Lingyun Sun

dblp:55/4732 · DBLP profile ↗
← Back
128ranked-venue papers
9as first author
117since 2021 · last 2026
0000-0002-5561-0493ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Human-computer interaction and ubiquitous computing · 69 · 7 first-author · 65 since 2021Graphics, computer vision, multimedia, augmented reality and games · 37 · 1 first-author · 33 since 2021Artificial intelligence and machine learning · 24 · 24 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 12 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Diffusion Distillation with Direct Preference Optimization for Efficient 3D LiDAR Scene Completion
abstract
The slow sampling speed of diffusion models hinders their application in 3D LiDAR scene completion. To address this, we propose Distillation-DPO, a novel framework that accelerates sampling through score distillation while simultaneously enhancing generation quality via preference alignment. Distillation-DPO follows a three-step procedure. First, the student model generates paired completion scenes with different initial noises. Second, using LiDAR scene evaluation metrics as preference, we construct winning and losing sample pairs. Third, as our core innovation, Distillation-DPO optimizes the student model by exploiting the difference in score functions between the teacher and student models on the paired completion scenes. This operation performs variational score distillation of the student model but simultaneously encourages the distilled student to prefer the winning samples over the losing ones. Extensive experiments demonstrate that Distillation-DPO achieves higher-quality scene completion than state-of-the-art diffusion models, while accelerating sampling by over 5-fold. To our knowledge, our work is the first to integrate the preference learning principle of DPO into the distillation of diffusion models, offering a new framework of preference-aligned distillation.
Shengyuan Zhang, Zejian Li, Ling Yang 0006, Pei Chen 0005, Anyang Wei, Perry Pengyun Gu, Lingyun Sun
AAAI10
2026 ThinkPersona: Thinking with Persona Graphs for Faithful Individualized Role-Playing
abstract
Yichen Cai, Pei Chen, Jiayang Li, Jingya Guo, Zejian Li, Changyuan Yang, Lingyun Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yichen Cai 0005, Pei Chen 0005, Jiayang Li 0003, Jingya Guo, Zejian Li, Chang-yuan Yang, Lingyun Sun
ACL (1)7
2026 IEvoAgent: Evolving Conversational Agent based on User Implicit Feedback
abstract
Yichen Cai, Jiayang Li, Junyuan Qiu, Jingya Guo, Weitao You, Changyuan Yang, Lingyun Sun, Pei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yichen Cai 0005, Jiayang Li 0003, Junyuan Qiu, Jingya Guo, Weitao You, Chang-yuan Yang, Lingyun Sun, Pei Chen 0005
ACL (1)7
2026 From Experts to Bases: Orthogonal Subspace Mixture for Continual Multimodal Instruction Tuning
abstract
Multimodal Continual Instruction Tuning (MCIT) is essential for adapting Multimodal Large Language Models (MLLMs) to dynamic data streams, yet preventing catastrophic forgetting remains a major challenge.Existing parameter-efficient approaches often face a dilemma: fixed architectures suffer from knowledge interference, while dynamic strategies incur inefficient capacity expansion, limiting scalability.We propose MoBLoRA (Mixture-of-Bases LoRA), a novel framework for MCIT.Motivated by our geometric analysis revealing subspace redundancy across sequential tasks, MoBLoRA shifts the paradigm from expert selection to subspace mixing: it decomposes adaptation weights into a globally shared pool of orthonormal bases to capture task-invariant knowledge, and lightweight mixing matrices to encode task-specific variations.This design effectively decouples knowledge accumulation from task reconstruction.Experiments on standard benchmarks show MoBLoRA significantly outperforms state-of-the-art methods while maintaining superior parameter efficiency. 1
Pei Chen 0005, Xilai Wang, Qixu Shi, Zejian Li, Lingyun Sun
ACL (1)5
2026 WeavePrint: A Generative Method for Woven-like Additive Manufacturing Based on Parametric Weave Structures
abstract
This paper presents WeavePrint, a parametric and multi-material additive manufacturing method for woven-like structures. By fusing traditional weaving logic with computational generation, WeavePrint overcomes limitations in pattern programmability, mechanical tunability, and build size. A parametric generator creates plain, twill, satin, and image-based jacquard patterns, while supporting curved-surface mapping and continuous vertical roll-to-roll printing for scalable production. Systematic tensile and compression tests quantify how overlap length, filament width, and multi-material combinations influence inter-layer adhesion and global mechanics. We define four motion primitives: bending, twisting, curved extension-contraction, and hinged extension-contraction, implemented through straight, diagonal, and curved weaves to produce predictable deformations. Demonstrations in wearable supports, robotic components, and rehabilitation devices highlight its broad potential in human-computer interaction. By unifying parametric modeling with multi-material continuous fabrication, WeavePrint provides a scalable route to programmable, anisotropic, and dynamically responsive interactive fabrics.
Jiacheng Cao, Zhaojia Yang, Manman Fan, Tianshu Dong, Jiaji Li, Lingyun Sun, Guanyun Wang
CHI8
2026 Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to Eye
abstract
Despite advances in multimodal AI, current vision-based assistants often remain inefficient in collaborative tasks. We identify two key gulfs: a communication gulf, where users must translate rich parallel intentions into verbal commands due to the channel mismatch, and an understanding gulf, where AI struggles to interpret subtle embodied cues. To address these, we propose Eye2Eye, a framework that leverages first-person perspective as a channel for human-AI cognitive alignment. It integrates three components: (1) joint attention coordination for fluid focus alignment, (2) revisable memory to maintain evolving common ground, and (3) reflective feedback allowing users to clarify and refine AI’s understanding. We implement this framework in an AR prototype and evaluate it through a user study and a post-hoc pipeline evaluation. Results show that Eye2Eye significantly reduces task completion time and interaction load while increasing trust, demonstrating its components work in concert to improve collaboration.
Zhuyu Teng, Pei Chen 0005, Yichen Cai 0005, Ruoqing Lu, Zhaoqu Jiang, Jiayang Li 0003, Weitao You, Lingyun Sun
CHI8
2026 PoemPalette: Facilitating Poetry Creative Exploration and Foundational Understanding through the Ideorealm Alignment of Paintings and Poems
abstract
The “Ideorealm Alignment of Paintings and Poems (IA-PP)” theory rooted in Chinese classical aesthetics offers a perspective for exploring poetry’s deep connotations. This study presents PoemPalette, a novel IA-PP creative-exploration tool that integrates Generative Artificial Intelligence (GAI) to guide poetry enthusiasts in actively constructing an ideorealm for the poetic painting they envision, informed by a formative study with six experts. We extract the core symbols of poetry, transform them into Scene Graph (SG), and generate images for users to freely compose, enabling IA-PP creative exploration. The system incorporates Large Language Model (LLM) agents to enhance the foundational understanding of poetry. In a controlled experiment on Chinese poetry and Japanese haiku with 60 participants, we analyze which interaction mechanisms most contribute to foundational understanding and creative outcomes, compared with both AI and non-AI baselines. Situated within East Asian poetry traditions, this study introduces cultural theories to guide the design of AI co-creation tools, using a graph-based interface of interpretable intermediate representations.
Ying Zhang 0076, Kaixin Jia, Hong Jian Zhang, Kewen Zhu, Chenye Meng, Jiesi Zhang, Zejian Li, Pei Chen 0005, Lingyun Sun
CHI9
2026 ImmersiProtor: A Collaborative Mixed-Prototype Tool Integrating Spatial Augmented Reality and Component-layered Generation
abstract
Conceptual design is a critical stage in product development, which is a co-design process involving multidisciplinary collaboration based on prototypes. In this paper, we aim to propose a novel prototype paradigm that combines the distinct strengths of generative artificial intelligence (GAI) and spatial augmented reality (SAR), leveraging the expressive potential of SAR and the creative potential of GAI for co-design. To achieve this, we initially conducted a formative study with designers to explore how these technologies could be effectively combined to facilitate co-design. Based on our findings, we introduce ImmersiProtor, a prototype tool integrating multi-view SAR and component-layered GAI for co-design. On one hand, ImmersiProtor allows design team members to freely create and modify physical prototypes while automatically generating multi-view and high-fidelity renderings that are projected onto the surfaces of the physical prototype using SAR technology, enabling immersive communication and intuitive evaluation. On the other hand, ImmersiProtor introduces a component-layered generation and collaboration mode, offering both personal and shared team component resources. It ensures that individual team members can explore ideas independently without interference, while also supporting concept integration, evaluation, and iteration. We implemented ImmersiProtor, which involves a web-based application and an SAR design space. We conducted a user study to verify ImmersiProtor’s usability in supporting prototype and collaboration. Our results highlighted ImmersiProtor’s inherent strengths in enhancing intuition, promoting collaboration, and strengthening GAI controllability. We also explored the effect of mixed interaction on design and critically discuss its best practices for the HCI community.
Chaoyi Lin, Weitao You, Lingyun Sun, Pei Chen 0005
CHI5
2026 3DInkGen: Extending Traditional Ink-Painting Artistry with Generative 3D Creation for Novices
abstract
Ink painting, renowned for its aesthetics and historical significance, plays a vital role in global art. Further, 3D ink art extends this tradition into spatial forms, enriching digital media like animation and games. However, existing methods for 3D ink creation demand expertise in both 3D modeling and ink aesthetics, limiting novice participation and 3D ink application. Through formative research with four experts, including ink painting artists and 3D designers, we summarize the core challenge: how to preserve the expressive pattern of ink paintings while constructing 3D structures. To tackle this challenge, we introduce 3DInkGen, a system transforming 2D ink elements into editable 3D compositions. 3DInkGen follows a four-stage workflow: element extraction, form generation, 3D reconstruction, and style transfer. A user study with sixteen novices showed 3DInkGen lowers technical barriers and enables intuitive 3D composition. The four experts believe novice-created works captured the artistic style of ink painting while maintain 3D structure of elements.
Jiesi Zhang, Ying Zhang 0076, Zejian Li, Changle Xie, Huanghuang Deng, Lingyun Sun
CHI6
2026 Cross-ancestry information transfer framework improves protein abundance prediction and protein-trait association identification
abstract
Genetics-informed proteome-wide association studies (PWASs) provide an effective way to uncover proteomic mechanisms underlying complex diseases. PWAS relies on an ancestry-matched reference panel to model the impact of genetically determined protein expression on phenotype. However, reference panels from underrepresented populations remain relatively limited. We developed a multi-ancestry framework to enhance protein prediction in these populations by integrating diverse information-sharing strategies into a Multi-Ancestry Best-performing Model (MABM). Results indicated that MABM increased the prediction performance with higher performance observed in both cross-validation and an external dataset. Leveraging the Biobank Japan, we identified three times as many significant PWAS associations using MABM as using Lasso model. Notably, 47.5% of the MABM specific associations were reproduced in independent East Asian datasets with concordant effect sizes. Furthermore, MABM enhanced decision-making in gene/protein prioritization for functional validation for complex traits by validating well-established associations and uncovering novel trait-related candidates. The benefits of MABM were further validated in additional ancestries and demonstrated in brain tissue-based PWAS, underscoring its broad applicability. Our findings close critical gaps in multi-omics research among underrepresented populations and facilitate trait-relevant protein discovery in underrepresented populations.
Wenli Zhai, Lingyun Sun, Wenwei Fang, Yidan Dong, Chunxiao Cheng, Yuanjiao Liu, Jiadong Ji, An Pan, Eric R. Gamazon, Xiong-Fei Pan
Briefings Bioinform.2
2026 CoRemix: Supporting Online Learning in Scratch Community with Visual Flowchart and Generative AI
abstract
Online programming communities give novices places to explore computing through user-generated projects, but limited structure can hinder a steadily challenging learning path. Beginners often struggle to interpret key events and relationships in projects, connect them to core concepts, and remix practices. We present CoRemix, a generative-AI community support system that uses visual flowcharts to clarify project logic. CoRemix introduces a prompting pipeline paired with a visual-textual scaffold that guides learners in constructing flowcharts. We further incorporate static project analysis and retrieval-augmented generation (RAG) to raise the precision of large-language-model outputs. In technical evaluations, static analysis and RAG improved response quality. In a user study, CoRemix outperformed a baseline in helping learners understand complex projects, strengthen computing-concept skills, and report better learning experiences within online communities. These gains include clearer event sequencing, improved identification of relationships across sprites and scripts, stronger remix strategies, and higher perceived scaffolding for progressive challenge.
Yunnong Chen, Yishu Shen, Ruiyi Liu, Lingyun Sun, Liuqing Chen 0002
Int. J. Hum. Comput. Interact.5
2026 ULMGNN: Fragmented layer grouping in GUI designs through graph learning based on multimodal information
Yunnong Chen, Shuhong Xiao, Jiazhi Li 0002, Lingyun Sun, Liuqing Chen 0002
Neurocomputing5
2026 Visionary Co-Driver: Enhancing Driver Perception of Potential Risks With LLM and HUD
abstract
Drivers’ perception of risky situations has always been a challenge in driving. Existing risk-detection methods excel at identifying collisions but face challenges in assessing the behavior of road users in non-collision situations. This paper introduces Visionary Co-Driver, a system that leverages large language models (LLMs) to identify non-collision roadside risks and alert drivers based on their eye movements. Specifically, the system combines video processing algorithms and LLMs to identify potentially risky road users. These risks are dynamically indicated on an adaptive heads-up display interface to enhance drivers’ attention. A user study with 41 drivers confirms that Visionary Co-Driver improves drivers’ risk perception and supports their recognition of roadside risks.
Wei Xiang 0008, Ziyue Lei, Lingyun Sun
IEEE Trans. Intell. Transp. Syst.8
2025 Integrating Sequence and Image Modeling in Irregular Medical Time Series Through Self-Supervised Learning
abstract
Medical time series are often irregular and face significant missingness, posing challenges for data analysis and clinical decision-making. Existing methods typically adopt a single modeling perspective, either treating series data as sequences or transforming them into image representations for further classification. In this paper, we propose a joint learning framework that incorporates both sequence and image representations. We also design three self-supervised learning strategies to facilitate the fusion of sequence and image representations, capturing a more generalizable joint representation. The results indicate that our approach outperforms seven other state-of-the-art models in three representative real-world clinical datasets. We further validate our approach by simulating two major types of real-world missingness through leave-sensors-out and leave-samples-out techniques. The results demonstrate that our approach is more robust and significantly surpasses other baselines in terms of classification performance.
Liuqing Chen 0002, Shuhong Xiao, Shixian Ding, Shanhai Hu, Lingyun Sun
AAAI5
2025 Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-Learning
abstract
Dynamic Music Emotion Recognition (DMER) aims to predict the emotion of different moments in music, playing a crucial role in music information retrieval. The existing DMER methods struggle to capture long-term dependencies when dealing with sequence data, which limits their performance. Furthermore, these methods often overlook the influence of individual differences on emotion perception, even though everyone has their own personalized emotional perception in the real world. Motivated by these issues, we explore more effective sequence processing methods and introduce the Personalized DMER (PDMER) problem, which requires models to predict emotions that align with personalized perception. Specifically, we propose a Dual-Scale Attention-Based Meta-Learning (DSAML) method. This method fuses features from a dual-scale feature extractor and captures both short and long-term dependencies using a dual-scale attention transformer, improving the performance in traditional DMER. To achieve PDMER, we design a novel task construction strategy that divides tasks by annotators. Samples in a task are annotated by the same annotator, ensuring consistent perception. Leveraging this strategy alongside meta-learning, DSAML can predict personalized perception of emotions with just one personalized annotation sample. Our objective and subjective experiments demonstrate that our method can achieve state-of-the-art performance in both traditional DMER and PDMER.
Dengming Zhang, Weitao You, Lingyun Sun, Pei Chen 0005
AAAI4
2025 GVMGen: A General Video-to-Music Generation Model with Hierarchical Attentions
abstract
Composing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative diversity, primarily due to inadequate feature alignment methods and insufficient datasets. In this study, we present General Video-to-Music Generation model (GVMGen), designed for generating high-related music to the video input. Our model employs hierarchical attentions to extract and align video features with music in both spatial and temporal dimensions, ensuring the preservation of pertinent features while minimizing redundancy. Remarkably, our method is versatile, capable of generating multi-style music from different video inputs, even in zero-shot scenarios. We also propose an evaluation model along with two novel objective metrics for assessing video-music alignment. Additionally, we have compiled a large-scale dataset comprising diverse types of video-music pairs. Experimental results demonstrate that GVMGen surpasses previous models in terms of music-video correspondence, music quality generative diversity, and application universality.
Heda Zuo, Weitao You, Junxian Wu 0003, Shihong Ren, Pei Chen 0005, Mingxu Zhou, Yujia Lu, Lingyun Sun
AAAI8
2025 CoExploreDS: Framing and Advancing Collaborative Design Space Exploration Between Human and AI
Pei Chen 0005, Zhuoyi Cheng, Yichen Cai 0005, Jiayang Li 0003, Weitao You, Lingyun Sun
CHI7
2025 I-Card: A Generative AI-Supported Intelligent Design Method Card Deck
abstract
A design method card deck helps designers understand and provoke thinking by presenting each method in a simple format and allow designers to switch between methods seamlessly by maintaining the same simple format across the deck. However, recent observations have shown designers hesitate to use a card deck due to the lack of support, while other tools have provided identified support with generative AI. Through a formative study, we identified the specific support designers need when applying the design method cards and intentions in integrating generative AI. Accordingly, we developed the intelligent design method card deck, I-Card, which integrates generative AI to provide applicable design methods, design knowledge and data support, and interactive and dynamic support. A user study demonstrates that I-Card improved the design efficiency and applicability by offering personalized guidance, enhanced decision-making with comprehensive data generation and provided more design inspiration via interactive support.
Liuqing Chen 0002, Wengteng Cheang, Zhaojun Jiang, Yuan Xu 0027, Zebin Cai, Lingyun Sun, Peter R. N. Childs, Preben Hansen, Haoyu Zuo
CHI6
2025 Voice by the Non-sighted: Practices and Challenges of Audiobook Voice Actors with Blind and Low Vision in China
Shi Chen 0005, Jingao Zhang, Suqi Lou, Wei Xiang 0008, Lingyun Sun
CHI6
2025 TH-Wood: Developing Thermo-Hygro-Coordinating Driven Wood Actuators to Enhance Human-Nature Interaction
abstract
CHI ’25, Yokohama, Japan
Guanyun Wang, Yangweizhe Zheng, Qianzi Zhen, Yang Zhang 0041, Jiaji Li, Yue Yang 0005, Ye Tao 0001, Shijian Luo, Lingyun Sun
CHI12
2025 FusionProtor: A Mixed-Prototype Tool for Component-level Physical-to-Virtual 3D Transition and Simulation
Pei Chen 0005, Xuelong Xie, Zhaoqu Jiang, Zejian Li, Lingyun Sun
CHI8
2025 IEDS: Exploring an Intelli-Embodied Design Space Combining Designer, AR, and GAI to Support Industrial Conceptual Design
Pei Chen 0005, Zhaoqu Jiang, Xuelong Xie, Weitao You, Lingyun Sun
CHI8
2025 Ink Restorer: Virtual Restoration of Ancient Chinese Paintings Inheriting Traditional Restoration Processes
Ying Zhang 0076, Zejian Li, Jiesi Zhang, Kewen Zhu, Qi Liu 0076, Huanghuang Deng, Lingyun Sun
CHI9
2025 Distilling Diffusion Models to Efficient 3D LiDAR Scene Completion
abstract
Diffusion models have been applied to 3D LiDAR scene completion due to their strong training stability and high completion quality. However, the slow sampling speed limits the practical application of diffusion-based scene completion models since autonomous vehicles require an efficient perception of surrounding environments. This paper proposes a novel distillation method tailored for 3D Li- DAR scene completion models, dubbed ScoreLiDAR, which achieves efficient yet high-quality scene completion. Score- LiDAR enables the distilled model to sample in significantly fewer steps after distillation. To improve completion quality, we also introduce a novel Structural Loss, which encourages the distilled model to capture the geometric structure of the 3D LiDAR scene. The loss contains a scene-wise term constraining the holistic structure and a point-wise term constraining the key landmark points and their relative configuration. Extensive experiments demonstrate that ScoreLiDAR significantly accelerates the completion time from 30.55 to 5.37 seconds per frame (>5x) on SemanticKITTI and achieves superior performance compared to state-of-the-art 3D LiDAR scene completion models. Our model and code are publicly available on https://github.com/happyw1nd/ScoreLiDAR.
Shengyuan Zhang, Ling Yang 0006, Zejian Li, Chenye Meng, Tianrun Chen, Anyang Wei, Perry Pengyun Gu, Lingyun Sun
ICCV10
2025 Distribution Backtracking Builds A Faster Convergence Trajectory for Diffusion Distillation
abstract
Accelerating the sampling speed of diffusion models remains a significant challenge. Recent score distillation methods distill a heavy teacher model into a student generator to achieve one-step generation, which is optimized by calculating the difference between two score functions on the samples generated by the student model. However, there is a score mismatch issue in the early stage of the score distillation process, since existing methods mainly focus on using the endpoint of pre-trained diffusion models as teacher models, overlooking the importance of the convergence trajectory between the student generator and the teacher model. To address this issue, we extend the score distillation process by introducing the entire convergence trajectory of the teacher model and propose $\textbf{Dis}$tribution $\textbf{Back}$tracking Distillation ($\textbf{DisBack}$). DisBask is composed of two stages: $\textit{Degradation Recording}$ and $\textit{Distribution Backtracking}$. $\textit{Degradation Recording}$ is designed to obtain the convergence trajectory by recording the degradation path from the pre-trained teacher model to the untrained student generator. The degradation path implicitly represents the intermediate distributions between the teacher and the student, and its reverse can be viewed as the convergence trajectory from the student generator to the teacher model. Then $\textit{Distribution Backtracking}$ trains the student generator to backtrack the intermediate distributions along the path to approximate the convergence trajectory of the teacher model. Extensive experiments show that DisBack achieves faster and better convergence than the existing distillation method and achieves comparable or better generation performance, with an FID score of 1.38 on the ImageNet 64$\times$64 dataset. DisBack is easy to implement and can be generalized to existing distillation methods to boost performance.
Shengyuan Zhang, Ling Yang 0006, Zejian Li, Chenye Meng, Chang-yuan Yang, Guang Yang 0022, Lingyun Sun
ICLR9
2025 Hand by Hand: LLM Driving EMS Assistant for Operational Skill Learning
abstract
Operational skill learning, inherently physical and reliant on hands-on practice and kinesthetic feedback, has yet to be effectively replicated in large language model (LLM)-supported training. Current LLM training assistants primarily generate customized textual feedback, neglecting the crucial kinesthetic modality. This gap derives from the textual and uncertain nature of LLMs, compounded by concerns on user acceptance of LLM driven body control. To bridge this gap and realize the potential of collaborative human-LLM action, this work explores human experience of LLM driven kinesthetic assistance. Specifically, we introduced an "Align-Analyze-Adjust" strategy and developed FlightAxis, a tool that integrates LLM with Electrical Muscle Stimulation (EMS) for flight skill acquisition, a representative operational skill domain. FlightAxis learns flight skills from manuals and guides forearm movements during simulated flight tasks. Our results demonstrate high user acceptance of LLM-mediated body control and significantly reduced task completion times. Crucially, trainees reported that this kinesthetic assistance enhanced their awareness of operation flaws and fostered increased engagement in the training process, rather than relieving perceived load. This work demonstrated the potential of kinesthetic LLM training in operational skill acquisition.
Wei Xiang 0008, Ziyue Lei, Haoyuan Che, Fangyuan Ye, Xueting Wu, Lingyun Sun
IJCAI6
2025 Intoner: For Chinese Poetry Intoning Synthesis
abstract
Chinese Poetry Intoning, with improvised melodies devoid of fixed musical scores, is crucial for emotional expression and prosodic rendition. However, this cultural heritage faces challenges in propagation due to scant audio records and a scarcity of domain experts. Existing text-to-speech models lack the ability to generate melodious audio, while singing-voice-synthesis models rely on predetermined musical scores, which are all unsuitable for intoning synthesis. Hence, we introduce Chinese Poetry Intoning Synthesis (PIS) as a novel task to reproduce intoning audio and preserve this age-old cultural art. Corresponding to this task, we summarize three-level principles from poetry metrical patterns and construct a diffusion PIS model Intoner based on them. We also collect a multi-style Chinese poetry intoning dataset of text-audio pairs accompanied by feature annotations. Experimental results show that our model effectively learns diverse intoning styles and contents which can synthesize more melodious and vibrant intoning audio. To the best of our knowledge, we are the first to work on poetry intoning synthesis task.
Heda Zuo, Liyao Sun, Zeyu Lai, Weitao You, Pei Chen 0005, Lingyun Sun
IJCAI6
2025 An Exploratory Study on How AI Awareness Impacts Human-AI Design Collaboration
Zhuoyi Cheng, Pei Chen 0005, Wenzheng Song, Zhuoshu Li, Lingyun Sun
IUI6
2025 Controllable Video-to-Music Generation with Multiple Time-Varying Conditions
abstract
Music enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations.
Junxian Wu 0003, Weitao You, Heda Zuo, Dengming Zhang, Pei Chen 0005, Lingyun Sun
ACM Multimedia6
2025 Inversion-DPO: Precise and Efficient Post-Training for Diffusion Models
abstract
Recent advancements in diffusion models (DMs) have been propelled by alignment methods that post-train models to better conform to human preferences. However, these approaches typically require computation-intensive training of a base model and a reward model, which not only incurs substantial computational overhead but may also compromise model accuracy and training efficiency. To address these limitations, we propose Inversion-DPO, a novel alignment framework that circumvents reward modeling by reformulating Direct Preference Optimization (DPO) with DDIM inversion for DMs. Our method conducts intractable posterior sampling in Diffusion-DPO with the deterministic inversion from winning and losing samples to noise and thus derive a new post-training paradigm. This paradigm eliminates the need for auxiliary reward models or inaccurate appromixation, significantly enhancing both precision and efficiency of training. We apply Inversion-DPO to a basic task of text-to-image generation and a challenging task of compositional image generation. Extensive experiments show substantial performance improvements achieved by Inversion-DPO compared to existing post-training methods and highlight the ability of the trained generative models to generate high-fidelity compositionally coherent images. For the post-training of compostitional image geneation, we curate a paired dataset consisting of 11,140 images with complex structural annotations and comprehensive scores, designed to enhance the compositional capabilities of generative models. Inversion-DPO explores a new avenue for efficient, high-precision alignment in diffusion models, advancing their applicability to complex realistic generation tasks. Our code is available at https://github.com/MIGHTYEZ/Inversion-DPO
Zejian Li, Yize Li 0001, Chenye Meng, Zhongni Liu, Ling Yang 0006, Shengyuan Zhang, Guang Yang 0022, Chang-yuan Yang, Lingyun Sun
ACM Multimedia10
2025 Spatial-Temporal Decomposition and Alignment in Controllable Video-to-Music Generation
abstract
Achieving high-quality output alongside enhanced controllability is crucial in video-to-music generation, especially for optimizing user experience in real-life application scenarios. Most existing studies emphasize generative quality, but often overlooking the vital aspect of controllability. Therefore, the generated music cannot be easily fine-tuned or modified to meet users' expectations. In this paper, we delve into the spatial-temporal decomposition and alignment in controllable video-to-music generation. We first introduce a novel video-music decomposition and transformation approach in both spatial and temporal domain, and enhance the cross-modal correspondence through feature alignment and flow-matching based alignment. Furthermore, our method attains unsupervised controllability during training via feature-free guidance. Experimental results demonstrate that our model achieves state-of-the-art results in overall generative quality. Moreover, its controllability significantly outperforms existing models, making it exceptionally well-suited to accommodate users' flexible and diverse control requirements.
Weitao You, Heda Zuo, Junxian Wu 0003, Dengming Zhang, Zhibin Zhou 0002, Lingyun Sun
ACM Multimedia6
2025 Driver Assistant: Persuading Drivers to Adjust Secondary Tasks Using Large Language Models
abstract
Level 3 automated driving systems allows drivers to engage in secondary tasks while diminishing their perception of risk. In the event of an emergency necessitating driver intervention, the system will alert the driver with a limited window for reaction and imposing a substantial cognitive burden. To address this challenge, this study employs a Large Language Model (LLM) to assist drivers in maintaining an appropriate attention on road conditions through a "humanized" persuasive advice. Our tool leverages the road conditions encountered by Level 3 systems as triggers, proactively steering driver behavior via both visual and auditory routes. Empirical study indicates that our tool is effective in sustaining driver attention with reduced cognitive load and coordinating secondary tasks with takeover behavior. Our work provides insights into the potential of using LLMs to support drivers during multi-task automated driving.
Wei Xiang 0008, Muchen Li, Manling Zheng, Hanfei Zhu, Mengyun Jiang, Lingyun Sun
SMC7
2025 KiriInflate: Fabricating Cross-Scale Inflatables with Large Contraction and Tunable Stretchability for Tangible Interaction
Yue Yang 0005, Bolan Yao, Lingyun Sun, Ye Tao 0001, Guanyun Wang
UIST7
2025 SCENIC: A Location-based System to Foster Cognitive Development in Children During Car Rides
Liuqing Chen 0002, Yaxuan Song, Ke Lyu, Shuhong Xiao, Yilang Shen, Lingyun Sun
UIST6
2025 Touch-n-Curl: Designing and Constructing Skeletal Form through 3D Printing Flattened Zipper Assembly
Deying Pan, Fanqi Zhou, Yitao Fan, Tianshu Dong, Fanke Qi, Yongbo Ni, Ye Tao 0001, Lingyun Sun, Guanyun Wang
UIST8
2025 "This is My Fault", Really? Understanding Blind and Low-Vision People's Perception of Hallucination in Large Vision Language Models
Yilin Tang, Yuyang Fang, Tianle Wang 0013, Lingyun Sun, Liuqing Chen 0002
UIST4
2025 MyWay: a 3D and audio-enhanced transportation learning kit for the visually impaired teenagers
Qionghui Cai, Kuangqi Zhu, Chenyi Dai, Lingyun Sun, Ye Tao 0001, Guanyun Wang
CCF Trans. Pervasive Comput. Interact.7
2025 GPSdesign: Integrating Generative AI with Problem-Solution Co-Evolution Network to Support Product Conceptual Design
abstract
In conceptual design, designers often face the challenge of navigating vast design spaces to define ambiguous problems and generate feasible solutions. Recent advancements in generative artificial intelligence (GenAI) offer new opportunities to support this process. However, formative research revealed that designers struggle to simultaneously advance both problem and solution spaces when using GenAI in conceptual design, leading to increased communication load and diminished solution practicality. This study explores the integration of GenAI with the problem-solution co-evolution model to facilitate the construction of a structured design space. We propose a GenAI-supported method for expanding and evaluating the design space and developed the GPSdesign system based on this method. Compared with a baseline system, GPSdesign fosters greater design space divergence, retrospection, and structured construction, while improving design efficiency and solution quality.
Pei Chen 0005, Yexinrui Wu, Zhuoshu Li, Mingxu Zhou, Weitao You, Lingyun Sun
Int. J. Hum. Comput. Interact.8
2025 MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AI
abstract
Programming is essential in K-12 education and fosters computational thinking skills. Given the complexity of programming and the advanced skills it requires, previous research has introduced user-friendly tools to support young learners. However, our interviews with six programming educators revealed that current tools often fail to reflect classroom learning objectives, offer flexible guidance, and foster creativity. Therefore, we introduced MindScratch, a multimodal generative AI (GAI)-powered visual programming support tool. MindScratch aims to balance structured classroom activities with free programming creation, supporting students in completing creative programming projects based on teacher-set learning objectives while also providing programming scaffolding. The results indicate that, compared to the baseline, MindScratch more effectively helps students achieve high-quality projects aligned with learning objectives. It also enhances students’ computational thinking and thinking. Overall, we believe that GAI-driven educational tools like MindScratch offer students a focused and engaging learning experience.
Yunnong Chen, Shuhong Xiao, Yaxuan Song, Zejian Li, Lingyun Sun, Liuqing Chen 0002
Int. J. Hum. Comput. Interact.5
2025 RealtimeGen: An Intervenable AI Image Generation System for Commercial Digital Art Asset Creators
abstract
Recent advances in artificial intelligence-generated content (AIGC) have led to the rapid generation of high-quality images. AIGC has attracted the attention of commercial digital art asset creators. Traditional artist-led processes contrast with current AI tools that often reduce creators to passive roles. This study examines the integration of AI image generation into commercial digital art, emphasizing the importance of preserving creators’ creative autonomy. Our formative study (S1) involved interviews with commercial digital art creators, highlighting a need for greater control and transparency in AI-assisted painting. In response, we developed RealtimeGen, an integrated tool that merges human creativity with AI’s capabilities, allowing creators to intervene in the generative process. A user study (S2) comparing RealtimeGen with the popular AIGC tool Stable Diffusion was also carried out. The results showed its enhanced user experience and workflow compatibility. Our work contributes to understanding and improving AI-assisted painting workflows for commercial creators, offering them greater creative agency.
Zejian Li, Ying Zhang 0076, Shengzhe Zhou, Qi Liu 0076, Jiesi Zhang, Shuyao Chen, Lingyun Sun
Int. J. Hum. Comput. Interact.9
2025 Play With Morphing Food: Supporting Children-Food Interaction With an Interactive Cooking Toolkit
abstract
To support children’s food interaction and enhance their understanding of food through morphing food technology, we develop a design exploration through the Research through Design (RtD) methodology. Our exploration integrates four stages: (1) defining design objectives through empathy with stakeholders, (2) investigating morphing food materials to understand their deformation mechanisms, (3) designing and iteratively developing tools based on user feedback, and (4) conducting a workshop-based evaluation. Our design outcome is a toolkit, comprising a morphing food library, trigger tools, and instructional interfaces. The workshop showed that through interaction with morphing food, children learned not only scientific principles but also developed culinary skills, as well as the diversity of food forms and functions. We discussed the detailed findings, insights, and implications for future design.
Guanyun Wang, Yilin Shao, Boyu Feng, Mengge Wang, Xiaojing Zhou, Zhengke Li, Yue Yang 0005, Kuangqi Zhu, Yanan Wang 0005, Lingyun Sun, Ye Tao 0001
Int. J. Hum. Comput. Interact.11
2025 VAEnvGen: A Real-Time Virtual Agent Environment Generation System Based on Large Language Models
abstract
Environment plays an important role in non-verbal communication for human-virtual agent interaction. Existing research explores the influence of an agent’s appearance and attributes to enhance human-virtual agent communication. However, there is no common practice for dynamically adjusting the surrounding environments of the virtual agent. In this paper, we introduce a real-time virtual agent environment generation system (VAEnvGen), which contributes to the field by enhancing users’ content perception and improving task performance through dynamic environment adjustment. The system dynamically analyzes both the appropriate communication environment and filters the key information according to the current context. Leveraging Large Language Models, it generates a pseudo-3D background space to create an engaging atmosphere and a dynamic foreground content space for vivid key information display, thereby significantly enhancing content perception. For widespread adoption and flexibility, VAEnvGen is developed as a web application. We further evaluate the impact of VAEnvGen on content perception, user attention, and subjective satisfaction through a mixed-design user study with 50 participants. Quantitative and qualitative results reveal significant improvements in content perception, task completion time, and user satisfaction when using VAEnvGen. The system effectively redistributes user attention from subtitles and the virtual agent itself to the dynamically generated background and key foreground information, leading to a more immersive and less fatiguing user experience.
Jingyu Wu, Pengchen Chen, Shi Chen 0005, Wei Xiang 0008, Lingyun Sun
Int. J. Hum. Comput. Interact.5
2025 Using a Configurational Approach to Examine the Impacts of Vehicle Appearance Perception on Pedestrian Acceptance of the External Human-Machine Interfaces on Autonomous Vehicles
abstract
The interaction between autonomous vehicles (AVs) and pedestrians has gained significant attention, leading to the exploration of external human-machine interfaces (eHMIs) equipped on AVs to facilitate effective communication. While existing research suggests that perceptions of vehicle appearances may influence interactions between pedestrians and AVs, a comprehensive study on the eHMIs related to AV appearance remains lacking. Therefore, we conducted a virtual reality (VR) experiment to investigate how AV appearances affect pedestrians’ acceptance regarding Awareness, Intent, and Harmony during interactions with various eHMIs. Leveraging the fuzzy set qualitative comparative analysis (fsQCA) method, we identified specific combinations of AV appearances and eHMIs that yield either high or low-performance interactions. For example, our findings reveal that text displays exhibit high performance in terms of awareness on AVs that are aggressive and ordinary. Furthermore, we distilled design guidelines to provide actionable suggestions for the design of eHMIs, fostering the acceptance of AVs among pedestrians.
Zhibin Zhou 0002, Yitao Fan, Wenan Li, Hao Jiang 0046, Weitao You, Lingyun Sun
Int. J. Hum. Comput. Interact.6
2025 Image generation evaluation: a comprehensive survey of human and automatic evaluations
abstract
Image generation models have made remarkable progress, and image evaluation is crucial for explaining and driving the development of these models. Previous studies have extensively explored human and automatic evaluations of image generation. Herein, these studies are comprehensively surveyed, specifically for two main parts: evaluation protocols and evaluation methods. First, 10 image generation tasks are summarized with focus on their differences in evaluation aspects. Based on this, a novel protocol is proposed to cover human and automatic evaluation aspects required for various image generation tasks. Second, the review of automatic evaluation methods in the past five years is highlighted. To our knowledge, this paper presents the first comprehensive summary of human evaluation, encompassing evaluation methods, tools, details, and data analysis methods. Finally, the challenges and potential directions for image generation evaluation are discussed. We hope that this survey will help researchers develop a systematic understanding of image generation evaluation, stay updated with the latest advancements in the field, and encourage further research.
Qi Liu 0076, Shuanglin Yang, Zejian Li, Lefan Hou, Chenye Meng, Ying Zhang 0076, Lingyun Sun
Frontiers Inf. Technol. Electron. Eng.7
2025 Exploring the role of Mixed Reality on Design Representations to Enhance User-Involved Co-Design Communication
abstract
As users transition from passive subjects to active partners in the co-design process, they bring unique insights based on their experiences, collaboratively envisioning a better future with designers. However, unlike designers who are adept at various forms of representation, most users lack advanced modeling or sketching skills to concretely present the three-dimensional (3D) forms or dynamic features of a design proposal. This hinders user expression and increases the cognitive load on designers, thereby reducing communication efficiency in the co-design process. Mixed Reality (MR) technology enables users to depict 3D information in real physical space using natural gestures. This means that MR can provide a low-learning-cost concrete expression method without compromising traditional communication methods. This study explores the role of MR in enhancing communication between designers and users during the early stages of design. A formative study was conducted to identify four key requirements, which informed the development of the DuoMR system. DuoMR supports designers and users in expressing design ideas through gesture modeling in a collaborative MR space. Results from the user study and practical case study show that DuoMR effectively reduces cognitive load and enhances mutual understanding during the co-design process.
Pei Chen 0005, Kexing Wang, Lianyan Liu, Xuanhui Liu, Zhuyu Teng, Lingyun Sun
Proc. ACM Hum. Comput. Interact.7
2025 Img2CAD: Conditioned 3-D CAD Model Generation From Single Image With Structured Visual Geometry
abstract
In this article, we propose Img2CAD, the first approach to our knowledge that uses 2-D image inputs to generate computer-aided design (CAD) models with editable parameters. Unlike existing artificial intelligence (AI) methods for 3-D model generation using text or image inputs often rely on mesh-based representations, which are incompatible with CAD tools and lack editability and fine control, Img2CAD enables seamless integration between AI-based 3-D reconstruction and CAD software. We have identified an innovative intermediate representation called structured visual geometry, characterized by vectorized wireframes extracted from objects. This representation significantly enhances the performance of generating conditioned CAD models. In addition, we introduce two new datasets to further support research in this area:a big cad model dataset (ABC)-mono, the largest known dataset comprising over 200 000 3-D CAD models with rendered images, andKOCAD, the first dataset featuring real-world captured objects alongside their ground truth CAD models, supporting further research in conditioned CAD model generation.
Tianrun Chen, Chunan Yu, Yuanqi Hu, Jing Li 0145, Tao Xu 0048, Runlong Cao, Lanyun Zhu, Ying Zang, Yong Zhang 0030, Zejian Li, Lingyun Sun
IEEE Trans. Ind. Informatics11
2025 DesignManager: An Agent-Powered Copilot for Designers to Integrate AI Design Tools into Creative Workflows
abstract
Creative design is an inherently complex and iterative process characterized by continuous exploration, evaluation, and refinement. While recent advances in generative AI have demonstrated remarkable potential in supporting specific design tasks, there remains a critical gap in understanding how these technologies can enhance the holistic design process rather than just isolated stages. This paper introduces DesignManager, a novel AI-powered design support system that aims to transform how designers collaborate with AI throughout their creative workflow. Through a formative study examining designers' current practices with generative AI, we identified key challenges and opportunities in integrating AI into the creative design process. Based on these insights, we developed DesignManager as an interactive copilot system that provides node-based visualization of design evolution, enabling designers to track, modify, and branch their design processes while maintaining meaningful dialogue-based collaboration. The system offers two collaboration modes: DesignManager-guiding and Designer-guiding. Designers can engage in conversational interactions with the DesignManager to obtain design inspiration and tool recommendations, and proactively advance the design progress. The system employs an agent framework to manage decoupled contextual information emerged during the design process, facilitating deep understanding of designers' needs and providing context-aware assistance. Our technical evaluation validated the effectiveness of context decoupling and the use of agent framework, while the open-ended user study with experts demonstrated that DesignManager successfully supports intuitive intention expression, flexible process control, and deeper creative articulation. This work contributes to the understanding of how AI can evolve from task-specific tools to collaborative partners in creative design processes.
Weitao You, Yinyu Lu, Zirui Ma, Nan Li 0085, Mingxu Zhou, Pei Chen 0005, Lingyun Sun
ACM Trans. Graph.8
2025 Measuring Human Perception of Airflow for Natural Motion Simulation in Virtual Reality
abstract
Airflow is recognized as an effective method for inducing the illusion of self-motion (vection) and reducing motion sickness in virtual reality. However, the quantitative relationship between virtual motion and the airflow perceived as consistent with it has not been fully explored. To address this gap, this study conducted three experiments. In Experiment 1, we carried out a series of cross-modal matching tasks to establish the relationship between the speed of virtual motion and the airflow speed perceived as consistent with it, revealing a strong linear correlation. In Experiment 2, we introduced the concept of an "Airflow Gradient" to simulate the bodily sensation of curvilinear motion and examined the relationship between the radius and angular velocity of the motion and the difference in airflow speed between the left and right sides. The results indicated a linear relationship between the radius and the left-right airflow speed difference, while the angular velocity showed a near-quadratic pattern, similar to the centripetal acceleration formula. Based on these findings, Experiment 3 developed a dynamic airflow scheme and compared it with constant airflow and no-airflow conditions during locomotion tasks in a complex urban environment. The results demonstrated that dynamic airflow, which ensures consistency between visual and bodily vection, further reduces motion sickness, enhances presence, and provides a more natural and consistent virtual motion experience.
Yu Cai 0014, Sanyi Jin, Daiwei Yang, Han Tu, Preben Hansen, Lingyun Sun, Liuqing Chen 0002
IEEE Trans. Vis. Comput. Graph.7
2024 KiPneu: Designing a Constructive Pneumatic Platform for Biomimicry Learning in STEAM Education
abstract
Biomimicry, a methodology adapted from nature, always inspires optimum solutions and innovative technologies in human history. To get children interested in, excited about, and inspired by biomimicry, we introduce KiPneu, a robotic platform that facilitates biomimicry education through hands-on, solution-oriented learning and a digital learning environment. KiPneu allows children to mimic flexible animal locomotion, like fish swimming or worm squirming, using low-cost building blocks and non-electrical pneumatic actuators. We provide five types of non-electrical tangible valves to adjust robot motion characteristics, such as direction and speed, through engaging tangible programming. Additionally, to facilitate the whole learning process, KiPneu comes with interactive instructional interface that visualize and simulate the pneumatic system. To validate KiPneu’s educational efficacy, we conducted a three-day workshop with 21 children aged 5-12. Pre-and-post surveys revealed KiPneu not only enhanced their understanding of animal locomotion mechanisms but also spurred interest in creative construction using acquired knowledge.
Guanyun Wang, Chenda Zheng, Yanbo Fu, Kuangqi Zhu, Fuyi Lai, Likang Zhang, Muyi Ren, Yanpei Zheng, Boyi Lian, Qi Wang 0075, Shijian Luo, Fangtian Ying, Lingyun Sun, Ye Tao 0001
Conference on Designing Interactive Systems17
2024 Reducing Spatial Fitting Error in Distillation of Denoising Diffusion Models
abstract
Denoising Diffusion models have exhibited remarkable capabilities in image generation. However, generating high-quality samples requires a large number of iterations. Knowledge distillation for diffusion models is an effective method to address this limitation with a shortened sampling process but causes degraded generative quality. Based on our analysis with bias-variance decomposition and experimental observations, we attribute the degradation to the spatial fitting error occurring in the training of both the teacher and student model in the distillation. Accordingly, we propose Spatial Fitting-Error Reduction Distillation model (SFERD). SFERD utilizes attention guidance from the teacher model and a designed semantic gradient predictor to reduce the student's fitting error. Empirically, our proposed model facilitates high-quality sample generation in a few function evaluations. We achieve an FID of 5.31 on CIFAR-10 and 9.39 on ImageNet 64x64 with only one step, outperforming existing diffusion methods. Our study provides a new perspective on diffusion distillation by highlighting the intrinsic denoising ability of models.
Shengzhe Zhou, Zejian Li, Shengyuan Zhang, Lefan Hou, Chang-yuan Yang, Guang Yang 0022, Lingyun Sun
AAAI8
2024 BIDTrainer: An LLMs-driven Education Tool for Enhancing the Understanding and Reasoning in Bio-inspired Design
abstract
Bio-inspired design (BID) fosters innovations in engineering. Learning BID is crucial for developing multidisciplinary innovation skills of designers and engineers. Current BID education aims to enhance learners’ understanding and analogical reasoning skills. However, it often heavily relies on the teachers’ expertise. When learners pursue independent learning using some educational tools, they face challenges in understanding and reasoning practice within this multidisciplinary field. Additionally, evaluating their learning outcomes comprehensively becomes problematic. Addressing these challenges, we introduce a LLMs-driven BID education method based on a structured ontology and three strategies: enhancing understanding through LLMs-enpowered "learning by asking", assisting reasoning by providing hints and feedback, and assessing learning outcomes through benchmarking against existing BID cases. Implementing the method, we developed BIDTrainer, a BID education tool. User studies indicate that learners using BIDTrainer understood BID knowledge better, reason faster with higher interactivity than the baseline, and BIDTrainer assessed the learning outcomes consistent with experts.
Liuqing Chen 0002, Zhaojun Jiang, Duowei Xia, Zebin Cai, Lingyun Sun, Peter R. N. Childs, Haoyu Zuo
CHI5
2024 ChatScratch: An AI-Augmented System Toward Autonomous Visual Programming Learning for Children Aged 6-12
abstract
As Computational Thinking (CT) continues to permeate younger age groups in K-12 education, established CT platforms such as Scratch face challenges in catering to these younger learners, particularly those in the elementary school (ages 6-12). Through formative investigation with Scratch experts, we uncover three key obstacles to children’s autonomous Scratch learning: artist’s block in project planning, bounded creativity in asset creation, and inadequate coding guidance during implementation. To address these barriers, we introduce ChatScratch, an AI-augmented system to facilitate autonomous programming learning for young children. ChatScratch employs structured interactive storyboards and visual cues to overcome artist’s block, integrates digital drawing and advanced image generation technologies to elevate creativity, and leverages Scratch-specialized Large Language Models (LLMs) for professional coding guidance. Our study shows that, compared to Scratch, ChatScratch efficiently fosters autonomous programming learning, and contributes to the creation of high-quality, personally meaningful Scratch projects for children.
Liuqing Chen 0002, Shuhong Xiao, Yunnong Chen, Yaxuan Song, Lingyun Sun
CHI6
2024 IntelliTex: Fabricating Low-cost and Washable Functional Textiles using A Double-coating Process
abstract
We present IntelliTex, a low-cost and highly accessible double-coating fabrication method for washable and reusable functional textiles with customized input functionalities. Specifically, off-the-shelf textiles are firstly coated with conductive carbon black using pen ink, which endows textiles with rich sensing capabilities, such as pressure, stretch, slide, and temperature. Secondly, textiles are coated with polyurethane to enhance the sensing stability over wash cycles for good reusability. To support user customization, we enrich the design space of double-coating by exploring various coating methods and diverse textiles to be coated. We further contribute a comprehensive library of input components and an online document to make our approach accessible to novice users. Finally, five application examples and a user study showcase the versatile functionalities and user accessibility of our method, with which we hope to support designers, makers, and researchers to easily create functional textiles ready to use in everyday life.
Yuecheng Peng, Danchang Yan, Yue Yang 0005, Ye Tao 0001, Lingyun Sun, Guanyun Wang
CHI7
2024 Touch-n-Go: Designing and Fabricating Touch Fastening Structures by FDM 3D Printing
abstract
Touch fastening structures are widely used to quickly assemble and disassemble an object with multiple parts. However, such structures are under-explored in the context of additive manufacturing for personal fabrication. We proposed Touch-n-Go, a method for designing touch-fastening structures with customizable mechanical properties such as holding capacities or shearing strength. Additionally, the customization of fastener patterns enables both static and dynamic connections, and the dynamic connections grant the freedom of rotation and translation. To facilitate the customization process, we developed a design tool that allows the integration of fastening structures on the surface of a 3D-printed object. Furthermore, we validated the fastening properties of Touch-n-Go through a series of experiments, and the result exhibits performances that match or even surpass off-the-shelf fasteners. Finally, we demonstrated the implementation of Touch-n-Go through a collection of applications.
Lingyun Sun, Deying Pan, Hongyi Hu, Junzhe Ji, Yue Tao, Shanghua Lou, Boyi Lian, Yitao Fan, Ye Tao 0001, Guanyun Wang
CHI1
2024 EmoEden: Applying Generative Artificial Intelligence to Emotional Learning for Children with High-Function Autism
abstract
Children with high-functioning autism (HFA) often face challenges in emotional recognition and expression, leading to emotional distress and social difficulties. Conversational agents developed for HFA children in previous studies show limitations in children's learning effectiveness due to the conversational agents’ inability to dynamically generate personalized and contextual content. Recent advanced generative Artificial Intelligence techniques, with the capability to generate substantial diverse and high-quality texts and visual content, offer an opportunity for personalized assistance in emotional learning for HFA children. Based on the findings of our formative study, we integrated large language models and text-to-image models to develop a tool named EmoEden supporting children with HFA. Over a 22-day study involving six HFA children, it is observed that EmoEden effectively engaged children and improved their emotional recognition and expression abilities. Additionally, we identified the advantages and potential risks of applying generative AI to assist HFA children in emotional learning.
Yilin Tang, Liuqing Chen 0002, Yu Cai 0014, Yao Du 0002, Lingyun Sun
CHI8
2024 SimUser: Generating Usability Feedback by Simulating Various Users Interacting with Mobile Applications
abstract
The conflict between the rapid iteration demand of prototyping and the time-consuming nature of user tests has led researchers to adopt AI methods to identify usability issues. However, these AI-driven methods concentrate on evaluating the feasibility of a system, while often overlooking the influence of specified user characteristics and usage contexts. Our work proposes a tool named SimUser based on large language models (LLMs) with the Chain-of-Thought structure and user modeling method. It generates usability feedback by simulating the interaction between users and applications, which is influenced by user characteristics and contextual factors. The empirical study (48 human users and 21 designers) validated that in the context of a simple smartwatch interface, SimUser could generate heuristic usability feedback with the similarity varying from 35.7% to 100% according to the user groups and usability category. Our work provides insights into simulating users by LLM to improve future design activities.
Wei Xiang 0008, Hanfei Zhu, Suqi Lou, Xinli Chen, Zhenghua Pan, Yuping Jin, Shi Chen 0005, Lingyun Sun
CHI8
2024 SnapInflatables: Designing Inflatables with Snap-through Instability for Responsive Interaction
abstract
Snap-through instability, like the rapid closure of the Venus flytrap, is gaining attention in robotics and HCI. It offers rapid shape reconfiguration, self-sensing, actuation, and enhanced haptic feedback. However, conventional snap-through structures face limitations in fabrication efficiency, scale, and tunability. We introduce SnapInflatables, enabling safe, multi-scale interaction with adjustable sensitivity and force reactions, utilizing the snap-through instability of inflatables. We designed six types of heat-sealing structures enabling versatile snap-through passive motion of inflatables with diverse reaction and trigger directions. A block structure enables ultra-sensitive states for rapid energy release and force amplification. The motion range is facilitated by geometry parameters, while force feedback properties are tunable through internal pressure settings. Based on experiments, we developed a design tool for creating desired inflatable snap-through shapes and motions, offering previews and inflation simulations. Example applications, including a self-locking medical stretcher, interactive animals, a bounce button, and a large-scale light demonstrate enhanced passive interaction with inflatables.
Yue Yang 0005, Zhuoyi Zhang, Yanchen Shen, Kuangqi Zhu, Junzhe Ji, Yongbo Ni, Jiayi Wu 0008, Qi Wang 0075, Jiang Wu 0019, Lingyun Sun, Ye Tao 0001, Guanyun Wang
CHI15
2024 Rapid 3D Model Generation with Intuitive 3D Input
abstract
With the emergence of AR/VR, 3D models are in tremendous demand. However, conventional 3D modeling with Computer-Aided Design software requires much expertise and is difficult for novice users. We find that AR/VR devices, in addition to serving as effective display mediums, can offer a promising potential as an intuitive 3D model creation tool, especially with the assistance of AI generative models. Here, we propose Deep3DVRSketch, the first 3D model generation network that inputs 3D VR sketches from novice users and generates highly consistent 3D models in multiple categories within seconds, irrespective of the users' drawing abilities. We also contribute KO3D+, the largest 3D sketch-shape dataset. Our method pre-trains a conditional diffusion model on quality 3D data, then fine-tunes an encoder to map 3D sketches onto the generator's manifold using an adaptive curriculum strategy for limited ground truths. In our experiment, our approach achieves state-of-the-art performance in both model quality and fidelity with real-world input from novice users, and users can even draw and obtain very detailed geometric structures. In our user study, users were able to complete the 3D modeling tasks over 10 times faster using our approach compared to conventional CAD software tools. We believe that our Deep3DVRSketch and KO3D+ dataset can offer a promising solution for future 3D modeling in metaverse era. Check the project page at http://research.kokoni3d.com/Deep3DVRSketch.
Tianrun Chen, Chaotao Ding, Shangzhan Zhang, Chunan Yu, Ying Zang, Zejian Li, Sida Peng, Lingyun Sun
CVPR8
2024 EGFE: End-to-end Grouping of Fragmented Elements in UI Designs with Multimodal Learning
abstract
When translating UI design prototypes to code in industry, automatically generating code from design prototypes can expedite the development of applications and GUI iterations. However, in design prototypes without strict design specifications, UI components may be composed of fragmented elements. Grouping these fragmented elements can greatly improve the readability and maintainability of the generated code. Current methods employ a two-stage strategy that introduces hand-crafted rules to group fragmented elements. Unfortunately, the performance of these methods is not satisfying due to visually overlapped and tiny UI elements. In this study, we propose EGFE, a novel method for automatically End-to-end Grouping Fragmented Elements via UI sequence prediction. To facilitate the UI understanding, we innovatively construct a Transformer encoder to model the relationship between the UI elements with multi-modal representation learning. The evaluation on a dataset of 4606 UI prototypes collected from professional UI designers shows that our method outperforms the state-of-the-art baselines in the precision (by 29.75%), recall (by 31.07%), and F1-score (by 30.39%) at edit distance threshold of 4. In addition, we conduct an empirical study to assess the improvement of the generated front-end code. The results demonstrate the effectiveness of our method on a real software engineering application. Our end-to-end fragmented elements grouping method creates opportunities for improving UI-related software engineering tasks.
Liuqing Chen 0002, Yunnong Chen, Shuhong Xiao, Yaxuan Song, Lingyun Sun, Yankun Zhen, Yanfang Chang
ICSE5
2024 AutoSpark: Supporting Automobile Appearance Design Ideation with Kansei Engineering and Generative AI
abstract
Rapid creation of novel product appearance designs that align with consumer emotional requirements poses a significant challenge. Text-to-image models, with their excellent image generation capabilities, have demonstrated potential in providing inspiration to designers. However, designers still encounter issues including aligning emotional needs, expressing design intentions, and comprehending generated outcomes in practical applications. To address these challenges, we introduce AutoSpark, an interactive system that integrates Kansei Engineering and generative AI to provide creativity support for designers in creating automobile appearance designs that meet emotional needs. AutoSpark employs a Kansei Engineering engine powered by generative AI and a semantic network to assist designers in emotional need alignment, design intention expression, and prompt crafting. It also facilitates designers’ understanding and iteration of generated results through fine-grained image-image similarity comparisons and text-image relevance assessments. The design-thinking map within its interface aids in managing the design process. Our user study indicates that AutoSpark effectively aids designers in producing designs that are more aligned with emotional needs and of higher quality compared to a baseline system, while also enhancing the designers’ experience in the human-AI co-creation process.
Liuqing Chen 0002, Qianzhi Jing, Yixin Tsang, Qianyi Wang, Ruocong Liu, Duowei Xia, Yunzhan Zhou, Lingyun Sun
UIST8
2024 MagneDot: Integrated Fabrication and Actuation Methods of Dot-Based Magnetic Shape Displays
abstract
This paper presents MagneDot, a novel method for making interactive magnetic shape displays through an integrated fabrication process. Magnetic soft materials can potentially create fast, responsive morphing structures for interactions. However, novice users and designers typically do not have access to sophisticated equipment and materials or cannot afford heavy labor to create interactive objects based on this material. Modified from an open-source 3D printer, the fabrication system of MagneDot integrates the processes of mold-making, pneumatic extrusion, magnetization, and actuation, using cost-effective materials only. By providing a design tool, MagneDot allows users to generate G-code for fabricating and actuating displays of various morphing effects. Finally, a series of design examples demonstrate the possibilities of shape displays enabled by MagneDot.
Lingyun Sun, Yitao Fan, Boyu Feng, Deying Pan, Yiwen Ren, Qi Wang 0075, Ye Tao 0001, Guanyun Wang
UIST1
2024 X-Hair: 3D Printing Hair-like Structures with Multi-form, Multi-property and Multi-function
abstract
In this paper, we present X-Hair, a method that enables 3D-printed hair with various forms, properties, and functions. We developed a two-step suspend printing strategy to fabricate hair-like structures in different forms (e.g. fluff, bristle, barb) by adjusting parameters including Extrusion Length Ratio and Total Length. Moreover, a design tool is also established for users to customize hair-like structures with various properties (e.g. pointy, stiff, soft) on imported 3D models, which virtually shows the results for previewing and generates G-code files for 3D printing. We demonstrate the design space of X-Hair and evaluate the properties of them with different parameters. Through a series of applications with hair-like structures, we validate X-hair’s practical usage of biomimicry, decoration, heat preservation, adhesion, and haptic interaction.
Guanyun Wang, Junzhe Ji, Yunkai Xu, Xiaojing Zhou, Boyu Feng, Lingyun Sun, Ye Tao 0001, Jiaji Li
UIST10
2024 ProtoDreamer: A Mixed-prototype Tool Combining Physical Model and Generative AI to Support Conceptual Design
abstract
Prototyping serves as a critical phase in the industrial conceptual design process, enabling exploration of problem space and identification of solutions. Recent advancements in large-scale generative models have enabled AI to become a co-creator in this process. However, designers often consider generative AI challenging due to the necessity to follow computer-centered interaction rules, diverging from their familiar design materials and languages. Physical prototype is a commonly used design method, offering unique benefits in prototype process, such as intuitive understanding and tangible testing. In this study, we propose ProtoDreamer, a mixed-prototype tool that synergizes generative AI with physical prototype to support conceptual design. ProtoDreamer allows designers to construct preliminary prototypes using physical materials, while AI recognizes these forms and vocal inputs to generate diverse design alternatives. This tool empowers designers to tangibly interact with prototypes, intuitively convey design intentions to AI, and continuously draw inspiration from the generated artifacts. An evaluation study confirms ProtoDreamer’s utility and strengths in time efficiency, creativity support, defects exposure, and detailed thinking facilitation.
Pei Chen 0005, Xuelong Xie, Chaoyi Lin, Lianyan Liu, Zhuoshu Li, Weitao You, Lingyun Sun
UIST8
2024 Supporting Text Entry in Virtual Reality with Large Language Models
abstract
Text entry in virtual reality (VR) often faces challenges in terms of efficiency and task loads. Prior research has explored various solutions, including specialized keyboard layouts, tracked physical devices, and hands-free interaction. Yet, these efforts often fall short of replicating the efficiency of real-world text entry, or introduce additional spatial and device constraints. This study leverages the extensive capabilities of large language models (LLMs) in context perception and text prediction to enhance text entry efficiency by reducing users’ manual keystrokes. Three LLM-assisted text entry methods - Simplified Spelling, Content Prediction, and Keyword-to-Sentence Generation - are introduced, aligning with user cognition and the contextual predictability of English text at word, grammatical structure, and sentence levels. Through user experiments encompassing various text entry tasks on an Oculus-based VR prototype, these methods demonstrate a 16.4%, 49.9%, 43.7% reduction in manual keystrokes, translating to efficiency gains of 21.4%,74.0%, 76.3%, respectively. Importantly, these methods do not increase manual corrections compared to manual typing, while significantly reducing physical, mental, and temporal loads and enhancing overall usability. Long-term observations further reveal users’ strategies for using these LLM-assisted methods, showing that users’ proficiency with the methods can reinforce their positive effects on text entry efficiency.
Liuqing Chen 0002, Yu Cai 0014, Ruyue Wang, Shixian Ding, Yilin Tang, Preben Hansen, Lingyun Sun
VR7
2024 AskNatureNet: A divergent thinking tool based on bio-inspired design knowledge
abstract
Divergent thinking is a process in design by exploring multiple possible solutions, is crucial in the early stages of design to break fixation and expand the design ideation. Design-by-Analogy promotes divergent thinking, by studying solutions have solved similar problems and using this knowledge to make inferences and solve problems in new and unfamiliar situations. Bio-inspired design (BID) is a form of design by analogy and its knowledge provides diverse sources for analogy, making BID knowledge as a potential source for divergent thinking. Existing BID database has focused on collecting BID cases and facilitating the retrieval of biological knowledge. Despite its success, applying BID knowledge into divergent thinking still encounters challenge, as the association between source domain and target domain are always limited within a single case. In this work, a novel approach is proposed to support divergent thinking from three subsequent phases: encoding, retrieval and mapping. Specifically, biological knowledge is encoded in a triple form by employing a large language model (LLM) to extract key information from a well-known BID knowledge base. The created triples are implemented in a semantic network to facilitate bidirectional retrieval modes: problem-driven and solution-driven, as well as mapping for divergent thinking. The mapping algorithm calculates the semantic similarity between nodes in the semantic network based on their attributes in three progressive steps by following the paradigm of divergent thinking. The proposed approach is implemented as tool called AskNatureNet,1 which supports divergent thinking by retrieving and mapping knowledge in a visualized interactive semantic network. An ideation case study on evaluating the effectiveness of AskNatureNet shows that our tool is capable of supporting divergent thinking efficiently.
Liuqing Chen 0002, Zebin Cai, Zhaojun Jiang, Jianxi Luo, Lingyun Sun, Peter R. N. Childs, Haoyu Zuo
Adv. Eng. Informatics5
2024 The Doctrine of the Mean: Chinese Calligraphy with Moderate Visual Complexity Elicits High Aesthetic Preference
abstract
Chinese calligraphy is a symbol of cultural heritage; it is widely used in visual design (e.g., film posters and modern fashion) for its aesthetic value. In human-computer interaction, visual design complexity strongly influences users’ preferences. Indeed, the relevance of visual complexity to aesthetic preference has been confirmed. However, the visual complexity of calligraphy artworks has seldom been investigated and the process of how visual complexity affects aesthetic preference remains vague. Therefore, in this work we proposed a computational method to evaluate complexity and conducted several perception studies. Results showed that layout features and calligraphy style (regular script, running script, and cursive script) affected visual complexity. The level of visual complexity (low, medium, and high) affected aesthetic preference but calligraphy style did not. Furthermore, Chinese calligraphy with moderate visual complexity evokes strong aesthetic preference. The present findings can help designers redesign websites and interfaces for high aesthetic preference and can provide insights for developing the theoretical design of calligraphic art for advanced interaction in cultural heritage, such as in interactive systems for teaching calligraphy based on aesthetic cognition.
Kaixin Han, Weitao You, Shuhui Shi, Huanghuang Deng, Lingyun Sun
Int. J. Hum. Comput. Interact.5
2024 The Impression of Round and Square: Chinese Calligraphy Aesthetics in Modern Type Design
abstract
Round and square are important concepts in Chinese aesthetics. The legacy of Chinese calligraphy aesthetics has led the way to modern design. In HCI, interface typefaces influence users’ reading experiences. Users’ aesthetic preferences may influence their reading experiences as well. The investigation of how human aesthetic preferences are affected by character visual features (e.g., character size and character weight) in type design has received substantial attention. However, it is difficult to accurately study human aesthetic preferences through these visual features in Chinese calligraphy-type design. How Chinese character visual features affect human aesthetic preferences is not clear. Inspired by graphic design, different visual features (e.g., color and texture) can cause changes in impression expressions (e.g., modern and classical), thus affecting human aesthetic preferences. In this article, we introduced a character impression analysis method to investigate how human aesthetic preferences are affected by Chinese character visual features. First, we proposed two impression expressions of character (round and square) from the concepts of Chinese aesthetics; Then, for these two impression expressions, we modeled four visual features of calligraphy characters from the principles of Gestalt psychology. Finally, we conducted two studies to investigate how visual features affect impression expressions and explored the relationship between aesthetic preferences. The results clearly explain how character visual features affect human aesthetic preferences. The present study can give designers a deep understanding of how visual features affect human aesthetic preferences based on character impression analysis. It can guide designers to make better type design decisions to achieve a high aesthetic preference for users. Furthermore, the findings of the present study can potentially provide new insights into Latin type design.
Kaixin Han, Weitao You, Shuhui Shi, Lingyun Sun
Int. J. Hum. Comput. Interact.4
2024 CNAMD Corpus: A Chinese Natural Audiovisual Multimodal Database of Conversations for Social Interactive Agents
abstract
Impressive progress has been made in developing companion Socially Interactive Agents (SIAs) that provide companionship and reduce loneliness. However, recent works focus on analyzing multimodal feedback in Answer part but ignore Question part. Furthermore, research on SIAs is primarily based on English, which poses a challenge for Chinese SIAs because of cultural differences between English and Chinese. Therefore, we introduce a Chinese Natural Audiovisual Multimodal Database (CNAMD) corpus, the first and largest freely available Chinese multimodal database for multi-person interaction, containing 48 hours of videos and annotations across eight modalities. Using CNAMD, we analyze the characteristics of vocal-verbal, audio, behavioral, and multimodal combinations during questioning, test the performance of six baselines on three tasks, and propose improvements for processing daily Chinese data. The present findings will help designers consider Chinese customs and language when designing Chinese SIAs, making them more suitable for the Chinese cultural context and users.
Jingyu Wu, Shi Chen 0005, Wei Xiang 0008, Lingyun Sun, Hongzeng Zhang, Yanxu Li
Int. J. Hum. Comput. Interact.4
2024 HierVid: Lowering the Barriers to Entry of Interactive Video Making with a Hierarchical Authoring System
abstract
Interactive videos have been applied to various areas due to their engagement potential and efficiency improvement of information communication. However, creating interactive videos can be challenging because of a lack of novice-oriented guidance in current platforms, and the logic-building process when authoring interactive videos. To address these challenges, we obtained insights from four creativity support tool designers, proposed a series of hierarchical interactive video structures based on existing narrative structures, and presented the HierVid system. The system is designed as a Template-Module-Unit Mode-based hierarchical authoring platform grounded on three design requirements, and we conducted two user studies to evaluate HierVid. The results showed that novice users could get started to use and understand the functions easily, and the system allowed users to use and explore freely, with an enhanced efficiency compared to the bilibili platform. In conclusion, our research and design of HierVid offer guidance and support for novice users, making interactive video authoring quicker and more accessible.
Weitao You, Zhuoyi Cheng, Zirui Ma, Guang Yang 0022, Zhibin Zhou 0002, Lingyun Sun
Int. J. Hum. Comput. Interact.6
2024 Elicitation and Evaluation of Hand-based Interaction Language for 3D Conceptual Design in Mixed Reality
Lingyun Sun, Pei Chen 0005, Zhaoqu Jiang, Xuelong Xie, Zihong Zhou, Xuanhui Liu
Int. J. Hum. Comput. Stud.1
2024 Element-conditioned GAN for graphic layout generation
Liuqing Chen 0002, Qianzhi Jing, Yunzhan Zhou, Zhaoxing Li, Lei Shi 0003, Lingyun Sun
Neurocomputing6
2024 Deep3DSketch-im: rapid high-fidelity AI 3D model generation by single freehand sketches
abstract
The rise of artificial intelligence generated content (AIGC) has been remarkable in the language and image fields, but artificial intelligence (AI) generated three-dimensional (3D) models are still under-explored due to their complex nature and lack of training data. The conventional approach of creating 3D content through computer-aided design (CAD) is labor-intensive and requires expertise, making it challenging for novice users. To address this issue, we propose a sketch-based 3D modeling approach, Deep3DSketch-im, which uses a single freehand sketch for modeling. This is a challenging task due to the sparsity and ambiguity. Deep3DSketch-im uses a novel data representation called the signed distance field (SDF) to improve the sketch-to-3D model process by incorporating an implicit continuous field instead of voxel or points, and a specially designed neural network that can capture point and local features. Extensive experiments are conducted to demonstrate the effectiveness of the approach, achieving state-of-the-art (SOTA) performance on both synthetic and real datasets. Additionally, users show more satisfaction with results generated by Deep3DSketch-im, as reported in a user study. We believe that Deep3DSketch-im has the potential to revolutionize the process of 3D modeling by providing an intuitive and easy-to-use solution for novice users.
Tianrun Chen, Runlong Cao, Zejian Li, Ying Zang, Lingyun Sun
Frontiers Inf. Technol. Electron. Eng.5
2024 Recent advances in artificial intelligence generated content
abstract
人工智能生成内容(AIGC)是近年来人工智能(AI)领域一个研究热点,它有望取代人类以较低成本高效率执行内容生成工作,如音乐、绘画、多模态内容生成、新闻文章、总结报告、股评摘要,以至元宇宙中的内容生成和数字人。AIGC为未来AI发展和实现提供了一条新的技术路径。 在此背景下,《信息与电子工程前沿(英文)》期刊组织了一期关于AIGC最新进展的特刊。本期特刊关注AIGC理论、算法、应用及相关领域。通过吸引高质量论文,我们希望帮助学术界和工业界研究人员更深入了解AIGC背后的基本理论及其潜在应用,激励更多研究人员加入并推进AIGC领域的研究。因此,我们就以下主题(但不限于)征集论文:(1)AI生成音乐;(2)AI生成绘画;(3)AI对话模型;(4)AI新闻摘要;(5)AI与元宇宙;(6)AI与数字人;(7)AI图像编辑;(8)AI生成短视频;(9)AI生成多媒体内容;(10)ChatGPT相关工作。经严格评审,选出12篇论文,包括1篇评论、1篇观点、3篇综述、6篇研究和1篇通讯。我们将其划分为3个主要部分:ChatGPT、扩散模型、提示学习和多模态。 总体而言,本期特刊涵盖了与AIGC开发和应用相关的广泛研究主题,包括人工智能图像/文本生成、三维内容创建、以用户为中心的图形设计、特定风格的音乐生成,以及与因果表征学习、高阶扩散模型相关的工作。此外,还详细调研了概率扩散模型、提示学习和ChatGPT。 最后,感谢所有作者对本期特刊的支持,特别感谢所有评审人对专刊投稿富有见地的意见和有益建议。
Junping Zhang, Lingyun Sun, Junbin Gao, Jiebo Luo 0001, Jingdong Wang 0001
Frontiers Inf. Technol. Electron. Eng.2
2024 LanT: finding experts for digital calligraphy character restoration
Kaixin Han, Weitao You, Huanghuang Deng, Lingyun Sun, Jinyu Song, Zijin Hu, Heyang Yi
Multim. Tools Appl.4
2024 Correction: LanT: finding experts for digital calligraphy character restoration
Kaixin Han, Weitao You, Huanghuang Deng, Lingyun Sun, Jinyu Song, Zijin Hu, Heyang Yi
Multim. Tools Appl.4
2024 Reality3DSketch: Rapid 3D Modeling of Objects From Single Freehand Sketches
abstract
The emerging trend of AR/VR places great demands on 3D content. However, most existing software requires expertise and is difficult for novice users to use. In this paper, we aim to create sketch-based modeling tools for user-friendly 3D modeling. We introduce Reality3DSketch with a novel application of an immersive 3D modeling experience, in which a user can capture the surrounding scene using a monocular RGB camera and can draw a single sketch of an object in the real-time reconstructed 3D scene. A 3D object is generated and placed in the desired location, enabled by our novel neural network with the input of a single sketch. Our neural network can predict the pose of a drawing and can turn a single sketch into a 3D model with view and structural awareness, which addresses the challenge of sparse sketch input and view ambiguity. We conducted extensive experiments synthetic and real-world datasets and achieved state-of-the-art (SOTA) results in both sketch view estimation and 3D modeling performance. According to our user study, our method of performing 3D modeling in a scene is$>$5x faster than conventional methods. Users are also more satisfied with the generated 3D model than the results of existing methods.
Tianrun Chen, Chaotao Ding, Lanyun Zhu, Ying Zang, Yiyi Liao, Zejian Li, Lingyun Sun
IEEE Trans. Multim.7
2024 Automatic Generation of Interactive Nonlinear Video for Online Apparel Shopping Navigation
abstract
We present an automatic generation pipeline of interactive nonlinear video for online apparel shopping navigation. Our approach was inspired by Google's “Messy Middle” theory, which suggests that people mentally are faced with two tasks—exploration and evaluation—before purchasing online. Given a set of apparel product presentation videos, our navigation UI organizes them to optimize users' product exploration and automatically generates interactive videos for users' product evaluation. To support automatic methods, we proposed a video clustering similarity ($\operatorname{CSIM}$) and a camera movement similarity ($\operatorname{MSIM}$), as well as a comparative video generation algorithm for product recommendation, presentation, and comparison. To evaluate our pipeline's effectiveness, we conducted several user studies. The results showed that our pipeline can help users complete the consumption process more efficiently, making it easier for them to understand and choose a product.
Weitao You, Juntao Ji, Lingyun Sun, Chang-yuan Yang, Mi Yu, Shi Chen 0005
IEEE Trans. Multim.3
2024 A Hybrid Prototype Method Combining Physical Models and Generative Artificial Intelligence to Support Creativity in Conceptual Design
abstract
Conceptual design is an essential stage in the design process, and its ultimate success largely depends on designers’ creativity. Both physical and digital prototypes are commonly adopted by designers to support ideation and creativity, providing intuitive perception and rapid iteration, respectively. In recent advancements, large-scale generation models are able to offer data-enabled creativity support by generating high-quality solutions comparable to human designers. This opens up an imaginary space for designers and brings new possibilities for design tools. In this study, we proposed a hybrid prototype method that synergistically combines physical models and generative artificial intelligence (AI) in the conceptual design stage. Correspondingly, we developed a hybrid prototype system to implement the proposed method. We conducted a comparative user study with 45 designers who completed a design task using the physical prototype method, standalone generative AI and the hybrid prototype method, respectively. Our results verified the effectiveness of the hybrid prototype method and investigated its mechanism in supporting creativity. Finally, we discussed the application value and optimisation space of the hybrid prototype method.
Pei Chen 0005, Xuelong Xie, Zhaoqu Jiang, Zihong Zhou, Lingyun Sun
ACM Trans. Comput. Hum. Interact.6
2024 Hearing with the eyes: modulating lyrics typography for music visualization
Kaixin Han, Weitao You, Shuhui Shi, Lingyun Sun
Vis. Comput.4
2023 Preserving Structural Consistency in Arbitrary Artist and Artwork Style Transfer
abstract
Deep generative models are effective in style transfer. Previous methods learn one or several specific artist-style from a collection of artworks. These methods not only homogenize the artist-style of different artworks of the same artist but also lack generalization for the unseen artists. To solve these challenges, we propose a double-style transferring module (DSTM). It extracts different artist-style and artwork-style from different artworks (even untrained) and preserves the intrinsic diversity between different artworks of the same artist. DSTM swaps the two styles in the adversarial training and encourages realistic image generation given arbitrary style combinations. However, learning style from single artwork can often cause over-adaption to it, resulting in the introduction of structural features of style image. We further propose an edge enhancing module (EEM) which derives edge information from multi-scale and multi-level features to enhance structural consistency. We broadly evaluate our method across six large-scale benchmark datasets. Empirical results show that our method achieves arbitrary artist-style and artwork-style extraction from a single artwork, and effectively avoids introducing the style image’s structural features. Our method improves the state-of-the-art deception rate from 58.9% to 67.2% and the average FID from 48.74 to 42.83.
Jingyu Wu, Lefan Hou, Zejian Li, Jun Liao 0001, Li Liu 0001, Lingyun Sun
AAAI6
2023 "I Never Envy Anyone, for I Have Already Built a Kingdom With My Fingertips": Exploring Teenagers' Experience in Chat-based Cosplay Community
abstract
This paper reports an interview study about the practice of teenagers’ chat-based cosplay in China. Findings reveal the four primary motivations of the participants and their main practice in chat-based cosplay. We found that adolescents perceived character presentation and portrayal as a central aspect of chat-based cosplay and they devoted significant effort to refine their characters to achieve higher character consistency. We highlighted the positive feedback loop between social relationships and story creation in chat-based cosplay community. In addition, we identified the influence and negative experiences on adolescents in the chat-based cosplay community.
Yaohua Bu, Suqi Lou, Shi Chen 0005, Lingyun Sun, Chang-yuan Yang
IDC5
2023 EdibleToy: Empowering Children to Create Their Own Meals with a DIY Wafer Paper Kit
abstract
Existing methods in human-computer interaction to enhance children’s eating habits predominantly rely on digital interactive technologies, which pose the risk of increasing sensory stimulation and diverting children’s attention away from the food itself. Drawing inspiration from shape-changing food research, we propose an approach that combines deformable wafer paper for food preparation. We summarize the principles of wafer paper controllable deformation and develop a toolkit to facilitate its use. We support children in creating personalized, transformable food items using this method, aiming to provide a playful, convenient, and safe food-making experience tailored for children, thereby enhancing children’s mealtime engagement and habits.
Yilin Shao, Boyu Feng, Yingpin Chen, Yue Yang 0005, Yanan Wang 0005, Ye Tao 0001, Lingyun Sun, Guanyun Wang
IDC9
2023 Layout Generation for Various Scenarios in Mobile Shopping Applications
abstract
Layout is essential for the product listing pages (PLPs) in mobile shopping applications. To clearly convey the information that consumers require and to achieve specific functions, PLPs layouts often have many variations driven by scenarios. In this work, we study the PLPs layout design for different scenarios and propose a design space to guide the large-scale creation of PLPs. We propose LayoutVQ-VAE, a novel model specialized in generating layouts with internal and external constraints. LayoutVQ-VAE differs from previous methods as it learns a discrete latent representation of layout and can model the relationship between layout representation and scenarios without applying heuristics. Experiments on publicly available benchmarks for different layout types validate that our method performs comparably or favorably against the state-of-the-art methods. Case studies show that the proposed approach including the design space and model is effective in producing large-scale high-quality PLPs layouts for mobile shopping platforms.
Qianzhi Jing, Yixin Tsang, Liuqing Chen 0002, Lingyun Sun, Yankun Zhen, Yichun Du
CHI5
2023 All-in-One Print: Designing and 3D Printing Dynamic Objects Using Kinematic Mechanism Without Assembly
abstract
The field of Human-Computer-Interaction (HCI) has been consistently utilizing kinematic mechanisms to create tangible dynamic interfaces and objects. However, the design and fabrication of these mechanisms are challenging due to complex spatial structures, step-by-step assembly processes, and unstable joint connections resulting from the inevitable matching errors within separated parts. In this paper, we propose an integrated fabrication method for one-step FDM 3D printing (FDM3DP) kinematic mechanisms to create dynamic objects without additional post-processing. We describe the Arch-printing and Support-bridges method, which we call All-in-One Print, that compiles given arbitrary solid 3D models into printable kinematic models as G-Code for FDM3DP. To expand the design space, we investigate a series of motion structures (e.g., rotate, slide, and screw) with multi-stabilities and develop a design tool to help users quickly design such dynamic objects. We also demonstrate various application cases, including physical interfaces, toys with interactive aesthetics and daily items with internalized functions.
Jiaji Li, Junzhe Ji, Deying Pan, Yitao Fan, Kuangqi Zhu, Yue Yang 0005, Lingyun Sun, Ye Tao 0001, Guanyun Wang
CHI9
2023 4Doodle: 4D Printing Artifacts Without 3D Printers
abstract
4D printing encodes transformability over time, which empowers users to create artifacts by on-demand deformation. The creative process of 4D printing shape-changing artifacts can be challenging because of its discontinuous fabrication steps, such as digital designing, specific path planning, automatic printing and manual triggering. We hypothesize that switching from typical 4D printing reliant on 3D printers to a more “handcrafted” method can allow users to understand and continuously reflect upon the artifact and its transformability. Towards this vision, we introduce 4Doodle, a hybrid craft approach that integrates unique deformation controllability and five techniques for freehand 4D printing, using a 3D pen. To tackle the shape-changing challenges of uncertain hands-on fabrication, we develop a mixed reality system to help novices master the manual skills of 4D printing. We also demonstrate a series of 4D printed artifacts with fully human intervention. Finally, our user study shows that 4Doodle lowers the skill-acquisition barrier associated with handcrafting 4D printed artifacts, and it has great potential for creative production and spatial ability.
Ye Tao 0001, Junzhe Ji, Linlin Cai, Hongmei Xia, Jinghai He, Yitao Fan, Shengzhang Pan, Jinghua Xu, Cheng Yang 0014, Lingyun Sun, Guanyun Wang
CHI12
2023 PneuFab: Designing Low-Cost 3D-Printed Inflatable Structures for Blow Molding Artifacts
abstract
Access to computer-aided fabrication tools, such as 3D printing, empowers various craft techniques to democratize the creation of artifacts. To afford new blow molding techniques in the field of Human-Computer Interaction, we make efforts to simplify this challenging handy fabrication and enrich the design space of blow molding by taking advantage of the thermoformability and heat deformability of 3D printed thermoplastics. We propose a novel and democratized blow molding technique, PneuFab, enabled by FDM 3D-printed custom structures and temporal triggering methods. Then we implement and evaluate a design tool that allows users to play with parameters and preview the resulting forms until achieving their desired shapes. Showcasing design spaces including artifacts with complex geometries and tunable stiffness, we hope to expand access and dive into what more the digital blow molding fabrication can be.
Guanyun Wang, Kuangqi Zhu, Lingchuan Zhou, Mengyan Guo, Deying Pan, Yue Yang 0005, Jiaji Li, Jiang Wu 0019, Ye Tao 0001, Lingyun Sun
CHI12
2023 E-Orthosis: Augmenting Off-the-Shelf Orthoses with Electronics
abstract
Orthoses with electronic functions have emerged as a promising medical product in response to the increasing demand for rehabilitation training, therapy assistance, and health monitoring. However, fabricating this “smart orthosis” often requires long development cycles and exorbitant prices. We introduce E-Orthosis, an integrated fabrication approach with construction toolkits for healthcare professionals to quickly embed electronics in off-the-shelf orthoses with customized functions cost-effectively and time-efficiently. Specifically, we develop components with magnets and pogo pins to support rapid attachment and sustainable use, and textile-based electrodes with snap installation to improve the wearing experience. We also provide a circuit iron tool to apply circuit traces on complex surfaces of orthoses directly and a hot punch tool to embed magnet ports and electrodes. Three application examples, technical evaluations, and expert reviews demonstrate the functionality of E-Orthosis and the potential for democratizing rapid-developed and low-cost smart orthoses for patients.
Yue Yang 0005, Yitao Fan, Yilin Shao, Kuangqi Zhu, Jiaji Li, Qi Wang 0075, Lingyun Sun, Ye Tao 0001, Guanyun Wang
CHI10
2023 Deep3DSketch: 3D Modeling from Free-Hand Sketches with View- and Structural-Aware Adversarial Training
abstract
This work aims to investigate the problem of 3D modeling using single free-hand sketches, which is one of the most natural ways we humans express ideas. Although sketch-based 3D modeling can drastically make the 3D modeling process more accessible, the sparsity and ambiguity of sketches bring significant challenges for creating high-fidelity 3D models that reflect the creators’ ideas. In this work, we propose a view-and structural-aware deep learning approach, Deep3DSketch, which tackles the ambiguity and fully uses sparse information of sketches, emphasizing the structural information. Specifically, we introduced random pose sampling on both 3D shapes and 2D silhouettes, and an adversarial training scheme with an effective progressive discriminator to facilitate learning of the shape structures. Extensive experiments demonstrated the effectiveness of our approach, which outperforms existing methods – with state-of-the-art (SOTA) performance on both synthetic and real datasets.
Tianrun Chen, Chenglong Fu 0003, Lanyun Zhu, Papa Mao, Ying Zang, Lingyun Sun
ICASSP7
2023 Efficient Emotional Adaptation for Audio-Driven Talking-Head Generation
abstract
Audio-driven talking-head synthesis is a popular research topic for virtual human-related applications. However, the inflexibility and inefficiency of existing methods, which necessitate expensive end-to-end training to transfer emotions from guidance videos to talking-head predictions, are significant limitations. In this work, we propose the Emotional Adaptation for Audio-driven Talking-head (EAT) method, which transforms emotion-agnostic talking-head models into emotion-controllable ones in a cost-effective and efficient manner through parameter-efficient adaptations. Our approach utilizes a pretrained emotion-agnostic talking-head transformer and introduces three lightweight adaptations (the Deep Emotional Prompts, Emotional Deformation Network, and Emotional Adaptation Module) from different perspectives to enable precise and realistic emotion controls. Our experiments demonstrate that our approach achieves state-of-the-art performance on widely-used benchmarks, including LRW and MEAD. Additionally, our parameter-efficient adaptations exhibit remarkable generalization ability, even in scenarios where emotional training videos are scarce or nonexistent. Project website: https://yuangan.github.io/eat/
Yuan Gan, Zongxin Yang, Xihang Yue, Lingyun Sun, Yi Yang 0001
ICCV4
2023 Learning Object Consistency and Interaction in Image Generation from Scene Graphs
abstract
This paper is concerned with synthesizing images conditioned on a scene graph (SG), a set of object nodes and their edges of interactive relations. We divide existing works into image-oriented and code-oriented methods. In our analysis, the image-oriented methods do not consider object interaction in spatial hidden feature. On the other hand, in empirical study, the code-oriented methods lose object consistency as their generated images miss certain objects in the input scene graph. To alleviate these two issues, we propose Learning Object Consistency and Interaction (LOCI). To preserve object consistency, we design a consistency module with a weighted augmentation strategy for objects easy to be ignored and a matching loss between scene graphs and image codes. To learn object interaction, we design an interaction module consisting of three kinds of message propagation between the input scene graph and the learned image code. Experiments on COCO-stuff and Visual Genome datasets show our proposed method alleviates the ignorance of objects and outperforms the state-of-the-art on visual fidelity of generated images and objects.
Yangkang Zhang, Chenye Meng, Zejian Li, Pei Chen 0005, Guang Yang 0022, Chang-yuan Yang, Lingyun Sun
IJCAI7
2023 Cultural Self-Adaptive Multimodal Gesture Generation Based on Multiple Culture Gesture Dataset
abstract
Co-speech gesture generation is essential for multimodal chatbots and agents. Previous research extensively studies the relationship between text, audio, and gesture. Meanwhile, to enhance cross-culture communication, culture-specific gestures are crucial for chatbots to learn cultural differences and incorporate cultural cues. However, culture-specific gesture generation faces two challenges: lack of large-scale, high-quality gesture datasets that include diverse cultural groups, and lack of generalization across different cultures. Therefore, in this paper, we first introduce a Multiple Culture Gesture Dataset (MCGD), the largest freely available gesture dataset to date. It consists of ten different cultures, over 200 speakers, and 10,000 segmented sequences. We further propose a Cultural Self-adaptive Gesture Generation Network (CSGN) that takes multimodal relationships into consideration while generating gestures using a cascade architecture and learnable dynamic weight. The CSGN adaptively generates gestures with different cultural characteristics without the need to retrain a new network. It extracts cultural features from the multimodal inputs or a cultural style embedding space with a designated culture. We broadly evaluate our method across four large-scale benchmark datasets. Empirical results show that our method achieves multiple cultural gesture generation and improves comprehensiveness of multimodal inputs. Our method improves the state-of-the-art average FGD from 53.7 to 48.0 and culture deception rate (CDR) from 33.63% to 39.87%.
Jingyu Wu, Shi Chen 0005, Shuyu Gan, Chang-yuan Yang, Lingyun Sun
ACM Multimedia6
2023 Deep3DSketch+: Rapid 3D Modeling from Single Free-Hand Sketches
Tianrun Chen, Chenglong Fu 0003, Ying Zang, Lanyun Zhu, Papa Mao, Lingyun Sun
MMM (2)7
2023 Novel 3D-Aware Composition Images Synthesis for Object Display with Diffusion Model
abstract
Designing attractive images for object display can be a time-consuming and skill-intensive process. The emergence of advanced algorithms, particularly the Diffusion Model, has made it possible to synthesize attractive images using AI. However, the existing diffusion models are mostly used to generate entire images and lack control over specific objects for object display. Here, to the best of our knowledge, we pioneers to extend the application of the diffusion model to synthesize novel images for specific objects. By encoding the input images of objects into NeRF representation and synthesizing the desired backgrounds using diffusion models with the input of rendered object images and text prompts, our method can generate 3D aware object display images at arbitrary angles and arbitrary backgrounds. We have conducted extensive experiments to demonstrate that our method is capable of generating high-quality and photo-realistic images, which are > 6 times faster than the conventional photomontage approach. Moreover, our generated images have higher compositional scores, image quality scores, and aesthetics scores in our user experiments. By significantly reducing the need for human effort and producing higher quality generated images, our approach opens up exciting possibilities for creating versatile novel images of specific objects.
Tianrun Chen, Tao Xu 0048, Yiyu Ye, Papa Mao, Ying Zang, Lingyun Sun
SMC6
2023 Supporting Crowd Workers in Ideation Tasks Through Information Gathering and Reflective Activity
abstract
Crowdsourcing is widely used to solve creative problems of ideation in HCI domain. To improve the creativity of crowdsourcing outcomes, researchers have proposed multiple approaches to support crowd workers’ idea proposal step. Other steps in designers’ creative process, such as information gathering and reflective activity, also impact idea creativity, while their effects in crowdsourcing scenarios remain unexplored. Therefore, referring to the creativity research, this study proposed an approach that optimized crowdsourcing tasks to help crowd workers perform information gathering and reflective activity. Two experiments involving 427 workers were conducted to test the effects of our approach. Results showed that instructing crowd workers to gather information and reflect on their ideas positively affected the creativity of crowdsourcing outcomes. This study provides inspiration for the optimization of crowdsourcing tasks and offers insights to the researchers who focus on using crowd power to achieve the ideation process.
Lingyun Sun, Wei-yue Gao, Wei Xiang 0008
Int. J. Hum. Comput. Interact.1
2023 UI layers merger: merging UI layers via visual learning and boundary prior
abstract
With the fast-growing graphical user interface (GUI) development workload in the Internet industry, some work attempted to generate maintainable front-end code from GUI screenshots. It can be more suitable for using user interface (UI) design drafts that contain UI metadata. However, fragmented layers inevitably appear in the UI design drafts, which greatly reduces the quality of the generated code. None of the existing automated GUI techniques detects and merges the fragmented layers to improve the accessibility of generated code. In this paper, we propose UI layers merger (UILM), a vision-based method that can automatically detect and merge fragmented layers into UI components. Our UILM contains the merging area detector (MAD) and a layer merging algorithm. The MAD incorporates the boundary prior knowledge to accurately detect the boundaries of UI components. Then, the layer merging algorithm can search for the associated layers within the components’ boundaries and merge them into a whole. We present a dynamic data augmentation approach to boost the performance of MAD. We also construct a large-scale UI dataset for training the MAD and testing the performance of UILM. Experimental results show that the proposed method outperforms the best baseline regarding merging area detection and achieves decent layer merging accuracy. A user study on a real application also confirms the effectiveness of our UILM.
Yunnong Chen, Yankun Zhen, Chu-ning Shi, Jiazhi Li 0002, Liuqing Chen 0002, Zejian Li, Lingyun Sun, Yanfang Chang
Frontiers Inf. Technol. Electron. Eng.7
2023 What makes virtual intimacy...intimate? Understanding the Phenomenon and Practice of Computer-Mediated Paid Companionship
abstract
Virtual romance service (VRS), as a notable commodification of intimacy, is currently emerging in China. Such service is not similar to the kind of intimacy that fans and idols generate through parasocial relationships, but behaves as the direct dyadic intimacy between service providers (virtual lovers) and buyers (customers). To gain a deep understanding of computer-mediated paid companionship, we study emerging user behaviors in VRS through a mixed-method study, including a survey (N = 178) and a follow-up semi-structured interview (N = 22) with both virtual lovers and customers to learn about their motivations, perceptions, and how virtual lovers provide online paid companionship to meet customers' emotional needs. We found three behavioral strategies of virtual lovers and the fact that they provide service in surface and deep acting and real feeling. Customers see VRS as a way to obtain affective benefits with reduced affective cost. We also found that VRS customers paid for the tangible benefits of an idealized romantic partner, rather than long-term commitment and emotional investment, and we identified key characteristics that VRS reduces from intimate relationships that fit its pay-per-use feature. We conclude by discussing the nature of virtual lovers and design implications for computer-mediated paid companionship.
Shi Chen 0005, Lingyun Sun, Chang-yuan Yang
Proc. ACM Hum. Comput. Interact.3
2023 Werewolf-XL: A Database for Identifying Spontaneous Affect in Large Competitive Group Interactions
abstract
Affective computing and natural human-computer interaction, which would be capable of interpreting and responding intelligently to the social cues of interaction in crowds, are more needed than ever as an individual's affective experience is often related to others in group activities. To develop the next-generation intelligent interactive systems, we require numerous human facial expressions with accurate annotations. However, existing databases usually consider nonspontaneous human behavior (posed or induced), individual or dyadic setting, and a single type of emotion annotation. To address this need, we created the Werewolf-XL database, which contains a total of 890 minutes of spontaneous audio-visual recordings of 129 subjects in a group interaction of nine individuals playing a conversational role-playing game called Werewolf. We provide 131,688 individual utterance-level video clips with internal self-assessment of 18 non-prototypical emotional categories and external assessment of pleasure, arousal, and dominance, including 14,632 speakers' samples and the rest of listeners' samples. Besides, the results of the annotation agreement analysis show fair reliability and validity. Role information and outcomes of the game are also recorded. Furthermore, we provided extensive benchmarks of unimodal and multimodal emotional recognition results. The database is made publicly available.
Xinda Wu, Xinhang Xie, Hui Zhang 0064, Lingyun Sun
IEEE Trans. Affect. Comput.7
2022 PneuMesh: Pneumatic-driven Truss-based Shape Changing System
abstract
From transoceanic bridges to large-scale installations, truss structures have been known for their structural stability and shape complexity. In addition to the advantages of static trusses, truss structures have a large degree of freedom to change shape when equipped with rotatable joints and retractable beams. However, it is difficult to design a complex motion and build a control system for large numbers of trusses. In this paper, we present PneuMesh, a novel truss-based shape-changing system that is easy to design and build but still able to achieve a range of tasks. PneuMesh accomplishes this by introducing an air channel connection strategy and reconfigurable constraint design that drastically decreases the number of control units without losing the complexity of shape-changing. We develop a design tool with real-time simulation to assist users in designing the shape and motion of truss-based shape-changing robots and devices. A design session with seven participants demonstrates that PneuMesh empowers users to design and build truss structures with a wide range of shapes and various functional motions.
Jianzhe Gu, Yuyu Lin, Jiaji Li, Lingyun Sun, Fangtian Ying, Guanyun Wang, Lining Yao
CHI6
2022 Few-Shot Incremental Learning for Label-to-Image Translation
abstract
Label-to-image translation models generate images from semantic label maps. Existing models depend on large volumes of pixel-level annotated samples. When given new training samples annotated with novel semantic classes, the models should be trained from scratch with both learned and new classes. This hinders their practical applications and motivates us to introduce an incremental learning strategy to the label-to-image translation scenario. In this paper, we introduce a few-shot incremental learning method for label-to-image translation. It learns new classes one by one from a few samples of each class. We propose to adopt semantically-adaptive convolution filters and normalization. When incrementally trained on a novel semantic class, the model only learns a few extra parameters of class-specific modulation. Such design avoids catastrophic forgetting of already-learned semantic classes and enables label-to-image translation of scenes with increasingly rich content. Furthermore, to facilitate few-shot learning, we propose a modulation transfer strategy for better initialization. Extensive experiments show that our method outperforms existing related methods in most cases and achieves zero forgetting.
Pei Chen 0005, Yangkang Zhang, Zejian Li, Lingyun Sun
CVPR4
2022 Recognizing Cognitive Load by a Hybrid Spatio-Temporal Causal Model from Multivariate Physiological Data
Zirui Yong, Guoxin Su, Xiaohu Li, Lingyun Sun, Zejian Li, Li Liu 0001
ECML/PKDD (6)4
2022 Enrichment of Product Presentation Video: Methods and Impacts on User Experience
abstract
Product presentation video (PPV) is a genre of short-form video in online retailing that presents product features and facilitates online shopping. To improve PPVs, quality and provide a better experience for viewers, PPV producers, most of whom are online retailers and non-professionals in video production, have made efforts to enrich these short-form videos. Despite the great demand for PPV enrichment, these methods have not been systematically explored, and their impacts on user experience remain unclear, impeding the improvement of PPVs. This study combined qualitative and quantitative methods to explore the impacts of PPV enrichment methods on user experience. As an exploratory study, we focused on PPVs of female fashion products on Chinese e-commerce platforms. We collected 240 PPVs and summarized ten enrichment methods from them accordingly. A questionnaire-based experiment, including 48 participants, was then conducted to explore the impacts of these methods on the experience of PPVs. Results indicated that eight out of ten methods effectively improved PPV experience from multiple dimensions. This study brings insights for exploring PPV enrichment from the perspective of user experience and provides support for PPV production process.
Wei-yue Gao, Wei Xiang 0008, Xuanhui Liu, Lingyun Sun
QoMEX5
2022 Impacts of Presenting Extra Information in Short Videos via Text and Voice on User Experience
abstract
Short video is an increasingly prevalent medium in online shopping environments to present products. To cope with the great demand for short videos rising from the enormous number and the rapid update of online products, computer-supported video production is becoming a trend. The optimization of short videos considering user experience is essential. Currently, using text and voice to integrate extra information into short videos is a potential and promising approach for optimizing computer-supported video production, while the effects of these elements on user experience remain unclear. In this study, we conducted a questionnaire-based experiment including 580 participants to explore the impacts of presenting extra information in short videos via text and voice on multi-dimensional user experience. Results indicated that these two elements positively impacted user experience from different dimensions. Gender differences were also found in this study. Based on experimental results, we provided suggestions to support the use of text and voice elements in short video production considering user experience.
Wei-yue Gao, Wei Xiang 0008, Xuanhui Liu, Xueyou Wang, Lingyun Sun
QoMEX5
2022 X-Bridges: Designing Tunable Bridges to Enrich 3D Printed Objects' Deformation and Stiffness
abstract
Bridges are unique structures appeared in fused deposition modeling (FDM) that make rigid prints flexible but not fully explored. This paper presents X-Bridges, an end-to-end workflow that allows novice users to design tunable bridges that can enrich 3D printed objects' deformable and physical properties. Specifically, we firstly provide a series of deformation primitives (e.g. bend, twist, coil, compress and stretch) with three versions of stiffness (loose, elastic, stable) based on parametrized bridging experiments. Embedding the printing parameters, a design tool is developed to modify the imported 3D model, evaluate optimized printing parameters for bridges, preview shape-changing process, and generate the G-code file for 3D printing. Finally, we demonstrate the design space of X-Bridges through a set of applications that enable foldable, resilient, and interactive shape-changing objects.
Lingyun Sun, Jiaji Li, Junzhe Ji, Deying Pan, Kuangqi Zhu, Yitao Fan, Yue Yang 0005, Ye Tao 0001, Guanyun Wang
UIST1
2022 Designing for trust: a set of design principles to increase trust in chatbot
Yunsan Guo, Runfan Wu, Lingyun Sun
CCF Trans. Pervasive Comput. Interact.5
2022 Few-shot font style transfer with multiple style encoders
Yonglin Wu, Yonggen Ling, Lingyun Sun, Yingming Li
Sci. China Inf. Sci.7
2022 Transparent-AI Blueprint: Developing a Conceptual Tool to Support the Design of Transparent AI Agents
abstract
With the increasing prevalence of artificial intelligence (AI) agents, the transparency of agents has become vital in addressing interaction issues (e.g., trust, usefulness, and understandability). However, determining the transparency of AI agents requires a systematic consideration of complex related factors, including stakeholders, algorithms, context, etc. Thus, in our study, we presented an overview of studies on the transparency of AI agents through multiple-stage bibliometric analysis, and identified an ontological framework of the key concepts relevant to transparent AI. We then built a Transparent-AI Blueprint prototype which is a diagram that visualizes the ontological framework of design concepts. In the subsequent pilot test, we updated Blueprint to the final version, and validated it in a workshop. Our work structurally summarized the design concepts related to the transparency of AI agents, and proposed a useful and practical conceptual design tool that effectively guides designers to operationalize the transparency of AI agents.
Zhibin Zhou 0002, Zhuoshu Li, Lingyun Sun
Int. J. Hum. Comput. Interact.4
2022 USIS: A unified semantic image synthesis model trained on a single or multiple samples
Pei Chen 0005, Zejian Li, Yangkang Zhang, Lingyun Sun
Neurocomputing5
2022 DPC-FSC: An approach of fuzzy semantic cells to density peaks clustering
abstract
Density peaks clustering (DPC) algorithm is a succinct and efficient density-based clustering approach to data analysis. It computes the local density and the relative distance for objects to seek cluster centers and form clusters. However, it is difficult to estimate an appropriate local density by a rule of thumb; therefore, the performance of DPC may be poor on real datasets in practice. Thus, this study proposes a novel method for density peaks clustering based on fuzzy semantic cells. Specifically, each object is coarsened into a fuzzy point with the form of the fuzzy semantic cell model. Based on this model, a local density metric is defined and estimated via the principle of justifiable granularity . Our local density estimation can be converted into an optimization problem . In addition, a relative semantic distance is also introduced which concerns the distance between fuzzy semantic cells. The relative semantic distance is more informative for selecting cluster centers in the decision graph. The experimental results show that our method not only exhibits higher performance but also provides a clearer decision graph to select cluster centers.
Lingyun Sun
Inf. Sci.2
2022 Visual knowledge guided intelligent generation of Chinese seal carving
abstract
We digitally reproduce the process of resource collaboration, design creation, and visual presentation of Chinese seal-carving art. We develop an intelligent seal-carving art-generation system (Zhejiang University Intelligent Seal-Carving System, http://www.next.zju.edu.cn/seal/ ; the website of the seal-carving search and layout system is http://www.next.zju.edu.cn/seal/search_app/ ) to deal with the difficulty in using a visual knowledge guided computational art approach. The knowledge base in this study is the Qiushi Seal-Carving Database, which consists of open datasets of images of seal characters and seal stamps. We propose a seal character generation method based on visual knowledge, guided by the database and expertise. Furthermore, to create the layout of the seal, we propose a deformation algorithm to adjust the seal characters and calculate layout parameters from the database and knowledge to achieve an intelligent structure. Experimental results show that this method and system can effectively deal with the difficulties in the generation of seal carving. Our work provides theoretical and applied references for the rebirth and innovation of seal-carving art.
Yehang Yin, Lingyun Sun, Huanghuang Deng, Yunhe Pan
Frontiers Inf. Technol. Electron. Eng.6
2022 EDVAM: a 3D eye-tracking dataset for visual attention modeling in a virtual museum
abstract
Predicting visual attention facilitates an adaptive virtual museum environment and provides a context-aware and interactive user experience. Explorations toward development of a visual attention mechanism using eye-tracking data have so far been limited to 2D cases, and researchers are yet to approach this topic in a 3D virtual environment and from a spatiotemporal perspective. We present the first 3D Eye-tracking Dataset for Visual Attention modeling in a virtual Museum, known as the EDVAM. In addition, a deep learning model is devised and tested with the EDVAM to predict a user’s subsequent visual attention from previous eye movements. This work provides a reference for visual attention modeling and context-aware interaction in the context of virtual museums.
Yunzhan Zhou, Tian Feng 0001, Shihui Shuai, Lingyun Sun, Henry Been-Lirn Duh
Frontiers Inf. Technol. Electron. Eng.5
2021 Shing: A Conversational Agent to Alert Customers of Suspected Online-payment Fraud with Empathetical Communication Skills
abstract
Alerting customers on suspected online-payment fraud and persuade them to terminate transactions is increasingly requested with the rapid growth of digital finance worldwide. We explored the feasibility of using a conversational agent (CA) to fulfill this request. Shing, a voice-based CA, proactively initializes and repairs the conversation with empathetical communication skills in order to alert customers when a suspected online-payment fraud is detected, collects important information for fraud scrutiny and persuades customers to terminate the transaction once the fraud is confirmed. We evaluated our system by comparing it with a rule-based CA with regards to customer response and perceptions in a real-world context where our systems took 144,795 phone calls in total in which 83,019 (57.3%) natural breakdowns happened. Results showed that more customers stopped risky transactions after conversing with Shing. They seemed more willing to converse with Shing for more dialogue turns and provide transaction details. Our work presents practical implications for the design of proactive CA.
Jingya Guo, Jiajing Guo, Chang-yuan Yang, Yanjing Wu, Lingyun Sun
CHI5
2021 FlexTruss: A Computational Threading Method for Multi-material, Multi-form and Multi-use Prototyping
abstract
3D printing, as a rapid prototyping technique, usually fabricates objects that are difficult to modify physically. This paper presents FlexTruss, a design and construction pipeline based on the assembly of modularized truss-shaped objects fabricated with conventional 3D printers and assembled by threading. To create an end-to-end system, a parametric design tool with an optimal Euler path calculation method is developed, which can support both inverse and forward design workflow and multi-material construction of modular parts. In addition, the assembly of truss modules by threading is evaluated with a series of application cases to demonstrate the affordance of FlexTruss. We believe that FlexTruss extends the design space of 3D printing beyond typically hard and fixed forms, and it will provide new capabilities for designers and researchers to explore the use of such flexible truss structures in human-object interaction.
Lingyun Sun, Jiaji Li, Yue Yang 0005, Danli Luo, Jianzhe Gu, Lining Yao, Ye Tao 0001, Guanyun Wang
CHI1
2021 ShrinCage: 4D Printing Accessories that Self-Adapt
abstract
3D printing technology makes Do-It-Yourself and reforming everyday objects a reality. However, designing and fabricating attachments that can seamlessly adapt existing objects to extended functionality is a laborious process, which requires accurate measuring, modeling, manufacturing, and assembly. This paper presents ShrinCage, a 4D printing system that allows novices to easily create shrinkable adaptations to fit and fasten existing objects. Specifically, the design tool presented in this work aid in the design of attachment that adapts to irregular morphologies, which accommodates the variations in measurements and fabrication, subsequently simplifying the modeling and assembly processes. We further conduct mechanical tests and user studies to evaluate the availability and feasibility of this method. Numerous application examples created by ShrinCage prove that it can be adopted by aesthetic modification, assistive technology, repair, upcycling, and augmented 3D printing.
Lingyun Sun, Yue Yang 0005, Jiaji Li, Danli Luo, Lining Yao, Ye Tao 0001, Guanyun Wang
CHI1
2021 Image Synthesis from Layout with Locality-Aware Mask Adaption
abstract
This paper is concerned with synthesizing images conditioned on a layout (a set of bounding boxes with object categories). Existing works construct a layout-mask-image pipeline. Object masks are generated separately and mapped to bounding boxes to form a whole semantic segmentation mask (layout-to-mask), with which a new image is generated (mask-to-image). However, overlapped boxes in layouts result in overlapped object masks, which reduces the mask clarity and causes confusion in image generation. We hypothesize the importance of generating clean and semantically clear semantic masks. The hypothesis is supported by the finding that the performance of state-of-the-art LostGAN decreases when input masks are tainted. Motivated by this hypothesis, we propose Locality-Aware Mask Adaption (LAMA) module to adapt overlapped or nearby object masks in the generation. Experimental results show our proposed model with LAMA outperforms existing approaches regarding visual fidelity and alignment with input layouts. On COCO-stuff in 256×256, our method improves the state-of-the-art FID score from 41.65 to 31.12 and the SceneFID from 22.00 to 18.64.
Zejian Li, Jingyu Wu, Immanuel Koh, Lingyun Sun
ICCV5
2021 A ML-Based Stock Trading Model for Profit Predication
Jimmy Ming-Tai Wu, Lingyun Sun, Gautam Srivastava 0001, Jerry Chun-Wei Lin
IEA/AIE (2)2
2021 Gated neural network framework for interactive character control
Xin Wang 0204, Xiaotao Jiang, Gloria Rumbidzai Regedzai, Haohao Meng, Lingyun Sun
Multim. Tools Appl.5
2021 F³A-GAN: Facial Flow for Face Animation With Generative Adversarial Networks
abstract
Formulated as a conditional generation problem, face animation aims at synthesizing continuous face images from a single source image driven by a set of conditional face motion. Previous works mainly model the face motion as conditions with 1D or 2D representation (e.g., action units, emotion codes, landmark), which often leads to low-quality results in some complicated scenarios such as continuous generation and large-pose transformation. To tackle this problem, the conditions are supposed to meet two requirements, i.e., motion information preserving and geometric continuity. To this end, we propose a novel representation based on a 3D geometric flow, termed facial flow, to represent the natural motion of the human face at any pose. Compared with other previous conditions, the proposed facial flow well controls the continuous changes to the face. After that, in order to utilize the facial flow for face editing, we build a synthesis framework generating continuous images with conditional facial flows. To fully take advantage of the motion information of facial flows, a hierarchical conditional framework is designed to combine the extracted multi-scale appearance features from images and motion features from flows in a hierarchical manner. The framework then decodes multiple fused features back to images progressively. Experimental results demonstrate the effectiveness of our method compared to other state-of-the-art methods.
Xintian Wu, Qihang Zhang, Yiming Wu 0005, Lingyun Sun, Xi Li 0001
IEEE Trans. Image Process.6
2020 ML Lifecycle Canvas: Designing Machine Learning-Empowered UX with Material Lifecycle Thinking
abstract
As a particular type of artificial intelligence technology, machine learning (ML) is widely used to empower user experience (UX). However, designers, especially the novice designers, struggle to integrate ML into familiar design activities because of its ever-changing and growable nature. This paper proposes a design method called Material Lifecycle Thinking (MLT) that considers ML as a design material with its own lifecycle. MLT encourages designers to regard ML, users, and scenarios as three co-creators who cooperate in creating ML-empowered UX. We have developed ML Lifecycle Canvas (Canvas), a conceptual design tool that incorporates visual representations of the co-creators and ML lifecycle. Canvas guides designers to organize essential information for the application of MLT. By involving design students in the “research through design” process, the development of Canvas was iterated through its application to design projects. MLT and Canvas have been evaluated in design workshops, with completed proposals and evaluation results demonstrating that our work is a solid step forward in bridging the gap between UX and ML.
Zhibin Zhou 0002, Lingyun Sun, Xuanhui Liu, Qing Gong
Hum. Comput. Interact.2
2020 Automatic synthesis of advertising images according to a specified style
abstract
Images are widely used by companies to advertise their products and promote awareness of their brands. The automatic synthesis of advertising images is challenging because the advertising message must be clearly conveyed while complying with the style required for the product, brand, or target audience. In this study, we proposed a data-driven method to capture individual design attributes and the relationships between elements in advertising images with the aim of automatically synthesizing the input of elements into an advertising image according to a specified style. To achieve this multi-format advertisement design, we created a dataset containing 13 280 advertising images with rich annotations that encompassed the outlines and colors of the elements, in addition to the classes and goals of the advertisements. Using our probabilistic models, users guided the style of synthesized advertisements via additional constraints (e.g., context-based keywords). We applied our method to a variety of design tasks, and the results were evaluated in several perceptual studies, which showed that our method improved users’ satisfaction by 7.1% compared to designs generated by nonprofessional students, and that more users preferred the coloring results of our designs to those generated by the color harmony model and Colormind.
Weitao You, Hao Jiang 0046, Zhi-Yuan Yang, Chang-yuan Yang, Lingyun Sun
Frontiers Inf. Technol. Electron. Eng.5
2020 A music-driven system for generating apparel display video
Hui Zhang 0064, Yingping Cao, Xiaoyi Huang, Chang-yuan Yang, Lingyun Sun
Multim. Tools Appl.7
2019 A Survey of Users' Expectations Towards On-body Companion Robots
abstract
Being as a robotic companion is an extensive application of on-body robots; yet, as an emerging type of robots, few previous works focus on the design of on-body companion robots from the users' perspective, remaining users' expectations towards this type of robots unclear. To assist designers in the design process of on-body companion robots, we surveyed users' expectations towards on-body companion robots (n=215) by a questionnaire constituting of questions on factors that may affect robot acceptance, including robot functionality, robot appearance, and robot social ability. Based on the survey results, we stated design guidelines for the design of on-body companion robots supporting designers with insights into users. To demonstrate how to design on-body companion robots based on our findings, we organized a workshop with experienced designers to develop a conceptual on-body companion robot, and they proposed Bubo, an example prototype of on-body companion robot.
Hao Jiang 0046, Siyuan Lin, Veerajagadheswar Prabakaran, Mohan Rajesh Elara, Lingyun Sun
Conference on Designing Interactive Systems5
2019 SmartPaint: a co-creative drawing system based on generative adversarial networks
abstract
Artificial intelligence (AI) has played a significant role in imitating and producing large-scale designs such as e-commerce banners. However, it is less successful at creative and collaborative design outputs. Most humans express their ideas as rough sketches, and lack the professional skills to complete pleasing paintings. Existing AI approaches have failed to convert varied user sketches into artistically beautiful paintings while preserving their semantic concepts. To bridge this gap, we have developed SmartPaint, a co-creative drawing system based on generative adversarial networks (GANs), enabling a machine and a human being to collaborate in cartoon landscape painting. SmartPaint trains a GAN using triples of cartoon images, their corresponding semantic label maps, and edge detection maps. The machine can then simultaneously understand the cartoon style and semantics, along with the spatial relationships among the objects in the landscape images. The trained system receives a sketch as a semantic label map input, and automatically synthesizes its edge map for stable handling of varied sketches. It then outputs a creative and fine painting with the appropriate style corresponding to the human’s sketch. Experiments confirmed that the proposed SmartPaint system successfully generates high-quality cartoon paintings.
Lingyun Sun, Pei Chen 0005, Wei Xiang 0008, Wei-yue Gao
Frontiers Inf. Technol. Electron. Eng.1
2018 The PMEmo Dataset for Music Emotion Recognition
abstract
Music Emotion Recognition (MER) has recently received considerable attention. To support the MER research which requires large music content libraries, we present the PMEmo dataset containing emotion annotations of 794 songs as well as the simultaneous electrodermal activity (EDA) signals. A Music Emotion Experiment was well-designed for collecting the affective-annotated music corpus of high quality, which recruited 457 subjects.
Hui Zhang 0064, Chang-yuan Yang, Lingyun Sun
ICMR5
2018 Crowdsourcing intelligent design
abstract
Design intelligence, namely, artificial intelligence to solve creative problems and produce creative ideas, has improved rapidly with the new generation artificial intelligence. However, existing methods are more skillful in learning from data and have limitations in creating original ideas different from the training data. Crowdsourcing offers a promising method to produce creative designs by combining human inspiration and machines’ computational ability. We propose a crowdsourcing intelligent design method called ‘flexible crowdsourcing design’. Design ideas produced through crowdsourcing design can be unreliable and inconsistent because they rely solely on selection among participants’ submissions of ideas. In contrast, the flexible crowdsourcing design method employs a cultivation procedure that integrates the ideas from crowd participants and cultivates these ideas to improve design quality at the same time. We introduce a series of studies to show how flexible crowdsourcing design can produce original design ideas consistently. Specifically, we will describe the typical procedure of flexible crowdsourcing design, the refined crowdsourcing tasks, the factors that affect the idea development process, the method for calculating idea development potential, and two applications of the flexible crowdsourcing design method. Finally, it summarizes the design capabilities enabled by crowdsourcing intelligent design. This method enhances the performance of crowdsourcing design and supports the development of design intelligence.
Wei Xiang 0008, Lingyun Sun, Weitao You, Chang-yuan Yang
Frontiers Inf. Technol. Electron. Eng.2
2018 A one-to-many conditional generative adversarial network framework for multiple image-to-image translations
Chunlei Chai, Jing Liao 0001, Ning Zou, Lingyun Sun
Multim. Tools Appl.4
2018 Cross-objects user interfaces for video interaction in virtual reality museum context
Lingyun Sun, Yunzhan Zhou, Preben Hansen, Weidong Geng
Multim. Tools Appl.1
2016 Odor emoticon: An olfactory application that conveys emotions
Wei Xiang 0008, Shi Chen 0005, Lingyun Sun, Shiwei Cheng 0001, V. Michael Bove Jr.
Int. J. Hum. Comput. Stud.3
2015 Gaze-Based Annotations for Reading Comprehension
abstract
We study eye gaze movement behavior during paper reading and generate a series of annotations from a user's reading features: gray shading to indicate reading speed, borders to indicate frequency of re-reading, and lines to indicate transitions between sections of a document. Through a user study, we validate that our SocialReading system that shares teachers' gaze data for an academic paper can improve students' reading comprehension of that paper.
Shiwei Cheng 0001, Lingyun Sun, Kirsten Yee, Anind K. Dey
CHI3