EDBT 2026 Demo / reviewers in the wild / expert
Pei Chen 0005
dblp:98/4148-5
· DBLP profile ↗
27ranked-venue papers
6as first author
26since 2021 · last 2026
0000-0003-0962-6459ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Human-computer interaction and ubiquitous computing · 14 · 3 first-author · 14 since 2021Artificial intelligence and machine learning · 10 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion Distillation with Direct Preference Optimization for Efficient 3D LiDAR Scene CompletionabstractThe slow sampling speed of diffusion models hinders their application in 3D LiDAR scene completion. To address this, we propose Distillation-DPO, a novel framework that accelerates sampling through score distillation while simultaneously enhancing generation quality via preference alignment. Distillation-DPO follows a three-step procedure. First, the student model generates paired completion scenes with different initial noises. Second, using LiDAR scene evaluation metrics as preference, we construct winning and losing sample pairs. Third, as our core innovation, Distillation-DPO optimizes the student model by exploiting the difference in score functions between the teacher and student models on the paired completion scenes. This operation performs variational score distillation of the student model but simultaneously encourages the distilled student to prefer the winning samples over the losing ones. Extensive experiments demonstrate that Distillation-DPO achieves higher-quality scene completion than state-of-the-art diffusion models, while accelerating sampling by over 5-fold. To our knowledge, our work is the first to integrate the preference learning principle of DPO into the distillation of diffusion models, offering a new framework of preference-aligned distillation. Shengyuan Zhang, Zejian Li, Ling Yang 0006, Pei Chen 0005, Anyang Wei, Perry Pengyun Gu, Lingyun Sun |
AAAI | 5 |
| 2026 | ThinkPersona: Thinking with Persona Graphs for Faithful Individualized Role-PlayingabstractYichen Cai, Pei Chen, Jiayang Li, Jingya Guo, Zejian Li, Changyuan Yang, Lingyun Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yichen Cai 0005, Pei Chen 0005, Jiayang Li 0003, Jingya Guo, Zejian Li, Chang-yuan Yang, Lingyun Sun |
ACL (1) | 2 |
| 2026 | IEvoAgent: Evolving Conversational Agent based on User Implicit FeedbackabstractYichen Cai, Jiayang Li, Junyuan Qiu, Jingya Guo, Weitao You, Changyuan Yang, Lingyun Sun, Pei Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yichen Cai 0005, Jiayang Li 0003, Junyuan Qiu, Jingya Guo, Weitao You, Chang-yuan Yang, Lingyun Sun, Pei Chen 0005 |
ACL (1) | 8 |
| 2026 | From Experts to Bases: Orthogonal Subspace Mixture for Continual Multimodal Instruction TuningabstractMultimodal Continual Instruction Tuning (MCIT) is essential for adapting Multimodal Large Language Models (MLLMs) to dynamic data streams, yet preventing catastrophic forgetting remains a major challenge.Existing parameter-efficient approaches often face a dilemma: fixed architectures suffer from knowledge interference, while dynamic strategies incur inefficient capacity expansion, limiting scalability.We propose MoBLoRA (Mixture-of-Bases LoRA), a novel framework for MCIT.Motivated by our geometric analysis revealing subspace redundancy across sequential tasks, MoBLoRA shifts the paradigm from expert selection to subspace mixing: it decomposes adaptation weights into a globally shared pool of orthonormal bases to capture task-invariant knowledge, and lightweight mixing matrices to encode task-specific variations.This design effectively decouples knowledge accumulation from task reconstruction.Experiments on standard benchmarks show MoBLoRA significantly outperforms state-of-the-art methods while maintaining superior parameter efficiency. 1 Pei Chen 0005, Xilai Wang, Qixu Shi, Zejian Li, Lingyun Sun |
ACL (1) | 1 |
| 2026 | Does Sycophancy Change Decisions? Effect of LLM Sycophancy on AI-Assisted Decision-MakingabstractLarge language models are increasingly integrated into everyday and professional decision making, yet often exhibit sycophantic behavior by aligning with users’ views or preferences. While sycophancy can enhance interaction, its influence on users’ decisions remain unclear given different styles and task risks. We examine three forms of sycophancy—opinion agreement, direct praise, and self-deprecation—in two contrasting contexts: a low-risk speed-dating prediction task and a high-risk ETF investment task. In a 4×2 mixed-design online study (N = 106), we compare non-sycophantic AI with sycophantic variants on decision outcomes and confidence changes. Results show that sycophancy influences decision patterns in type-dependent ways. Specifically, opinion agreement reinforces initial decisions and self-deprecation boosts confidence. Interviews further indicate that users value supportive AI but question its objectivity when praise becomes excessive. These findings reveal the multifaceted effects of AI sycophancy and offer design implications for balancing support and credibility in human–AI interaction. Zejian Li, Jiaman Pan, Qi Liu 0076, Yuning Xi, Yixiang Zhou, Yike Jin, Rongjie Mao, Pei Chen 0005 |
CHI | 8 |
| 2026 | Seeing Eye to Eye: Enabling Cognitive Alignment Through Shared First-Person Perspective in Human-AI Collaboration: Seeing Eye to EyeabstractDespite advances in multimodal AI, current vision-based assistants often remain inefficient in collaborative tasks. We identify two key gulfs: a communication gulf, where users must translate rich parallel intentions into verbal commands due to the channel mismatch, and an understanding gulf, where AI struggles to interpret subtle embodied cues. To address these, we propose Eye2Eye, a framework that leverages first-person perspective as a channel for human-AI cognitive alignment. It integrates three components: (1) joint attention coordination for fluid focus alignment, (2) revisable memory to maintain evolving common ground, and (3) reflective feedback allowing users to clarify and refine AI’s understanding. We implement this framework in an AR prototype and evaluate it through a user study and a post-hoc pipeline evaluation. Results show that Eye2Eye significantly reduces task completion time and interaction load while increasing trust, demonstrating its components work in concert to improve collaboration. Zhuyu Teng, Pei Chen 0005, Yichen Cai 0005, Ruoqing Lu, Zhaoqu Jiang, Jiayang Li 0003, Weitao You, Lingyun Sun |
CHI | 2 |
| 2026 | PoemPalette: Facilitating Poetry Creative Exploration and Foundational Understanding through the Ideorealm Alignment of Paintings and PoemsabstractThe “Ideorealm Alignment of Paintings and Poems (IA-PP)” theory rooted in Chinese classical aesthetics offers a perspective for exploring poetry’s deep connotations. This study presents PoemPalette, a novel IA-PP creative-exploration tool that integrates Generative Artificial Intelligence (GAI) to guide poetry enthusiasts in actively constructing an ideorealm for the poetic painting they envision, informed by a formative study with six experts. We extract the core symbols of poetry, transform them into Scene Graph (SG), and generate images for users to freely compose, enabling IA-PP creative exploration. The system incorporates Large Language Model (LLM) agents to enhance the foundational understanding of poetry. In a controlled experiment on Chinese poetry and Japanese haiku with 60 participants, we analyze which interaction mechanisms most contribute to foundational understanding and creative outcomes, compared with both AI and non-AI baselines. Situated within East Asian poetry traditions, this study introduces cultural theories to guide the design of AI co-creation tools, using a graph-based interface of interpretable intermediate representations. Ying Zhang 0076, Kaixin Jia, Hong Jian Zhang, Kewen Zhu, Chenye Meng, Jiesi Zhang, Zejian Li, Pei Chen 0005, Lingyun Sun |
CHI | 8 |
| 2026 | ImmersiProtor: A Collaborative Mixed-Prototype Tool Integrating Spatial Augmented Reality and Component-layered GenerationabstractConceptual design is a critical stage in product development, which is a co-design process involving multidisciplinary collaboration based on prototypes. In this paper, we aim to propose a novel prototype paradigm that combines the distinct strengths of generative artificial intelligence (GAI) and spatial augmented reality (SAR), leveraging the expressive potential of SAR and the creative potential of GAI for co-design. To achieve this, we initially conducted a formative study with designers to explore how these technologies could be effectively combined to facilitate co-design. Based on our findings, we introduce ImmersiProtor, a prototype tool integrating multi-view SAR and component-layered GAI for co-design. On one hand, ImmersiProtor allows design team members to freely create and modify physical prototypes while automatically generating multi-view and high-fidelity renderings that are projected onto the surfaces of the physical prototype using SAR technology, enabling immersive communication and intuitive evaluation. On the other hand, ImmersiProtor introduces a component-layered generation and collaboration mode, offering both personal and shared team component resources. It ensures that individual team members can explore ideas independently without interference, while also supporting concept integration, evaluation, and iteration. We implemented ImmersiProtor, which involves a web-based application and an SAR design space. We conducted a user study to verify ImmersiProtor’s usability in supporting prototype and collaboration. Our results highlighted ImmersiProtor’s inherent strengths in enhancing intuition, promoting collaboration, and strengthening GAI controllability. We also explored the effect of mixed interaction on design and critically discuss its best practices for the HCI community. Chaoyi Lin, Weitao You, Lingyun Sun, Pei Chen 0005 |
CHI | 6 |
| 2025 | Personalized Dynamic Music Emotion Recognition with Dual-Scale Attention-Based Meta-LearningabstractDynamic Music Emotion Recognition (DMER) aims to predict the emotion of different moments in music, playing a crucial role in music information retrieval. The existing DMER methods struggle to capture long-term dependencies when dealing with sequence data, which limits their performance. Furthermore, these methods often overlook the influence of individual differences on emotion perception, even though everyone has their own personalized emotional perception in the real world. Motivated by these issues, we explore more effective sequence processing methods and introduce the Personalized DMER (PDMER) problem, which requires models to predict emotions that align with personalized perception. Specifically, we propose a Dual-Scale Attention-Based Meta-Learning (DSAML) method. This method fuses features from a dual-scale feature extractor and captures both short and long-term dependencies using a dual-scale attention transformer, improving the performance in traditional DMER. To achieve PDMER, we design a novel task construction strategy that divides tasks by annotators. Samples in a task are annotated by the same annotator, ensuring consistent perception. Leveraging this strategy alongside meta-learning, DSAML can predict personalized perception of emotions with just one personalized annotation sample. Our objective and subjective experiments demonstrate that our method can achieve state-of-the-art performance in both traditional DMER and PDMER. Dengming Zhang, Weitao You, Lingyun Sun, Pei Chen 0005 |
AAAI | 5 |
| 2025 | GVMGen: A General Video-to-Music Generation Model with Hierarchical AttentionsabstractComposing music for video is essential yet challenging, leading to a growing interest in automating music generation for video applications. Existing approaches often struggle to achieve robust music-video correspondence and generative diversity, primarily due to inadequate feature alignment methods and insufficient datasets. In this study, we present General Video-to-Music Generation model (GVMGen), designed for generating high-related music to the video input. Our model employs hierarchical attentions to extract and align video features with music in both spatial and temporal dimensions, ensuring the preservation of pertinent features while minimizing redundancy. Remarkably, our method is versatile, capable of generating multi-style music from different video inputs, even in zero-shot scenarios. We also propose an evaluation model along with two novel objective metrics for assessing video-music alignment. Additionally, we have compiled a large-scale dataset comprising diverse types of video-music pairs. Experimental results demonstrate that GVMGen surpasses previous models in terms of music-video correspondence, music quality generative diversity, and application universality. Heda Zuo, Weitao You, Junxian Wu 0003, Shihong Ren, Pei Chen 0005, Mingxu Zhou, Yujia Lu, Lingyun Sun |
AAAI | 5 |
| 2025 | CoExploreDS: Framing and Advancing Collaborative Design Space Exploration Between Human and AI
Pei Chen 0005, Zhuoyi Cheng, Yichen Cai 0005, Jiayang Li 0003, Weitao You, Lingyun Sun |
CHI | 1 |
| 2025 | FusionProtor: A Mixed-Prototype Tool for Component-level Physical-to-Virtual 3D Transition and Simulation
Pei Chen 0005, Xuelong Xie, Zhaoqu Jiang, Zejian Li, Lingyun Sun |
CHI | 2 |
| 2025 | IEDS: Exploring an Intelli-Embodied Design Space Combining Designer, AR, and GAI to Support Industrial Conceptual Design
Pei Chen 0005, Zhaoqu Jiang, Xuelong Xie, Weitao You, Lingyun Sun |
CHI | 2 |
| 2025 | Intoner: For Chinese Poetry Intoning SynthesisabstractChinese Poetry Intoning, with improvised melodies devoid of fixed musical scores, is crucial for emotional expression and prosodic rendition. However, this cultural heritage faces challenges in propagation due to scant audio records and a scarcity of domain experts. Existing text-to-speech models lack the ability to generate melodious audio, while singing-voice-synthesis models rely on predetermined musical scores, which are all unsuitable for intoning synthesis. Hence, we introduce Chinese Poetry Intoning Synthesis (PIS) as a novel task to reproduce intoning audio and preserve this age-old cultural art. Corresponding to this task, we summarize three-level principles from poetry metrical patterns and construct a diffusion PIS model Intoner based on them. We also collect a multi-style Chinese poetry intoning dataset of text-audio pairs accompanied by feature annotations. Experimental results show that our model effectively learns diverse intoning styles and contents which can synthesize more melodious and vibrant intoning audio. To the best of our knowledge, we are the first to work on poetry intoning synthesis task. Heda Zuo, Liyao Sun, Zeyu Lai, Weitao You, Pei Chen 0005, Lingyun Sun |
IJCAI | 5 |
| 2025 | An Exploratory Study on How AI Awareness Impacts Human-AI Design Collaboration
Zhuoyi Cheng, Pei Chen 0005, Wenzheng Song, Zhuoshu Li, Lingyun Sun |
IUI | 2 |
| 2025 | Controllable Video-to-Music Generation with Multiple Time-Varying ConditionsabstractMusic enhances video narratives and emotions, driving demand for automatic video-to-music (V2M) generation. However, existing V2M methods relying solely on visual features or supplementary textual inputs generate music in a black-box manner, often failing to meet user expectations. To address this challenge, we propose a novel multi-condition guided V2M generation framework that incorporates multiple time-varying conditions for enhanced control over music generation. Our method uses a two-stage training strategy that enables learning of V2M fundamentals and audiovisual temporal synchronization while meeting users' needs for multi-condition control. In the first stage, we introduce a fine-grained feature selection module and a progressive temporal alignment attention mechanism to ensure flexible feature alignment. For the second stage, we develop a dynamic conditional fusion module and a control-guided decoder module to integrate multiple conditions and accurately guide the music composition process. Extensive experiments demonstrate that our method outperforms existing V2M pipelines in both subjective and objective evaluations, significantly enhancing control and alignment with user expectations. Junxian Wu 0003, Weitao You, Heda Zuo, Dengming Zhang, Pei Chen 0005, Lingyun Sun |
ACM Multimedia | 5 |
| 2025 | GPSdesign: Integrating Generative AI with Problem-Solution Co-Evolution Network to Support Product Conceptual DesignabstractIn conceptual design, designers often face the challenge of navigating vast design spaces to define ambiguous problems and generate feasible solutions. Recent advancements in generative artificial intelligence (GenAI) offer new opportunities to support this process. However, formative research revealed that designers struggle to simultaneously advance both problem and solution spaces when using GenAI in conceptual design, leading to increased communication load and diminished solution practicality. This study explores the integration of GenAI with the problem-solution co-evolution model to facilitate the construction of a structured design space. We propose a GenAI-supported method for expanding and evaluating the design space and developed the GPSdesign system based on this method. Compared with a baseline system, GPSdesign fosters greater design space divergence, retrospection, and structured construction, while improving design efficiency and solution quality. Pei Chen 0005, Yexinrui Wu, Zhuoshu Li, Mingxu Zhou, Weitao You, Lingyun Sun |
Int. J. Hum. Comput. Interact. | 1 |
| 2025 | Exploring the role of Mixed Reality on Design Representations to Enhance User-Involved Co-Design CommunicationabstractAs users transition from passive subjects to active partners in the co-design process, they bring unique insights based on their experiences, collaboratively envisioning a better future with designers. However, unlike designers who are adept at various forms of representation, most users lack advanced modeling or sketching skills to concretely present the three-dimensional (3D) forms or dynamic features of a design proposal. This hinders user expression and increases the cognitive load on designers, thereby reducing communication efficiency in the co-design process. Mixed Reality (MR) technology enables users to depict 3D information in real physical space using natural gestures. This means that MR can provide a low-learning-cost concrete expression method without compromising traditional communication methods. This study explores the role of MR in enhancing communication between designers and users during the early stages of design. A formative study was conducted to identify four key requirements, which informed the development of the DuoMR system. DuoMR supports designers and users in expressing design ideas through gesture modeling in a collaborative MR space. Results from the user study and practical case study show that DuoMR effectively reduces cognitive load and enhances mutual understanding during the co-design process. Pei Chen 0005, Kexing Wang, Lianyan Liu, Xuanhui Liu, Zhuyu Teng, Lingyun Sun |
Proc. ACM Hum. Comput. Interact. | 1 |
| 2025 | DesignManager: An Agent-Powered Copilot for Designers to Integrate AI Design Tools into Creative WorkflowsabstractCreative design is an inherently complex and iterative process characterized by continuous exploration, evaluation, and refinement. While recent advances in generative AI have demonstrated remarkable potential in supporting specific design tasks, there remains a critical gap in understanding how these technologies can enhance the holistic design process rather than just isolated stages. This paper introduces DesignManager, a novel AI-powered design support system that aims to transform how designers collaborate with AI throughout their creative workflow. Through a formative study examining designers' current practices with generative AI, we identified key challenges and opportunities in integrating AI into the creative design process. Based on these insights, we developed DesignManager as an interactive copilot system that provides node-based visualization of design evolution, enabling designers to track, modify, and branch their design processes while maintaining meaningful dialogue-based collaboration. The system offers two collaboration modes: DesignManager-guiding and Designer-guiding. Designers can engage in conversational interactions with the DesignManager to obtain design inspiration and tool recommendations, and proactively advance the design progress. The system employs an agent framework to manage decoupled contextual information emerged during the design process, facilitating deep understanding of designers' needs and providing context-aware assistance. Our technical evaluation validated the effectiveness of context decoupling and the use of agent framework, while the open-ended user study with experts demonstrated that DesignManager successfully supports intuitive intention expression, flexible process control, and deeper creative articulation. This work contributes to the understanding of how AI can evolve from task-specific tools to collaborative partners in creative design processes. Weitao You, Yinyu Lu, Zirui Ma, Nan Li 0085, Mingxu Zhou, Pei Chen 0005, Lingyun Sun |
ACM Trans. Graph. | 7 |
| 2024 | ProtoDreamer: A Mixed-prototype Tool Combining Physical Model and Generative AI to Support Conceptual DesignabstractPrototyping serves as a critical phase in the industrial conceptual design process, enabling exploration of problem space and identification of solutions. Recent advancements in large-scale generative models have enabled AI to become a co-creator in this process. However, designers often consider generative AI challenging due to the necessity to follow computer-centered interaction rules, diverging from their familiar design materials and languages. Physical prototype is a commonly used design method, offering unique benefits in prototype process, such as intuitive understanding and tangible testing. In this study, we propose ProtoDreamer, a mixed-prototype tool that synergizes generative AI with physical prototype to support conceptual design. ProtoDreamer allows designers to construct preliminary prototypes using physical materials, while AI recognizes these forms and vocal inputs to generate diverse design alternatives. This tool empowers designers to tangibly interact with prototypes, intuitively convey design intentions to AI, and continuously draw inspiration from the generated artifacts. An evaluation study confirms ProtoDreamer’s utility and strengths in time efficiency, creativity support, defects exposure, and detailed thinking facilitation. Pei Chen 0005, Xuelong Xie, Chaoyi Lin, Lianyan Liu, Zhuoshu Li, Weitao You, Lingyun Sun |
UIST | 2 |
| 2024 | StyleFactory: Towards Better Style Alignment in Image Creation through Style-Strength-Based Control and EvaluationabstractGenerative AI models have been widely used for image creation. However, generating images that are well-aligned with users’ personal styles on aesthetic features (e.g., color and texture) can be challenging due to the poor style expression and interpretation between humans and models. Through a formative study, we observed that participants showed a clear subjective perception of the desired style and variations in its strength, which directly inspired us to develop style-strength-based control and evaluation. Building on this, we present StyleFactory, an interactive system that helps users achieve style alignment. Our interface enables users to rank images based on their strengths in the desired style and visualizes the strength distribution of other images in that style from the model’s perspective. In this way, users can evaluate the understanding gap between themselves and the model, and define well-aligned personal styles for image creation through targeted iterations. Our technical evaluation and user study demonstrate that StyleFactory accurately generates images in specific styles, effectively facilitates style alignment in image creation workflow, stimulates creativity, and enhances the user experience in human-AI interactions. Mingxu Zhou, Dengming Zhang, Weitao You, Chenghao Pan, Tianyu Lao, Pei Chen 0005 |
UIST | 9 |
| 2024 | Elicitation and Evaluation of Hand-based Interaction Language for 3D Conceptual Design in Mixed Reality
Lingyun Sun, Pei Chen 0005, Zhaoqu Jiang, Xuelong Xie, Zihong Zhou, Xuanhui Liu |
Int. J. Hum. Comput. Stud. | 3 |
| 2024 | A Hybrid Prototype Method Combining Physical Models and Generative Artificial Intelligence to Support Creativity in Conceptual DesignabstractConceptual design is an essential stage in the design process, and its ultimate success largely depends on designers’ creativity. Both physical and digital prototypes are commonly adopted by designers to support ideation and creativity, providing intuitive perception and rapid iteration, respectively. In recent advancements, large-scale generation models are able to offer data-enabled creativity support by generating high-quality solutions comparable to human designers. This opens up an imaginary space for designers and brings new possibilities for design tools. In this study, we proposed a hybrid prototype method that synergistically combines physical models and generative artificial intelligence (AI) in the conceptual design stage. Correspondingly, we developed a hybrid prototype system to implement the proposed method. We conducted a comparative user study with 45 designers who completed a design task using the physical prototype method, standalone generative AI and the hybrid prototype method, respectively. Our results verified the effectiveness of the hybrid prototype method and investigated its mechanism in supporting creativity. Finally, we discussed the application value and optimisation space of the hybrid prototype method. Pei Chen 0005, Xuelong Xie, Zhaoqu Jiang, Zihong Zhou, Lingyun Sun |
ACM Trans. Comput. Hum. Interact. | 2 |
| 2023 | Learning Object Consistency and Interaction in Image Generation from Scene GraphsabstractThis paper is concerned with synthesizing images conditioned on a scene graph (SG), a set of object nodes and their edges of interactive relations. We divide existing works into image-oriented and code-oriented methods. In our analysis, the image-oriented methods do not consider object interaction in spatial hidden feature. On the other hand, in empirical study, the code-oriented methods lose object consistency as their generated images miss certain objects in the input scene graph. To alleviate these two issues, we propose Learning Object Consistency and Interaction (LOCI). To preserve object consistency, we design a consistency module with a weighted augmentation strategy for objects easy to be ignored and a matching loss between scene graphs and image codes. To learn object interaction, we design an interaction module consisting of three kinds of message propagation between the input scene graph and the learned image code. Experiments on COCO-stuff and Visual Genome datasets show our proposed method alleviates the ignorance of objects and outperforms the state-of-the-art on visual fidelity of generated images and objects. Yangkang Zhang, Chenye Meng, Zejian Li, Pei Chen 0005, Guang Yang 0022, Chang-yuan Yang, Lingyun Sun |
IJCAI | 4 |
| 2022 | Few-Shot Incremental Learning for Label-to-Image TranslationabstractLabel-to-image translation models generate images from semantic label maps. Existing models depend on large volumes of pixel-level annotated samples. When given new training samples annotated with novel semantic classes, the models should be trained from scratch with both learned and new classes. This hinders their practical applications and motivates us to introduce an incremental learning strategy to the label-to-image translation scenario. In this paper, we introduce a few-shot incremental learning method for label-to-image translation. It learns new classes one by one from a few samples of each class. We propose to adopt semantically-adaptive convolution filters and normalization. When incrementally trained on a novel semantic class, the model only learns a few extra parameters of class-specific modulation. Such design avoids catastrophic forgetting of already-learned semantic classes and enables label-to-image translation of scenes with increasingly rich content. Furthermore, to facilitate few-shot learning, we propose a modulation transfer strategy for better initialization. Extensive experiments show that our method outperforms existing related methods in most cases and achieves zero forgetting. Pei Chen 0005, Yangkang Zhang, Zejian Li, Lingyun Sun |
CVPR | 1 |
| 2022 | USIS: A unified semantic image synthesis model trained on a single or multiple samples
Pei Chen 0005, Zejian Li, Yangkang Zhang, Lingyun Sun |
Neurocomputing | 1 |
| 2019 | SmartPaint: a co-creative drawing system based on generative adversarial networksabstractArtificial intelligence (AI) has played a significant role in imitating and producing large-scale designs such as e-commerce banners. However, it is less successful at creative and collaborative design outputs. Most humans express their ideas as rough sketches, and lack the professional skills to complete pleasing paintings. Existing AI approaches have failed to convert varied user sketches into artistically beautiful paintings while preserving their semantic concepts. To bridge this gap, we have developed SmartPaint, a co-creative drawing system based on generative adversarial networks (GANs), enabling a machine and a human being to collaborate in cartoon landscape painting. SmartPaint trains a GAN using triples of cartoon images, their corresponding semantic label maps, and edge detection maps. The machine can then simultaneously understand the cartoon style and semantics, along with the spatial relationships among the objects in the landscape images. The trained system receives a sketch as a semantic label map input, and automatically synthesizes its edge map for stable handling of varied sketches. It then outputs a creative and fine painting with the appropriate style corresponding to the human’s sketch. Experiments confirmed that the proposed SmartPaint system successfully generates high-quality cartoon paintings. Lingyun Sun, Pei Chen 0005, Wei Xiang 0008, Wei-yue Gao |
Frontiers Inf. Technol. Electron. Eng. | 2 |