VLDB 2026 Research / reviewers in the wild / expert
Zejian Li
dblp:211/6506
· DBLP profile ↗
41ranked-venue papers
8as first author
36since 2021 · last 2026
0000-0001-5313-2742ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 20 · 5 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 16 · 4 first-author · 13 since 2021Human-computer interaction and ubiquitous computing · 8 · 2 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diffusion Distillation with Direct Preference Optimization for Efficient 3D LiDAR Scene CompletionabstractThe slow sampling speed of diffusion models hinders their application in 3D LiDAR scene completion. To address this, we propose Distillation-DPO, a novel framework that accelerates sampling through score distillation while simultaneously enhancing generation quality via preference alignment. Distillation-DPO follows a three-step procedure. First, the student model generates paired completion scenes with different initial noises. Second, using LiDAR scene evaluation metrics as preference, we construct winning and losing sample pairs. Third, as our core innovation, Distillation-DPO optimizes the student model by exploiting the difference in score functions between the teacher and student models on the paired completion scenes. This operation performs variational score distillation of the student model but simultaneously encourages the distilled student to prefer the winning samples over the losing ones. Extensive experiments demonstrate that Distillation-DPO achieves higher-quality scene completion than state-of-the-art diffusion models, while accelerating sampling by over 5-fold. To our knowledge, our work is the first to integrate the preference learning principle of DPO into the distillation of diffusion models, offering a new framework of preference-aligned distillation. Shengyuan Zhang, Zejian Li, Ling Yang 0006, Pei Chen 0005, Anyang Wei, Perry Pengyun Gu, Lingyun Sun |
AAAI | 3 |
| 2026 | ThinkPersona: Thinking with Persona Graphs for Faithful Individualized Role-PlayingabstractYichen Cai, Pei Chen, Jiayang Li, Jingya Guo, Zejian Li, Changyuan Yang, Lingyun Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yichen Cai 0005, Pei Chen 0005, Jiayang Li 0003, Jingya Guo, Zejian Li, Chang-yuan Yang, Lingyun Sun |
ACL (1) | 5 |
| 2026 | From Experts to Bases: Orthogonal Subspace Mixture for Continual Multimodal Instruction TuningabstractMultimodal Continual Instruction Tuning (MCIT) is essential for adapting Multimodal Large Language Models (MLLMs) to dynamic data streams, yet preventing catastrophic forgetting remains a major challenge.Existing parameter-efficient approaches often face a dilemma: fixed architectures suffer from knowledge interference, while dynamic strategies incur inefficient capacity expansion, limiting scalability.We propose MoBLoRA (Mixture-of-Bases LoRA), a novel framework for MCIT.Motivated by our geometric analysis revealing subspace redundancy across sequential tasks, MoBLoRA shifts the paradigm from expert selection to subspace mixing: it decomposes adaptation weights into a globally shared pool of orthonormal bases to capture task-invariant knowledge, and lightweight mixing matrices to encode task-specific variations.This design effectively decouples knowledge accumulation from task reconstruction.Experiments on standard benchmarks show MoBLoRA significantly outperforms state-of-the-art methods while maintaining superior parameter efficiency. 1 Pei Chen 0005, Xilai Wang, Qixu Shi, Zejian Li, Lingyun Sun |
ACL (1) | 4 |
| 2026 | Does Sycophancy Change Decisions? Effect of LLM Sycophancy on AI-Assisted Decision-MakingabstractLarge language models are increasingly integrated into everyday and professional decision making, yet often exhibit sycophantic behavior by aligning with users’ views or preferences. While sycophancy can enhance interaction, its influence on users’ decisions remain unclear given different styles and task risks. We examine three forms of sycophancy—opinion agreement, direct praise, and self-deprecation—in two contrasting contexts: a low-risk speed-dating prediction task and a high-risk ETF investment task. In a 4×2 mixed-design online study (N = 106), we compare non-sycophantic AI with sycophantic variants on decision outcomes and confidence changes. Results show that sycophancy influences decision patterns in type-dependent ways. Specifically, opinion agreement reinforces initial decisions and self-deprecation boosts confidence. Interviews further indicate that users value supportive AI but question its objectivity when praise becomes excessive. These findings reveal the multifaceted effects of AI sycophancy and offer design implications for balancing support and credibility in human–AI interaction. Zejian Li, Jiaman Pan, Qi Liu 0076, Yuning Xi, Yixiang Zhou, Yike Jin, Rongjie Mao, Pei Chen 0005 |
CHI | 1 |
| 2026 | PoemPalette: Facilitating Poetry Creative Exploration and Foundational Understanding through the Ideorealm Alignment of Paintings and PoemsabstractThe “Ideorealm Alignment of Paintings and Poems (IA-PP)” theory rooted in Chinese classical aesthetics offers a perspective for exploring poetry’s deep connotations. This study presents PoemPalette, a novel IA-PP creative-exploration tool that integrates Generative Artificial Intelligence (GAI) to guide poetry enthusiasts in actively constructing an ideorealm for the poetic painting they envision, informed by a formative study with six experts. We extract the core symbols of poetry, transform them into Scene Graph (SG), and generate images for users to freely compose, enabling IA-PP creative exploration. The system incorporates Large Language Model (LLM) agents to enhance the foundational understanding of poetry. In a controlled experiment on Chinese poetry and Japanese haiku with 60 participants, we analyze which interaction mechanisms most contribute to foundational understanding and creative outcomes, compared with both AI and non-AI baselines. Situated within East Asian poetry traditions, this study introduces cultural theories to guide the design of AI co-creation tools, using a graph-based interface of interpretable intermediate representations. Ying Zhang 0076, Kaixin Jia, Hong Jian Zhang, Kewen Zhu, Chenye Meng, Jiesi Zhang, Zejian Li, Pei Chen 0005, Lingyun Sun |
CHI | 7 |
| 2026 | 3DInkGen: Extending Traditional Ink-Painting Artistry with Generative 3D Creation for NovicesabstractInk painting, renowned for its aesthetics and historical significance, plays a vital role in global art. Further, 3D ink art extends this tradition into spatial forms, enriching digital media like animation and games. However, existing methods for 3D ink creation demand expertise in both 3D modeling and ink aesthetics, limiting novice participation and 3D ink application. Through formative research with four experts, including ink painting artists and 3D designers, we summarize the core challenge: how to preserve the expressive pattern of ink paintings while constructing 3D structures. To tackle this challenge, we introduce 3DInkGen, a system transforming 2D ink elements into editable 3D compositions. 3DInkGen follows a four-stage workflow: element extraction, form generation, 3D reconstruction, and style transfer. A user study with sixteen novices showed 3DInkGen lowers technical barriers and enables intuitive 3D composition. The four experts believe novice-created works captured the artistic style of ink painting while maintain 3D structure of elements. Jiesi Zhang, Ying Zhang 0076, Zejian Li, Changle Xie, Huanghuang Deng, Lingyun Sun |
CHI | 3 |
| 2026 | M-Control: Improving text-image consistency via Mask-Guided ControlNet
Wei Li 0183, Zejian Li, Yongxing He |
Comput. Vis. Image Underst. | 4 |
| 2026 | Corrigendum to "M-Control: Improving text-image consistency via Mask-Guided ControlNet" [Comput. Vis. Image Underst. 270 (2026) 104832]
Zejian Li, Yongxing He |
Comput. Vis. Image Underst. | 4 |
| 2026 | SCoRE: Standardized Human Evaluation Provides a Reliable Measure for Semantic Consistency of Text-to-Image Generation
Zejian Li, Qi Liu 0076, Jiaman Pan, Lefan Hou, Xiangfei Hu, Jiarui Ma, Shengyuan Zhang, Jiesi Zhang, Xuetao Tian, Xiaoming Deng 0001 |
Int. J. Comput. Vis. | 1 |
| 2026 | Let Human Sketches Help: Empowering the Challenging Image Segmentation Task With Freehand SketchesabstractSketches, with their expressive potential, enable humans to convey the essence of an object through a rough contour. This work leverages expressive power for the first time to improve segmentation performance in challenging tasks such as camouflaged object detection (COD). We propose a sketch guided interactive segmentation framework that allows users to intuitively annotate objects with freehand sketches rather than relying on traditional bounding boxes or points commonly used in models such as the SAM. Our method introduces dedicated network architectural enhancements and a novel sketch augmentation strategy to fully exploit sketch input, leading to significant accuracy gains compared with text- or box-based annotations. Furthermore, our model's output can directly train other neural networks, achieving performance comparable to that of pixel-level annotations while reducing the annotation time by up to 120× and thereby lowering the barrier for large-scale dataset creation and model training. To support future research, werelease KOSCamo+, the first freehand sketch dataset for COD, along with code and a labeling tool. These contributions open promising avenues for expanding sketch-based interaction to broader segmentation tasks and exploring multimodal annotation strategies that combine sketches, text, and other lightweight user inputs. Ying Zang, Runlong Cao, Jianqi Zhang, Yidong Han, Ziyue Cao, Didi Zhu, Zejian Li, Lanyun Zhu, Deyi Ji, Tianrun Chen |
IEEE Trans. Multim. | 8 |
| 2025 | CharacterCritique: Supporting Children's Development of Critical Thinking through Multi-Agent Interaction in Story Reading
Jiangyu Pan, Duola Jin, Jingao Zhang, Jiacheng Cao, Chao Zhang 0082, Zejian Li, Preben Hansen, Shouqian Sun, Xianyue Qiao |
CHI | 7 |
| 2025 | FusionProtor: A Mixed-Prototype Tool for Component-level Physical-to-Virtual 3D Transition and Simulation
Pei Chen 0005, Xuelong Xie, Zhaoqu Jiang, Zejian Li, Lingyun Sun |
CHI | 6 |
| 2025 | Ink Restorer: Virtual Restoration of Ancient Chinese Paintings Inheriting Traditional Restoration Processes
Ying Zhang 0076, Zejian Li, Jiesi Zhang, Kewen Zhu, Qi Liu 0076, Huanghuang Deng, Lingyun Sun |
CHI | 2 |
| 2025 | Distilling Diffusion Models to Efficient 3D LiDAR Scene CompletionabstractDiffusion models have been applied to 3D LiDAR scene completion due to their strong training stability and high completion quality. However, the slow sampling speed limits the practical application of diffusion-based scene completion models since autonomous vehicles require an efficient perception of surrounding environments. This paper proposes a novel distillation method tailored for 3D Li- DAR scene completion models, dubbed ScoreLiDAR, which achieves efficient yet high-quality scene completion. Score- LiDAR enables the distilled model to sample in significantly fewer steps after distillation. To improve completion quality, we also introduce a novel Structural Loss, which encourages the distilled model to capture the geometric structure of the 3D LiDAR scene. The loss contains a scene-wise term constraining the holistic structure and a point-wise term constraining the key landmark points and their relative configuration. Extensive experiments demonstrate that ScoreLiDAR significantly accelerates the completion time from 30.55 to 5.37 seconds per frame (>5x) on SemanticKITTI and achieves superior performance compared to state-of-the-art 3D LiDAR scene completion models. Our model and code are publicly available on https://github.com/happyw1nd/ScoreLiDAR. Shengyuan Zhang, Ling Yang 0006, Zejian Li, Chenye Meng, Tianrun Chen, Anyang Wei, Perry Pengyun Gu, Lingyun Sun |
ICCV | 4 |
| 2025 | Distribution Backtracking Builds A Faster Convergence Trajectory for Diffusion DistillationabstractAccelerating the sampling speed of diffusion models remains a significant challenge. Recent score distillation methods distill a heavy teacher model into a student generator to achieve one-step generation, which is optimized by calculating the difference between two score functions on the samples generated by the student model.
However, there is a score mismatch issue in the early stage of the score distillation process, since existing methods mainly focus on using the endpoint of pre-trained diffusion models as teacher models, overlooking the importance of the convergence trajectory between the student generator and the teacher model.
To address this issue, we extend the score distillation process by introducing the entire convergence trajectory of the teacher model and propose $\textbf{Dis}$tribution $\textbf{Back}$tracking Distillation ($\textbf{DisBack}$). DisBask is composed of two stages: $\textit{Degradation Recording}$ and $\textit{Distribution Backtracking}$.
$\textit{Degradation Recording}$ is designed to obtain the convergence trajectory by recording the degradation path from the pre-trained teacher model to the untrained student generator.
The degradation path implicitly represents the intermediate distributions between the teacher and the student, and its reverse can be viewed as the convergence trajectory from the student generator to the teacher model.
Then $\textit{Distribution Backtracking}$ trains the student generator to backtrack the intermediate distributions along the path to approximate the convergence trajectory of the teacher model.
Extensive experiments show that DisBack achieves faster and better convergence than the existing distillation method and achieves comparable or better generation performance, with an FID score of 1.38 on the ImageNet 64$\times$64 dataset.
DisBack is easy to implement and can be generalized to existing distillation methods to boost performance. Shengyuan Zhang, Ling Yang 0006, Zejian Li, Chenye Meng, Chang-yuan Yang, Guang Yang 0022, Lingyun Sun |
ICLR | 3 |
| 2025 | Vividportraits: Face Parsing Guided Portrait AnimationabstractPortrait animation aims to transfer the facial expressions and movements of a target character onto a reference character. This task presents two main challenges: accurately transferring motion and expressions while fully preserving the identity features of the reference portrait. We introduce Vividportraits, a diffusion-based model designed to effectively meet these objectives. In contrast to existing methods that rely on sparse representations such as facial landmarks, our approach leverages facial parsing maps for motion guidance, enabling a more precise conveyance of subtle expressions. A random scaling technique is applied during training to prevent the model from internalizing identity-specific features from the driving images. Furthermore, we perform foreground-background segmentation on the reference portrait to reduce data redundancy. The long-video generation process is refined to improve consistency across sequences. Our model, exclusively trained on public datasets, demonstrates superior performance relative to current state-of-the-art methods, achieving a notable 8% improvement in expression metric. More visual results are available on the anonymous website https://www.vividportraits.cn. Xuze Tian, Jinshan Zhang 0001, Boxi Wu 0001, Meng Xi 0002, Zejian Li, Jianwei Yin |
ICMR | 6 |
| 2025 | Inversion-DPO: Precise and Efficient Post-Training for Diffusion ModelsabstractRecent advancements in diffusion models (DMs) have been propelled by alignment methods that post-train models to better conform to human preferences. However, these approaches typically require computation-intensive training of a base model and a reward model, which not only incurs substantial computational overhead but may also compromise model accuracy and training efficiency. To address these limitations, we propose Inversion-DPO, a novel alignment framework that circumvents reward modeling by reformulating Direct Preference Optimization (DPO) with DDIM inversion for DMs. Our method conducts intractable posterior sampling in Diffusion-DPO with the deterministic inversion from winning and losing samples to noise and thus derive a new post-training paradigm. This paradigm eliminates the need for auxiliary reward models or inaccurate appromixation, significantly enhancing both precision and efficiency of training. We apply Inversion-DPO to a basic task of text-to-image generation and a challenging task of compositional image generation. Extensive experiments show substantial performance improvements achieved by Inversion-DPO compared to existing post-training methods and highlight the ability of the trained generative models to generate high-fidelity compositionally coherent images. For the post-training of compostitional image geneation, we curate a paired dataset consisting of 11,140 images with complex structural annotations and comprehensive scores, designed to enhance the compositional capabilities of generative models. Inversion-DPO explores a new avenue for efficient, high-precision alignment in diffusion models, advancing their applicability to complex realistic generation tasks. Our code is available at https://github.com/MIGHTYEZ/Inversion-DPO Zejian Li, Yize Li 0001, Chenye Meng, Zhongni Liu, Ling Yang 0006, Shengyuan Zhang, Guang Yang 0022, Chang-yuan Yang, Lingyun Sun |
ACM Multimedia | 1 |
| 2025 | ObjCtrl: Object-based Control Relaxation for Conditional Text-to-Image GenerationabstractConditional text-to-image diffusion models enhance the controllability of text-to-image generation by incorporating additional visual conditions. However, they often encounter two main challenges when dealing with complex visual conditions (namely, including multiple different objects): semantic leakage among objects and conflicts between visual inputs and text descriptions. To address these issues, we propose an innovative object-level conditional image generation method. It associates visual features with object semantic information, ensuring that generated objects are accurately positioned in their expected locations within the visual inputs. To address semantic leakage, we design an Object-level Structure Controller (OSC) module. This module utilizes an attention mechanism to fuse bounding box annotations, object prompts, and visual conditional inputs, allowing the model to learn essential object-level structural features. Besides, we propose an Object-level Control Relaxation (OCR) module to predict object-level scale features, which can reconcile conflicts between object semantics and visual features. Finally, the scaled backbone features are fused with structural features to form the final output features. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in terms of text-image alignment, structural similarity, and spatial fidelity. Zejian Li, Wei Li 0183, Chengyu Lin 0003 |
ACM Multimedia | 2 |
| 2025 | TransEC-GAN: A Transformer-Enhanced IDS for Robust Detection and Privacy in Industrial CPSabstractDespite the widespread deployment of various Intrusion Detection Systems (IDSs) in industrial Cyber-Physical Systems (CPS), significant challenges such as class imbalance, zero-day attacks, and privacy vulnerabilities persist. These issues underline the critical need for a more robust IDS solution that not only improves detection and generalization capabilities across different scenarios but also ensures stringent data privacy. In this paper, a novel IDS solution tailored for industrial CPS is proposed, designated as a Transformer-enhanced External Classifier-Generative Adversarial Network (TransEC-GAN). This innovative model extends the External Classifier-Generative Adversarial Network (EC-GAN) by integrating Transformer encoders, which leverage Wasserstein distance and label conditioning to enable stable gradient descent within the semi-supervised learning environment. To further fortify the system, Adaptive Differential Privacy (ADP) is incorporated, dynamically adjusting privacy settings to effectively prevent adversaries from exploiting sensitive information. Additionally, the proposed TransEC-GAN features a meticulously designed two-stage detection architecture that proficiently distinguishes between In-Distribution (InD) and Out-Of-Distribution (OOD) samples, enhancing its ability to identify and react to novel and evolving threats. Comprehensive experimental evaluations and theoretical analysis validate that the proposed TransEC-GAN not only safeguards against privacy breaches but also excels in detecting a wide array of attack types in industrial CPS settings. Junwei Liang 0004, Zejian Li, Muhammad Sadiq |
WCNC | 2 |
| 2025 | MindScratch: A Visual Programming Support Tool for Classroom Learning Based on Multimodal Generative AIabstractProgramming is essential in K-12 education and fosters computational thinking skills. Given the complexity of programming and the advanced skills it requires, previous research has introduced user-friendly tools to support young learners. However, our interviews with six programming educators revealed that current tools often fail to reflect classroom learning objectives, offer flexible guidance, and foster creativity. Therefore, we introduced MindScratch, a multimodal generative AI (GAI)-powered visual programming support tool. MindScratch aims to balance structured classroom activities with free programming creation, supporting students in completing creative programming projects based on teacher-set learning objectives while also providing programming scaffolding. The results indicate that, compared to the baseline, MindScratch more effectively helps students achieve high-quality projects aligned with learning objectives. It also enhances students’ computational thinking and thinking. Overall, we believe that GAI-driven educational tools like MindScratch offer students a focused and engaging learning experience. Yunnong Chen, Shuhong Xiao, Yaxuan Song, Zejian Li, Lingyun Sun, Liuqing Chen 0002 |
Int. J. Hum. Comput. Interact. | 4 |
| 2025 | RealtimeGen: An Intervenable AI Image Generation System for Commercial Digital Art Asset CreatorsabstractRecent advances in artificial intelligence-generated content (AIGC) have led to the rapid generation of high-quality images. AIGC has attracted the attention of commercial digital art asset creators. Traditional artist-led processes contrast with current AI tools that often reduce creators to passive roles. This study examines the integration of AI image generation into commercial digital art, emphasizing the importance of preserving creators’ creative autonomy. Our formative study (S1) involved interviews with commercial digital art creators, highlighting a need for greater control and transparency in AI-assisted painting. In response, we developed RealtimeGen, an integrated tool that merges human creativity with AI’s capabilities, allowing creators to intervene in the generative process. A user study (S2) comparing RealtimeGen with the popular AIGC tool Stable Diffusion was also carried out. The results showed its enhanced user experience and workflow compatibility. Our work contributes to understanding and improving AI-assisted painting workflows for commercial creators, offering them greater creative agency. Zejian Li, Ying Zhang 0076, Shengzhe Zhou, Qi Liu 0076, Jiesi Zhang, Shuyao Chen, Lingyun Sun |
Int. J. Hum. Comput. Interact. | 1 |
| 2025 | Image generation evaluation: a comprehensive survey of human and automatic evaluationsabstractImage generation models have made remarkable progress, and image evaluation is crucial for explaining and driving the development of these models. Previous studies have extensively explored human and automatic evaluations of image generation. Herein, these studies are comprehensively surveyed, specifically for two main parts: evaluation protocols and evaluation methods. First, 10 image generation tasks are summarized with focus on their differences in evaluation aspects. Based on this, a novel protocol is proposed to cover human and automatic evaluation aspects required for various image generation tasks. Second, the review of automatic evaluation methods in the past five years is highlighted. To our knowledge, this paper presents the first comprehensive summary of human evaluation, encompassing evaluation methods, tools, details, and data analysis methods. Finally, the challenges and potential directions for image generation evaluation are discussed. We hope that this survey will help researchers develop a systematic understanding of image generation evaluation, stay updated with the latest advancements in the field, and encourage further research. Qi Liu 0076, Shuanglin Yang, Zejian Li, Lefan Hou, Chenye Meng, Ying Zhang 0076, Lingyun Sun |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2025 | Img2CAD: Conditioned 3-D CAD Model Generation From Single Image With Structured Visual GeometryabstractIn this article, we propose Img2CAD, the first approach to our knowledge that uses 2-D image inputs to generate computer-aided design (CAD) models with editable parameters. Unlike existing artificial intelligence (AI) methods for 3-D model generation using text or image inputs often rely on mesh-based representations, which are incompatible with CAD tools and lack editability and fine control, Img2CAD enables seamless integration between AI-based 3-D reconstruction and CAD software. We have identified an innovative intermediate representation called structured visual geometry, characterized by vectorized wireframes extracted from objects. This representation significantly enhances the performance of generating conditioned CAD models. In addition, we introduce two new datasets to further support research in this area:a big cad model dataset (ABC)-mono, the largest known dataset comprising over 200 000 3-D CAD models with rendered images, andKOCAD, the first dataset featuring real-world captured objects alongside their ground truth CAD models, supporting further research in conditioned CAD model generation. Tianrun Chen, Chunan Yu, Yuanqi Hu, Jing Li 0145, Tao Xu 0048, Runlong Cao, Lanyun Zhu, Ying Zang, Yong Zhang 0030, Zejian Li, Lingyun Sun |
IEEE Trans. Ind. Informatics | 10 |
| 2025 | A Hierarchical Encrypted Compression Scheme for Intra-Vehicle NetworkabstractThe CAN bus is the most widely used bus for intra-vehicle communication due to its high transmission stability, excellent real-time communication capability, and relatively low cost. As the number of ECUs grows, the CAN bus load increases and thus raises the possibility of data transmission delays and errors. Message compression based on the differential algorithm has been proposed to reduce the CAN bus load. However, current works do not consider the security problems of the CAN bus. Attackers can manage to acquire the original messages before compression and disturb the message statistics to decrease compression rate by injecting malicious frames. In this paper, we propose a secure compression mechanism for the intra-vehicle network, including an improved compression algorithm, a stream key distribution scheme, and a hierarchical encryption scheme. Formal verification results show that the proposed scheme can achieve mutual authentication, message confidentiality and integrity, resist replay attacks, and support secure compression. Evaluations using real vehicle data on 16 MHz boards show the average communication overhead can be reduced by 46.38% compared to the original messages. Performance analysis results show our scheme can reduce computational overhead on compression by 31.82% and 19.43% on decompression compared to related schemes. Jin Cao 0001, Zejian Li, Ben Niu 0001, Kwok-Yan Lam, Chihung Chi, Hui Li 0006 |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | Reducing Spatial Fitting Error in Distillation of Denoising Diffusion ModelsabstractDenoising Diffusion models have exhibited remarkable capabilities in image generation. However, generating high-quality samples requires a large number of iterations. Knowledge distillation for diffusion models is an effective method to address this limitation with a shortened sampling process but causes degraded generative quality. Based on our analysis with bias-variance decomposition and experimental observations, we attribute the degradation to the spatial fitting error occurring in the training of both the teacher and student model in the distillation. Accordingly, we propose Spatial Fitting-Error Reduction Distillation model (SFERD). SFERD utilizes attention guidance from the teacher model and a designed semantic gradient predictor to reduce the student's fitting error. Empirically, our proposed model facilitates high-quality sample generation in a few function evaluations. We achieve an FID of 5.31 on CIFAR-10 and 9.39 on ImageNet 64x64 with only one step, outperforming existing diffusion methods. Our study provides a new perspective on diffusion distillation by highlighting the intrinsic denoising ability of models. Shengzhe Zhou, Zejian Li, Shengyuan Zhang, Lefan Hou, Chang-yuan Yang, Guang Yang 0022, Lingyun Sun |
AAAI | 2 |
| 2024 | Rapid 3D Model Generation with Intuitive 3D InputabstractWith the emergence of AR/VR, 3D models are in tremendous demand. However, conventional 3D modeling with Computer-Aided Design software requires much expertise and is difficult for novice users. We find that AR/VR devices, in addition to serving as effective display mediums, can offer a promising potential as an intuitive 3D model creation tool, especially with the assistance of AI generative models. Here, we propose Deep3DVRSketch, the first 3D model generation network that inputs 3D VR sketches from novice users and generates highly consistent 3D models in multiple categories within seconds, irrespective of the users' drawing abilities. We also contribute KO3D+, the largest 3D sketch-shape dataset. Our method pre-trains a conditional diffusion model on quality 3D data, then fine-tunes an encoder to map 3D sketches onto the generator's manifold using an adaptive curriculum strategy for limited ground truths. In our experiment, our approach achieves state-of-the-art performance in both model quality and fidelity with real-world input from novice users, and users can even draw and obtain very detailed geometric structures. In our user study, users were able to complete the 3D modeling tasks over 10 times faster using our approach compared to conventional CAD software tools. We believe that our Deep3DVRSketch and KO3D+ dataset can offer a promising solution for future 3D modeling in metaverse era. Check the project page at http://research.kokoni3d.com/Deep3DVRSketch. Tianrun Chen, Chaotao Ding, Shangzhan Zhang, Chunan Yu, Ying Zang, Zejian Li, Sida Peng, Lingyun Sun |
CVPR | 6 |
| 2024 | A Message-based Lightweight Session Key Distribution Scheme for Intra-Vehicle NetworkabstractModern vehicles are equipped with ECU nodes and intra-vehicle buses. Among these, the CAN bus stands out as the most widely utilized intra-vehicle bus due to its affordability and straightforward deployment. However, the CAN bus suffers from significant security vulnerabilities, such as the absence of access control, identity authentication, message encryption, and authentication. In this paper, we propose a lightweight message-based key distribution scheme aimed at addressing these vulnerabilities. Our scheme facilitates mutual authentication during key distribution and assigns a unique key to each class of message. Formal verification using the Scyther tool demonstrates that our protocol achieves mutual authentication and effectively mitigates several protocol attacks, including replay, tampering, and manin-the-middle attacks. We evaluate our scheme using Arduino UNO boards. Performance analysis indicates that our scheme exhibits superior security capabilities compared to other related schemes and outperforms them in terms of communication and computation overheads. Jin Cao 0001, Zejian Li, Yurong Luo, Kwok-Yan Lam, Chihung Chi, Hui Li 0006 |
GLOBECOM | 3 |
| 2024 | Deep3DSketch-im: rapid high-fidelity AI 3D model generation by single freehand sketchesabstractThe rise of artificial intelligence generated content (AIGC) has been remarkable in the language and image fields, but artificial intelligence (AI) generated three-dimensional (3D) models are still under-explored due to their complex nature and lack of training data. The conventional approach of creating 3D content through computer-aided design (CAD) is labor-intensive and requires expertise, making it challenging for novice users. To address this issue, we propose a sketch-based 3D modeling approach, Deep3DSketch-im, which uses a single freehand sketch for modeling. This is a challenging task due to the sparsity and ambiguity. Deep3DSketch-im uses a novel data representation called the signed distance field (SDF) to improve the sketch-to-3D model process by incorporating an implicit continuous field instead of voxel or points, and a specially designed neural network that can capture point and local features. Extensive experiments are conducted to demonstrate the effectiveness of the approach, achieving state-of-the-art (SOTA) performance on both synthetic and real datasets. Additionally, users show more satisfaction with results generated by Deep3DSketch-im, as reported in a user study. We believe that Deep3DSketch-im has the potential to revolutionize the process of 3D modeling by providing an intuitive and easy-to-use solution for novice users. Tianrun Chen, Runlong Cao, Zejian Li, Ying Zang, Lingyun Sun |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2024 | Reality3DSketch: Rapid 3D Modeling of Objects From Single Freehand SketchesabstractThe emerging trend of AR/VR places great demands on 3D content. However, most existing software requires expertise and is difficult for novice users to use. In this paper, we aim to create sketch-based modeling tools for user-friendly 3D modeling. We introduce Reality3DSketch with a novel application of an immersive 3D modeling experience, in which a user can capture the surrounding scene using a monocular RGB camera and can draw a single sketch of an object in the real-time reconstructed 3D scene. A 3D object is generated and placed in the desired location, enabled by our novel neural network with the input of a single sketch. Our neural network can predict the pose of a drawing and can turn a single sketch into a 3D model with view and structural awareness, which addresses the challenge of sparse sketch input and view ambiguity. We conducted extensive experiments synthetic and real-world datasets and achieved state-of-the-art (SOTA) results in both sketch view estimation and 3D modeling performance. According to our user study, our method of performing 3D modeling in a scene is$>$5x faster than conventional methods. Users are also more satisfied with the generated 3D model than the results of existing methods. Tianrun Chen, Chaotao Ding, Lanyun Zhu, Ying Zang, Yiyi Liao, Zejian Li, Lingyun Sun |
IEEE Trans. Multim. | 6 |
| 2023 | Preserving Structural Consistency in Arbitrary Artist and Artwork Style TransferabstractDeep generative models are effective in style transfer. Previous methods learn one or several specific artist-style from a collection of artworks. These methods not only homogenize the artist-style of different artworks of the same artist but also lack generalization for the unseen artists. To solve these challenges, we propose a double-style transferring module (DSTM). It extracts different artist-style and artwork-style from different artworks (even untrained) and preserves the intrinsic diversity between different artworks of the same artist. DSTM swaps the two styles in the adversarial training and encourages realistic image generation given arbitrary style combinations. However, learning style from single artwork can often cause over-adaption to it, resulting in the introduction of structural features of style image. We further propose an edge enhancing module (EEM) which derives edge information from multi-scale and multi-level features to enhance structural consistency. We broadly evaluate our method across six large-scale benchmark datasets. Empirical results show that our method achieves arbitrary artist-style and artwork-style extraction from a single artwork, and effectively avoids introducing the style image’s structural features. Our method improves the state-of-the-art deception rate from 58.9% to 67.2% and the average FID from 48.74 to 42.83. Jingyu Wu, Lefan Hou, Zejian Li, Jun Liao 0001, Li Liu 0001, Lingyun Sun |
AAAI | 3 |
| 2023 | Learning Object Consistency and Interaction in Image Generation from Scene GraphsabstractThis paper is concerned with synthesizing images conditioned on a scene graph (SG), a set of object nodes and their edges of interactive relations. We divide existing works into image-oriented and code-oriented methods. In our analysis, the image-oriented methods do not consider object interaction in spatial hidden feature. On the other hand, in empirical study, the code-oriented methods lose object consistency as their generated images miss certain objects in the input scene graph. To alleviate these two issues, we propose Learning Object Consistency and Interaction (LOCI). To preserve object consistency, we design a consistency module with a weighted augmentation strategy for objects easy to be ignored and a matching loss between scene graphs and image codes. To learn object interaction, we design an interaction module consisting of three kinds of message propagation between the input scene graph and the learned image code. Experiments on COCO-stuff and Visual Genome datasets show our proposed method alleviates the ignorance of objects and outperforms the state-of-the-art on visual fidelity of generated images and objects. Yangkang Zhang, Chenye Meng, Zejian Li, Pei Chen 0005, Guang Yang 0022, Chang-yuan Yang, Lingyun Sun |
IJCAI | 3 |
| 2023 | UI layers merger: merging UI layers via visual learning and boundary priorabstractWith the fast-growing graphical user interface (GUI) development workload in the Internet industry, some work attempted to generate maintainable front-end code from GUI screenshots. It can be more suitable for using user interface (UI) design drafts that contain UI metadata. However, fragmented layers inevitably appear in the UI design drafts, which greatly reduces the quality of the generated code. None of the existing automated GUI techniques detects and merges the fragmented layers to improve the accessibility of generated code. In this paper, we propose UI layers merger (UILM), a vision-based method that can automatically detect and merge fragmented layers into UI components. Our UILM contains the merging area detector (MAD) and a layer merging algorithm. The MAD incorporates the boundary prior knowledge to accurately detect the boundaries of UI components. Then, the layer merging algorithm can search for the associated layers within the components’ boundaries and merge them into a whole. We present a dynamic data augmentation approach to boost the performance of MAD. We also construct a large-scale UI dataset for training the MAD and testing the performance of UILM. Experimental results show that the proposed method outperforms the best baseline regarding merging area detection and achieves decent layer merging accuracy. A user study on a real application also confirms the effectiveness of our UILM. Yunnong Chen, Yankun Zhen, Chu-ning Shi, Jiazhi Li 0002, Liuqing Chen 0002, Zejian Li, Lingyun Sun, Yanfang Chang |
Frontiers Inf. Technol. Electron. Eng. | 6 |
| 2022 | Few-Shot Incremental Learning for Label-to-Image TranslationabstractLabel-to-image translation models generate images from semantic label maps. Existing models depend on large volumes of pixel-level annotated samples. When given new training samples annotated with novel semantic classes, the models should be trained from scratch with both learned and new classes. This hinders their practical applications and motivates us to introduce an incremental learning strategy to the label-to-image translation scenario. In this paper, we introduce a few-shot incremental learning method for label-to-image translation. It learns new classes one by one from a few samples of each class. We propose to adopt semantically-adaptive convolution filters and normalization. When incrementally trained on a novel semantic class, the model only learns a few extra parameters of class-specific modulation. Such design avoids catastrophic forgetting of already-learned semantic classes and enables label-to-image translation of scenes with increasingly rich content. Furthermore, to facilitate few-shot learning, we propose a modulation transfer strategy for better initialization. Extensive experiments show that our method outperforms existing related methods in most cases and achieves zero forgetting. Pei Chen 0005, Yangkang Zhang, Zejian Li, Lingyun Sun |
CVPR | 3 |
| 2022 | Recognizing Cognitive Load by a Hybrid Spatio-Temporal Causal Model from Multivariate Physiological Data
Zirui Yong, Guoxin Su, Xiaohu Li, Lingyun Sun, Zejian Li, Li Liu 0001 |
ECML/PKDD (6) | 5 |
| 2022 | USIS: A unified semantic image synthesis model trained on a single or multiple samples
Pei Chen 0005, Zejian Li, Yangkang Zhang, Lingyun Sun |
Neurocomputing | 2 |
| 2021 | Image Synthesis from Layout with Locality-Aware Mask AdaptionabstractThis paper is concerned with synthesizing images conditioned on a layout (a set of bounding boxes with object categories). Existing works construct a layout-mask-image pipeline. Object masks are generated separately and mapped to bounding boxes to form a whole semantic segmentation mask (layout-to-mask), with which a new image is generated (mask-to-image). However, overlapped boxes in layouts result in overlapped object masks, which reduces the mask clarity and causes confusion in image generation. We hypothesize the importance of generating clean and semantically clear semantic masks. The hypothesis is supported by the finding that the performance of state-of-the-art LostGAN decreases when input masks are tainted. Motivated by this hypothesis, we propose Locality-Aware Mask Adaption (LAMA) module to adapt overlapped or nearby object masks in the generation. Experimental results show our proposed model with LAMA outperforms existing approaches regarding visual fidelity and alignment with input layouts. On COCO-stuff in 256×256, our method improves the state-of-the-art FID score from 41.65 to 31.12 and the SceneFID from 22.00 to 18.64. Zejian Li, Jingyu Wu, Immanuel Koh, Lingyun Sun |
ICCV | 1 |
| 2020 | FET-GAN: Font and Effect Transfer via K-shot Adaptive Instance NormalizationabstractText effect transfer aims at learning the mapping between text visual effects while maintaining the text content. While remarkably successful, existing methods have limited robustness in font transfer and weak generalization ability to unseen effects. To address these problems, we propose FET-GAN, a novel end-to-end framework to implement visual effects transfer with font variation among multiple text effects domains. Our model achieves remarkable results both on arbitrary effect transfer between texts and effect translation from text to graphic objects. By a few-shot fine-tuning strategy, FET-GAN can generalize the transfer of the pre-trained model to the new effect. Through extensive experimental validation and comparison, our model advances the state-of-the-art in the text effect transfer task. Besides, we have collected a font dataset including 100 fonts of more than 800 Chinese and English characters. Based on this dataset, we demonstrated the generalization ability of our model by the application that complements the font library automatically by few-shot samples. This application is significant in reducing the labor cost for the font designer. Wei Li 0183, Yongxing He, Yanwei Qi, Zejian Li |
AAAI | 4 |
| 2019 | Learning Disentangled Representation with Pairwise IndependenceabstractUnsupervised disentangled representation learning is one of the foundational methods to learn interpretable factors in the data. Existing learning methods are based on the assumption that disentangled factors are mutually independent and incorporate this assumption with the evidence lower bound. However, our experiment reveals that factors in real-world data tend to be pairwise independent. Accordingly, we propose a new method based on a pairwise independence assumption to learn the disentangled representation. The evidence lower bound implicitly encourages mutual independence of latent codes so it is too strong for our assumption. Therefore, we introduce another lower bound in our method. Extensive experiments show that our proposed method gives competitive performances as compared with other state-of-the-art methods. Zejian Li, Wei Li 0183, Yongxing He |
AAAI | 1 |
| 2019 | A review of design intelligence: progress, problems, and challengesabstractDesign intelligence is an important branch of artificial intelligence (AI), focusing on the intelligent models and algorithms in creativity and design. In the context of AI 2.0, studies on design intelligence have developed rapidly. We summarize mainly the current emerging framework of design intelligence and review the state-of-the-art techniques of related topics, including user needs analysis, ideation, content generation, and design evaluation. Specifically, the models and methods of intelligence-generated content are reviewed in detail. Finally, we discuss some open problems and challenges for future research in design intelligence. Jiangjie Huang, Meng-ting Yao, Wei Li 0183, Yongxing He, Zejian Li |
Frontiers Inf. Technol. Electron. Eng. | 7 |
| 2018 | Unsupervised Disentangled Representation Learning with Analogical RelationsabstractLearning the disentangled representation of interpretable generative factors of data is one of the foundations to allow artificial intelligence to think like people. In this paper, we propose the analogical training strategy for the unsupervised disentangled representation learning in generative models. The analogy is one of the typical cognitive processes, and our proposed strategy is based on the observation that sample pairs in which one is different from the other in one specific generative factor show the same analogical relation. Thus, the generator is trained to generate sample pairs from which a designed classifier can identify the underlying analogical relation. In addition, we propose a disentanglement metric called the subspace score, which is inspired by subspace learning methods and does not require supervised information. Experiments show that our proposed training strategy allows the generative models to find the disentangled factors, and that our methods can give competitive performances as compared with the state-of-the-art methods. Zejian Li, Yongxing He |
IJCAI | 1 |
| 2018 | Comparative density peaks clustering
Zejian Li |
Expert Syst. Appl. | 1 |