Zongjian Li

dblp:227/9294 · DBLP profile ↗
← Back
14ranked-venue papers
3as first author
11since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 4 · 3 since 2021Systems, architecture and hardware · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 LibreFace 2.0: A Generalizable Facial Expression Analysis Toolkit Leveraging Synthetic Data
abstract
Facial expression analysis is central to social AI and human-computer interaction. However, existing toolkits often struggle to generalize across diverse demographics, largely due to the limited diversity of training data for tasks such as action unit (AU) detection, which typically require costly per-frame annotations. In this work, we introduce LibreFace 2.0, a toolkit that leverages recent advances in face generation and motion retargeting to enrich AU datasets with broader demographic coverage. Specifically, we employ stable diffusion to synthesize a wide range of identities spanning age, gender, race and facial attributes and retarget AU motions from annotated datasets onto these generated identities. Training on this large-scale, demographically diverse dataset yields consistent improvements in benchmark performance and enhances fairness across demographic groups. Beyond AU detection and intensity estimation, LibreFace 2.0 also supports facial expression recognition and gaze estimation through lightweight models that achieve competitive accuracy with substantially fewer parameters, enabling efficient inference. Our work provides a scalable approach to achieving fairer face analysis in real-world applications. The code and the synthetic data will be released publicly at https://github.com/ihp-lab/LibreFace.
Xulang Guan, Ashutosh Chaubey, Maksim Siniukov, Annabelle Hsieh, Zongjian Li, Mohammad Soleymani 0001
FG5
2026 Bringing Diversity from Diffusion Models to Semantic-Guided Face Asset Generation
abstract
High-quality 3D face asset creation remains costly due to reliance on controlled capture setups and manual processing, limiting scalability and diversity. We introduce a fully automated, semantically controllable framework for generating PBR-ready 3D facial assets without requiring dedicated scans. Our pipeline begins with a diffusion-based data synthesis stage, where 2D portrait samples from a pre-trained diffusion model are converted into 44K textured 3D face reconstructions via our proposed geometry recovery and texture normalization algorithm, which aligns arbitrarily shaded outputs into clean albedo space. Using this dataset, we train a disentangled adversarial generator that maps semantic attributes (age, gender, ethnicity) to UV-space geometry and albedo, enabling both direct sampling and continuous latent editing while preserving identity. A refinement stage further produces PBR materials and secondary assets (eyeballs, teeth, gums). The resulting system supports controllable face generation and post-editing in real time and exports directly to standard rendering and animation pipelines. We evaluate each component extensively and provide a web-based interactive interface to showcase practical deployment.
Yunxuan Cai, Sitao Xiang, Zongjian Li, Haiwei Chen
ACM Trans. Graph.3
2025 AGFSync: Leveraging AI-Generated Feedback for Preference Optimization in Text-to-Image Generation
abstract
Text-to-Image (T2I) diffusion models have achieved remarkable success in image generation. Despite their progress, challenges remain in both prompt-following ability, image quality and lack of high-quality datasets, which are essential for refining these models. As acquiring labeled data is costly, we introduce AGFSync, a framework that enhances T2I diffusion models through Direct Preference Optimization (DPO) in a fully AI-driven approach. AGFSync utilizes Vision-Language Models (VLM) to assess image quality across style, coherence, and aesthetics, generating feedback data within an AI-driven loop. By applying AGFSync to leading T2I models such as SD v1.4, v1.5, and SDXL-base, our extensive experiments on the TIFA dataset demonstrate notable improvements in VQA scores, aesthetic evaluations, and performance on the HPS v2 benchmark, consistently outperforming the base models. AGFSync's method of refining T2I diffusion models paves the way for scalable alignment techniques.
Jingkun An, Yinghao Zhu, Zongjian Li, Enshen Zhou, Xijie Huang, Bohua Chen, Yemin Shi 0001, Chengwei Pan
AAAI3
2025 WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
abstract
Video Variational Autoencoder (VAE) encodes videos into a low-dimensional latent space, becoming a key component of most Latent Video Diffusion Models (LVDMs) to reduce model training costs. However, as the resolution and duration of generated videos increase, the encoding cost of Video VAEs becomes a limiting bottleneck in training LVDMs. Moreover, the block-wise inference method adopted by most LVDMs can lead to discontinuities of latent space when processing long-duration videos. The key to addressing the computational bottleneck lies in decomposing videos into distinct components and efficiently encoding the critical information. Wavelet transform can decompose videos into multiple frequency-domain components and improve the efficiency significantly, we thus propose Wavelet Flow VAE (WF-VAE), an autoencoder that leverages multi-level wavelet transform to facilitate low-frequency energy flow into latent representation. Furthermore, we introduce a method called Causal Cache, which maintains the integrity of latent space during block-wise inference. Compared to state-of-the-art video VAEs, WF-VAE demonstrates superior performance in both PSNR and LPIPS metrics, achieving 2× higher throughput and 4× lower memory consumption while maintaining competitive reconstruction quality. Our code and models are available at https://github.com/PKU-YuanGroup/WF-VAE.
Zongjian Li, Bin Lin 0014, Liuhan Chen, Xinhua Cheng, Shenghai Yuan 0002, Li Yuan 0007
CVPR1
2025 OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
abstract
Variational Autoencoder (VAE), compressing videos into latent representations, is a crucial preceding component of Latent Video Diffusion Models (LVDMs). However, most LVDMs utilize 2D image VAE, which only compresses video spatially. This will lead to temporally redundant representations, reducing the efficiency of LVDMs. To eliminate this issue, we propose an omni-dimension compression VAE, named OD-VAE, which can temporally and spatially compress videos based on 3D-Causal-CNN architecture. To obtain a better trade-off between video reconstruction quality and compression speed, we further introduce and analyze four model variants of OD-VAE. In addition, a novel initialization method is designed to train our OD-VAE more efficiently, and a novel inference strategy is proposed to enable OD-VAE to handle videos of arbitrary length with limited GPU memory. Comprehensive experiments on video reconstruction and LVDM-based video generation demonstrate the effectiveness and efficiency of our proposed methods. The source code and models are available at here.
Liuhan Chen, Zongjian Li, Bin Lin 0014, Qian Wang 0062, Shenghai Yuan 0002, Xinhua Cheng, Li Yuan 0007
ICME2
2025 Multimodal Behavioral Characterization of Dyadic Alliance in Support Groups
Kevin Hyekang Joo, Zongjian Li, Yunwen Wang, Yuanfeixue Nan, Mina J. Kian, Shriya Upadhyay, Maja J. Mataric, Lynn C. Miller, Mohammad Soleymani 0001
ICMI2
2025 ImgEdit: A Unified Image Editing Dataset and Benchmark
abstract
Recent advancements in generative models have enabled high-fidelity text-to-image generation. However, open-source image-editing models still lag behind their proprietary counterparts, primarily due to limited high-quality data and insufficient benchmarks.To overcome these limitations, we introduce ImgEdit, a large-scale, high-quality image-editing dataset comprising one million carefully curated edit pairs, which contain both novel and complex single-turn edits, as well as challenging multi-turn tasks.To ensure the data quality, we employ a multi-stage pipeline that integrates a cutting-edge vision-language model, a detection model, a segmentation model, alongside task-specific in-painting procedures and strict post-processing. ImgEdit surpasses existing datasets in both task novelty and data quality.Using ImgEdit, we train ImgEdit-E1, an editing model using Vision Language Model to process the reference image and editing prompt, which outperforms existing open-source models on multiple tasks, highlighting the value of ImgEdit and model design.For comprehensive evaluation, we introduce ImgEdit-Bench, a benchmark designed to evaluate image editing performance in terms of instruction adherence, editing quality, and detail preservation.It includes a basic testsuite, a challenging single-turn suite, and a dedicated multi-turn suite. We evaluate both open-source and proprietary models, as well as ImgEdit-E1, providing deep analysis and actionable insights into the current behavior of image-editing models.
Xianyi He, Zongjian Li, Bin Lin 0014, Shenghai Yuan 0002, Zhiyuan Yan 0002, Bohan Hou, Li Yuan 0007
NeurIPS3
2024 Build Your Own Robot Friend: An Open-Source Learning Module for Accessible and Engaging AI Education
abstract
As artificial intelligence (AI) is playing an increasingly important role in our society and global economy, AI education and literacy have become necessary components in college and K-12 education to prepare students for an AI-powered society. However, current AI curricula have not yet been made accessible and engaging enough for students and schools from all socio-economic backgrounds with different educational goals. In this work, we developed an open-source learning module for college and high school students, which allows students to build their own robot companion from the ground up. This open platform can be used to provide hands-on experience and introductory knowledge about various aspects of AI, including robotics, machine learning (ML), software engineering, and mechanical engineering. Because of the social and personal nature of a socially assistive robot companion, this module also puts a special emphasis on human-centered AI, enabling students to develop a better understanding of human-AI interaction and AI ethics through hands-on learning activities. With open-source documentation, assembling manuals and affordable materials, students from different socio-economic backgrounds can personalize their learning experience based on their individual educational goals. To evaluate the student-perceived quality of our module, we conducted a usability testing workshop with 15 college students recruited from a minority-serving institution. Our results indicate that our AI module is effective, easy-to-follow, and engaging, and it increases student interest in studying AI/ML and robotics in the future. We hope that this work will contribute toward accessible and engaging AI education in human-AI interaction for college and high school students.
Zhonghao Shi, Amy O'Connell, Zongjian Li, Siqi Liu 0012, Jennifer Ayissi, Guy Hoffman, Mohammad Soleymani 0001, Maja J. Mataric
AAAI3
2024 LibreFace: An Open-Source Toolkit for Deep Facial Expression Analysis
abstract
Facial expression analysis is an important tool for human-computer interaction. In this paper, we introduce LibreFace, an open-source toolkit for facial expression analysis. This open-source toolbox offers real-time and offline analysis of facial behavior through deep learning models, including facial action unit (AU) detection, AU intensity estimation, and facial expression recognition. To accomplish this, we employ several techniques, including the utilization of a large-scale pre-trained network, feature-wise knowledge distillation, and task-specific fine-tuning. These approaches are designed to effectively and accurately analyze facial expressions by leveraging visual information, thereby facilitating the implementation of real-time interactive applications. In terms of Action Unit (AU) intensity estimation, we achieve a Pearson Correlation Coefficient (PCC) of 0.63 on DISFA, which is 7% higher than the performance of OpenFace 2.0 [4] while maintaining highly-efficient inference that runs two times faster than OpenFace 2.0 [4]. Despite being compact, our model also demonstrates competitive performance to state-of-the-art facial expression analysis methods on AffecNet, FFHQ, and RAF-DB. Our code will be released at https://github.com/ihp-lab/LibreFace
Di Chang, Yufeng Yin 0002, Zongjian Li, Minh Tran 0004, Mohammad Soleymani 0001
WACV3
2023 DIVIS: Digital Interactive Victim Intake Simulator
abstract
The Digital Interactive Victim Intake Simulator ("DIVIS") is an interactive, agent-based simulated training tool that has been deployed at the U.S. Army's Sexual Harassment/Assault Response Prevention Program ("SHARP") Academy since May 2021. The system allows student Sexual Assault Response Coordinators ("SARCs") and Victim Advocates ("VAs") to practice critical interpersonal intake skills needed when conducting the initial interview of a survivor of military sexual assault. Currently the system includes two scenarios -- one with a male victim and a second with a female victim -- with two more scenarios under development. Each victim exhibits one of a possible three different emotional vectors, (e.g., angry, ashamed or defensive). Scenarios can run multiple times, giving trainees the ability to navigate through various potential story paths based on how they engage with the victim during each session.
Alesia Gainer, Allison Aptaker, Ron Artstein, David Cobbins, Mark G. Core, Carla Gordon, Anton Leuski, Zongjian Li, Chirag Merchant, David Nelson, Mohammad Soleymani 0001, David R. Traum
IVA8
2022 Re-architecting the virtual human toolkit: towards an interoperable platform for embodied conversational agent research and development
abstract
The research and development (R&D) of intelligent virtual agents (IVAs) is inherently complex. We aim to manage this complexity by combining the best aspects of academic and commercial approaches into a principled R&D platform that emphasizes interoperability, ex-tendability, re-use, and support for multiple hardware targets. This IVA platform, the Virtual Human Toolkit 2.0, is a re-architecture of our earlier work and combines a modular message passing architecture with that of a microservices architecture. This paper discusses our approach, design decisions, lessons learned, and current status of this ongoing effort. We illustrate the strengths of the architecture, how best to use commodity AI cloud services in one's own work, and how to port legacy stand-alone software to a web service.
Arno Hartholt, Edward Fast, Zongjian Li, Kevin Kim, Andrew Leeds, Sharon Mozgai
IVA3
2020 OpenSense: A Platform for Multimodal Data Acquisition and Behavior Perception
abstract
Automatic multimodal acquisition and understanding of social signals is an essential building block for natural and effective human-machine collaboration and communication. This paper introduces OpenSense, a platform for real-time multimodal acquisition and recognition of social signals. OpenSense enables precisely synchronized and coordinated acquisition and processing of human behavioral signals. Powered by the Microsoft's Platform for Situated Intelligence, OpenSense supports a range of sensor devices and machine learning tools and encourages developers to add new components to the system through straightforward mechanisms for component integration. This platform also offers an intuitive graphical user interface to build application pipelines from existing components. OpenSense is freely available for academic research.
Kalin Stefanov, Baiyu Huang, Zongjian Li, Mohammad Soleymani 0001
ICMI3
2019 Mining the rank of universities with Wikipedia
Zongjian Li, Cong Li 0009, Xiang Li 0010
Sci. China Inf. Sci.1
2017 A novel discrete sinusoidal signal integrator (DSSI) based adaptive frequency controller for single-phase grid-connected inverter
abstract
A novel discrete sinusoidal signal integrator (DSSI) structure is proposed in this paper for single grid-tied inverters. The DSSI provides an adjustable gain and bandwidth at the expected frequency and a highly attenuated gain at other frequencies which make the DSSI based controllers obtain the zero steady-state error tracking performance of arbitrary frequencies. The performance of the conventional proportional resonant (PR) or quasi-PR controllers is limited by the discrete process. The DSSI based controller is implemented with difference equations directly without any discrete process and the parameters of DSSI are explicit and concise, which is convenient to implement the adaptive frequency control without resource-consuming. The experimental results provide a real-time comparison among the PR, PI and DSSI based controllers, which validate the theoretical analysis.
Zongjian Li, Jun Wang 0054, Xi Jiang 0005, Z. John Shen
IECON1