EDBT 2026 Demo / reviewers in the wild / expert
Sijing Wu
dblp:73/11354
· DBLP profile ↗
21ranked-venue papers
6as first author
21since 2021 · last 2026
0009-0000-7753-1596ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 18 · 6 first-author · 18 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Q-Agent: An MLLM-Driven Framework for Universal Visual Quality Assessment
Peihang Chen, Huiyu Duan, Zitong Xu, Yuqin Cao, Sijing Wu, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
QoMEX | 7 |
| 2026 | MI3S: A multimodal large language model assisted quality assessment framework for AI-generated talking heads
Yingjie Zhou 0003, Sijing Wu, Jun Jia, Yanwei Jiang, Wei Sun 0029, Xiaohong Liu 0001, Xiongkuo Min, Guangtao Zhai |
Inf. Process. Manag. | 3 |
| 2026 | DHQA-4D: A large-scale dataset and LMM-based metric for dynamic 4D digital human quality assessment
Sijing Wu, Yucheng Zhu, Huiyu Duan, Wei Sun 0029, Xiongkuo Min, Guangtao Zhai |
Pattern Recognit. | 2 |
| 2026 | Surveillance Facial Image Quality Assessment: A Multi-Dimensional Dataset and Lightweight ModelabstractSurveillance facial images are often captured under unconstrained conditions, resulting in severe quality degradation due to factors such as low resolution, motion blur, occlusion, and poor lighting. Although recent face restoration techniques applied to surveillance cameras can significantly enhance visual quality, they often compromise fidelity (i.e., identity-preserving features), which directly conflicts with the primary objective of surveillance images -- reliable identity verification. Existing facial image quality assessment (FIQA) predominantly focus on either visual quality or recognition-oriented evaluation, thereby failing to jointly address visual quality and fidelity, which are critical for surveillance applications. To bridge this gap, we propose the first comprehensive study on surveillance facial image quality assessment (SFIQA), targeting the unique challenges inherent to surveillance scenarios. Specifically, we first construct SFIQA-Bench, a multi-dimensional quality assessment benchmark for surveillance facial images, which consists of 5,004 surveillance facial images captured by three widely deployed surveillance cameras in real-world scenarios. A subjective experiment is conducted to collect six dimensional quality ratings, including noise, sharpness, colorfulness, contrast, fidelity and overall quality, covering the key aspects of SFIQA. Furthermore, we propose SFIQA-Assessor, a lightweight multi-task FIQA model that jointly exploits complementary facial views through cross-view feature interaction, and employs learnable task tokens to guide the unified regression of multiple quality dimensions. The experiment results on the proposed dataset show that our method achieves the best performance compared with the state-of-the-art general image quality assessment (IQA) and FIQA methods, validating its effectiveness for real-world surveillance applications. Yanwei Jiang, Wei Sun 0029, Yingjie Zhou 0003, Yuqin Cao, Jun Jia, Sijing Wu, Dandan Zhu 0001, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2026 | AGHI-QA: A Subjective-Aligned Dataset and Metric for AI-Generated Human ImagesabstractThe rapid development of text-to-image (T2I) generation approaches has attracted extensive interest in evaluating the quality of generated images, leading to the development of various quality assessment methods for general-purpose T2I outputs. However, existing image quality assessment (IQA) methods are limited to providing global quality scores, failing to deliver fine-grained perceptual evaluations for structurally complex subjects like humans, which is a critical challenge considering the frequent anatomical and textural distortions in AI-generated human images (AGHIs). To address this gap, we introduce AGHI-QA, a large-scale benchmark specifically designed for quality assessment of AGHIs. The dataset comprises 4, 000 images generated from 400 carefully crafted text prompts using 10 state-of-the-art T2I models. We conduct a systematic subjective study to collect multidimensional annotations, including perceptual quality scores, text-image correspondence scores, visible and distorted body part labels. Based on AGHI-QA, we evaluate the strengths and weaknesses of current T2I methods in generating human images from multiple dimensions. Furthermore, we propose AGHI-Assessor, a novel quality metric that integrates the large multimodal model (LMM) with domain-specific human features for precise quality prediction and identification of visible and distorted body parts in AGHIs. Extensive experimental results demonstrate that AGHI-Assessor showcases state-of-the-art performance, significantly outperforming existing IQA methods in multidimensional quality assessment and surpassing leading LMMs in detecting structural distortions in AGHIs. Sijing Wu, Wei Sun 0029, Yucheng Zhu, Huiyu Duan, Xiongkuo Min, Guangtao Zhai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2026 | SingingHead: A Large-Scale 4D Dataset for Singing Head AnimationabstractSinging, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often overlooked in the field of audio-driven 3D facial animation due to the lack of singing head datasets and the domain gap between singing and talking in rhythm and amplitude. To this end, we collect a large-scale high-quality multi-modal singing head dataset,SingingHead, which consists of more than 27 hours of synchronized singing video, 3D facial motion, singing audio, and background music from 76 individuals and 8 types of music. Along with the SingingHead dataset, we benchmark existing audio-driven 3D facial animation methods and 2D talking head methods on the singing task. Existing 3D facial animation methods and 2D talking head methods fail to produce satisfactory singing results. Focusing on the 3D singing head animation, we first utilize the proposed singing-specific dataset to retrain the 3D facial animation methods, resulting in substantial performance improvements. Besides, considering the absence of background music and the slow generation speed of existing methods, we propose a simple but efficient non-autoregressive VAE-based framework with background music as an input signal to generate diverse and accurate 3D singing facial motions in real time. Extensive experiments demonstrate the significance of the SingingHead dataset in promoting the development of singing head animation. The dataset is released for research purposes at:https://wsj-sjtu.github.io/SingingHead/. Sijing Wu, Weitian Zhang, Jun Jia, Yucheng Zhu, Yichao Yan, Guangtao Zhai, Xiaokang Yang 0001 |
IEEE Trans. Multim. | 1 |
| 2025 | Disentangled Clothed Avatar Generation with Layered RepresentationabstractClothed avatar generation has wide applications in virtual and augmented reality, filmmaking, and more. Previous methods have achieved success in generating diverse digital avatars, however, generating avatars with disentangled components (\eg, body, hair, and clothes) has long been a challenge. In this paper, we propose LayerAvatar, the first feed-forward diffusion-based method for generating component-disentangled clothed avatars. To achieve this, we first propose a layered UV feature plane representation, where components are distributed in different layers of the Gaussian-based UV feature plane with corresponding semantic labels. This representation supports high-resolution and real-time rendering, as well as expressive animation including controllable gestures and facial expressions. Based on the well-designed representation, we train a single-stage diffusion model and introduce constrain terms to address the severe occlusion problem of the innermost human body layer. Extensive experiments demonstrate the impressive performances of our method in generating disentangled clothed avatars, and we further explore its applications in component transfer. The project page is available at: https://olivia23333.github.io/LayerAvatar/ Weitian Zhang, Yichao Yan, Sijing Wu, Manwen Liao, Xiaokang Yang 0001 |
ICCV | 3 |
| 2025 | Multi-Dimensional Text-to-Face Image Quality Assessment Using LLM: Database and MethodabstractWith the rise of Text-to-Image (T2I) models, generating face images from text prompts has emerged as a prominent research area. However, evaluating the quality of these generated face images, particularly with respect to fine-grained facial attributes, remains a significant challenge. To address this, we introduce the Fine -grained Text-to-Face Image Quality Assessment (FineTFIQA) database, which is designed to evaluate the ability of T2I models to generate fine-grained face images. To the best of our knowledge, this database is the largest of its kind, containing 7,218 face images generated from 1,000 text prompts that cover 111 distinct facial attributes. A large group of subjects was invited to assess the quality of text-to-face images on four evaluation dimensions: perceptual quality, human likeness, attractiveness, and consistency. Additionally, we develop the Multi-Dimensional Text-to-Face Image Quality Assessment (MDTFIQA) method based on the Large Language Model (LLM), which combines both face image features and text features to evaluate generated images on all evaluation dimensions. Extensive experimental results demonstrate that traditional face image assessment methods and general image quality assessment methods are inadequate for accurately evaluating generated text-to-face images. Our method significantly outperforms these existing methods on all evaluation dimensions, proving to be an effective method for assessing the quality of generated text-to-face images. Xiongkuo Min, Jinliang Han, Yuqin Cao, Sijing Wu, Yunze Dou, Guangtao Zhai |
ACM Multimedia | 5 |
| 2025 | RGC-VQA: An Exploration Database for Robotic-Generated Video Quality AssessmentabstractAs camera-equipped robotic platforms become increasingly integrated into daily life, robotic-generated videos have begun to appear on streaming media platforms, enabling us to envision a future where humans and robots coexist. We innovatively propose the concept of Robotic-Generated Content (RGC) to term these videos generated from egocentric perspective of robots. The perceptual quality of RGC videos is critical in human-robot interaction scenarios, and RGC videos exhibit unique distortions and visual requirements that differ markedly from those of professionally-generated content (PGC) videos and user-generated content (UGC) videos. However, dedicated research on quality assessment of RGC videos is still lacking. To address this gap and to support broader robotic applications, we establish the first Robotic-Generated Content Database (RGCD), which contains a total of 2,100 videos drawn from three robot categories and sourced from diverse platforms. A subjective VQA experiment is conducted subsequently to assess human visual perception of robotic-generated videos. Finally, we conduct a benchmark experiment to evaluate the performance of 11 state-of-the-art VQA models on our database. Experimental results reveal significant limitations in existing VQA models when applied to complex, robotic-generated content, highlighting a critical need for RGC-specific VQA models. Our RGCD is publicly available at: https://github.com/IntMeGroup/RGC-VQA. Jianing Jin, Jiangyong Ying, Huiyu Duan, Sijing Wu, Yushuo Zheng, Xiongkuo Min, Guangtao Zhai |
ACM Multimedia | 5 |
| 2025 | HVEval: Towards Unified Evaluation of Human-Centric Video Generation and UnderstandingabstractHuman-centric videos play a significant role in the pervasive video content of modern life. However, the capabilities of text-to-video (T2V) generation models and video-to-text (V2T) understanding models for human-centric videos remain largely unexplored. To this end, we present HVEval, the first comprehensive evaluation dataset focusing on human-centric videos, which consists of 20,000 videos, 60k MOS annotations across 3 dimensions (i.e., spatial quality, temporal quality, and text-video correspondence), and 20k category-specific Q&A pairs. Based on the HVEval dataset, this paper aims to answer three questions: (1) can today's T2V models effectively generate human-centric videos following the given prompts? (2) how effective are today's V2T LMMs in understanding and evaluating human-centric videos? (3) are current VQA metrics good enough for evaluating human-centric videos? Comprehensive evaluations of 24 T2V models, 20 LMMs, and 18 VQA metrics reveal their limitations in fine-grained text-controlled generation and human-aligned perception and understanding, highlighting the significant potential of our dataset and benchmarks to advance research in human-centric video generation and understanding. Sijing Wu, Huiyu Duan, Yanwei Jiang, Yucheng Zhu, Guangtao Zhai |
ACM Multimedia | 1 |
| 2025 | FVQ: A Large-Scale Dataset and an LMM-based Method for Face Video Quality AssessmentabstractFace video quality assessment (FVQA) deserves to be explored in addition to general video quality assessment (VQA), as face videos are the primary content on social media platforms and human visual system (HVS) is particularly sensitive to human faces. However, FVQA is rarely explored due to the lack of large-scale FVQA datasets. To fill this gap, we present the first large-scale in-the-wild FVQA dataset, FVQ-20K, which contains 20,000 in-the-wild face videos together with corresponding mean opinion score (MOS) annotations. Along with the FVQ-20K dataset, we further propose a specialized FVQA method named FVQ-Rater to achieve human-like rating and scoring for face video, which is the first attempt to explore the potential of large multimodal models (LMMs) for the FVQA task. Concretely, we elaborately extract multi-dimensional features including spatial features, temporal features, and face-specific features (i.e., portrait features and face embeddings) to provide comprehensive visual information, and take advantage of the LoRA-based instruction tuning technique to achieve quality-specific fine-tuning, which shows superior performance on both FVQ-20K and CFVQA datasets. Extensive experiments and comprehensive analysis demonstrate the significant potential of the FVQ-20K dataset and FVQ-Rater method in promoting the development of FVQA. The code and dataset will be released at: https://github.com/wsj-sjtu/FVQ. Sijing Wu, Ziwen Xu, Huiyu Duan, Wei Sun 0029, Guangtao Zhai |
ACM Multimedia | 1 |
| 2025 | LMME3DHF: Benchmarking and Evaluating Multimodal 3D Human Face Generation with LMMsabstractThe rapid advancement in generative artificial intelligence have enabled the creation of 3D human faces (HFs) for applications including media production, virtual reality, security, healthcare, and game development, etc. However, assessing the quality and realism of these AI-generated 3D human faces remains a significant challenge due to the subjective nature of human perception and innate perceptual sensitivity to facial features. To this end, we conduct a comprehensive study on the quality assessment of AI-generated 3D human faces. We first introduce Gen3DHF, a large-scale benchmark comprising 2,000 videos of AI-Generated 3D Human Faces along with 4,000 Mean Opinion Scores (MOS) collected across two dimensions, i.e., quality and authenticity, 2,000 distortion-aware saliency maps and distortion descriptions. Based on Gen3DHF, we propose LMME3DHF, a Large Multimodal Model (LMM)-based metric for Evaluating 3DHF capable of quality and authenticity score prediction, distortion-aware visual question answering, and distortion-aware saliency prediction. Experimental results show that LMME3DHF achieves state-of-the-art performance, surpassing existing methods in both accurately predicting quality scores for AI-generated 3D human faces and effectively identifying distortion-aware salient regions and distortion types, while maintaining strong alignment with human perceptual judgments. Both the Gen3DHF database and the LMME3DHF will be released upon the publication. Woo Yi Yang, Sijing Wu, Huiyu Duan, Guangtao Zhai, Xiongkuo Min |
ACM Multimedia | 3 |
| 2025 | Ges-QA: A Multidimensional Quality Assessment Dataset for Audio-to-3D Gesture GenerationabstractThe Audio-to-3D-Gesture (A2G) task exhibits significant potential across domains including virtual reality, computer graphics, and 3D animation production. However, current evaluation metrics, such as Fréchet Gesture Distance or Beat Constancy, fail at reflecting the human preference of the generated 3D gestures. To cope with this problem, exploring human preference and an objective quality assessment metric for AI-generated 3D human gestures is becoming increasingly significant. In this paper, we introduce the Ges-QA dataset, which includes 1,400 samples with multidimensional scores for gesture quality and audio-gesture consistency. Moreover, we collect binary classification labels to determine whether the generated gestures match the emotions of the audio. Equipped with our Ges-QA dataset, we propose a multi-modal transformer-based neural network with 3 branches for video, audio and 3D skeleton modalities, which can score A2G contents in multiple dimensions. Comparative experimental results and ablation studies demonstrate that Ges-QAer yields state-of-the-art performance on our dataset. Zhilin Gao, Sijing Wu, Yuqin Cao, Huiyu Duan, Guangtao Zhai |
VCIP | 3 |
| 2025 | Hybrid attention multi-scale feature aggregation for efficient nuclei segmentation and classification in H&E-stained images
Xingpeng Zhang, Qiuli Wang 0001, Sijing Wu |
Multim. Syst. | 5 |
| 2025 | Knowledge Integration for Grounded Situation Recognition
Jiaming Lei, Sijing Wu, Lin Li 0065, Lei Chen 0082, Jun Xiao 0001, Yi Yang 0001, Long Chen 0016 |
Pattern Recognit. | 2 |
| 2024 | UniProcessor: A Text-Induced Unified Low-Level Image Processor
Huiyu Duan, Xiongkuo Min, Sijing Wu, Wei Shen 0002, Guangtao Zhai |
ECCV (67) | 3 |
| 2024 | HQ-Avatar: Towards High-Quality 3D Avatar Generation via Point-based RepresentationabstractDespite the flourishing of 3D object generation, generating high-quality digital avatars with detailed geometry and texture that are free to animate remains a challenging task. Existing avatar generation techniques often suffer from limitations such as low-quality geometry and blurry texture. Thus, we propose HQ-Avatar, a novel method for generating animatable avatars with high-quality geometry and texture. We enhance the geometry quality by proposing an importance sampling strategy and capturing intricate details through learned normal maps. To achieve high-quality texture, we present a neural point-based avatar representation, which enables high-resolution rendering results at a resolution of 10242, allowing detailed supervision. Extensive experiments on THuman2.0 dataset demonstrate the superiority of our method over state-of-the-art techniques in generating high-quality avatars. Furthermore, we show the applicability of our method by employing it as 3D priors to simplify the human avatar reconstruction process from scans or even single images. Code is available at https://github.com/olivia23333/HQ-Avatar Weitian Zhang, Sijing Wu, Yichao Yan, Ben Xue, Wenhan Zhu, Xiaokang Yang 0001 |
ICME | 2 |
| 2024 | MMHead: Towards Fine-grained Multi-modal 3D Facial Animationabstract3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation, especially text-guided 3D facial animation is rarely explored due to the lack of multi-modal 3D facial animation dataset. To fill this gap, we first construct a large-scale multi-modal 3D facial animation dataset, MMHead, which consists of 49 hours of 3D facial motion sequences, speech audios, and rich hierarchical text annotations. Each text annotation contains abstract action and emotion descriptions, fine-grained facial and head movements (i.e., expression and head pose) descriptions, and three possible scenarios that may cause such emotion. Concretely, we integrate five public 2D portrait video datasets, and propose an automatic pipeline to 1) reconstruct 3D facial motion sequences from monocular videos; and 2) obtain hierarchical text annotations with the help of AU detection and ChatGPT. Based on the MMHead dataset, we establish benchmarks for two new tasks: text-induced 3D talking head animation and text-to-3D facial motion generation. Moreover, a simple but efficient VQ-VAE-based method named MM2Face is proposed to unify the multi-modal information and generate diverse and plausible 3D facial motions, which achieves competitive results on both benchmarks. Extensive experiments and comprehensive analysis demonstrate the significant potential of our dataset and benchmarks in promoting the development of multi-modal 3D facial animation. The dataset will be released at: https://wsj-sjtu.github.io/MMHead/. Sijing Wu, Yichao Yan, Huiyu Duan, Ziwei Liu 0002, Guangtao Zhai |
ACM Multimedia | 1 |
| 2024 | MVBind: Self-Supervised Music Recommendation for Videos via Embedding Space BindingabstractRecent years have witnessed the rapid development of short videos, which usually contain both visual and audio modalities. Background music is important to the short videos, which can significantly influence the emotions of the viewers. However, at present, the background music of short videos is generally chosen by the video producer, and there is a lack of automatic music recommendation methods for short videos. This paper introduces MVBind, an innovative Music-Video embedding space Binding model for cross-modal retrieval. MVBind operates as a self-supervised approach, acquiring inherent knowledge of intermodal relationships directly from data, without the need of manual annotations. Additionally, to compensate the lack of a corresponding musical-visual pair dataset for short videos, we construct a dataset, SVM-10K (Short Video with Music-10K), which mainly consists of meticulously selected short videos. On this dataset, MVBind manifests significantly improved performance compared to other baseline methods. The database and code are available at: https://github.com/IntMeGroup/MVBind. Jiajie Teng, Huiyu Duan, Yucheng Zhu, Sijing Wu, Guangtao Zhai |
VCIP | 4 |
| 2023 | GANHead: Towards Generative Animatable Neural Head AvatarsabstractTo bring digital avatars into people's lives, it is highly demanded to efficiently generate complete, realistic, and animatable head avatars. This task is challenging, and it is difficult for existing methods to satisfy all the requirements at once. To achieve these goals, we propose GANHead (Generative Animatable Neural Head Avatar), a novel generative head model that takes advantages of both the fine-grained control over the explicit expression parameters and the realistic rendering results of implicit representations. Specifically, GANHead represents coarse geometry, fine-gained details and texture via three networks in canonical space to obtain the ability to generate complete and realistic head avatars. To achieve flexible animation, we define the deformation filed by standard linear blend skinning (LBS), with the learned continuous pose and expression bases and LBS weights. This allows the avatars to be directly animated by FLAME [22] parameters and generalize well to unseen poses and expressions. Compared to state-of-the-art (SOTA) methods, GANHead achieves superior performance on head avatar generation and raw scan fitting. Sijing Wu, Yichao Yan, Yuhao Cheng, Wenhan Zhu, Ke Gao 0012, Guangtao Zhai |
CVPR | 1 |
| 2021 | Accurate Compensation Makes the World More Clear for the Visually ImpairedabstractVisual impairment is one of the most serious social and public health problems in the world, therefore, it is of great theoretical and practical significance to study the image enhancement algorithms for the visually impaired, which is the basis for the development of assistive devices. In this paper, a general deep learning based image enhancement framework for the visually impaired is proposed, which can be used to enhance images to compensate for any visually impaired symptom that can be modeled. Take central vision loss as an example, we first model the central vision loss based on the contrast sensitivity function (CSF) specified by clinical indicator Pelli-Robson score and logMAR visual acuity, and then use the proposed framework to generate an image enhancement method aiming at compensating for the central vision loss. Both the simulation experiment and the patient experiment show the superiority of the proposed image enhancement method designed for the central vision loss, which also validates the effectiveness of the proposed framework. Sijing Wu, Huiyu Duan, Xiongkuo Min, Danyang Tu, Guangtao Zhai |
ICIP | 1 |