Peipei Li 0002

dblp:60/3675-2 · also Pei-Pei Li 0002 · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
22since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 16 · 6 first-author · 12 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 11 since 2021Security and privacy · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 T2Agent: A Tool-augmented Multimodal Misinformation Detection Agent with Monte Carlo Tree Search
abstract
Real-world multimodal misinformation often arises from mixed forgery sources, requiring dynamic reasoning and adaptive verification. However, existing methods mainly rely on static pipelines and limited tool usage, limiting their ability to handle such complexity and diversity. To address this challenge, we propose T2Agent, a novel misinformation detection agent that incorporates an extensible toolkit with Monte Carlo Tree Search (MCTS). The toolkit consists of modular tools such as web search, forgery detection, and consistency analysis. Each tool is described using standardized templates, enabling seamless integration and future expansion. To avoid inefficiency from using all tools simultaneously, a greedy search-based selector is proposed to identify a task-relevant subset. This subset then serves as the action space for MCTS to dynamically collect evidence and perform multi-source verification. To better align MCTS with the multi-source nature of misinformation detection, T2Agent extends traditional MCTS with multi-source verification, which decomposes the task into coordinated subtasks targeting different forgery sources. A dual reward mechanism containing a reasoning trajectory score and a confidence score is further proposed to encourage a balance between exploration across mixed forgery sources and exploitation for more reliable evidence. We conduct ablation studies to confirm the effectiveness of the tree search mechanism and tool usage. Extensive experiments further show that T2Agent consistently outperforms existing baselines on challenging mixed-source multimodal misinformation benchmarks, demonstrating its strong potential as a training-free detector.
Xing Cui, Yueying Zou, Zekun Li 0001, Peipei Li 0002, Xuannan Liu, Huaibo Huang
AAAI4
2025 Can Machines Understand Composition? Dataset and Benchmark for Photographic Image Composition Embedding and Understanding
abstract
With the rapid growth of social media and digital photography, visually appealing images have become essential for effective communication and emotional engagement. Among the factors influencing aesthetic appeal, composition—the arrangement of visual elements within a frame—plays a crucial role. In recent years, specialized models for photographic composition have achieved impressive results across various aesthetic tasks. Meanwhile, rapidly advancing multimodal large language models (MLLMs) have excelled in several visual perception tasks. However, their ability to embed and understand compositional information remains underexplored, primarily due to the lack of suitable evaluation datasets. To address this gap, we introduce the Photographic Image Composition Dataset (PICD), a large-scale dataset consisting of 36,857 images categorized into 24 composition categories across 355 diverse scenes. We demonstrate the advantages of PICD over existing datasets in terms of data scale, composition category, label quality, and scene diversity. Building on PICD, we establish benchmarks to evaluate the composition embedding capabilities of specialized models and the compositional understanding ability of MLLMs. To enable efficient and effective evaluation, we propose a novel Composition Discrimination Accuracy (CDA) metric. Our evaluation highlights the limitations of current models and provides insights into directions for improving their ability to embed and understand composition.
Zhaoran Zhao, Peng Lu 0007, Peipei Li 0002, Xuannan Liu, Shiyi Chen, Wenhao Guo 0003
CVPR4
2025 Eye Movements as Images: A Multimodal Framework for Eye Movements Representation
abstract
Eye movements are increasingly popular for enhancing natural language processing and modeling individual states. Although specialized methods have been developed to represent eye movements for various tasks, effectively modeling the complex dynamics of eye movements and the heterogeneity with stimulus text remains challenging. This paper proposes a text-guided eye movement representation framework that introduces a novel perspective by converting raw eye movement sequences into line graph images and encoding them with a powerful pre-trained vision transformer. To address the disparities between eye movements and text, we guide their temporal alignment using human reading order and combine Canonical Correlation Analysis with Optimal Transport to fuse the two modalities. This approach not only significantly simplifies the design of specialized models but also has the potential to become a universal representation for eye movements. Experimental results on six different domain tasks show that the proposed method achieves state-of-the-art performance. We release the source code at https://github.com/wulalahalala/VLEM.
Dongsen Zhang, Peipei Li 0002, Zekun Li 0001, Yiwei Ru, Huijia Wu, Zhaofeng He 0001
ICASSP2
2025 MMFakeBench: A Mixed-Source Multimodal Misinformation Detection Benchmark for LVLMs
abstract
Current multimodal misinformation detection (MMD) methods often assume a single source and type of forgery for each sample, which is insufficient for real-world scenarios where multiple forgery sources coexist. The lack of a benchmark for mixed-source misinformation has hindered progress in this field. To address this, we introduce MMFakeBench, the first comprehensive benchmark for mixed-source MMD. MMFakeBench includes 3 critical sources: textual veracity distortion, visual veracity distortion, and cross-modal consistency distortion, along with 12 sub-categories of misinformation forgery types. We further conduct an extensive evaluation of 6 prevalent detection methods and 15 Large Vision-Language Models (LVLMs) on MMFakeBench under a zero-shot setting. The results indicate that current methods struggle under this challenging and realistic mixed-source MMD setting. Additionally, we propose MMD-Agent, a novel approach to integrate the reasoning, action, and tool-use capabilities of LVLM agents, significantly enhancing accuracy and generalization. We believe this study will catalyze future research into more realistic mixed-source multimodal misinformation and provide a fair evaluation of misinformation detection methods.
Xuannan Liu, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Shuhan Xia, Xing Cui, Linzhi Huang, Weihong Deng, Zhaofeng He 0001
ICLR3
2025 SpineBench: Benchmarking Multimodal LLMs for Spinal Pathology Analysis
abstract
With the increasing integration of Multimodal Large Language Models (MLLMs) into the medical field, comprehensive evaluation of their performance in various medical domains becomes critical. However, existing benchmarks primarily assess general medical tasks, inadequately capturing performance in nuanced areas like the spine, which relies heavily on visual input. To address this, we introduce SpineBench, a comprehensive Visual Question Answering (VQA) benchmark designed for fine-grained analysis and evaluation of MLLMs in the spinal domain. SpineBench comprises 64,878 QA pairs from 40,263 spine images, covering 11 spinal diseases through two critical clinical tasks: spinal disease diagnosis and spinal lesion localization, both in multiple-choice format. SpineBench is built by integrating and standardizing image-label pairs from open-source spinal disease datasets, and samples challenging hard negative options for each VQA pair based on visual similarity (similar but not the same disease), simulating real-world challenging scenarios. We evaluate 12 leading MLLMs on SpineBench. The results reveal that these models exhibit poor performance in spinal tasks, highlighting limitations of current MLLM in the spine domain and guiding future improvements in spinal medicine applications. SpineBench is publicly available at https://zhangchenghanyu.github.io/SpineBench.github.io/.
Chenghanyu Zhang, Zekun Li 0001, Peipei Li 0002, Xing Cui, Shuhan Xia, Weixiang Yan, Yiqiao Zhang, Qianyu Zhuang
ACM Multimedia3
2025 Video-SafetyBench: A Benchmark for Safety Evaluation of Video LVLMs
abstract
The increasing deployment of Large Vision-Language Models (LVLMs) raises safety concerns under potential malicious inputs. However, existing multimodal safety evaluations primarily focus on model vulnerabilities exposed by static image inputs, ignoring the temporal dynamics of video that may induce distinct safety risks. To bridge this gap, we introduce Video-SafetyBench, the first comprehensive benchmark designed to evaluate the safety of LVLMs under video-text attacks. It comprises 2,264 video-text pairs spanning 48 fine-grained unsafe categories, each pairing a synthesized video with either a harmful query, which contains explicit malice, or a benign query, which appears harmless but triggers harmful behavior when interpreted alongside the video. To generate semantically accurate videos for safety evaluation, we design a controllable pipeline that decomposes video semantics into subject images (what is shown) and motion text (how it moves), which jointly guide the synthesis of query-relevant videos. To effectively evaluate uncertain or borderline harmful outputs, we propose RJScore, a novel LLM-based metric that incorporates the confidence of judge models and human-aligned decision threshold calibration. Extensive experiments show that benign-query video composition achieves average attack success rates of 67.2%, revealing consistent vulnerabilities to video-induced attacks. We believe Video-SafetyBench will catalyze future research into video-based safety evaluation and defense strategies.
Xuannan Liu, Zekun Li 0001, Zheqi He, Peipei Li 0002, Shuhan Xia, Xing Cui, Huaibo Huang, Xi Yang 0023, Ran He 0001
NeurIPS4
2025 A*-Thought: Efficient Reasoning via Bidirectional Compression for Low-Resource Settings
abstract
Large Reasoning Models (LRMs) achieve superior performance by extending the thought length. However, a lengthy thinking trajectory leads to reduced efficiency. Most of the existing methods are stuck in the assumption of overthinking and attempt to reason efficiently by compressing the Chain-of-Thought, but this often leads to performance degradation. To address this problem, we introduce A*-Thought, an efficient tree search-based unified framework designed to identify and isolate the most essential thoughts from the extensive reasoning chains produced by these models. It formulates the reasoning process of LRMs as a search tree, where each node represents a reasoning span in the giant reasoning space. By combining the A* search algorithm with a cost function specific to the reasoning path, it can efficiently compress the chain of thought and determine a reasoning path with high information density and low cost. In addition, we also propose a bidirectional importance estimation mechanism, which further refines this search process and enhances its efficiency beyond uniform sampling. Extensive experiments on several advanced math tasks show that A*-Thought effectively balances performance and efficiency over a huge search space. Specifically, A*-Thought can improve the performance of QwQ-32B by 2.39$\times$ with low-budget and reduce the length of the output token by nearly 50\% with high-budget. The proposed method is also compatible with several other LRMs, demonstrating its generalization capability. The code can be accessed at: https://github.com/AI9Stars/AStar-Thought.
Xiaoang Xu, Shuo Wang 0013, Zhenghao Liu 0001, Huijia Wu, Peipei Li 0002, Zhiyuan Liu 0001, Maosong Sun 0001, Zhaofeng He 0001
NeurIPS6
2025 CAHS-Attack: CLIP-Aware Heuristic Search Attack Method for Stable Diffusion
abstract
Diffusion models exhibit notable fragility when faced with adversarial prompts, and strengthening attack capabilities is crucial for uncovering such vulnerabilities and building more robust generative systems. Existing works often rely on white-box access to model gradients or hand-crafted prompt engineering, which is infeasible in real-world deployments due to restricted access or poor attack effect. In this paper, we propose CAHS-Attack, a CLIP-Aware Heuristic Search attack method. CAHS-Attack integrates Monte Carlo Tree Search (MCTS) to perform fine-grained suffix optimization, leveraging a constrained genetic algorithm to preselect high-potential adversarial prompts as root nodes, and retaining the most semantically disruptive outcome at each simulation rollout for efficient local search. Extensive experiments demonstrate that our method achieves state-of-the-art attack performance across both short and long prompts of varying semantics. Furthermore, we find that the fragility of SD models can be attributed to the inherent vulnerability of their CLIP-based text encoders, suggesting a fundamental security risk in current text-to-image pipelines.
Shuhan Xia, Hui Ouyang, Yadong Shang, Dongxiao Zhao, Peipei Li 0002
TrustCom6
2025 AdvCloak: Customized adversarial cloak for privacy protection
Xuannan Liu, Yaoyao Zhong, Xing Cui, Yuhang Zhang 0016, Peipei Li 0002, Weihong Deng
Pattern Recognit.5
2025 Toward Real-World Remote Sensing Image Super-Resolution: A New Benchmark and an Efficient Model
abstract
Super-resolution (SR) is a fundamental and crucial task in remote sensing. It can improve low-resolution (LR) remote sensing images and has potential benefits for downstream tasks such as remote sensing object detection and recognition. Existing remote sensing image SR (RSISR) methods are trained on simulated paired datasets, in which LR images are obtained by a simple and uniform (i.e., bicubic) degradation from corresponding high-resolution (HR) images. However, since this simulated degradation usually deviates from the real degradation, the performance of the trained model is limited when applied to real scenarios. To address this issue, we construct a novel real-world RSISR (RRSISR) dataset to model the real-world degradation, which exploits the imaging characteristics of the spectral camera to capture paired LR-HR images of the same scene. To ensure the precise alignment of the paired images, algorithms such as image registration and geometric correction are utilized. In addition, considering the vast amount of data involved in the RSISR task and its requirement for higher efficiency, we divide the image into patches with different restoration difficulties and propose a reference table-based patch exiting (RPE) method to efficiently reduce the computation of SR. Specifically, this method incorporates a predictor to estimate the performance of the current layer and a lookup table to decide whether to exit. Extensive experiments show that models trained on the proposed RRSISR dataset produce more realistic images than models with simulated datasets and generalize well to other satellites. We also demonstrate the efficiency of our RPE.
Jia Wang 0038, Liuyu Xiang, Jiaochong Xu, Peipei Li 0002, Qizhi Xu, Zhaofeng He 0001
IEEE Trans. Geosci. Remote. Sens.5
2025 FDNet: A Frequency-Aware Decomposition Network for Robust Face Super-Resolution Against Adversarial Attacks
abstract
Face super-resolution (FSR) is a crucial step in the face analysis pipeline, achieving remarkable progress by applying deep neural networks (DNNs). However, DNN-based FSR models are not robust enough and may suffer significant performance degradation due to subtle adversarial perturbations. In addition, the high-frequency details of images restored by existing models are insufficient, especially at large upsampling factors. In this paper, we propose a frequency-aware decomposition network (FD-Net) for robust face super-resolution, which aims to defend against adversarial attacks and obtain face images with fidelity. Observing that the noise introduced by adversarial attacks is often intricately mixed with the high-frequency information of the input image, we decompose and process the features of different frequencies separately to eliminate harmful perturbations and enhance high-frequency information. Specifically, by leveraging the frequency-aware capability of empirical mode decomposition (EMD), we propose an EMD-based multi-branch structure. The framework implicitly compels different branches to adaptively extract features from distinct frequency bands, limiting the adversarial noise into decoupled components restricted to specific branches. It also improves the recovery of high-frequency information, which is conducive to producing more credible results. Furthermore, we introduce a high-frequency noise suppressor capable of randomly eliminating imperceptible noise in the high-frequency components. Quantitative and qualitative results demonstrate the superior robustness of our proposed method against adversarial attacks, showing better fidelity in image reconstruction compared to state-of-the-art FSR methods, especially for upscaling factors of 8 and 16.
Jia Wang 0038, Peipei Li 0002, Liuyu Xiang, Rui Wang 0124, Zhaofeng He 0001
IEEE Trans. Inf. Forensics Secur.2
2024 INSTASTYLE: Inversion Noise of a Stylized Image is Secretly a Style Adviser
Xing Cui, Zekun Li 0001, Peipei Li 0002, Huaibo Huang, Xuannan Liu, Zhaofeng He 0001
ECCV (51)3
2024 I3FDM: IRIS Inpainting Via Inverse Fusion of Diffusion Models
abstract
Iris images captured in real-world scenarios are often occluded, which leads to a significant degradation for the iris recognition system. Therefore, it is necessary to propose an effective iris inpainting method. While generative adversarial network (GAN)-based image inpainting methods have shown promise, they often suffer from issues such as mode collapse and training instability. Recently, denoising diffusion probabilistic model (DDPM) has surpassed GAN in terms of image quality while maintaining stable training. Combining DDPM and the characteristics of iris image, I3FDM (Iris Inpainting via Inverse Fusion of Diffusion Models), a method that iteratively modifies intermediate variables in the generation process based on a given occluded image. Since these modifications introduce semantic differences, we introduce an inverse fusion module to enhance the performance of iris inpainting. I3FDM enables the processing of various types of occluded images using an unconditional DDPM without the need for additional learning. Extensive experimental results qualitatively and quantitatively demonstrate that our method produces iris images with richer texture information and improves the performance of iris recognition.
Chenyang Li 0011, Peipei Li 0002, Zhaofeng He 0001
ICASSP3
2024 Exploring 3D-aware Lifespan Face Aging via Disentangled Shape-Texture Representations
abstract
Existing face aging methods often focus on modeling either texture aging or using an entangled shape-texture representation to achieve face aging. However, shape and texture are two distinct factors that mutually affect the human face aging process. In this paper, we propose 3D-STD, a novel 3D-aware Shape-Texture Disentangled face aging network that explicitly disentangles the facial image into shape and texture representations using 3D face reconstruction. Additionally, to facilitate high-fidelity texture synthesis, we propose a novel texture generation method based on Empirical Mode Decomposition (EMD). Extensive qualitative and quantitative experiments show that our method achieves state-of-the-art performance in terms of shape and texture transformation. Moreover, our method supports producing plausible 3D face aging results, which is rarely accomplished by current methods.
Qianrui Teng, Rui Wang 0124, Xing Cui, Peipei Li 0002, Zhaofeng He 0001
ICME4
2024 FKA-Owl: Advancing Multimodal Fake News Detection through Knowledge-Augmented LVLMs
abstract
The massive generation of multimodal fake news involving both text and images exhibits substantial distribution discrepancies, prompting the need for generalized detectors. However, the insulated nature of training restricts the capability of classical detectors to obtain open-world facts. While Large Vision-Language Models (LVLMs) have encoded rich world knowledge, they are not inherently tailored for combating fake news and struggle to comprehend local forgery details. In this paper, we propose FKA-Owl, a novel framework that leverages forgery-specific knowledge to augment LVLMs, enabling them to reason about manipulations effectively. The augmented forgery-specific knowledge includes semantic correlation between text and images, and artifact trace in image manipulation. To inject these two kinds of knowledge into the LVLM, we design two specialized modules to establish their representations, respectively. The encoded knowledge embeddings are then incorporated into LVLMs. Extensive experiments on the public benchmark demonstrate that FKA-Owl achieves superior cross-domain performance compared to previous methods. Code is publicly available at https://liuxuannan.github.io/FKA_Owl.github.io/.
Xuannan Liu, Peipei Li 0002, Huaibo Huang, Zekun Li 0001, Xing Cui, Lixiong Qin, Weihong Deng, Zhaofeng He 0001
ACM Multimedia2
2024 Localize, Understand, Collaborate: Semantic-Aware Dragging via Intention Reasoner
abstract
Flexible and accurate drag-based editing is a challenging task that has recently garnered significant attention. Current methods typically model this problem as automatically learning "how to drag" through point dragging and often produce one deterministic estimation, which presents two key limitations: 1) Overlooking the inherently ill-posed nature of drag-based editing, where multiple results may correspond to a given input, as illustrated in Fig.1; 2) Ignoring the constraint of image quality, which may lead to unexpected distortion. To alleviate this, we propose LucidDrag, which shifts the focus from "how to drag" to "what-then-how" paradigm. LucidDrag comprises an intention reasoner and a collaborative guidance sampling mechanism. The former infers several optimal editing strategies, identifying what content and what semantic direction to be edited. Based on the former, the latter addresses "how to drag" by collaboratively integrating existing editing guidance with the newly proposed semantic guidance and quality guidance. Specifically, semantic guidance is derived by establishing a semantic editing direction based on reasoned intentions, while quality guidance is achieved through classifier guidance using an image fidelity discriminator. Both qualitative and quantitative comparisons demonstrate the superiority of LucidDrag over previous methods.
Xing Cui, Peipei Li 0002, Zekun Li 0001, Xuannan Liu, Yueying Zou, Zhaofeng He 0001
NeurIPS2
2024 Bidirectional Knowledge Reconfiguration for Lightweight Point Cloud Analysis
abstract
Point cloud analysis faces computational system overhead, limiting its application on mobile or edge devices. Directly employing small models may result in a significant drop in performance since it is difficult for a small model to adequately capture local structure and global shape information simultaneously, which are essential clues for point cloud analysis. This paper explores feature distillation for lightweight point cloud models. To mitigate the semantic gap between the lightweight student and the cumbersome teacher, we propose bidirectional knowledge reconfiguration (BKR) to distill informative contextual knowledge from the teacher to the student. Specifically, a top-down knowledge reconfiguration and a bottom-up knowledge reconfiguration are developed to inherit diverse local structure information and consistent global shape knowledge from the teacher, respectively. However, due to the farthest point sampling in most point cloud models, the intermediate features between teacher and student are misaligned, deteriorating the feature distillation performance. To eliminate it, we propose a feature mover's distance (FMD) loss based on optimal transportation, which can measure the distance between unordered point cloud features effectively. Extensive experiments conducted on shape classification, part segmentation, and semantic segmentation benchmarks demonstrate the universality and superiority of our method.
Peipei Li 0002, Xing Cui, Yibo Hu 0001, Man Zhang 0005, Ting Yao 0003, Tao Mei 0001
IEEE Trans. Multim.1
2023 Pluralistic Aging Diffusion Autoencoder
abstract
Face aging is an ill-posed problem because multiple plausible aging patterns may correspond to a given input. Most existing methods often produce one deterministic estimation. This paper proposes a novel CLIP-driven Pluralistic Aging Diffusion Autoencoder (PADA) to enhance the diversity of aging patterns. First, we employ diffusion models to generate diverse low-level aging details via a sequential denoising reverse process. Second, we present Probabilistic Aging Embedding (PAE) to capture diverse high-level aging patterns, which represents age information as probabilistic distributions in the common CLIP latent space. A text-guided KL-divergence loss is designed to guide this learning. Our method can achieve pluralistic face aging conditioned on open-world aging texts and arbitrary unseen face images. Qualitative and quantitative experiments demonstrate that our method can generate more diverse and high-quality plausible aging results.
Peipei Li 0002, Rui Wang 0124, Huaibo Huang, Ran He 0001, Zhaofeng He 0001
ICCV1
2023 Generative Iris Prior Embedded Transformer for Iris Restoration
abstract
Iris restoration from complexly degraded iris images, aiming to improve iris recognition performance, is a challenging problem. Due to the complex degradation, directly training a convolutional neural network (CNN) without prior cannot yield satisfactory results. In this work, we propose a generative iris prior embedded Transformer model (Gformer), in which we build a hierarchical encoder-decoder network employing Transformer block and generative iris prior. First, we tame Transformer blocks to model long-range dependencies in target images. Second, we pretrain an iris generative adversarial network (GAN) to obtain the rich iris prior, and incorporate it into the iris restoration process with our iris feature modulator. Our experiments demonstrate that the proposed Gformer outperforms state-of-the-art methods. Besides, iris recognition performance has been significantly improved after applying Gformer.
Jia Wang 0038, Peipei Li 0002, Liuyu Xiang, Peigang Li, Zhaofeng He 0001
ICME3
2023 Sensing Micro-Motion Human Patterns using Multimodal mmRadar and Video Signal for Affective and Psychological Intelligence
abstract
Affective and psychological perception are pivotal in human-machine interaction and essential domains within artificial intelligence. Existing physiological signal-based affective and psychological datasets primarily rely on contact-based sensors, potentially introducing extraneous affectives during the measurement process. Consequently, creating accurate non-contact affective and psychological perception datasets is crucial for overcoming these limitations and advancing affective intelligence. In this paper, we introduce the Remote Multimodal Affective and Psychological (ReMAP) dataset, for the first time, apply head micro-tremor (HMT) signals for affective and psychological perception. ReMAP features 68 participants and comprises two sub-datasets. The stimuli videos utilized for affective perception undergo rigorous screening to ensure the efficacy and universality of affective elicitation. Additionally, we propose a novel remote affective and psychological perception framework, leveraging multimodal complementarity and interrelationships to enhance affective and psychological perception capabilities. Extensive experiments demonstrate HMT as a "small yet powerful" physiological signal in psychological perception. Our method outperforms existing state-of-the-art approaches in remote affective recognition and psychological perception. The ReMAP dataset is publicly accessible at https://remap-dataset.github.io/ReMAP.
Yiwei Ru, Peipei Li 0002, Muyi Sun, Yunlong Wang 0003, Kunbo Zhang, Qi Li 0005, Zhaofeng He 0001, Zhenan Sun
ACM Multimedia2
2023 Learning-to-Rank Meets Language: Boosting Language-Driven Ordering Alignment for Ordinal Classification
abstract
We present a novel language-driven ordering alignment method for ordinal classification. The labels in ordinal classification contain additional ordering relations, making them prone to overfitting when relying solely on training data. Recent developments in pre-trained vision-language models inspire us to leverage the rich ordinal priors in human language by converting the original task into a vision-language alignment task. Consequently, we propose L2RCLIP, which fully utilizes the language priors from two perspectives. First, we introduce a complementary prompt tuning technique called RankFormer, designed to enhance the ordering relation of original rank prompts. It employs token-level attention with residual-style prompt blending in the word embedding space. Second, to further incorporate language priors, we revisit the approximate bound optimization of vanilla cross-entropy loss and restructure it within the cross-modal embedding space. Consequently, we propose a cross-modal ordinal pairwise loss to refine the CLIP feature space, where texts and images maintain both semantic alignment and ordering alignment. Extensive experiments on three ordinal classification tasks, including facial age estimation, historical color image (HCI) classification, and aesthetic assessment demonstrate its promising performance.
Rui Wang 0124, Peipei Li 0002, Huaibo Huang, Chunshui Cao, Ran He 0001, Zhaofeng He 0001
NeurIPS2
2023 Theme-Aware Aesthetic Distribution Prediction With Full-Resolution Photographs
abstract
Aesthetic quality assessment (AQA) is a challenging task due to complex aesthetic factors. Currently, it is common to conduct AQA using deep neural networks (DNNs) that require fixed-size inputs. The existing methods mainly transform images by resizing, cropping, and padding or use adaptive pooling to alternately capture the aesthetic features from fixed-size inputs. However, these transformations potentially damage aesthetic features. To address this issue, we propose a simple but effective method to accomplish full-resolution image AQA by combining image padding with region of image (RoM) pooling. Padding turns inputs into the same size. RoM pooling pools image features and discards extra padded features to eliminate the side effects of padding. In addition, the image aspect ratios are encoded and fused with visual features to remedy the shape information loss of RoM pooling. Furthermore, we observe that the same image may receive different aesthetic evaluations under different themes, which we call the theme criterion bias. Hence, a theme-aware model that uses theme information to guide model predictions is proposed. Finally, we design an attention-based feature fusion module to effectively use both the shape and theme information. Extensive experiments prove the effectiveness of the proposed method over state-of-the-art methods.
Gengyun Jia, Peipei Li 0002, Ran He 0001
IEEE Trans. Neural Networks Learn. Syst.2
2020 Hierarchical Face Aging Through Disentangled Latent Characteristics
Peipei Li 0002, Huaibo Huang, Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun
ECCV (3)1
2020 Dual-Structure Disentangling Variational Generation for Data-Limited Face Parsing
abstract
Deep learning based face parsing methods have attained state-of-the-art performance in recent years. Their superior performance heavily depends on the large-scale annotated training data. However, it is expensive and time-consuming to construct a large-scale pixel-level manually annotated dataset for face parsing. To alleviate this issue, we propose a novel Dual-Structure Disentangling Variational Generation (D2VG) network. Benefiting from the interpretable factorized latent disentanglement in VAE, D2VG can learn a joint structural distribution of facial image and its corresponding parsing map. Owing to these, it can synthesize large-scale paired face images and parsing maps from a standard Gaussian distribution. Then, we adopt both manually annotated and synthesized data to train a face parsing model in a supervised way. Since there are inaccurate pixel-level labels in synthesized parsing maps, we introduce a coarseness-tolerant learning algorithm, to effectively handle these noisy or uncertain labels. In this way, we can significantly boost the performance of face parsing. Extensive quantitative and qualitative results on HELEN, CelebAMask-HQ and LaPa demonstrate the superiority of our methods.
Peipei Li 0002, Yinglu Liu, Hailin Shi, Xiang Wu 0001, Yibo Hu 0001, Ran He 0001, Zhenan Sun
ACM Multimedia1
2020 Deep label refinement for age estimation
Peipei Li 0002, Yibo Hu 0001, Xiang Wu 0001, Ran He 0001, Zhenan Sun
Pattern Recognit.1
2019 M2FPA: A Multi-Yaw Multi-Pitch High-Quality Dataset and Benchmark for Facial Pose Analysis
abstract
Facial images in surveillance or mobile scenarios often have large view-point variations in terms of pitch and yaw angles. These jointly occurred angle variations make face recognition challenging. Current public face databases mainly consider the case of yaw variations. In this paper, a new large-scale Multi-yaw Multi-pitch high-quality database is proposed for Facial Pose Analysis (M2FPA), including face frontalization, face rotation, facial pose estimation and pose-invariant face recognition. It contains 397,544 images of 229 subjects with yaw, pitch, attribute, illumination and accessory. M2FPA is the most comprehensive multi-view face database for facial pose analysis. Further, we provide an effective benchmark for face frontalization and pose-invariant face recognition on M2FPA with several state-of-the-art methods, including DR-GAN, TP-GAN and CAPG-GAN. We believe that the new database and benchmark can significantly push forward the advance of facial pose analysis in real-world applications. Moreover, a simple yet effective parsing guided discriminator is introduced to capture the local consistency during GAN optimization. Extensive quantitative and qualitative results on M2FPA and Multi-PIE demonstrate the superiority of our face frontalization method. Baseline results for both face synthesis and face recognition from state-of-the-art methods demonstrate the challenge offered by this new database.
Peipei Li 0002, Xiang Wu 0001, Yibo Hu 0001, Ran He 0001, Zhenan Sun
ICCV1
2019 Global and Local Consistent Wavelet-Domain Age Synthesis
abstract
Age synthesis is a challenging task due to the complicated and non-linear transformation in the human aging process. Aging information is usually reflected in local facial parts, such as wrinkles at the eye corners. However, these local facial parts contribute less in previous GAN-based methods for age synthesis. To address this issue, we propose a wavelet-domain global and local consistent age generative adversarial network (WaveletGLCA-GAN), in which one global specific network and three local specific networks are integrated together to capture both global topology information and local texture details of human faces. Different from the most existing methods that modeling age synthesis in image domain, we adopt wavelet transform to depict the textual information in frequency domain. Moreover, five types of losses are adopted: 1) adversarial loss aims to generate realistic wavelets; 2) identity preserving loss aims to better preserve identity information; 3) age preserving loss aims to enhance the accuracy of age synthesis; 4) pixel-wise loss aims to preserve the background information of the input face; and 5) the total variation regularization aims to remove ghosting artifacts. Our method is evaluated on three face aging datasets, including CACD2000, Morph, and FG-NET. Qualitative and quantitative experiments show the superiority of the proposed method over other state-of-the-arts.
Peipei Li 0002, Yibo Hu 0001, Ran He 0001, Zhenan Sun
IEEE Trans. Inf. Forensics Secur.1
2018 Global and Local Consistent Age Generative Adversarial Networks
abstract
Age progression/regression is a challenging task due to the complicated and non-linear transformation in human aging process. Many researches have shown that both global and local facial features are essential for face representation [1], but previous GAN based methods mainly focused on the global feature in age synthesis. To utilize both global and local facial information, we propose a Global and Local Consistent Age Generative Adversarial Network (GLCA-GAN). In our generator, a global network learns the whole facial structure and simulates the aging trend of the whole face, while three crucial facial patches are progressed or regressed by three local networks aiming at imitating subtle changes of crucial facial subregions. To preserve most of the details in age-attribute-irrelevant areas, our generator learns the residual face. Moreover, we employ an identity preserving loss to better preserve the identity information, as well as age preserving loss to enhance the accuracy of age synthesis. A pixel loss is also adopted to preserve detailed facial information of the input face. Our proposed method is evaluated on three face aging datasets, i.e., CACD dataset, Morph dataset and FG-NET dataset. Experimental results show appealing performance of the proposed method by comparing with the state-of-the-art.
Peipei Li 0002, Yibo Hu 0001, Qi Li 0005, Ran He 0001, Zhenan Sun
ICPR1