EDBT 2026 Demo / reviewers in the wild / expert
Jianfu Zhang 0003
dblp:78/3993-3
· DBLP profile ↗
51ranked-venue papers
10as first author
38since 2021 · last 2026
0000-0002-2673-5860ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 8 first-author · 26 since 2021Artificial intelligence and machine learning · 28 · 4 first-author · 22 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | High-Quality Full-Head 3D Avatar Generation from Any Single Portrait ImageabstractIn this work, we introduce a novel high-fidelity full-head 3D avatar generation method from a single image, regardless of perspective, style, expression, or accessories. Prior works often fail to preserve consistent head geometry and facial details, primarily due to their limited capacity in modeling fine-grained facial textures and maintaining identity information. To address these challenges, we construct a new high-quality dataset containing 227 sequences of digital human portraits captured from 96 different perspectives, totalling 21,792 frames, featuring high-quality facial texture details. To further improve performance, we propose a novel multi-view diffusion named ID-TS diffusion model, which integrate identity and expression information into the two-stage multi-view diffusion process. The low-resolution stage ensures structural consistency of heads across multiple views, while the high-resolution stage preserves facial detail fidelity and coherence. Finally, we propose an enhanced feed-forward Gaussian avatar reconstruction method that optimizes the network on multi-view images of each single subject, significantly improving 3D facial texture details. Extensive experiments show that our method demonstrates robust performance across challenging scenarios, while showcasing broad applicability across numerous downstream tasks. Yujie Gao 0001, Chencheng Wang, Xianbing Sun, Jiahui Zhan, Wentao Wang 0009, Yiyi Zhang 0002, Haohua Zhao 0001, Liqing Zhang 0001, Jianfu Zhang 0003 |
AAAI | 9 |
| 2025 | WildFake: A Large-Scale and Hierarchical Dataset for AI-Generated Images DetectionabstractThe development of text-to-image generative models has enabled the creation of images so realistic that distinguishing between AI-generated images and real photos is becoming a challenge. This progress offers new possibilities but also raises concerns over privacy, authenticity, and security. Detecting AI-generated images is crucial to prevent misuse. To assess the generalizability and robustness of AI-generated image detection, we present a large-scale dataset, referred to as WildFake. This dataset features cutting-edge image generators, a wide variety of generator categories, and generators for various applications, organized in a hierarchical framework. WildFake collects fake images from the open-source community, enriching its diversity with a broad range of image classes and image styles. Its design significantly improves the effectiveness of detection algorithms, making it a valuable resource for enhancing AI-generated image detection in practical applications. Our evaluations offer insights into the performance of generative models at various levels, showcasing WildFake's unique hierarchical structure's benefits. Yan Hong 0001, Jianming Feng, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003 |
AAAI | 7 |
| 2025 | Hyperspectral Image Classification Based on Local Low-Rank RepresentationabstractHyperspectral image classification is an attractive and challenging task due to the difficulty to acquire large labeled datasets, and its susceptibility to natural environmental influences. Deep learning-based methods can effectively enhance the efficiency of hyperspectral feature learning. Tensor-based approaches can effectively help preserving the inherent structural information. This paper introduces a novel tensor-based framework for hyperspectral image classification. It leverages the ‘bot-tleneck’ structure of an AutoEncoder to extract low-dimensional representations in a self-supervised manner. Additionally, local low-rank constraints in the embedding space facilitate the distribution of features across a low-rank manifold. Experiments conducted on real hyperspectral datasets demonstrate that the proposed method yields superior classification performance. Jianting Wang, Haohua Zhao 0001, Jianfu Zhang 0003, Liqing Zhang 0001 |
CSCWD | 3 |
| 2025 | Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and ModalitiesabstractReconstructing pedestrian motion from dynamic sensors, with a focus on pedestrian intention, is crucial for advancing autonomous driving safety. However, this task is challenging due to data limitations arising from technical complexities, safety, and cost concerns. We introduce the Pedestrian Motion Reconstruction (PMR) dataset, which focuses on pedestrian intention to reconstruct behavior using multiple perspectives and modalities. PMR is developed from a mixed reality platform that combines real-world realism with the extensive, accurate labels of simulations, thereby reducing costs and risks. It captures the intricate dynamics of pedestrian interactions with objects and vehicles, using different modalities for a comprehensive understanding of human-vehicle interaction. Analyses show that PMR can naturally exhibit pedestrian intent and simulate extreme cases. PMR features a vast collection of data from 54 subjects interacting across 12 urban settings with 7 objects, encompassing 12,138 sequences with diverse weather conditions and vehicle speeds. This data provides a rich foundation for modeling pedestrian intent through multi-view and multi-modal insights. We also conduct comprehensive benchmark assessments across different modalities to thoroughly evaluate pedestrian motion reconstruction methods. Yiyi Zhang 0002, Xinhao Hu, Li Niu 0002, Jianfu Zhang 0003, Yasushi Makihara, Yasushi Yagi, Wenlong Liao, Junchi Yan, Liqing Zhang 0001 |
ICLR | 5 |
| 2025 | Towards Explainable Fake Image Detection with Multi-Modal Large Language ModelsabstractProgress in image generation raises significant public security concerns. We argue that fake image detection should not operate as a "black box". Instead, an ideal approach must ensure both strong generalization and transparency. Recent progress in Multi-modal Large Language Models (MLLMs) offers new opportunities for reasoning-based AI-generated image detection. In this work, we evaluate the capabilities of MLLMs in comparison to traditional detection methods and human evaluators, highlighting their strengths and limitations. Furthermore, we design six distinct prompts and propose a framework that integrates these prompts to develop a more robust, explainable, and reasoning-driven detection system. The code is available at https://github.com/Gennadiyev/mllm-defake. Yikun Ji, Yan Hong 0001, Jiahui Zhan, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Liqing Zhang 0001, Jianfu Zhang 0003 |
ACM Multimedia | 9 |
| 2025 | InterAnimate: Taming Region-Aware Diffusion Model for Realistic Human Interaction Animation
Yukang Lin, Yan Hong 0001, Zunnan Xu, Xindi Li, Chuanbiao Song, Ronghui Li, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Xiu Li 0001 |
ACM Multimedia | 12 |
| 2025 | GeneMAN: Generalizable Single-Image 3D Human Reconstruction from Multi-Source Human DataabstractGiven a single in-the-wild human photo, it remains a challenging task to reconstruct a high-fidelity 3D human model. Existing methods face difficulties including a) the varying body proportions captured by in-the-wild human images; b) diverse personal belongings within the shot; and c) ambiguities in human postures and inconsistency in human textures. In addition, the scarcity of high-quality human data intensifies the challenge. To address these problems, we propose a Generalizable image-to-3D huMAN reconstruction framework, dubbed GeneMAN, building upon a comprehensive multi-source collection of high-quality human data, including 3D scans, multi-view videos, single photos, and our generated synthetic human data. GeneMAN encompasses three key modules. 1) Without relying on parametric human models (e.g., SMPL), GeneMAN first trains a human-specific text-to-image diffusion model and a view-conditioned diffusion model, serving as GeneMAN 2D human prior and 3D human prior for reconstruction, respectively. 2) With the help of the pretrained human prior models, the Geometry Initialization-&-Sculpting pipeline is leveraged to recover high-quality 3D human geometry given a single image. 3) To achieve high-fidelity 3D human textures, GeneMAN employs the Multi-Space Texture Refinement pipeline, consecutively refining textures in the latent and the pixel spaces. Extensive experimental results demonstrate that GeneMAN could generate high-quality 3D human models from a single image input, outperforming prior state-of-the-art methods. Notably, GeneMAN could reveal much better generalizability in dealing with in-the-wild images, often yielding high-quality 3D human models in natural poses with common items, regardless of the body proportions in the input images. Wentao Wang 0009, Hang Ye 0002, Fangzhou Hong, Xue Yang 0005, Jianfu Zhang 0003, Yizhou Wang 0001, Ziwei Liu 0002, Liang Pan |
NeurIPS | 5 |
| 2025 | ID-MotionNet: Identity-Preserved 3D Skeleton Sequence Generation via Information Bottleneck Disentanglement
Jingyu Xie, Jianfu Zhang 0003, Liqing Zhang 0001, Yiyi Zhang 0002 |
PRCV (15) | 4 |
| 2025 | Defending adversarial attacks in Graph Neural Networks via tensor enhancement
Jianfu Zhang 0003, Yan Hong 0001, Dawei Cheng, Liqing Zhang 0001, Qibin Zhao |
Pattern Recognit. | 1 |
| 2024 | Assessing Image Inpainting via Re-Inpainting Self-Consistency EvaluationabstractImage inpainting, the task of reconstructing missing segments in corrupted images using available data, faces challenges in ensuring consistency and fidelity, especially under information-scarce conditions. Traditional evaluation methods, heavily dependent on the existence of unmasked reference images, inherently favor certain inpainting outcomes, introducing biases. Addressing this issue, we introduce an innovative evaluation paradigm that utilizes a self-supervised metric based on multiple re-inpainting passes. This approach, diverging from conventional reliance on direct comparisons in pixel or feature space with original images, emphasizes the principle of self-consistency to enable the exploration of various viable inpainting solutions, effectively reducing biases. Our extensive experiments across numerous benchmarks validate the alignment of our evaluation method with human judgment. Jianfu Zhang 0003, Yan Hong 0001, Yiyi Zhang 0002, Liqing Zhang 0001 |
CIKM | 2 |
| 2024 | ComFusion: Enhancing Personalized Generation by Instance-Scene Compositing and Fusion
Yan Hong 0001, Yuxuan Duan, Bo Zhang 0075, Haoxing Chen, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003 |
ECCV (44) | 8 |
| 2024 | Hierarchical Attacks on Large-Scale Graph Neural NetworksabstractIn this paper, we present a novel hierarchical approach to adversarial attacks targeting Graph Neural Networks (GNNs), tailored to overcome the complexities inherent in large-scale poisoning attacks. Traditional global attack strategies often fail to yield effective results on extensive graph structures. Our innovative method implements a divide-and-conquer tactic, clustering nodes based on their embeddings and forming coarse-grained graphs from these clusters. We initiate perturbations at this coarse level, gradually honing them in more detailed, finer-grained graphs, while keeping non-essential nodes grouped. By employing meta-gradients derived from these refined graphs, we pinpoint critical edges for perturbation, thereby vastly simplifying the process and reducing the intricacy involved in manipulating large-scale graphs. This hierarchical strategy not only enhances the efficacy of the attacks but also maintains operational efficiency across expansive network structures. Jianfu Zhang 0003, Yan Hong 0001, Dawei Cheng, Liqing Zhang 0001, Qibin Zhao |
ICASSP | 1 |
| 2024 | Arbitrary Style Transfer with Prototype-Based Channel AlignmentabstractStyle transfer aims to migrate the "style" from a style image to a content image. Despite the appealing results achieved by existing methods, few studies have considered the alignment of semantics or structures between the style image and the content image. To overcome this problem, we propose a novel network with two parallel branches: coarse-grained stylization branch and fine-grained decoration branch. In the stylization branch, we perform conventional AdaIN to produce globally stylized feature. In the decoration branch, we propose a ProtoType-based Channel Alignment module to align the channels between style feature and content feature, followed by Adaptive Group Transfer to produce locally stylized feature. Extensive experiments demonstrate that our proposed method outperforms state-of-the-art methods in terms of visual quality and efficiency. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003 |
ICASSP | 3 |
| 2024 | ProAug: Prototype-Based Augmentation for Long-Tailed Image ClassificationabstractReal-world data often exhibit long-tailed distributions with heavy class imbalance, which deteriorates the generalization performance of the classifier. To mitigate this problem, we propose a novel Prototype-based Augmentation framework (ProAug) to address the data scarcity issue by augmenting the feature space for tail classes. Our ProAug consists of a prototype construction branch and a dynamic augmentation branch. The prototype-based dictionary is optimized with category-aware margin loss to learn multi-center and discriminative prototypes for each category. In the dynamic augmentation branch, we aim to produce high-quality tail-class features by dynamically composing context-similar prototypes with an attention mechanism. Moreover, to further improve the reliability of prototypes and the quality of augmented features, a meta-update strategy is adopted to calibrate two branches of ProAug to boost performance. Extensive empirical results on CIFAR-LT-10/100, ImageNet-LT, and iNaturalist 2018 demonstrate the effectiveness of our method. Yan Hong 0001, Jianfu Zhang 0003, Zhongyi Sun 0002 |
ICASSP | 2 |
| 2024 | DomainGallery: Few-shot Domain-driven Image Generation by Attribute-centric FinetuningabstractThe recent progress in text-to-image models pretrained on large-scale datasets has enabled us to generate various images as long as we provide a text prompt describing what we want. Nevertheless, the availability of these models is still limited when we expect to generate images that fall into a specific domain either hard to describe or just unseen to the models. In this work, we propose DomainGallery, a few-shot domain-driven image generation method which aims at finetuning pretrained Stable Diffusion on few-shot target datasets in an attribute-centric manner. Specifically, DomainGallery features prior attribute erasure, attribute disentanglement, regularization and enhancement. These techniques are tailored to few-shot domain-driven generation in order to solve key issues that previous works have failed to settle. Extensive experiments are given to validate the superior performance of DomainGallery on a variety of domain-driven generation scenarios. Yuxuan Duan, Yan Hong 0001, Bo Zhang 0075, Jun Lan 0001, Huijia Zhu, Weiqiang Wang 0002, Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
NeurIPS | 7 |
| 2024 | An E-Commerce Dataset Revealing Variations during SalesabstractSince the development of artificial intelligence technology, E-Commerce has gradually become one of the world's largest commercial markets. Within this domain, sales events, which are based on sociological mechanisms, play a significant role. E-Commerce platforms frequently offer sales and promotions to encourage users to purchase items, leading to significant changes in live environments. Learning-To-Rank (LTR) is a crucial component of E-Commerce search and recommendations, and substantial efforts have been devoted to this area. However, existing methods often assume an independent and identically distributed data setting, which does not account for the evolving distribution of online systems beyond online finetuning strategies. This limitation can lead to inaccurate predictions of user behaviors during sales events, resulting in significant loss of revenue. In addition, models must readjust themselves once sales have concluded in order to eliminate any effects caused by the sales events, leading to further regret. To address these limitations, we introduce a long-term E-Commerce search data set specifically designed to incubate LTR algorithms during such sales events, with the objective of advancing the capabilities of E-Commerce search engines. Our investigation focuses on typical industry practices and aims to identify potential solutions to address these challenges. Jianfu Zhang 0003, Qingtao Yu, Guoliang Zhou, Yawei Sun, Guangda Huzhang, Yabo Ni, Anxiang Zeng, Han Yu 0001 |
SIGIR | 1 |
| 2024 | Multi-Model UNet: An Adversarial Defense Mechanism for Robust Visual TrackingabstractAbstract Currently, state-of-the-art object-tracking algorithms are facing a severe threat from adversarial attacks, which can significantly undermine their performance. In this research, we introduce MUNet, a novel defensive model designed for visual tracking. This model is capable of generating defensive images that can effectively counter attacks while maintaining a low computational overhead. To achieve this, we experiment with various configurations of MUNet models, finding that even a minimal three-layer setup significantly improves tracking robustness when the target tracker is under attack. Each model undergoes end-to-end training on randomly paired images, which include both clean and adversarial noise images. This training separately utilizes pixel-wise denoiser and feature-wise defender. Our proposed models significantly enhance tracking performance even when the target tracker is attacked or the target frame is clean. Additionally, MUNet can simultaneously share its parameters on both template and search regions. In experimental results, the proposed models successfully defend against top attackers on six benchmark datasets, including OTB100, LaSOT, UAV123, VOT2018, VOT2019, and GOT-10k. Performance results on all datasets show a significant improvement over all attackers, with a decline of less than 4.6% for every benchmark metric compared to the original tracker. Notably, our model demonstrates the ability to enhance tracking robustness in other blackbox trackers. Wattanapong Suttapak, Jianfu Zhang 0003, Haohua Zhao 0001, Liqing Zhang 0001 |
Neural Process. Lett. | 2 |
| 2023 | Memorization Weights for Instance Reweighting in Adversarial TrainingabstractAdversarial training is an effective way to defend deep neural networks (DNN) against adversarial examples. However, there are atypical samples that are rare and hard to learn, or even hurt DNNs' generalization performance on test data. In this paper, we propose a novel algorithm to reweight the training samples based on self-supervised techniques to mitigate the negative effects of the atypical samples. Specifically, a memory bank is built to record the popular samples as prototypes and calculate the memorization weight for each sample, evaluating the "typicalness" of a sample. All the training samples are reweigthed based on the proposed memorization weights to reduce the negative effects of atypical samples. Experimental results show the proposed method is flexible to boost state-of-the-art adversarial training methods, improving both robustness and standard accuracy of DNNs. Jianfu Zhang 0003, Yan Hong 0001, Qibin Zhao |
AAAI | 1 |
| 2023 | Amodal Instance Segmentation via Prior-Guided ExpansionabstractAmodal instance segmentation aims to infer the amodal mask, including both the visible part and occluded part of each object instance. Predicting the occluded parts is challenging. Existing methods often produce incomplete amodal boxes and amodal masks, probably due to lacking visual evidences to expand the boxes and masks. To this end, we propose a prior-guided expansion framework, which builds on a two-stage segmentation model (i.e., Mask R-CNN) and performs box-level (resp., pixel-level) expansion for amodal box (resp., mask) prediction, by retrieving regression (resp., flow) transformations from a memory bank of expansion prior. We conduct extensive experiments on KINS, D2SA, and COCOA cls datasets, which show the effectiveness of our method. Junjie Chen 0008, Li Niu 0002, Jianfu Zhang 0003, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
AAAI | 3 |
| 2023 | Critical Firms Prediction for Stemming Contagion Risk in Networked-Loans through Graph-Based Deep Reinforcement LearningabstractThe networked-loan is major financing support for Micro, Small and Medium-sized Enterprises (MSMEs) in some developing countries. But external shocks may weaken the financial networks' robustness; an accidental default may spread across the network and collapse the whole network. Thus, predicting the critical firms in networked-loans to stem contagion risk and prevent potential systemic financial crises is of crucial significance to the long-term health of inclusive finance and sustainable economic development. Existing approaches in the banking industry dismiss the contagion risk across loan networks and need extensive knowledge with sophisticated financial expertise. Regarding the issues, we propose a novel approach to predict critical firms for stemming contagion risk in the bank industry with deep reinforcement learning integrated with high-order graph message-passing networks. We demonstrate that our approach outperforms the state-of-the-art baselines significantly on the dataset from a large commercial bank. Moreover, we also conducted empirical studies on the real-world loan dataset for risk mitigation. The proposed approach enables financial regulators and risk managers to better track and understands contagion and systemic risk in networked-loans. The superior performance also represents a paradigm shift in addressing the modern challenges in financing support of MSMEs and sustainable economic development. Dawei Cheng, Zhibin Niu, Jianfu Zhang 0003, Yiyi Zhang 0002, Changjun Jiang 0002 |
AAAI | 3 |
| 2023 | Isometric Manifold Learning Using Hierarchical FlowabstractWe propose the Hierarchical Flow (HF) model constrained by isometric regularizations for manifold learning that combines manifold learning goals such as dimensionality reduction, inference, sampling, projection and density estimation into one unified framework. Our proposed HF model is regularized to not only produce embeddings preserving the geometric structure of the manifold, but also project samples onto the manifold in a manner conforming to the rigorous definition of projection. Theoretical guarantees are provided for our HF model to satisfy the two desired properties. In order to detect the real dimensionality of the manifold, we also propose a two-stage dimensionality reduction algorithm, which is a time-efficient algorithm thanks to the hierarchical architecture design of our HF model. Experimental results justify our theoretical analysis, demonstrate the superiority of our dimensionality reduction algorithm in cost of training time, and verify the effect of the aforementioned properties in improving performances on downstream tasks such as anomaly detection. Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
AAAI | 2 |
| 2023 | Taming the Power of Diffusion Models for High-Quality Virtual Try-On with Appearance FlowabstractVirtual try-on is a critical image synthesis task that aims to transfer clothes from one image to another while preserving the details of both humans and clothes. While many existing methods rely on Generative Adversarial Networks (GANs) to achieve this, flaws can still occur, particularly at high resolutions. Recently, the diffusion model has emerged as a promising alternative for generating high-quality images in various applications. However, simply using clothes as a condition for guiding the diffusion model to inpaint is insufficient to maintain the details of the clothes. To overcome this challenge, we propose an exemplar-based inpainting approach that leverages a warping module to guide the diffusion model's generation effectively. The warping module performs initial processing on the clothes, which helps to preserve the local details of the clothes. We then combine the warped clothes with clothes-agnostic person image and add noise as the input of diffusion model. Additionally, the warped clothes is used as local conditions for each denoising process to ensure that the resulting output retains as much detail as possible. Our approach, namely Diffusion-based Conditional Inpainting for Virtual Try-ON(DCI-VTON), effectively utilizes the power of the diffusion model, and the incorporation of the warping module helps to produce high-quality and realistic virtual try-on results. Experimental results on VITON-HD demonstrate the effectiveness and superiority of our method. Source code and trained models will be publicly released at: https://github.com/bcmi/DCI-VTON-Virtual-Try-On. Junhong Gou, Jianfu Zhang 0003, Jianlou Si, Chen Qian 0006, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2023 | Natural Image Matting with Attended Global Context
Yiyi Zhang 0002, Li Niu 0002, Yasushi Makihara, Jianfu Zhang 0003, Weijie Zhao 0003, Yasushi Yagi, Liqing Zhang 0001 |
J. Comput. Sci. Technol. | 4 |
| 2023 | Diverse image inpainting with disentangled uncertainty
Wentao Wang 0009, Li Niu 0002, Jianfu Zhang 0003, Haoyu Ling, Liqing Zhang 0001 |
Pattern Recognit. | 4 |
| 2022 | Shadow Generation for Composite Image in Real-World ScenesabstractImage composition targets at inserting a foreground object into a background image. Most previous image composition methods focus on adjusting the foreground to make it compatible with background while ignoring the shadow effect of foreground on the background. In this work, we focus on generating plausible shadow for the foreground object in the composite image. First, we contribute a real-world shadow generation dataset DESOBA by generating synthetic composite images based on paired real images and deshadowed images. Then, we propose a novel shadow generation network SGRNet, which consists of a shadow mask prediction stage and a shadow filling stage. In the shadow mask prediction stage, foreground and background information are thoroughly interacted to generate foreground shadow mask. In the shadow filling stage, shadow parameters are predicted to fill the shadow area. Extensive experiments on our DESOBA dataset and real composite images demonstrate the effectiveness of our proposed method. Our dataset and code are available at https://github.com/bcmi/Object-Shadow-Generation- Dataset-DESOBA. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003 |
AAAI | 3 |
| 2022 | Deep Image Harmonization by Bridging the Reality Gap
Junyan Cao, Wenyan Cong, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
BMVC | 4 |
| 2022 | Dual-path Image Inpainting with Auxiliary GAN InversionabstractDeep image inpainting can inpaint a corrupted image using a feed-forward inference, but still fails to handle large missing area or complex semantics. Recently, GAN inversion based inpainting methods propose to leverage semantic information in pretrained generator (e.g., StyleGAN) to solve the above issues. Different from feed-forward methods, they seek for a closest latent code to the corrupted image and feed it to a pretrained generator. However, inferring the latent code is either time-consuming or inaccurate. In this paper, we develop a dual-path inpainting network with inversion path and feed-forward path, in which inversion path provides auxiliary information to help feed-forward path. We also design a novel deformable fusion module to align the feature maps in two paths. Experiments on FFHQ and LSUN demonstrate that our method is effective in solving the aforementioned problems while producing more realistic results than state-of-the-art methods. Wentao Wang 0009, Li Niu 0002, Jianfu Zhang 0003, Xue Yang 0005, Liqing Zhang 0001 |
CVPR | 3 |
| 2022 | DeltaGAN: Towards Diverse Few-Shot Image Generation with Sample-Specific Delta
Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ECCV (16) | 3 |
| 2022 | SAFA: Sample-Adaptive Feature Augmentation for Long-Tailed Image Classification
Yan Hong 0001, Jianfu Zhang 0003, Zhongyi Sun 0002 |
ECCV (24) | 2 |
| 2022 | Efficient Learning-based Community-Preserving Graph GenerationabstractGraph generation is beneficial to comprehend the creation of meaningful structures of networks in a broad spec-trum of applications such as social networks and biological net-works. Recent studies tend to leverage deep learning techniques to learn the topology structures in graphs. However, we notice that the community structure, which is one of the most unique and prominent features of the graph, cannot be well captured by the existing graph generators. Moreover, the existing advanced deep learning-based graph generators are not efficient and scalable, which can only handle small graphs. In this paper, we propose a novel community-preserving generative adversarial network (CPGAN) for effective and efficient (scalable) graph simulation. We employ graph convolution networks in the encoder and share parameters in the generation process to transmit information about community structures and preserve the permutation-invariance in CPGAN. We conducted extensive experiments on benchmark datasets, including six sets of real-life graphs. The results demonstrate that CPGAN can achieve a good trade-off between efficiency (scalability) and graph simulation quality for real-life graph simulation compared with state-of-the-art baselines. Sheng Xiang 0001, Dawei Cheng, Jianfu Zhang 0003, Zhenwei Ma, Xiaoyang Wang 0002, Ying Zhang 0001 |
ICDE | 3 |
| 2022 | Few-shot Image Generation Using Discrete Content RepresentationabstractFew-shot image generation and few-shot image translation are two related tasks, both of which aim to generate new images for an unseen category with only a few images. In this work, we make the first attempt to adapt few-shot image translation method to few-shot image generation task. Few-shot image translation disentangles an image into style vector and content map. An unseen style vector can be combined with different seen content maps to produce different images. However, it needs to store seen images to provide content maps and the unseen style vector may be incompatible with seen content maps. To adapt it to few-shot image generation task, we learn a compact dictionary of local content vectors via quantizing continuous content maps into discrete content maps instead of storing seen images. Furthermore, we model the autoregressive distribution of discrete content map conditioned on style vector, which can alleviate the incompatibility between content map and style vector. Qualitative and quantitative results on three real datasets demonstrate that our model can produce images of higher diversity and fidelity for unseen categories than previous methods. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2022 | Diminishing-feature attack: The adversarial infiltration on visual tracking
Wattanapong Suttapak, Jianfu Zhang 0003, Liqing Zhang 0001 |
Neurocomputing | 2 |
| 2021 | Disentangled Information BottleneckabstractThe information bottleneck (IB) method is a technique for extracting information that is relevant for predicting the target random variable from the source random variable, which is typically implemented by optimizing the IB Lagrangian that balances the compression and prediction terms. However, the IB Lagrangian is hard to optimize, and multiple trials for tuning values of Lagrangian multiplier are required. Moreover, we show that the prediction performance strictly decreases as the compression gets stronger during optimizing the IB Lagrangian. In this paper, we implement the IB method from the perspective of supervised disentangling. Specifically, we introduce Disentangled Information Bottleneck (DisenIB) that is consistent on compressing source maximally without target prediction performance loss (maximum compression). Theoretical and experimental results demonstrate that our method is consistent on maximum compression, and performs well in terms of generalization, robustness to adversarial attack, out-of-distribution detection, and supervised disentangling. Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
AAAI | 3 |
| 2021 | Tensor Decomposition Via Core Tensor NetworksabstractTensor decomposition (TD) has shown promising performance in image completion and denoising. Existing methods always aim to decompose one tensor into latent factors or core tensors by optimizing a particular cost function based on a specific tensor model. These algorithms iteratively learn the optima from random initialization given any individual tensor, resulting in slow convergence and low efficiency. In this paper, we propose an efficient TD algorithm that aims to learn a global mapping from input tensors to latent core tensors, under the assumption that the mappings of multiple tensors might be shared or highly correlated. To this end, we train a deep neural network (DNN) to model the global mapping and then apply it to decompose a newly given tensor with high efficiency. Furthermore, the initial values of DNN are learned based on meta-learning methods. By leveraging the pretrained core tensor DNN, our proposed method enables us to perform TD efficiently and accurately. Experimental results demonstrate the significant improvements of our method over other TD methods in terms of speed and accuracy. Jianfu Zhang 0003, Zerui Tao, Liqing Zhang 0001, Qibin Zhao |
ICASSP | 1 |
| 2021 | Parallel Multi-Resolution Fusion Network for Image InpaintingabstractConventional deep image inpainting methods are based on auto-encoder architecture, in which the spatial details of images will be lost in the down-sampling process, leading to the degradation of generated results. Also, the structure information in deep layers and texture information in shallow layers of the auto-encoder architecture can not be well integrated. Differing from the conventional image inpainting architecture, we design a parallel multi-resolution inpainting network with multi-resolution partial convolution, in which low-resolution branches focus on the global structure while high-resolution branches focus on the local texture details. All these high- and low-resolution streams are in parallel and fused repeatedly with multi-resolution masked representation fusion so that the reconstructed images are semantically robust and textually plausible. Experimental results show that our method can effectively fuse structure and texture information, producing more realistic results than state-of-the-art methods. Wentao Wang 0009, Jianfu Zhang 0003, Li Niu 0002, Haoyu Ling, Xue Yang 0005, Liqing Zhang 0001 |
ICCV | 2 |
| 2021 | Bargainnet: Background-Guided Domain Translation for Image HarmonizationabstractGiven a composite image with inharmonious foreground and background, image harmonization aims to adjust the foreground to make it compatible with the background. Previous image harmonization methods mainly focus on learning the mapping from composite image to real image, while ignoring the crucial guidance role that background plays. In this work, we formulate image harmonization task as background-guided domain translation. Specifically, we use a domain code extractor to capture the background domain information to guide the foreground harmonization, which is regulated by well-tailored triplet losses. Extensive experiments on the benchmark dataset demonstrate the effectiveness of our proposed method. Code is available at https://github.com/bcmi/BargainNet. Wenyan Cong, Li Niu 0002, Jianfu Zhang 0003, Jing Liang 0007, Liqing Zhang 0001 |
ICME | 3 |
| 2021 | DANet: Deformable Alignment Network for Video Inpainting
Xutong Lu, Jianfu Zhang 0003 |
MMM (1) | 2 |
| 2021 | Person Re-Identification With Reinforced Attribute Attention SelectionabstractPerson re-identification (Re-ID) aims to match pedestrian images across various scenes in video surveillance. There are a few works using attribute information to boost Re-ID performance. Specifically, those methods leverage attribute information to boost Re-ID performance by introducing auxiliary tasks like verifying the image level attribute information of two pedestrian images or recognizing identity level attributes. Identity level attribute annotations cost less manpower and are well-fitted for person re-identification task compared with image-level attribute annotations. However, the identity attribute information may be very noisy due to incorrect attribute annotation or lack of discriminativeness to distinguish different persons, which is probably unhelpful for the Re-ID task. In this paper, we propose a novel Attribute Attentional Block (AAB), which can be integrated into any backbone network or framework. Our AAB adopts reinforcement learning to drop noisy attributes based on our designed reward and then utilizes aggregated attribute attention of the remaining attributes to facilitate the Re-ID task. Experimental results demonstrate that our proposed method achieves state-of-the-art results on three benchmark datasets. Jianfu Zhang 0003, Li Niu 0002, Liqing Zhang 0001 |
IEEE Trans. Image Process. | 1 |
| 2020 | A Proposal-Based Approach for Activity Image-to-Video RetrievalabstractActivity image-to-video retrieval task aims to retrieve videos containing the similar activity as the query image, which is a challenging task because videos generally have many background segments irrelevant to the activity. In this paper, we utilize R-C3D model to represent a video by a bag of activity proposals, which can filter out background segments to some extent. However, there are still noisy proposals in each bag. Thus, we propose an Activity Proposal-based Image-to-Video Retrieval (APIVR) approach, which incorporates multi-instance learning into cross-modal retrieval framework to address the proposal noise issue. Specifically, we propose a Graph Multi-Instance Learning (GMIL) module with graph convolutional layer, and integrate this module with classification loss, adversarial loss, and triplet loss in our cross-modal retrieval framework. Moreover, we propose geometry-aware triplet loss based on point-to-subspace distance to preserve the structural information of activity proposals. Extensive experiments on three widely-used datasets verify the effectiveness of our approach. Ruicong Xu, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
AAAI | 3 |
| 2020 | Exploiting Motion Information from Unlabeled Videos for Static Image Action RecognitionabstractStatic image action recognition, which aims to recognize action based on a single image, usually relies on expensive human labeling effort such as adequate labeled action images and large-scale labeled image dataset. In contrast, abundant unlabeled videos can be economically obtained. Therefore, several works have explored using unlabeled videos to facilitate image action recognition, which can be categorized into the following two groups: (a) enhance visual representations of action images with a designed proxy task on unlabeled videos, which falls into the scope of self-supervised learning; (b) generate auxiliary representations for action images with the generator learned from unlabeled videos. In this paper, we integrate the above two strategies in a unified framework, which consists of Visual Representation Enhancement (VRE) module and Motion Representation Augmentation (MRA) module. Specifically, the VRE module includes a proxy task which imposes pseudo motion label constraint and temporal coherence constraint on unlabeled videos, while the MRA module could predict the motion information of a static action image by exploiting unlabeled videos. We demonstrate the superiority of our framework based on four benchmark human action datasets with limited labeled data. Yiyi Zhang 0002, Li Niu 0002, Meichao Luo, Jianfu Zhang 0003, Dawei Cheng, Liqing Zhang 0001 |
AAAI | 5 |
| 2020 | DoveNet: Deep Image Harmonization via Domain VerificationabstractImage composition is an important operation in image processing, but the inconsistency between foreground and background significantly degrades the quality of composite image. Image harmonization, aiming to make the foreground compatible with the background, is a promising yet challenging task. However, the lack of high-quality publicly available dataset for image harmonization greatly hinders the development of image harmonization techniques. In this work, we contribute an image harmonization dataset iHarmony4 by generating synthesized composite images based on COCO (resp., Adobe5k, Flickr, day2night) dataset, leading to our HCOCO (resp., HAdobe5k, HFlickr, Hday2night) sub-dataset. Moreover, we propose a new deep image harmonization method DoveNet using a novel domain verification discriminator, with the insight that the foreground needs to be translated to the same domain as background. Extensive experiments on our constructed dataset demonstrate the effectiveness of our proposed method. Our dataset and code are available at https://github.com/bcmi/Image_Harmonization_Datasets. Wenyan Cong, Jianfu Zhang 0003, Li Niu 0002, Liu Liu 0022, Zhixin Ling, Weiyuan Li, Liqing Zhang 0001 |
CVPR | 2 |
| 2020 | Beyond Without Forgetting: Multi-Task Learning for Classification with Disjoint DatasetsabstractMulti-task Learning (MTL) for classification with disjoint datasets aims to explore MTL when one task only has one labeled dataset. In existing methods, for each task, the unlabeled datasets are not fully exploited to facilitate this task. Inspired by semi-supervised learning, we use unlabeled datasets with pseudo labels to facilitate each task. However, there are two major issues: 1) the pseudo labels are very noisy; 2) the unlabeled datasets and the labeled dataset for each task has considerable data distribution mismatch. To address these issues, we propose our MTL with Selective Augmentation (MTL-SA) method to select the training samples in unlabeled datasets with confident pseudo labels and close data distribution to the labeled dataset. Then, we use the selected training samples to add information and use the remaining training samples to preserve information. Extensive experiments on face-centric and human-centric applications demonstrate the effectiveness of our MTL-SA method. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ICME | 3 |
| 2020 | Matchinggan: Matching-Based Few-Shot Image GenerationabstractTo generate new images for a given category, most deep generative models require abundant training images from this category, which are often too expensive to acquire. To achieve the goal of generation based on only a few images, we propose matching-based Generative Adversarial Network (GAN) for few-shot generation, which includes a matching generator and a matching discriminator. Matching generator can match random vectors with a few conditional images from the same category and generate new images for this category based on the fused features. The matching discriminator extends conventional GAN discriminator by matching the feature of generated image with the fused feature of conditional images. Extensive experiments on three datasets demonstrate the effectiveness of our proposed method. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Liqing Zhang 0001 |
ICME | 3 |
| 2020 | Inductive Guided Filter: Real-Time Deep Matting with Weakly Annotated Masks on Mobile DevicesabstractRecently, significant progress has been achieved in deep image matting. Most of the classical image matting methods are time-consuming and require an ideal trimap which is difficult to attain in practice. An efficient image matting method based on a weakly annotated mask is in demand for mobile applications. In this paper, we propose a novel method called Inductive Guided Filter, which tackles the real-time general image matting task with weakly annotated masks on mobile devices. The Inductive Guided Filter exploits the gradient prior implicit in Guided Filter to reduce the computational burden tremendously in a deep learning manner. The use of Gabor loss is also proposed for complicated textures in image matting. Moreover, we create an image matting dataset MAT-2793 with a variety of foreground objects. Experimental results demonstrate that our proposed method massively reduces running time with robust accuracy. Yaoyi Li, Jianfu Zhang 0003, Weijie Zhao 0003, Hongtao Lu 0001 |
ICME | 2 |
| 2020 | F2GAN: Fusing-and-Filling GAN for Few-shot Image GenerationabstractIn order to generate images for a given category, existing deep generative models generally rely on abundant training images. However, extensive data acquisition is expensive and fast learning ability from limited data is necessarily required in real-world applications. Also, these existing methods are not well-suited for fast adaptation to a new category. Few-shot image generation, aiming to generate images from only a few images for a new category, has attracted some research interest. In this paper, we propose a Fusing-and-Filling Generative Adversarial Network (F2GAN) to generate realistic and diverse images for a new category with only a few images. In our F2GAN, a fusion generator is designed to fuse the high-level features of conditional images with random interpolation coefficients, and then fills in attended low-level details with non-local attention module to produce a new image. Moreover, our discriminator can ensure the diversity of generated images by a mode seeking loss and an interpolation regression loss. Extensive experiments on five datasets demonstrate the effectiveness of our proposed method for few-shot image generation. Yan Hong 0001, Li Niu 0002, Jianfu Zhang 0003, Weijie Zhao 0003, Liqing Zhang 0001 |
ACM Multimedia | 3 |
| 2020 | Image Editing via Segmentation Guided Self-Attention NetworkabstractImage editing is one of the most popular directions in computer vision. Recently, many methods have benefited from the advances in deep learning, showing promising performance in the image editing task by inpainting the editing areas. These methods take advantage of edge information as user guidance to generate the desired content. However, they are suffering from generating color discrepancy and inconsistent boundaries. In this letter, we propose a deep image editing method based on a self-attention network which copies information for each of the small patches from distant spatial locations. The proposed method smooths the image, computes segmentation maps, and utilizes the segmentation information for guiding the self-attention layers to explicitly leverage image features from surrounding areas with similar appearances. Experimental results show that the proposed method achieves better performance, is flexible for different purposes, and is fast for implementation. Jianfu Zhang 0003, Peiming Yang, Wentao Wang 0009, Yan Hong 0001, Liqing Zhang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2019 | Multi-Attribute Transfer via Disentangled RepresentationabstractRecent studies show significant progress in image-to-image translation task, especially facilitated by Generative Adversarial Networks. They can synthesize highly realistic images and alter the attribute labels for the images. However, these works employ attribute vectors to specify the target domain which diminishes image-level attribute diversity. In this paper, we propose a novel model formulating disentangled representations by projecting images to latent units, grouped feature channels of Convolutional Neural Network, to disassemble the information between different attributes. Thanks to disentangled representation, we can transfer attributes according to the attribute labels and moreover retain the diversity beyond the labels, namely, the styles inside each image. This is achieved by specifying some attributes and swapping the corresponding latent units to “swap” the attributes appearance, or applying channel-wise interpolation to blend different attributes. To verify the motivation of our proposed model, we train and evaluate our model on face dataset CelebA. Furthermore, the evaluation of another facial expression dataset RaFD demonstrates the generalizability of our proposed model. Jianfu Zhang 0003, Yaoyi Li, Weijie Zhao 0003, Liqing Zhang 0001 |
AAAI | 1 |
| 2019 | Clothes Keypoints Localization and Attribute Recognition via Prior KnowledgeabstractRich clothes datasets and high-quality annotations have driven recent advances in fashion clothes recognition. However, the existing approaches treat clothes as common images, ignoring the prior clothing knowledge such as spatial relations, symmetry, proportions, and key characteristics of clothes. In order to combine the semantic information with the advantages of deep learning, we propose Detection+, a model using the prior symmetric constraint to refine the keypoints located by any backbone detection networks. To deal with uncertainty in labelling clothing, we introduce a new loss to utilize all available data which contain "maybe" labels. Detection+ has reduced about 2.54% Normalized Error in FashionAI dataset and improved 3.2% AP in human keypoints dataset coco2017 compared to the Mask R-CNN baseline. A large number of experimental results show the proposed approach achieves better results in different recognition datasets (resp., FashionAI, and Deepfashion) with about (resp., 2.57% mAP, and 10% recall) improvements. Zhangxuan Gu, Jianfu Zhang 0003, Haohua Zhao 0001, Liqing Zhang 0001 |
ICME | 2 |
| 2019 | GAIN: Gradient Augmented Inpainting Network for Irregular HolesabstractImage inpainting, which aims to fill the missing holes of the images, is a challenging task because the holes may contain complicated structures or different possible layouts. Deep learning methods have shown promising performance in image inpainting but still, suffer from generating poor-structured artifacts when the holes are large and irregular. Some existing methods use edge inpainting to help image inpainting, with binary edge map obtained from image gradient. However, by only using the binary edge map, these methods discard the rich information in image gradient and thus leave some critical issues (e.g. , color discrepancy) unattended. In this paper, we propose Gradient Augmented Inpainting Network (GAIN), which uses image gradient information instead of edge information to facilitate image inpainting. Specifically, we formulate a multi-task learning framework which performs image inpainting and gradient inpainting simultaneously. A novel GAI-Block is designed to encourage the information fusion between the image feature map and the gradient feature map. Moreover, gradient information is also used to determine the filling priority, which can guide the network to construct more plausible semantic structures for the holes. Experimental results on public datasets CelebA-HQ and Places2 show that our proposed method outperforms state-of-the-art methods quantitatively and qualitatively. Jianfu Zhang 0003, Li Niu 0002, Dexin Yang, Liwei Kang, Yaoyi Li, Weijie Zhao 0003, Liqing Zhang 0001 |
ACM Multimedia | 1 |
| 2018 | Multi-Shot Pedestrian Re-Identification via Sequential Decision MakingabstractMulti-shot pedestrian re-identification problem is at the core of surveillance video analysis. It matches two tracks of pedestrians from different cameras. In contrary to existing works that aggregate single frames features by time series model such as recurrent neural network, in this paper, we propose an interpretable reinforcement learning based approach to this problem. Particularly, we train an agent to verify a pair of images at each time. The agent could choose to output the result (same or different) or request another pair of images to verify (unsure). By this way, our model implicitly learns the difficulty of image pairs, and postpone the decision when the model does not accumulate enough evidence. Moreover, by adjusting the reward for unsure action, we can easily trade off between speed and accuracy. In three open benchmarks, our method are competitive with the state-of-the-art methods while only using 3% to 6% images. These promising results demonstrate that our method is favorable in both efficiency and performance. Jianfu Zhang 0003, Naiyan Wang, Liqing Zhang 0001 |
CVPR | 1 |
| 2018 | Attention-Based Network for Cross-View Gait Recognition
Jianfu Zhang 0003, Haohua Zhao 0001, Liqing Zhang 0001 |
ICONIP (7) | 2 |