Fan Tang

dblp:46/6804 · DBLP profile ↗
← Back
78ranked-venue papers
4as first author
69since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 63 · 2 first-author · 55 since 2021Artificial intelligence and machine learning · 30 · 1 first-author · 29 since 2021Systems, architecture and hardware · 5 · 2 first-author · 4 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2026 Visual-Friendly Concept Protection via Selective Adversarial Perturbations
abstract
Personalized concept generation by tuning diffusion models with a few images raises potential legal and ethical concerns regarding privacy and intellectual property rights. Researchers attempt to prevent malicious personalization using adversarial perturbations. However, previous efforts have mainly focused on the effectiveness of protection while neglecting the visibility of perturbations. They utilize global adversarial perturbations, which introduce noticeable alterations to original images and significantly degrade visual quality. In this work, we propose the Visual-Friendly Concept Protection (VCPro) framework, which prioritizes the protection of key concepts chosen by the image owner through adversarial perturbations with lower perceptibility. To ensure these perturbations are as inconspicuous as possible, we introduce a relaxed optimization objective to identify the least perceptible yet effective adversarial perturbations, solved using the Lagrangian multiplier method. Qualitative and quantitative experiments validate that VCPro achieves a better trade-off between the visibility of perturbations and protection effectiveness, effectively prioritizing the protection of target concepts in images with less perceptible perturbations.
Xiaoyue Mi, Fan Tang, Juan Cao 0001, Peng Li 0030, Yang Liu 0005
AAAI2
2026 Task-oriented medical image super-resolution via target prior guidance
Yuxiang Meng, Dengwen Zhou, Shaoxin Li 0004, Feiyue Huang, Xingkun Xu, Lifeng Zhu, Yuchen Xu 0008, Fan Tang
Knowl. Based Syst.8
2026 Boosting cross-domain semi-supervised medical image segmentation with internal and external regularizations
abstract
Cross-domain medical image segmentation has been challenging due to the extraordinary cost of collecting sufficient data in various imaging conditions. Mainstream methodologies enhance generalizability by directly adapting large-scale vision foundation models for medical image segmentation tasks. However, these approaches frequently incur high memory costs and necessitate additional prompts for deployment. In this study, we utilize dark knowledge in pretrained segmentation models as external regularization to improve the model’s generalizability. Furthermore, an activation-restricted regularization item is proposed to eliminate the noise/errors within the generated pseudo labels. Experiments on medical cross-domain datasets demonstrate the SOTA performance (above 2.66% improvement) of the proposed method with no additional cost during inference.
Rui Wang 0177, Fan Tang, Feiyue Huang, Shaoxin Li 0004, Xinkun Xu, Yuchen Xu 0008, Lifeng Zhu, Weiming Dong
Pattern Recognit.2
2026 MoAnimate: Bridging the Motion-Oriented Latent Representation Gaps in Human Video Animation
abstract
Human animation strives to bring static characters to life. Existing methods produce high-quality outcomes for single-frame animation; however, they often fail to maintain satisfactory temporal consistency, especially in facial and hand movements. This limitation arises from commonly used motion modules that do not explicitly model inter-entity relationships. In this work, we introduce MoAnimate, a Motion-oriented Human Animation framework designed to improve inter-entity consistency. Specifically, we extract motion flows from driving videos and transfer them to align the shape of character. During initialization, we propose a motion-oriented latent refinement that optimizes low-frequency subbands to regulate the layout of visual objects along flow trajectories, while preserving random high-frequency subbands to accommodate appearance variations. During denoising, we further introduce a motion-oriented entity attention module to enable direct and efficient interaction among entities within a coordinated subspace. Extensive experiments demonstrate that our method significantly enhances temporal consistency, particularly the visual consistency of the entities.
Haipeng Fang, Sheng Tang, Ziyao Huang 0002, Juan Cao 0001, Fan Tang, Yongdong Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.6
2026 HeadRouter: A Training-free Image Editing Framework for MM-DiTs by Adaptively Routing Attention Heads
abstract
Diffusion Transformers (DiTs) have exhibited robust capabilities in image generation tasks. However, accurate text-guided image editing for multimodal DiTs (MM-DiTs) still poses a significant challenge. Unlike UNet-based structures that could utilize self/cross-attention maps for semantic editing, MM-DiTs inherently lack support for explicit and consistent incorporated text guidance, resulting in semantic misalignment between the edited results and texts. In this study, we disclose the sensitivity of different attention heads to different image semantics within MM-DiTs and introduce HeadRouter , a training-free image editing framework that edits the source image by adaptively routing the text guidance to different attention heads in MM-DiTs. Furthermore, we propose a dual-token refinement module to refine text/image token representations for precise semantic guidance and accurate region expression. Experiments on multiple benchmarks demonstrate HeadRouter’s performance in terms of editing fidelity and image quality. The code is available at https://github.com/ICTMCG/HeadRouter .
Fan Tang, Juan Cao 0001, Xiaoyu Kong, Yuxin Zhang 0006, Jintao Li 0001, Oliver Deussen, Tong-Yee Lee
ACM Trans. Graph.2
2026 Make-Your-Anchor+: Temporal Consistent 2D Avatar Generation via Video Diffusion Prior
abstract
Despite the remarkable process of talking-head-based avatar-creating solutions, directly generating anchor-style videos with full-body motions remains challenging. In this study, we propose Make-Your-Anchor+, a novel system necessitating only a one-minute video clip of an individual for training, subsequently enabling the automatic generation of anchor-style videos with precise torso and hand movements. Specifically, we finetune a proposed structure-guided diffusion model on input video to render 3D mesh conditions into human appearances. We adopt a two-stage training strategy for the diffusion model, effectively mapping movements with specific appearances to create digital avatars for online streamers, live shopping hosts, and other applications. To produce arbitrary long temporal video, we extract human motion information from video diffusion prior by adapting the frame-wise diffusion model to pretrained video diffusion weights with lower cost, and a simple yet effective batch-overlapped temporal denoising module is proposed to bypass the constraints on video length during inference. Finally, a novel identity-specific face enhancement module is introduced to improve the visual quality of facial regions in the output videos. Comparative experiments demonstrate the system's effectiveness and superiority in visual quality, temporal coherence, and identity preservation, outperforming SOTA diffusion/non-diffusion methods.
Ziyao Huang 0002, Fan Tang, Juan Cao 0001, Yong Zhang 0034, Xiaodong Cun, Yihang Bo, Jintao Li 0001, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2026 Interactive Visual Assessment for Text-to-Image Generation Models
abstract
Visual generation models have achieved remarkable progress in computer graphics applications but still face significant challenges in real-world deployment. Current assessment approaches for visual generation tasks typically follow an isolated three-phase framework: test input collection, model output generation, and user assessment. These fashions suffer from fixed coverage, evolving difficulty, and data leakage risks, limiting their effectiveness in comprehensively evaluating increasingly complex generation models. To address these limitations, we propose DyEval, an LLM-powered dynamic interactive visual assessment framework that facilitates collaborative evaluation between humans and generative models for text-to-image systems. DyEval features an intuitive visual interface that enables users to interactively explore and analyze model behaviors, while adaptively generating hierarchical, fine-grained, and diverse textual inputs to continuously probe the capability boundaries of the models based on their feedback. Additionally, to provide interpretable analysis for users to further improve tested models, we develop a contextual reflection module that mines failure triggers of test inputs and reflects model potential failure patterns, supporting in-depth analysis using the logical reasoning ability of LLM. Qualitative and quantitative experiments demonstrate that DyEval can effectively help users identify max up to 2.56 timesmore generation failures than conventional methods, and uncover complex and rare failure patterns, such as issues with pronoun generation and specific cultural context generation. Our framework provides valuable insights for improving generative models and has broad implications for advancing the reliability and capabilities of visual generation systems across various domains.
Xiaoyue Mi, Fan Tang, Juan Cao 0001, Qiang Sheng 0001, Ziyao Huang 0002, Peng Li 0030, Yang Liu 0005, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2026 AnchorCrafter: Animate Cyber-Anchors Selling Your Products via Human-Object Interacting Video Generation
abstract
The generation of anchor-style product promotion videos presents promising opportunities in e-commerce, advertising, and consumer engagement. Despite advancements in pose-guided human video generation, creating product promotion videos remains challenging. In addressing this challenge, we identify the integration of human-object interactions (HOI) into pose-guided human video generation as a core issue. To this end, we introduce AnchorCrafter, a novel diffusion-based system designed to generate 2D videos featuring a target human and a customized object, achieving high visual fidelity and controllable interactions. Specifically, we propose two key innovations: the HOI-appearance perception, which enhances object appearance recognition from arbitrary multi-view perspectives and disentangles object and human appearance, and the HOI-motion injection, which enables complex human-object interactions by overcoming challenges in object trajectory conditioning and inter-occlusion management. Extensive experiments show that our system improves object appearance preservation by 7.5%, and achieves the best video quality compared to existing state-of-the-art approaches. It also outperforms existing approaches in maintaining human motion consistency and high-quality video generation.
Ziyao Huang 0002, Juan Cao 0001, Yong Zhang 0034, Xiaodong Cun, Qing Shuai, Linchao Bao, Fan Tang
IEEE Trans. Vis. Comput. Graph.9
2025 Z-Magic: Zero-shot Multiple Attributes Guided Image Creator
abstract
The customization of multiple attributes has gained popularity with the rising demand for personalized content creation. Despite promising empirical results, the contextual coherence between different attributes has been largely overlooked. In this paper, we argue that subsequent attributes should follow the multivariable conditional distribution introduced by former attribute creation. In light of this, we reformulate multi-attribute creation from a conditional probability theory perspective and tackle the challenging zero-shot setting. By explicitly modeling the dependencies between attributes, we further enhance the coherence of generated images across diverse attribute combinations. Furthermore, we identify connections between multi-attribute customization and multi-task learning, effectively addressing the high computing cost encountered in multi-attribute synthesis. Extensive experiments demonstrate that Z-Magic outperforms existing models in zero-shot image generation, with broad implications for AI-driven design and creative applications.
Yingying Deng, Fan Tang, Weiming Dong
CVPR3
2025 Attend to Not Attended: Structure-then-Detail Token Merging for Post-training DiT Acceleration
abstract
Diffusion transformers have shown exceptional performance in visual generation but incur high computational costs. Token reduction techniques that compress models by sharing the denoising process among similar tokens have been introduced. However, existing approaches neglect the denoising priors of the diffusion models, leading to suboptimal acceleration and diminished image quality. This study proposes a novel concept: attend to prune feature redundancies in areas not attended by the diffusion process. We analyze the location and degree of feature redundancies based on the structure-then-detail denoising priors. Subsequently, we introduce SDTM, a structure-then-detail token merging approach that dynamically compresses feature redundancies. Specifically, we design dynamic visual token merging, compression ratio adjusting, and prompt reweighting for different stages. Served in a post-training way, the proposed method can be integrated seamlessly into any DiT architecture. Extensive experiments across various backbones, schedulers, and datasets showcase the superiority of our method, for example, it achieves 1.55× acceleration with negligible impact on image quality. Project page: https://github.com/ICTMCG/SDTM.
Haipeng Fang, Sheng Tang, Juan Cao 0001, Enshuo Zhang, Fan Tang, Tong-Yee Lee
CVPR5
2025 Beyond Words: Augmenting Discriminative Richness via Diffusions in Unsupervised Prompt Learning
abstract
Fine-tuning vision-language models (VLMs) with large amounts of unlabeled data has recently garnered significant interest. However, a key challenge remains the lack of high-quality pseudo-labeled data. Current pseudo-labeling strategies often struggle with mismatches between semantic and visual information, leading to sub-optimal performance of unsupervised prompt learning (UPL) methods. In this paper, we introduce a simple yet effective approach called Augmenting Discriminative Richness via Diffusions (AiR), toward learning a richer discriminating way to represent the class comprehensively and thus facilitate classification. Specifically, our approach includes a pseudo-label generation module that leverages high-fidelity synthetic samples to create an auxiliary classifier, which captures richer visual variation, bridging text-image-pair classification to a more robust image-image-pair classification. Additionally, we exploit the diversity of diffusion-based synthetic samples to enhance prompt learning, providing greater information for semantic-visual alignment. Extensive experiments on five public benchmarks, including RESISC45 and Flowers102, and across three learning paradigms-UL, SSL, and TRZSL-demonstrate that AiR achieves substantial and consistent performance improvements over state-of-the-art unsupervised prompt learning methods. Code is available.
Hairui Ren, Fan Tang, He Zhao 0001, Dandan Guo, Yi Chang 0001
CVPR2
2025 Adversarial Robust Memory-Based Continual Learner
abstract
Despite the remarkable advances that have been made in continual learning, the adversarial vulnerability of such methods has not been fully discussed. We delve into the adversarial robustness of memory-based continual learning algorithms and observe limited robustness improvement by directly applying adversarial training techniques. Preliminary studies reveal the twin challenges for building adversarial robust continual learners: accelerated forgetting in continual learning and gradient obfuscation in adversarial robustness. In this study, we put forward a novel adversarial robust memory-based continual learner that adjusts data logits to mitigate the forgetting of pasts caused by adversarial samples. Furthermore, we devise a gradient-based data selection mechanism to overcome the gradient obfuscation caused by limited stored data. The proposed approach can widely integrate with existing memory-based continual learning as well as adversarial training algorithms in a plug-and-play way. Extensive experiments on Split-CIFAR10/100 and Split-Tiny-ImageNet demonstrate the effectiveness of our approach, achieving up to 8.13% higher accuracy for adversarial data.
Xiaoyue Mi, Fan Tang, Zonghan Yang, Danding Wang, Juan Cao 0001, Peng Li 0030, Yang Liu 0005
ICCV2
2025 AMDANet: Attention-Driven Multi-Perspective Discrepancy Alignment for RGB-Infrared Image Fusion and Segmentation
abstract
The challenge of multimodal semantic segmentation lies in establishing semantically consistent and segmentable multimodal fusion features under conditions of significant visual feature discrepancies. Existing methods commonly construct cross-modal self-attention fusion frameworks or introduce additional multimodal fusion loss functions to establish fusion features. However, these approaches often overlook the challenge caused by feature discrepancies between modalities during the fusion process. To achieve precise segmentation, we propose an Attention-Driven Multimodal Discrepancy Alignment Network (AMDANet). AMDANet reallocates weights to reduce the saliency of discrepant features and utilizes low-weight features as cues to mitigate discrepancies between modalities, thereby achieving multimodal feature alignment. Furthermore, to simplify the feature alignment process, a semantic consistency inference mechanism is introduced to reveal the network’s inherent bias toward specific modalities, thereby compressing cross-modal feature discrepancies from the foundational level. Extensive experiments on the FMB, MFNet, and PST900 datasets demonstrate that AMDANet achieves mIoU improvements of 3.6%, 3.0%, and 1.6%, respectively, significantly outperforming state-of-the-art methods. The code is available at https://github.com/Zhonghaifeng6/AMDANet
Haifeng Zhong, Fan Tang, Zhuo Chen 0028, Hyung Jin Chang, Yixing Gao 0001
ICCV2
2025 Multi-Turn Consistent Image Editing
abstract
Many real-world applications, such as interactive photo retouching, artistic content creation, and product design, require flexible and iterative image editing. However, existing image editing methods primarily focus on achieving the desired modifications in a single step, which often struggles with ambiguous user intent, complex transformations, or the need for progressive refinements. As a result, these methods frequently produce inconsistent outcomes or fail to meet user expectations. To address these challenges, we propose a multi-turn image editing framework that enables users to iteratively refine their edits, progressively achieving more satisfactory results. Our approach leverages flow matching for accurate image inversion and a dual-objective Linear Quadratic Regulators (LQR) for stable sampling, effectively mitigating error accumulation. Additionally, by analyzing the layer-wise roles of transformers, we introduce a adaptive attention highlighting method that enhances editability while preserving multi-turn coherence. Extensive experiments demonstrate that our framework significantly improves edit success rates and visual fidelity compared to existing methods.
Zijun Zhou, Yingying Deng, Weiming Dong, Fan Tang
ICCV5
2025 FireFlow: Fast Inversion of Rectified Flow for Image Semantic Editing
abstract
Though Rectified Flows (ReFlows) with distillation offer a promising way for fast sampling, its fast inversion transforms images back to structured noise for recovery and following editing remains unsolved. This paper introduces FireFlow, an embarrassingly simple yet effective zero-shot approach that inherits the startling capacity of ReFlow-based models (such as FLUX) in generation while extending its capabilities to accurate inversion and editing in 8 steps. We first demonstrate that a carefully designed numerical solver is pivotal for ReFlow inversion, enabling accurate inversion and reconstruction with the precision of a second-order solver while maintaining the practical efficiency of a first-order Euler method. This solver achieves a $3\times$ runtime speedup compared to state-of-the-art ReFlow inversion and editing techniques while delivering smaller reconstruction errors and superior editing results in a training-free mode. The code is available at this-URL.
Yingying Deng, Changwang Mei, Fan Tang
ICML5
2025 DarkSeg: Infrared-Driven Semantic Segmentation for Garment Grasping Detection in Low-Light Conditions
abstract
Garment grasping in low-light environments is a critical challenge for domestic intelligent robots, yet existing research has not sufficiently addressed this issue. In low-light conditions, the scarcity of visual features due to insufficient illumination causes different categories of garments to exhibit ambiguous feature similarities, thereby hindering the robot’s ability to detect the categories of different garments. Although traditional methods can compensate for visual deficiencies in low-light scenarios by applying preprocessing strategies that fuse infrared multimodal features, their complex computational processes incur significant computational overhead. To address this limitation, we propose a low-light garment detection model based on the student-teacher model. The innovation of DarkSeg lies in its replacement of complex multimodal feature fusion with an indirect feature alignment mechanism between the student and teacher models, thereby circumventing high computational demands. Through feature alignment, DarkSeg enables the student model to learn illumination-invariant structural representations from the infrared features provided by the teacher model, effectively correcting structural deficiencies in low-light environments. Furthermore, to evaluate DarkSeg’s feasibility for low-light clothing grasping, we propose a depth-perceptive grasping strategy and build a low-light multimodal garment detection dataset, DarkClothes. Extensive experiments deploying DarkSeg on a Baxter robot demonstrate that DarkSeg achieves a 22% improvement in the grasping success rate while reducing the model parameters by 99.08 million compared to traditional methods, validating the practical viability of DarkSeg for robotic garment grasping in low-light conditions. The code and dataset are available at https://github.com/Zhonghaifeng6/Darkseg
Haifeng Zhong, Fan Tang, Hyung Jin Chang, Xingyu Zhu 0014, Yixing Gao 0001
IROS2
2025 HOMA: Towards Generic Human-Object Interaction in Multimodal Driven Human Animation with Weak Conditions
abstract
While recent advances in human-object interaction (HOI) video generation showcase promising capabilities for synthesizing coordinated human-object dynamics, existing methods remain constrained by their reliance on meticulously curated motion sequences and actor-specific data, thereby limiting practical scalability and user accessibility. Furthermore, generalization to novel object appearances and interaction scenarios remains understudied. To address these limitations, we propose HOMA, a weakly conditioned multimodal-driven HOI video generation framework that introduces sparse, decoupled motion guidance to enhance controllability and reduce dependency on stringent input conditions. Our approach encodes appearance and motion signals into the dual input space of a multimodal diffusion transformer (MMDiT), fusing them within a shared context space to enable temporally consistent and physically plausible interactions. To optimize learning efficiency and feature injection accuracy, we introduce a parameter-space HOI adapter initialized with pretrained MMDiT weights to preserve prior knowledge while enabling efficient adaptation. Additionally, we design a facial cross-attention adapter for audio-driven lip synchronization, ensuring anatomically accurate speech animation. Extensive experiments demonstrate that HOMA achieves state-of-the-art performance in interaction naturalness and generalization under weak supervision, outperforming existing methods by significant margins. We further illustrate HOMA’s versatility through diverse applications, including text-conditioned generation and interactive object manipulation, facilitated by a user-friendly demo interface. The project page is https://bone-11.github.io/homa-page/.
Ziyao Huang 0002, Juan Cao 0001, Yifeng Ma 0006, Zejing Rao, Qin Lin 0003, Qinglin Lu, Fan Tang
SIGGRAPH Asia12
2025 In-Context Brush: Zero-shot Customized Subject Insertion with Context-Aware Latent Space Manipulation
abstract
Recent advances in diffusion models have enhanced multimodal-guided visual generation, enabling customized subject insertion that seamlessly “brushes” user-specified objects into a given image guided by textual prompts. However, existing methods often struggle to insert customized subjects with high fidelity and align results with the user’s intent through textual prompts. In this work, we propose In-Context Brush, a zero-shot framework for customized subject insertion by reformulating the task within the paradigm of in-context learning. Without loss of generality, we formulate the object image and the textual prompts as cross-modal demonstrations, and the target image with the masked region as the query. The goal is to inpaint the target image with the subject aligning textual prompts without model tuning. Building upon a pretrained MMDiT-based inpainting network, we perform test-time enhancement via dual-level latent space manipulation: intra-head latent feature shifting within each attention head that dynamically shifts attention outputs to reflect the desired subject semantics and inter-head attention reweighting across different heads that amplifies prompt controllability through differential attention prioritization. Extensive experiments and applications demonstrate that our approach achieves superior identity preservation, text alignment, and image quality compared to existing state-of-the-art methods, without requiring dedicated training or additional data collection. Project page: https://yuci-gpt.github.io/In-Context-Brush/.
Fan Tang, Lin Gao 0004, Oliver Deussen, Hongbin Yan, Jintao Li 0001, Juan Cao 0001, Tong-Yee Lee
SIGGRAPH Asia2
2025 Modality-Consistent Prompt Tuning With Optimal Transport
abstract
Prompt tuning has been successfully used in leveraging the knowledge of Large-scale Vision-Language Pre-trained (VLP) models on downstream tasks. Most existing prompt tuning approaches learn prompts by maximizing the pairwise similarity. Although samples in different modalities might be relatively aligned pairwisely, such alignment does not fully utilize the information between samples, which can be less consistent on the modality level. In this paper, we propose a novel prompt tuning strategy by distributionally matching different modalities. Specifically, we minimize the distribution-wise distance between the image and text modalities with optimal transport (OT) theory. Simultaneously, we add a constraint on the learned transport plan during the modality matching to enhance the learning of vision and text prompts. Our proposed one can be applied to improve existing uni-modal and multi-modal prompt learning methods for being a plug-and-play method, which can generate modality-consistent representations. Experiments on eleven public datasets demonstrate that our proposed method has excellent performance, achieving substantial improvements on both uni-modal and multi-modal prompt tuning methods.
Hairui Ren, Fan Tang, Huangjie Zheng, He Zhao 0001, Dandan Guo, Yi Chang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2025 DiffStyler: Controllable Dual Diffusion for Text-Driven Image Stylization
abstract
Despite the impressive results of arbitrary image-guided style transfer methods, text-driven image stylization has recently been proposed for transferring a natural image into a stylized one according to textual descriptions of the target style provided by the user. Unlike the previous image-to-image transfer approaches, text-guided stylization progress provides users with a more precise and intuitive way to express the desired style. However, the huge discrepancy between cross-modal inputs/outputs makes it challenging to conduct text-driven image stylization in a typical feed-forward CNN pipeline. In this article, we present DiffStyler, a dual diffusion processing architecture to control the balance between the content and style of the diffused results. The cross-modal style information can be easily integrated as guidance during the diffusion process step-by-step. Furthermore, we propose a content image-based learnable noise on which the reverse denoising process is based, enabling the stylization results to better preserve the structure information of the content image. We validate the proposed DiffStyler beyond the baseline methods through extensive qualitative and quantitative experiments. The code is available at https://github.com/haha-lisa/Diffstyler.
Nisha Huang, Yuxin Zhang 0006, Fan Tang, Chongyang Ma, Weiming Dong, Changsheng Xu
IEEE Trans. Neural Networks Learn. Syst.3
2025 B4M: Breaking Low-Rank Adapter for Making Content-Style Customization
abstract
Personalized generation paradigms empower designers to customize visual intellectual property with the help of textual descriptions by adapting pre-trained text-to-image models on a few images. Recent studies focus on simultaneously customizing content and detailed visual style in images but often struggle with entangling the two. In this study, we reconsider the customization of content and style concepts from the perspective of parameter space construction. Unlike existing methods that utilize a shared parameter space for content and style learning, we propose a novel framework that separates the parameter space to facilitate individual learning of content and style by introducing “partly learnable projection” (PLP) matrices to separate the original adapters into divided sub-parameter spaces. A “ break-for-make ” customization learning pipeline based on PLP is proposed: we first break the original adapters into “up projection” and “down projection” for content and style concept under orthogonal prior and then make the entity parameter space by reconstructing the content and style PLP matrices by using Riemannian preconditioning to adaptively balance content and style learning. Experiments on various styles, including textures, materials, and artistic style, show that our method outperforms state-of-the-art single/multiple concept learning pipelines regarding content-style-prompt alignment. Code is available at https://github.com/ICTMCG/Break-for-make .
Fan Tang, Juan Cao 0001, Yuxin Zhang 0006, Oliver Deussen, Weiming Dong, Jintao Li 0001, Tong-Yee Lee
ACM Trans. Graph.2
2025 CreativeSynth: Cross-Art-Attention for Artistic Image Synthesis With Multimodal Diffusion
abstract
Although remarkable progress has been made in image style transfer, style is just one of the components of artistic paintings. Directly transferring extracted style features to natural images often results in outputs with obvious synthetic traces. This is because key painting attributes including layout, perspective, shape, and semantics often cannot be conveyed and expressed through style transfer. Large-scale pretrained text-to-image generation models have demonstrated their capability to synthesize a vast amount of high-quality images. However, even with extensive textual descriptions, it is challenging to fully express the unique visual properties and details of paintings. Moreover, generic models often disrupt the overall artistic effect when modifying specific areas, making it more complicated to achieve a unified aesthetic in artworks. Our main novel idea is to integrate multimodal semantic information as a synthesis guide into artworks, rather than transferring style to the real world. We also aim to reduce the disruption to the harmony of artworks while simplifying the guidance conditions. Specifically, we propose an innovative multi-task unified framework called CreativeSynth, based on the diffusion model with the ability to coordinate multimodal inputs. CreativeSynth combines multimodal features with customized attention mechanisms to seamlessly integrate real-world semantic content into the art domain through Cross-Art-Attention for aesthetic maintenance and semantic fusion. We demonstrate the results of our method across a wide range of different art categories, proving that CreativeSynth bridges the gap between generative models and artistic expression.
Nisha Huang, Weiming Dong, Yuxin Zhang 0006, Fan Tang, Ronghui Li, Chongyang Ma, Xiu Li 0001, Tong-Yee Lee, Changsheng Xu
IEEE Trans. Vis. Comput. Graph.4
2025 MotionCrafter: Plug-and-Play Motion Guidance for Diffusion Models
abstract
The essence of a video lies in the dynamic motions. While text-to-video generative diffusion models have made significant strides in creating diverse content, effectively controlling specific motions through text prompts remains a challenge. By utilizing user-specified reference videos, the more precise guidance for character actions, object movements, and camera movements can be achieved. This gives rise to the task of motion customization, where the primary challenge lies in effectively decoupling the appearance and motion within a video clip. To address this challenge, we introduce MotionCrafter, a novel one-shot instance-guided motion customization method that is suitable for both pre-trained text-to-video and text-to-image diffusion models. MotionCrafter employs a parallel spatial-temporal architecture that integrates the reference motion into the temporal component of the base model, while independently adjusting the spatial module for character or style control. To enhance the disentanglement of motion and appearance, we propose an innovative dual-branch motion disentanglement approach, which includes a motion disentanglement loss and an appearance prior enhancement strategy. To facilitate more efficient learning of motions, we further propose a novel timestep-layered tuning strategy that directs the diffusion model to focus on motion-level information. Through comprehensive quantitative and qualitative experiments, along with user preference tests, we demonstrate that MotionCrafter can successfully integrate dynamic motions while maintaining the coherence and quality of the base model, providing a wide range of appearance generation capabilities. MotionCrafter can be applied to various personalized backbones in the community to generate videos with a variety of artistic styles.
Yuxin Zhang 0006, Weiming Dong, Fan Tang, Nisha Huang, Chongyang Ma, Pengfei Wan 0001, Tong-Yee Lee, Changsheng Xu
IEEE Trans. Vis. Comput. Graph.3
2025 A Comprehensive Evaluation of Arbitrary Image Style Transfer Methods
abstract
Despite the remarkable process in the field of arbitrary image style transfer (AST), inconsistent evaluation continues to plague style transfer research. Existing methods often suffer from limited objective evaluation and inconsistent subjective feedback, hindering reliable comparisons among AST variants. In this study, we propose a multi-granularity assessment system that combines standardized objective and subjective evaluations. We collect a fine-grained dataset considering a range of image contexts such as different scenes, object complexities, and rich parsing information from multiple sources. Objective and subjective studies are conducted using the collected dataset. Specifically, we innovate on traditional subjective studies by developing an online evaluation system utilizing a combination of point-wise, pair-wise, and group-wise questionnaires. Finally, we bridge the gap between objective and subjective evaluations by examining the consistency between the results from the two studies. We experimentally evaluate CNN-based, flow-based, transformer-based, and diffusion-based AST methods by the proposed multi-granularity assessment system, which lays the foundation for a reliable and robust evaluation. Providing standardized measures, objective data, and detailed subjective feedback empowers researchers to make informed comparisons and drive innovation in this rapidly evolving field.
Zijun Zhou, Fan Tang, Yuxin Zhang 0006, Oliver Deussen, Juan Cao 0001, Weiming Dong, Xiangtao Li, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.2
2025 Attribute guided adversarial editing for face privacy protection
abstract
Nowadays, the proliferation of portraits or photographs containing human faces on the internet has created significant risks of illegal privacy collection and analysis by intelligent systems. Previous attempts to protect against unauthorized identification by face recognition models have primarily involved manipulating or adding adversarial perturbations to photos. However, it remains a challenge to balance privacy protection effectiveness and maintaining image visual quality. That is, to successfully attack real-world black-box face recognition models, significant manipulation is required for the source image, which will obviously damage the image visual quality. To address these issues, we propose an attribute-guided face identity protection (AG-FIP) approach that can protect facial privacy effectively without introducing meaningless or conspicuous artifacts into the source image. The proposed method involves mapping the images to latent space and subsequently implementing an adversarial attack through attribute editing. An attribute selection module followed by an attribute adversarially editing module is proposed to enhance the efficiency and effectiveness of adversarial attacks. Experimental results demonstrate that our approach outperforms SOTAs in terms of confusing black-box face recognition models, commercial face recognition APIs, and image visual quality.
Ziang Wang 0004, Fan Tang, Juan Cao 0001, Xirong Li 0001, Jintao Li 0001
Vis. Informatics3
2024 Music Style Transfer with Time-Varying Inversion of Diffusion Models
abstract
With the development of diffusion models, text-guided image style transfer has demonstrated great controllable and high-quality results. However, the utilization of text for diverse music style transfer poses significant challenges, primarily due to the limited availability of matched audio-text datasets. Music, being an abstract and complex art form, exhibits variations and intricacies even within the same genre, thereby making accurate textual descriptions challenging. This paper presents a music style transfer approach that effectively captures musical attributes using minimal data. We introduce a novel time-varying textual inversion module to precisely capture mel-spectrogram features at different levels. During inference, we utilize a bias-reduced stylization technique to get stable results. Experimental results demonstrate that our method can transfer the style of specific instruments, as well as incorporate natural sounds to compose melodies. Samples and code are available at https://lsfhuihuiff.github.io/MusicTI/.
Sifei Li, Yuxin Zhang 0006, Fan Tang, Chongyang Ma, Weiming Dong, Changsheng Xu
AAAI3
2024 Topology-preserving Adversarial Training for Alleviating Natural Accuracy Degradation
Xiaoyue Mi, Fan Tang, Yepeng Weng, Danding Wang, Juan Cao 0001, Sheng Tang, Peng Li 0030, Yang Liu 0005
BMVC2
2024 Z*: Zero-shot Style Transfer via Attention Reweighting
abstract
Despite the remarkable progress in image style transfer, formulating style in the context of art is inherently subjective and challenging. In contrast to existing methods, this study shows that vanilla diffusion models can directly extract style information and seamlessly integrate the generative prior into the content image without retraining. Specifically, we adopt dual denoising paths to represent content/style references in latent space and then guide the content image denoising process with style latent codes. We further reveal that the cross-attention mechanism in latent diffusion models tends to blend the content and style images, resulting in stylized outputs that deviate from the original content image. To overcome this limitation, we introduce a cross-attention reweighting strategy. Through theoretical analysis and experiments, we demonstrate the effectiveness and superiority of the diffusion-based zero-shot §_tyle transfer via attention reweighting, Z -STAR.
Yingying Deng, Fan Tang, Weiming Dong
CVPR3
2024 Make-Your-Anchor: A Diffusion-based 2D Avatar Generation Framework
abstract
Despite the remarkable process of talking-head-based avatar-creating solutions, directly generating anchor-style videos with full-body motions remains challenging. In this study, we propose Make-Your-Anchor, a novel system necessitating only a one-minute video clip of an individual for training, subsequently enabling the automatic generation of anchor-style videos with precise torso and hand movements. Specifically, we finetune a proposed structure-guided diffusion model on input video to render 3D mesh conditions into human appearances. We adopt a two-stage training strategy for the diffusion model, effectively binding movements with specific appearances. To produce arbitrary long temporal video, we extend the 2D U-Net in the frame-wise diffusion model to a 3D style without additional training cost, and a simple yet effective batch-overlapped temporal denoising module is proposed to bypass the constraints on video length during inference. Finally, a novel identity-specific face enhancement module is introduced to improve the visual quality of facial regions in the output videos. Comparative experiments demonstrate the effectiveness and superiority of the system in terms of visual quality, temporal coherence, and identity preservation, outperforming SOTA diffusion/non-diffusion methods. Project page: https://github.com/ICTMCG/Make-Your-Anchor.
Ziyao Huang 0002, Fan Tang, Yong Zhang 0034, Xiaodong Cun, Juan Cao 0001, Jintao Li 0001, Tong-Yee Lee
CVPR2
2024 U-VAP: User-specified Visual Appearance Personalization via Decoupled Self Augmentation
abstract
Concept personalization methods enable large text-to-image models to learn specific subjects (e.g., ob-jects/poses/3D models) and synthesize renditions in new contexts. Given that the image references are highly biased towards visual attributes, state-of-the-art personalization models tend to overfit the whole subject and cannot disentangle visual characteristics in pixel space. In this study, we proposed a more challenging setting, namely fine-grained visual appearance personalization. Different from existing methods, we allow users to provide a sentence describing the desired attributes. A novel decoupled self-augmentation strategy is proposed to generate target-related and non-target samples to learn user-specified visual attributes. These augmented data allow for refining the model's understanding of the target attribute while mitigating the impact of unrelated attributes. At the inference stage, adjustments are conducted on semantic space through the learned tar-get and non-target embeddings to further enhance the dis-entanglement of target attributes. Extensive experiments on various kinds of visual attributes with SOTA personalization methods show the ability of the proposed method to mimic target visual appearance in novel contexts, thus improving the controllability and flexibility of personalization. Project page: https://github.com/ICTMCG/U-VAP.
Kean Liu, Xiaoyue Mi, Fan Tang, Juan Cao 0001, Jintao Li 0001
CVPR4
2024 Lighting Image/Video Style Transfer Methods by Iterative Channel Pruning
abstract
Deploying style transfer methods on resource-constrained devices is challenging, which limits their real-world applicability. To tackle this issue, we propose using pruning techniques to accelerate various visual style transfer methods. We argue that typical pruning methods may not be well-suited for style transfer methods and present an iterative correlation-based channel pruning (ICCP) strategy for encoder-transform-decoder-based image/video style transfer models. The correlation-based channel regularization preserves the feature distributions for content and style references, and the iterative pruning strategy prevents layer collapse when pruning on the encoder-decoder structure. Experiments demonstrate that the proposed ICCP can generate visual competitive results compared to SOTA style transfer methods and significantly reduces the number of parameters (at least 70K) and inference time. Model is available at https://github.com/wukx-wukx/ICCP.
Kexin Wu, Fan Tang, Oliver Deussen, Thi Ngoc Hanh Le, Weiming Dong, Tong-Yee Lee
ICASSP2
2024 Revealing the Two Sides of Data Augmentation: An Asymmetric Distillation-based Win-Win Solution for Open-Set Recognition
Yunbing Jia, Xiaoyu Kong, Fan Tang, Yixing Gao 0001, Weiming Dong
IJCAI3
2024 Power Microservices Troubleshooting by Pretrained Language Model with Multi-source Data
abstract
Microservice has become the mainstream paradigm for developing cloud-native applications, but the intricate interdependencies between microservices and the vast amount of heterogeneous observable data (i.e. metrics, logs and traces) pose challenges for rapid troubleshooting. Several anomaly detection and root cause localization approaches that integrate multi-source data have been proposed. However, they are plagued with issues such as scarcity of high-quality data and insufficient model generalization. This is particularly evident when domain-specific models are trained from scratch for specific tasks. Recently, Large Language Models (LLMs) have shown outstanding capabilities in time series analysis, due to multi-source data generated by distributed microservices exhibit intrinsic spatio-temporal characteristics. In view of this, we propose LLM4MST, an LLM-empowered microservice troubleshooting model. We first unify and represent multi-source data by extracting service invocation graphs, and model dependencies between microservices by using a message-passing based graph neural network to generate graph-level sequences. The graph-level representation is then aligned with the LLM, and the LLM is fine-tuned to capture complex spatio-temporal patterns, generating a global vector that represents the state of microservice system within a timeslot. LLM4MST achieves accurate anomaly detection and root cause localization by jointly training the end-to-end model. Experiments on real datasets show that LLM4MST exhibits excellent performance in both full-sample and few-shot scenarios, demonstrating the powerful ability of LLMs in cross-domain knowledge transfer and few-shot learning.
Zhuang Lu, Fan Tang, Tong Li 0012, Jingguo Ge
ISPA4
2024 Dance-to-Music Generation with Encoder-based Textual Inversion
abstract
The seamless integration of music with dance movements is essential for communicating the artistic intent of a dance piece. This alignment also significantly improves the immersive quality of gaming experiences and animation productions. Although there has been remarkable advancement in creating high-fidelity music from textual descriptions, current methodologies mainly focus on modulating overall characteristics such as genre and emotional tone. They often overlook the nuanced management of temporal rhythm, which is indispensable in crafting music for dance, since it intricately aligns the musical beats with the dancers’ movements. Recognizing this gap, we propose an encoder-based textual inversion technique to augment text-to-music models with visual control, facilitating personalized music generation. Specifically, we develop dual-path rhythm-genre inversion to effectively integrate the rhythm and genre of a dance motion sequence into the textual space of a text-to-music model. Contrary to traditional textual inversion methods, which directly update text embeddings to reconstruct a single target object, our approach utilizes separate rhythm and genre encoders to obtain text embeddings for two pseudo-words, adapting to the varying rhythms and genres. We collect a new dataset called In-the-wild Dance Videos (InDV) and demonstrate that our approach outperforms state-of-the-art methods across multiple evaluation metrics. Furthermore, our method is able to adapt to changes in tempo and effectively integrates with the inherent text-guided generation capability of the pre-trained model. Our source code and demo videos are available at https://github.com/lsfhuihuiff/Dance-to-music_Siggraph_Asia_2024.
Sifei Li, Weiming Dong, Yuxin Zhang 0006, Fan Tang, Chongyang Ma, Oliver Deussen, Tong-Yee Lee, Changsheng Xu
SIGGRAPH Asia4
2024 Row-Column Separated Attention Based Low-Light Image/Video Enhancement
abstract
Abstract U‐Net structure is widely used for low‐light image/video enhancement. The enhanced images result in areas with large local noise and loss of more details without proper guidance for global information. Attention mechanisms can better focus on and use global information. However, attention to images could significantly increase the number of parameters and computations. We propose a Row–Column Separated Attention module (RCSA) inserted after an improved U‐Net. The RCSA module's input is the mean and maximum of the row and column of the feature map, which utilizes global information to guide local information with fewer parameters. We propose two temporal loss functions to apply the method to low‐light video enhancement and maintain temporal consistency. Extensive experiments on the LOL, MIT Adobe FiveK image, and SDSD video datasets demonstrate the effectiveness of our approach.
Chengqi Dong, Tuoshi Qi, Kexin Wu, Yixing Gao 0001, Fan Tang
Comput. Graph. Forum6
2024 Multi-Level Feature Exploration and Fusion Network for Prediction of IDH Status in Gliomas From MRI
abstract
Isocitrate dehydrogenase (IDH) is one of the most important genotypes in patients with glioma because it can affect treatment planning. Machine learning-based methods have been widely used for prediction of IDH status (denoted as IDH prediction). However, learning discriminative features for IDH prediction remains challenging because gliomas are highly heterogeneous in MRI. In this paper, we propose a multi-level feature exploration and fusion network (MFEFnet) to comprehensively explore discriminative IDH-related features and fuse different features at multiple levels for accurate IDH prediction in MRI. First, a segmentation-guided module is established by incorporating a segmentation task and is used to guide the network in exploiting features that are highly related to tumors. Second, an asymmetry magnification module is used to detect T2-FLAIR mismatch sign from image and feature levels. The T2-FLAIR mismatch-related features can be magnified from different levels to increase the power of feature representations. Finally, a dual-attention feature fusion module is introduced to fuse and exploit the relationships of different features from intra- and inter-slice feature fusion levels. The proposed MFEFnet is evaluated on a multi-center dataset and shows promising performance in an independent clinical dataset. The interpretability of the different modules is also evaluated to illustrate the effectiveness and credibility of the method. Overall, MFEFnet shows great potential for IDH prediction.
Jianyun Cao, Fan Tang, Meiyan Huang
IEEE J. Biomed. Health Informatics3
2024 Cross-Domain Mutual-Assistance Learning Framework for Fully Automated Diagnosis of Primary Tumor in Nasopharyngeal Carcinoma
abstract
Accurate T-staging of nasopharyngeal carcinoma (NPC) holds paramount importance in guiding treatment decisions and prognosticating outcomes for distinct risk groups. Regrettably, the landscape of deep learning-based techniques for T-staging in NPC remains sparse, and existing methodologies often exhibit suboptimal performance due to their neglect of crucial domain-specific knowledge pertinent to primary tumor diagnosis. To address these issues, we propose a new cross-domain mutual-assistance learning framework for fully automated diagnosis of primary tumor using H&N MR images. Specifically, we tackle primary tumor diagnosis task with the convolutional neural network consisting of a 3D cross-domain knowledge perception network (CKP net) for excavated cross-domain-invariant features emphasizing tumor intensity variations and internal tumor heterogeneity, and a multi-domain mutual-information sharing fusion network (M2SF net), comprising a dual-pathway domain-specific representation module and a mutual information fusion module, for intelligently gauging and amalgamating multi-domain, multi-scale T-stage diagnosis-oriented features. The proposed 3D cross-domain mutual-assistance learning framework not only embraces task-specific multi-domain diagnostic knowledge but also automates the entire process of primary tumor diagnosis. We evaluate our model on an internal and an external MR images dataset in a three-fold cross-validation paradigm. Exhaustive experimental results demonstrate that our method outperforms the other algorithms, and obtains promising performance for tumor segmentation and T-staging. These findings underscore its potential for clinical application, offering valuable assistance to clinicians in treatment decision-making and prognostication for various risk groups.
Xiuyu Dong, Kaifan Yang, Fan Tang, Wenjun Liao, Yu Zhang 0064, Shujun Liang
IEEE Trans. Medical Imaging4
2024 ${A^{2}Pt}$: Anti-Associative Prompt Tuning for Open Set Visual Recognition
abstract
Multi-modality pre-trained models (PTMs) have considerably boosted the performance on a broad range of computer vision topics. Still, they have not been explored purposefully in open set recognition (OSR) scenarios when applying PTMs to downstream recognition tasks. Directly fine/prompt tuning PTMs on closed-set classification tasks will inevitably suffer from data bias and always learn more or less target class-irrelevant cooccurring contextual information, which leads to over-confident predictions on unknown samples. In this paper, we propose a simple yet effective approach, termed Anti-Associative Prompt Tuning(A2Pt), toward learning compact and accurate class-related representation with few class-irrelevant associations from context using multi-modal priors. Specifically, a cross-modal guided activation module is adopted to refine the class-aware representation and suppress the associations from co-occurring contexts by involving text-modal information. We further design an anti-association calibration module to obtain compact class-aware and class-irrelevant representations, respectively, by introducing two additional object functions. Extensive experiments on publicly available benchmarks, including CIFAR series, Tiny-ImageNet, and ImageNet-21K-P, show that the proposed(A2Pt)achieves substantial and consistent performance gains compared with both SOTA OSR and PTM prompt tuning approaches.
Hairui Ren, Fan Tang, Xingjia Pan, Juan Cao 0001, Weiming Dong, Zhiwen Lin, Changsheng Xu
IEEE Trans. Multim.2
2024 Exploring the Temporal Consistency of Arbitrary Style Transfer: A Channelwise Perspective
abstract
Arbitrary image stylization by neural networks has become a popular topic, and video stylization is attracting more attention as an extension of image stylization. However, when image stylization methods are applied to videos, unsatisfactory results that suffer from severe flickering effects appear. In this article, we conducted a detailed and comprehensive analysis of the cause of such flickering effects. Systematic comparisons among typical neural style transfer approaches show that the feature migration modules for state-of-the-art (SOTA) learning systems are ill-conditioned and could lead to a channelwise misalignment between the input content representations and the generated frames. Unlike traditional methods that relieve the misalignment via additional optical flow constraints or regularization modules, we focus on keeping the temporal consistency by aligning each output frame with the input frame. To this end, we propose a simple yet efficient multichannel correlation network (MCCNet), to ensure that output frames are directly aligned with inputs in the hidden feature space while maintaining the desired style patterns. An inner channel similarity loss is adopted to eliminate side effects caused by the absence of nonlinear operations such as softmax for strict alignment. Furthermore, to improve the performance of MCCNet under complex light conditions, we introduce an illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/kongxiuxiu/MCCNetV2.
Xiaoyu Kong, Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Yongyong Chen, Zhenyu He 0001, Changsheng Xu
IEEE Trans. Neural Networks Learn. Syst.3
2024 Identity-Preserving Face Swapping via Dual Surrogate Generative Models
abstract
In this study, we revisit the fundamental setting of face-swapping models and reveal that only using implicit supervision for training leads to the difficulty of advanced methods to preserve the source identity. We propose a novel reverse pseudo-input generation approach to offer supplemental data for training face-swapping models, which addresses the aforementioned issue. Unlike the traditional pseudo-label-based training strategy, we assume that arbitrary real facial images could serve as the ground-truth outputs for the face-swapping network and try to generate corresponding input pair data. Specifically, we involve a source-creating surrogate that alters the attributes of the real image while keeping the identity, and a target-creating surrogate intends to synthesize attribute-preserved target images with different identities. Our framework, which utilizes proxy-paired data as explicit supervision to direct the face-swapping training process, partially fulfills a credible and effective optimization direction to boost the identity-preserving capability. We design explicit and implicit adaption strategies to better approximate the explicit supervision for face swapping. Quantitative and qualitative experiments on FF++, FFHQ, and wild images show that our framework could improve the performance of various face-swapping pipelines in terms of visual fidelity and ID preserving. Furthermore, we display applications with our method on re-aging, swappable attribute customization, cross-domain, and video face swapping. Code is available under https://github.com/ ICTMCG/CSCS.
Ziyao Huang 0002, Fan Tang, Yong Zhang 0034, Juan Cao 0001, Sheng Tang, Jintao Li 0001, Tong-Yee Lee
ACM Trans. Graph.2
2023 Adaptive Assignment for Geometry Aware Local Feature Matching
abstract
The detector-free feature matching approaches are currently attracting great attention thanks to their excellent performance. However, these methods still struggle at large-scale and viewpoint variations, due to the geometric inconsistency resulting from the application of the mutual nearest neighbour criterion (i.e., one-to-one assignment) in patch-level matching. Accordingly, we in-troduce AdaMatcher, which first accomplishes the feature correlation and co-visible area estimation through an elaborate feature interaction module, then performs adaptive assignment on patch-level matching while es-timating the scales between images, and finally refines the co-visible matches through scale alignment and sub-pixel regression module. Extensive experiments show that AdaMatcher outperforms solid baselines and achieves state-of-the-art results on many downstream tasks. Ad-ditionally, the adaptive assignment and sub-pixel refinement module can be used as a refinement network for other matching methods, such as SuperGlue, to boost their performance further. The code will be publicly available at https://github.com/AbyssGaze/AdaMatcher.
Dihe Huang, Yong Liu 0032, Shang Xu, Yikang Ding, Fan Tang, Chengjie Wang 0001
CVPR8
2023 Progressive Open Space Expansion for Open-Set Model Attribution
abstract
Despite the remarkable progress in generative technology, the Janus-faced issues of intellectual property protection and malicious content supervision have arisen. Efforts have been paid to manage synthetic images by attributing them to a set of potential source models. However, the closed-set classification setting limits the application in real-world scenarios for handling contents generated by arbitrary models. In this study, we focus on a challenging task, namely Open-Set Model Attribution (OSMA), to simultaneously attribute images to known models and identify those from unknown ones. Compared to existing openset recognition (OSR) tasks focusing on semantic novelty, OSMA is more challenging as the distinction between images from known and unknown models may only lie in visually imperceptible traces. To this end, we propose a Progressive Open Space Expansion (POSE) solution, which simulates open-set samples that maintain the same semantics as closed-set samples but embedded with different imperceptible traces. Guided by a diversity constraint, the open space is simulated progressively by a set of lightweight augmentation models. We consider three real-world scenarios and construct an OSMA benchmark dataset, including unknown models trained with different random seeds, architectures, and datasets from known ones. Extensive experiments on the dataset demonstrate POSE is superior to both existing model attribution methods and off-the-shelf OSR methods. Github: https://github.com/ICTMCG/POSE
Tianyun Yang, Danding Wang, Fan Tang, Juan Cao 0001, Sheng Tang
CVPR3
2023 Inversion-based Style Transfer with Diffusion Models
abstract
The artistic style within a painting is the means of expression, which includes not only the painting material, colors, and brushstrokes, but also the high-level attributes, including semantic elements and object shapes. Previous arbitrary example-guided artistic image generation methods often fail to control shape changes or convey elements. Pre-trained text-to-image synthesis diffusion probabilistic models have achieved remarkable quality but often require extensive textual descriptions to accurately portray the attributes of a particular painting. The uniqueness of an artwork lies in the fact that it cannot be adequately explained with normal language. Our key idea is to learn the artistic style directly from a single painting and then guide the synthesis without providing complex textual descriptions. Specifically, we perceive style as a learnable textual description of a painting. We propose an inversion-based style transfer method (InST), which can efficiently and accurately learn the key information of an image, thus capturing and transferring the artistic style of a painting. We demonstrate the quality and efficiency of our method on numerous paintings of various artists and styles. Codes are available at https://github.com/zyxElsa/InST.
Yuxin Zhang 0006, Nisha Huang, Fan Tang, Chongyang Ma, Weiming Dong, Changsheng Xu
CVPR3
2023 Towards harmonized regional style transfer and manipulation for facial images
abstract
Regional facial image synthesis conditioned on a semantic mask has achieved great attention in the field of computational visual media. However, the appearances of different regions may be inconsistent with each other after performing regional editing. In this paper, we focus on harmonized regional style transfer for facial images. A multi-scale encoder is proposed for accurate style code extraction. The key part of our work is a multi-region style attention module. It adapts multiple regional style embeddings from a reference image to a target image, to generate a harmonious result. We also propose style mapping networks for multi-modal style synthesis. We further employ an invertible flow model which can serve as mapping network to fine-tune the style code by inverting the code to latent space. Experiments on three widely used face datasets were used to evaluate our model by transferring regional facial appearance between datasets. The results show that our model can reliably perform style transfer and multi-modal manipulation, generating output comparable to the state of the art.
Fan Tang, Yong Zhang 0034, Tieru Wu, Weiming Dong
Comput. Vis. Media2
2023 Bias oriented unbiased data augmentation for cross-bias representation learning
Fan Tang, Juan Cao 0001, Xirong Li 0001, Danding Wang
Multim. Syst.2
2023 CrossRectify: Leveraging disagreement for semi-supervised object detection
Chengcheng Ma, Xingjia Pan, Qixiang Ye, Fan Tang, Weiming Dong, Changsheng Xu
Pattern Recognit.4
2023 Semantic-Context Graph Network for Point-Based 3D Object Detection
abstract
Point-based indoor 3D object detection has received increasing attention with the large demand for augmented reality, autonomous driving, and robot technology in the industry. However, the detection precision suffers from inputs with semantic ambiguity, i.e., shape symmetries, occlusion, and texture missing, which would lead that different objects appearing similar from different viewpoints and then confusing the detection model. Typical point-based detectors relieve this problem via learning proposal representations with both geometric and semantic information, while the entangled representation may cause a reduction in both semantic and spatial discrimination. In this paper, we focus on alleviating the confusion from entanglement and then enhancing the proposal representation by considering the proposal’s semantics and the context in one scene. A semantic-context graph network (SCGNet) is proposed, which mainly includes two modules: a category-aware proposal recoding module (CAPR) and a proposal context aggregation module (PCAg). To produce semantically clear features from entanglement representation, the CAPR module learns a high-level semantic embedding for each category to extract discriminative semantic clues. In view of further enhancing the proposal representation and leveraging the semantic clues, the PCAg module builds a graph to mine the most relevant context in the scene. With few bells and whistles, the SCGNet achieves SOTA performance and obtains consistent gains when applying to different backbones (0.9% ~ 2.4% on ScanNet V2 and 1.6% ~ 2.2% on SUN RGB-D for [email protected]). Code is available at https://github.com/dsw-jlu-rgzn/SCGNet.
Shuwei Dong, Xiaoyu Kong, Xingjia Pan, Fan Tang, Weiming Dong
IEEE Trans. Circuits Syst. Video Technol.4
2023 SPA2Net: Structure-Preserved Attention Activated Network for Weakly Supervised Object Localization
abstract
By exploring the localizable representations in deep CNN, weakly supervised object localization (WSOL) methods could determine the position of the object in each image just trained by the classification task. However, the partial activation problem caused by the discriminant function makes the network unable to locate objects accurately. To alleviate this problem, we propose Structure-Preserved Attention Activated Network (SPA2Net), a simple and effective one-stage WSOL framework to explore the ability of structure preservation of deep features. Different from traditional WSOL approaches, we decouple the object localization task from the classification branch to reduce their mutual influence by involving a localization branch which is online refined by a self-supervised structural-preserved localization mask. Specifically, we employ the high-order self-correlation as structural prior to enhance the perception of spatial interaction within convolutional features. By succinctly combining the structural prior with spatial attention, activations by SPA2Net will spread from part to the whole object during training. To avoid the structure-missing issue caused by the classification network, we furthermore utilize the restricted activation loss (RAL) to distinguish the difference between foreground and background in the channel dimension. In conjunction with the self-supervised localization branch, SPA2Net can directly predict the class-irrelevant localization map while prompting the network to pay more attention to the target region for accurate localization. Extensive experiments on two publicly available benchmarks, including CUB-200-2011 and ILSVRC, show that our SPA2Net achieves substantial and consistent performance gains compared with baseline approaches. The code and models are available at https://github.com/MsterDC/SPA2Net.
Dong Chen 0044, Xingjia Pan, Fan Tang, Weiming Dong, Changsheng Xu
IEEE Trans. Image Process.3
2023 SMNet: Synchronous Multi-Scale Low Light Enhancement Network With Local and Global Concern
abstract
Limited by objectively poor lighting conditions and hardware devices, low-light images with low visual quality and low visibility are inevitable in the real world. Accurate local details and reasonable global information play their essential and distinct roles in low-light image enhancement: local details contribute to fine textures, while global information is critical for a proper understanding of the global brightness level. In this paper, we focus on integrating local and global aspects to achieve high-quality low-light image enhancement by proposing the synchronous multi-scale low-light enhancement network (SMNet). A synchronous multi-scale representation learning structure and a global feature recalibration module are adopted in SMNet. Different from the traditional multi-scale feature learning architecture, SMNet carries out the multi-scale representation learning in a synchronous way: we first calculate the rough contextual representations in a top-down manner and then learn multi-scale representations in a bottom-up way to generate representations with rich local details. To acquire global brightness information, a global feature recalibration module (GFRM) is applied after the synchronous multi-scale representations to perceive and exploit proper global information by global pooling and projection to recalibrate channel weights globally. The synchronous multi-scale representation and GFRM compose the basic local-and-global block. Experimental results on mainstream real-world dataset LOL and synthetic dataset MIT-Adobe FiveK show that the proposed SMNet not only leads the way on objective metrics (0.41/2.31 improvement of PSNR on two datasets) but is also superior in subjective comparisons compared with typical SoTA methods. The code had already been uploaded tohttps://github.com/linshideng/SMNet.
Shideng Lin, Fan Tang, Weiming Dong, Xingjia Pan, Changsheng Xu
IEEE Trans. Multim.2
2023 ProSpect: Prompt Spectrum for Attribute-Aware Personalization of Diffusion Models
abstract
Personalizing generative models offers a way to guide image generation with user-provided references. Current personalization methods can invert an object or concept into the textual conditioning space and compose new natural sentences for text-to-image diffusion models. However, representing and editing specific visual attributes such as material, style, and layout remains a challenge, leading to a lack of disentanglement and editability. To address this problem, we propose a novel approach that leverages the step-by-step generation process of diffusion models, which generate images from low to high frequency information, providing a new perspective on representing, generating, and editing images. We develop the Prompt Spectrum Space P*, an expanded textual conditioning space, and a new image representation method called ProSpect. ProSpect represents an image as a collection of inverted textual token embeddings encoded from per-stage prompts, where each prompt corresponds to a specific generation stage (i.e., a group of consecutive steps) of the diffusion model. Experimental results demonstrate that P* and ProSpect offer better disentanglement and controllability compared to existing methods. We apply ProSpect in various personalized attribute-aware image generation applications, such as image-guided or text-driven manipulations of materials, style, and layout, achieving previously unattainable results from a single image input without fine-tuning the diffusion models. Our source code is available at https://github.com/zyxElsa/ProSpect.
Yuxin Zhang 0006, Weiming Dong, Fan Tang, Nisha Huang, Chongyang Ma, Tong-Yee Lee, Oliver Deussen, Changsheng Xu
ACM Trans. Graph.3
2023 A Unified Arbitrary Style Transfer Framework via Adaptive Contrastive Learning
abstract
This work presents Unified Contrastive Arbitrary Style Transfer (UCAST), a novel style representation learning and transfer framework, that can fit in most existing arbitrary image style transfer models, such as CNN-based, ViT-based, and flow-based methods. As the key component in image style transfer tasks, a suitable style representation is essential to achieve satisfactory results. Existing approaches based on deep neural networks typically use second-order statistics to generate the output. However, these hand-crafted features computed from a single image cannot leverage style information sufficiently, which leads to artifacts such as local distortions and style inconsistency. To address these issues, we learn style representation directly from a large number of images based on contrastive learning by considering the relationships between specific styles and the holistic style distribution. Specifically, we present an adaptive contrastive learning scheme for style transfer by introducing an input-dependent temperature. Our framework consists of three key components: a parallel contrastive learning scheme for style representation and transfer, a domain enhancement (DE) module for effective learning of style distribution, and a generative network for style transfer. Qualitative and quantitative evaluations show the results of our approach are superior to those obtained via state-of-the-art methods. The code is available at https://github.com/zyxElsa/CAST_pytorch .
Yuxin Zhang 0006, Fan Tang, Weiming Dong, Chongyang Ma, Tong-Yee Lee, Changsheng Xu
ACM Trans. Graph.2
2023 Balance-Aware Grid Collage for Small Image Collections
abstract
Grid collages (GClg) of small image collections are popular and useful in many applications, such as personal album management, online photo posting, and graphic design. In this article, we focus on how visual effects influence individual preferences through various arrangements of multiple images under such scenarios. A novel balance-aware metric is proposed to bridge the gap between multi-image joint presentation and visual pleasure. The metric merges psychological achievements into the field of grid collage. To capture user preference, a bonus mechanism related to a user-specified special location in the grid and uniqueness values of the subimages is integrated into the metric. An end-to-end reinforcement learning mechanism empowers the model without tedious manual annotations. Experiments demonstrate that our metric can evaluate the GClg visual balance in line with human subjective perception, and the model can generate visually pleasant GClg results, which is comparable to manual designs.
Fan Tang, Weiming Dong, Feiyue Huang, Tong-Yee Lee, Changsheng Xu
IEEE Trans. Vis. Comput. Graph.2
2022 ZINB-Based Graph Embedding Autoencoder for Single-Cell RNA-Seq Interpretations
abstract
Single-cell RNA sequencing (scRNA-seq) provides high-throughput information about the genome-wide gene expression levels at the single-cell resolution, bringing a precise understanding on the transcriptome of individual cells. Unfortunately, the rapidly growing scRNA-seq data and the prevalence of dropout events pose substantial challenges for cell type annotation. Here, we propose a single-cell model-based deep graph embedding clustering (scTAG) method, which simultaneously learns cell–cell topology representations and identifies cell clusters based on deep graph convolutional network. scTAG integrates the zero-inflated negative binomial (ZINB) model into a topology adaptive graph convolutional autoencoder to learn the low-dimensional latent representation and adopts Kullback–Leibler (KL) divergence for the clustering tasks. By simultaneously optimizing the clustering loss, ZINB loss, and the cell graph reconstruction loss, scTAG jointly optimizes cluster label assignment and feature learning with the topological structures preserved in an end-to-end manner. Extensive experiments on 16 single-cell RNA-seq datasets from diverse yet representative single-cell sequencing platforms demonstrate the superiority of scTAG over various state-of-the-art clustering methods.
Zhuohan Yu, Yifu Lu, Yunhe Wang 0002, Fan Tang, Ka-Chun Wong, Xiangtao Li
AAAI4
2022 StyTr2: Image Style Transfer with Transformers
abstract
The goal of image style transfer is to render an image with artistic features guided by a style reference while maintaining the original content. Owing to the locality in convolutional neural networks (CNNs), extracting and maintaining the global information of input images is difficult. Therefore, traditional neural style transfer methods face biased content representation. To address this critical issue, we take long-range dependencies of input images into account for image style transfer by proposing a transformer-based approach called StyTr2. In contrast with visual transformers for other vision tasks, StyTr2 contains two different transformer encoders to generate domain-specific sequences for content and style, respectively. Following the encoders, a multi-layer transformer decoder is adopted to stylize the content sequence according to the style sequence. We also analyze the deficiency of existing positional encoding methods and propose the content-aware positional encoding (CAPE), which is scale-invariant and more suitable for image style transfer tasks. Qualitative and quantitative experiments demonstrate the effectiveness of the proposed StyTr2 compared with state-of-the-art CNN-based and flow-based approaches. Code and models are available at https://github.com/diyiiyiii/StyTR-2.
Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Xingjia Pan, Changsheng Xu
CVPR2
2022 SIOD: Single Instance Annotated Per Category Per Image for Object Detection
abstract
Object detection under imperfect data receives great attention recently. Weakly supervised object detection (WSOD) suffers from severe localization issues due to the lack of instance-level annotation, while semi-supervised object detection (SSOD) remains challenging led by the inter-image discrepancy between labeled and unlabeled data. In this study, we propose the Single Instance annotated Object Detection (SIOD), requiring only one instance annotation for each existing category in an image. Degraded from inter-task (WSOD) or inter-image (SSOD) discrepancies to the intra-image discrepancy, SIOD provides more reliable and rich prior knowledge for mining the rest of unlabeled instances and trades off the annotation cost and performance. Under the SIOD setting, we propose a simple yet effective framework, termed Dual-Mining (DMiner), which consists of a Similarity-based Pseudo Label Generating module (SPLG) and a Pixel-level Group Contrastive Learning module (PGCL). SPLG firstly mines latent instances from feature representation space to alleviate the annotation missing problem. To avoid being misled by inaccurate pseudo labels, we propose PGCL to boost the tolerance to false pseudo labels. Extensive experiments on MS COCO verify the feasibility of the SIOD setting and the superiority of the proposed method, which obtains consistent and significant improvements compared to baseline methods and achieves comparable results with fully supervised object detection (FSOD) methods with only 40% instances annotated. Code is available at https://github.com/solicucu/SIOD.
Hanjun Li 0004, Xingjia Pan, Fan Tang, Wei-Shi Zheng 0001
CVPR4
2022 Draw Your Art Dream: Diverse Digital Art Synthesis with Multimodal Guided Diffusion
abstract
Digital art synthesis is receiving increasing attention in the multimedia community because of engaging the public with art effectively. Current digital art synthesis methods usually use single-modality inputs as guidance, thereby limiting the expressiveness of the model and the diversity of generated results. To solve this problem, we propose the multimodal guided artwork diffusion (MGAD) model, which is a diffusion-based digital artwork generation approach that utilizes multimodal prompts as guidance to control the classifier-free diffusion model. Additionally, the contrastive language-image pretraining (CLIP) model is used to unify text and image modalities. Extensive experimental results on the quality and quantity of the generated digital art paintings confirm the effectiveness of the combination of the diffusion model and multimodal guidance. Code is available at https://github.com/haha-lisa/MGAD-multimodal-guided-artwork-diffusion.
Nisha Huang, Fan Tang, Weiming Dong, Changsheng Xu
ACM Multimedia2
2022 Non-dominated sorting based multi-page photo collage
abstract
The development of social networking services (SNSs) revealed a surge in image sharing. The sharing mode of multi-page photo collage (MPC), which posts several image collages at a time, can often be observed on many social network platforms, which enables uploading images and arrangement in a logical order. This study focuses on the construction of MPC for an image collection and its formulation as an issue of joint optimization, which involves not only the arrangement in a single collage but also the arrangement among different collages. Novel balance-aware measurements, which merge graphic features and psychological achievements, are introduced. Non-dominated sorting genetic algorithm is adopted to optimize the MPC guided by the measurements. Experiments demonstrate that the proposed method can lead to diverse, visually pleasant, and logically clear MPC results, which are comparable to manually designed MPC results.
Fan Tang, Weiming Dong, Changsheng Xu
Comput. Vis. Media2
2022 Transformers in computational visual media: A survey
abstract
Transformers, the dominant architecture for natural language processing, have also recently attracted much attention from computational visual media researchers due to their capacity for long-range representation and high performance. Transformers are sequence-to-sequence models, which use a self-attention mechanism rather than the RNN sequential structure. Thus, such models can be trained in parallel and can represent global information. This study comprehensively surveys recent visual transformer works. We categorize them according to task scenario: backbone design, high-level vision, low-level vision and generation, and multimodal learning. Their key ideas are also analyzed. Differing from previous surveys, we mainly focus on visual transformer methods in low-level vision and generation. The latest works on backbone design are also reviewed in detail. For ease of understanding, we precisely describe the main contributions of the latest works in the form of tables. As well as giving quantitative comparisons, we also present image results for low-level vision and generation tasks. Computational costs and source code links for various important works are also given in this survey to assist further development.
Yifan Xu 0008, HuaPeng Wei, Minxuan Lin, Yingying Deng, Kekai Sheng, Mengdan Zhang, Fan Tang, Weiming Dong, Feiyue Huang, Changsheng Xu
Comput. Vis. Media7
2022 A Comparative Study of CNN- and Transformer-Based Visual Style Transfer
HuaPeng Wei, Yingying Deng, Fan Tang, Xingjia Pan, Weiming Dong
J. Comput. Sci. Technol.3
2021 Arbitrary Video Style Transfer via Multi-Channel Correlation
abstract
Video style transfer is attracting increasing attention from the artificial intelligence community because of its numerous applications, such as augmented reality and animation production. Relative to traditional image style transfer, video style transfer presents new challenges, including how to effectively generate satisfactory stylized results for any specified style while maintaining temporal coherence across frames. Towards this end, we propose a Multi-Channel Correlation network (MCCNet), which can be trained to fuse exemplar style features and input content features for efficient style transfer while naturally maintaining the coherence of input videos to output videos. Specifically, MCCNet works directly on the feature space of style and content domain where it learns to rearrange and fuse style features on the basis of their similarity to content features. The outputs generated by MCC are features containing the desired style patterns that can further be decoded into images with vivid style textures. Moreover, MCCNet is also designed to explicitly align the features to input and thereby ensure that the outputs maintain the content structures and the temporal continuity. To further improve the performance of MCCNet under complex light conditions, we also introduce illumination loss during training. Qualitative and quantitative evaluations demonstrate that MCCNet performs well in arbitrary video and image style transfer tasks. Code is available at https://github.com/diyiiyiii/MCCNet.
Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Changsheng Xu
AAAI2
2021 Unveiling the Potential of Structure Preserving for Weakly Supervised Object Localization
abstract
Weakly supervised object localization (WSOL) remains an open problem given the deficiency of finding object extent information using a classification network. Although prior works struggled to localize objects through various spatial regularization strategies, we argue that how to extract object structural information from the trained classification network is neglected. In this paper, we propose a two-stage approach, termed structure-preserving activation (SPA), toward fully leveraging the structure information incorporated in convolutional features for WSOL. First, a restricted activation module (RAM) is designed to alleviate the structure-missing issue caused by the classification network on the basis of the observation that the unbounded classification map and global average pooling layer drive the network to focus only on object parts. Second, we designed a post-process approach, termed self-correlation map generating (SCG) module to obtain structure-preserving localization maps on the basis of the activation maps acquired from the first stage. Specifically, we utilize the high-order self-correlation (HSC) to extract the inherent structural information retained in the learned model and then aggregate HSC of multiple points for precise object localization. Extensive experiments on two publicly available benchmarks including CUB-2002011 and ILSVRC show that the proposed SPA achieves substantial and consistent performance gains compared with baseline approaches. Code and models are available at github.com/Panxjia/SPA CVPR2021.
Xingjia Pan, Yingguo Gao, Zhiwen Lin, Fan Tang, Weiming Dong, Haolei Yuan, Feiyue Huang, Changsheng Xu
CVPR4
2021 DAE-GAN: Dynamic Aspect-aware GAN for Text-to-Image Synthesis
abstract
Text-to-image synthesis refers to generating an image from a given text description, the key goal of which lies in photo realism and semantic consistency. Previous methods usually generate an initial image with sentence embedding and then refine it with fine-grained word embedding. Despite the significant progress, the ‘aspect’ information (e.g., red eyes) contained in the text, referring to several words rather than a word that depicts ‘a particular part or feature of something’, is often ignored, which is highly helpful for synthesizing image details. How to make better utilization of aspect information in text-to-image synthesis still remains an unresolved challenge. To address this problem, in this paper, we propose a Dynamic Aspect-awarE GAN (DAE-GAN) that represents text information comprehensively from multiple granularities, including sentence-level, word-level, and aspect-level. Moreover, inspired by human learning behaviors, we develop a novel Aspect-aware Dynamic Re-drawer (ADR) for image refinement, in which an Attended Global Refinement (AGR) module and an Aspect-aware Local Refinement (ALR) module are alternately employed. AGR utilizes word-level embedding to globally enhance the previously generated image, while ALR dynamically employs aspect-level embedding to refine image details from a local perspective. Finally, a corresponding matching loss function is designed to ensure the text-image semantic consistency at different levels. Extensive experiments on two well-studied and publicly available datasets (i.e., CUB-200 and COCO) demonstrate the superiority and rationality of our method.
Shulan Ruan, Yong Zhang 0034, Kun Zhang 0015, Yanbo Fan, Fan Tang, Qi Liu 0003, Enhong Chen
ICCV5
2021 SiamCPN: Visual tracking with the Siamese center-prediction network
abstract
Object detection is widely used in object tracking; anchor-free object tracking provides an end-to-end single-object-tracking approach. In this study, we propose a new anchor-free network, the Siamese center-prediction network (SiamCPN). Given the presence of referenced object features in the initial frame, we directly predict the center point and size of the object in subsequent frames in a Siamese-structure network without the need for perframe post-processing operations. Unlike other anchor-free tracking approaches that are based on semantic segmentation and achieve anchor-free tracking by pixel-level prediction, SiamCPN directly obtains all information required for tracking, greatly simplifying the model. A center-prediction sub-network is applied to multiple stages of the backbone to adaptively learn from the experience of different branches of the Siamese net. The model can accurately predict object location, implement appropriate corrections, and regress the size of the target bounding box. Compared to other leading Siamese networks, SiamCPN is simpler, faster, and more efficient as it uses fewer hyperparameters. Experiments demonstrate that our method outperforms other leading Siamese networks on GOT-10K and UAV123 benchmarks, and is comparable to other excellent trackers on LaSOT, VOT2016, and OTB-100 while improving inference speed 1.5 to 2 times.
Dong Chen 0044, Fan Tang, Weiming Dong, Hanxing Yao, Changsheng Xu
Comput. Vis. Media2
2021 Task migration optimization for guaranteeing delay deadline with mobility consideration in mobile edge computing
Fan Tang, Chubo Liu, Kenli Li 0001, Zhuo Tang, Keqin Li 0001
J. Syst. Archit.1
2021 Exploring the Representativity of Art Paintings
abstract
Art painting evaluation is sophisticated for a novice with no or limited knowledge on art criticism, and history. In this study, we propose the concept ofrepresentativityto evaluate paintings instead of using professional concepts, such as genre, media, and style, which may be confusing to non-professionals. We define the concept of representativity to evaluate quantitatively the extent to which a painting can represent the characteristics of an artists creations. We begin by proposing a novel deep representation of art paintings, which is enhanced by style information through a weighted pooling feature fusion module. In contrast to existing feature extraction approaches, the proposed framework embeds painting styles, and authorship information, and learns specific artwork characteristics in a single framework. Subsequently, we propose a graph-based learning method for representativity learning, which considers intra-category, and extra-category information. In view of the significance of historical factors in the art domain, we introduce the creation time of a painting into the learning process. User studies demonstrate our approach helps the public effectively access the creation characteristics of artists through sorting paintings by representativity from highest to lowest.
Yingying Deng, Fan Tang, Weiming Dong, Chongyang Ma, Feiyue Huang, Oliver Deussen, Changsheng Xu
IEEE Trans. Multim.2
2021 VNE-HRL: A Proactive Virtual Network Embedding Algorithm Based on Hierarchical Reinforcement Learning
abstract
Virtual network embedding (VNE) that instantiates virtualized networks on a substrate infrastructure, is one of the key research problems for network virtualization. Most existing VNE approaches, however, focus on the current virtual network request (VNR) and treat all VNRs equally, which disregard the long-term impact and waste many resources on the process of embedding infeasible VNRs (i.e., VNRs that cannot be embedded completely). To address these problems, a proactive virtual network embedding algorithm based on hierarchical reinforcement learning, VNE-HRL, is proposed in this paper. Within our framework, the VNE task is performed by a two-level agent that considers both the long-term impact of a VNR and the short-term effect of an embedding action. For each processing, a high-level agent aims to select a currently feasible VNR with the maximum long-term reward from a window-based batch, and a low-level agent is assigned to embed the selected VNR on a substrate infrastructure by performing a series of embedding actions. Extensive simulation results indicate that our algorithm best performance on most metrics compared with existing state-of-the-art solutions, with up to 9.92% and 33.03% improvement on acceptance ratio and average revenue.
Jin Cheng 0008, Yulei Wu, Yeming Lin, Yuepeng E, Fan Tang, Jingguo Ge
IEEE Trans. Netw. Serv. Manag.5
2021 Distribution Aligned Multimodal and Multi-domain Image Stylization
abstract
Multimodal and multi-domain stylization are two important problems in the field of image style transfer. Currently, there are few methods that can perform multimodal and multi-domain stylization simultaneously. In this study, we propose a unified framework for multimodal and multi-domain style transfer with the support of both exemplar-based reference and randomly sampled guidance. The key component of our method is a novel style distribution alignment module that eliminates the explicit distribution gaps between various style domains and reduces the risk of mode collapse. The multimodal diversity is ensured by either guidance from multiple images or random style codes, while the multi-domain controllability is directly achieved by using a domain label. We validate our proposed framework on painting style transfer with various artistic styles and genres. Qualitative and quantitative comparisons with state-of-the-art methods demonstrate that our method can generate high-quality results of multi-domain styles and multimodal instances from reference style guidance or a random sampled style.
Minxuan Lin, Fan Tang, Weiming Dong, Xiao Li 0030, Changsheng Xu, Chongyang Ma
ACM Trans. Multim. Comput. Commun. Appl.2
2021 Distributed Task Migration Optimization in MEC by Extending Multi-Agent Deep Reinforcement Learning Approach
abstract
Closer to mobile users geographically, mobile edge computing (MEC) can provide some cloud-like capabilities to users more efficiently. This enables it possible for resource-limited mobile users to offload their computation-intensive and latency-sensitive tasks to MEC nodes. For its great benefits, MEC has drawn wide attention and extensive works have been done. However, few of them address task migration problem caused by distributed user mobility, which can't be ignored with quality of service (QoS) consideration. In this article, we study task migration problem and try to minimize the average completion time of tasks under migration energy budget. There are multiple independent users and the movement of each mobile user is memoryless with a sequential decision-making process, thus reinforcement learning algorithm based on Markov chain model is applied with low computation complexity. To further facilitate cooperation among users, we devise a distributed task migration algorithm based on counterfactual multi-agent (COMA) reinforcement learning approach to solve this problem. Extensive experiments are carried out to assess the performance of this distributed task migration algorithm. Compared with no migrating (NM) and single-agent actor-critic (AC) algorithms, the proposed distributed task migration algorithm can achieve up 30-50 percent reduction about average completion time.
Chubo Liu, Fan Tang, Yikun Hu 0001, Kenli Li 0001, Zhuo Tang, Keqin Li 0001
IEEE Trans. Parallel Distributed Syst.2
2021 Content-Based Visual Summarization for Image Collections
abstract
With the surge of images in the information era, people demand an effective and accurate way to access meaningful visual information. Accordingly, effective and accurate communication of information has become indispensable. In this article, we propose a content-based approach that automatically generates a clear and informative visual summarization based on design principles and cognitive psychology to represent image collections. We first introduce a novel method to make representative and nonredundant summarizations of image collections, thereby ensuring data cleanliness and emphasizing important information. Then, we propose a tree-based algorithm with a two-step optimization strategy to generate the final layout that operates as follows: (1) an initial layout is created by constructing a tree randomly based on the grouping results of the input image set; (2) the layout is refined through a coarse adjustment in a greedy manner, followed by gradient back propagation drawing on the training procedure of neural networks. We demonstrate the usefulness and effectiveness of our method via extensive experimental results and user studies. Our visual summarization algorithm can precisely and efficiently capture the main content of image collections better than alternative methods or commercial tools.
Xingjia Pan, Fan Tang, Weiming Dong, Chongyang Ma, Yiping Meng, Feiyue Huang, Tong-Yee Lee, Changsheng Xu
IEEE Trans. Vis. Comput. Graph.2
2020 Arbitrary Style Transfer via Multi-Adaptation Network
abstract
Arbitrary style transfer is a significant topic with research value and application prospect. A desired style transfer, given a content image and referenced style painting, would render the content image with the color tone and vivid stroke patterns of the style painting while synchronously maintaining the detailed content structure information. Style transfer approaches would initially learn content and style representations of the content and style references and then generate the stylized images guided by these representations. In this paper, we propose the multi-adaptation network which involves two self-adaptation (SA) modules and one co-adaptation (CA) module:the SA modules adaptively disentangle the content and style representations, i.e., content SA module uses position-wise self-attention to enhance content representation and style SA module uses channel-wise self-attention to enhance style representation; the CA module rearranges the distribution of style representation based on content representation distribution by calculating the local similarity between the disentangled content and style features in a non-local fashion. Moreover, a new disentanglement loss function enables our network to extract main style patterns and exact content structures to adapt to various input images, respectively. Various qualitative and quantitative experiments demonstrate that the proposed multi-adaptation network leads to better results than the state-of-the-art style transfer methods.
Yingying Deng, Fan Tang, Weiming Dong, Feiyue Huang, Changsheng Xu
ACM Multimedia2
2020 Destylization of text with decorative elements
abstract
Style text with decorative elements has a strong visual sense, and enriches our daily work, study and life. However, it introduces new challenges to text detection and recognition. In this study, we propose a text destylized framework, that can transform the stylized texts with decorative elements into a type that is easily distinguishable by a detection or recognition model. We arranged and integrate an existing stylistic text data set to train the destylized network. The new destylized data set contains English letters and Chinese characters. The proposed approach enables a framework to handle both Chinese characters and English letters without the need for additional networks. Experiments show that the method is superior to the state-of-the-art style-related models.
Fan Tang, Weiming Dong, Changsheng Xu
MMAsia2
2020 Self-Supervised Feature Augmentation for Large Image Object Detection
abstract
Input scale plays an important role in modern detection frameworks, and an optimal training scale for images exists empirically. However, the optimal one usually cannot be reached in facing extremely large images under the memory constraint. In this study, we explore the scale effect inside the object detection pipeline and find that feature upsampling with the introduction of high-resolution information benefits the detection. Compared with direct input upscaling, feature upsampling trades a small performance loss for a large amount of memory savings. From these observations, we propose a self-supervised feature augmentation network, which takes downsampled images as inputs and aims to generate comparable features with the ones when feeding upscaled images to networks. We present a guided feature upsampling module, which takes downsampled images as inputs, to learn upscaled feature representations with the supervision of real large features acquired from upscaled images. In a self-supervised learning manner, we can introduce detailed information of images to the network. For an efficient feature upsampling, we design a residualized sub-pixel convolution block based on a sub-pixel convolution layer, which involves considerable information in upsampling process. Experiments on Mapillary Vistas Dataset (MVD), Cityscapes, and COCO are conducted to demonstrate the effectiveness of our method. On the MVD and Cityscapes detection benchmarks, in which the images are extremely large, our method surpasses current approaches. On COCO, the proposed method obtains comparable results to existing methods but with higher efficiency.
Xingjia Pan, Fan Tang, Weiming Dong, Zhichao Song, Yiping Meng, Pengfei Xu 0013, Oliver Deussen, Changsheng Xu
IEEE Trans. Image Process.2
2020 Image Retargetability
abstract
Real-world applications could benefit from the ability to automatically retarget an image to different aspect ratios and resolutions while preserving its visually and semantically important content. However, not all images can be equally processed. This study introduces the notion of image retargetability to describe how well a particular image can be handled by content-aware image retargeting. We propose to learn a deep convolutional neural network to rank photo retargetability, in which the relative ranking of photo retargetability is directly modeled in the loss function. Our model incorporates the joint learning of meaningful photographic attributes and image content information, which can facilitate the regularization of the complicated retargetability rating problem. To train and analyze this model, we collect a dataset that contains retargetability scores and meaningful image attributes assigned by six expert raters. The experiments demonstrate that our unified model can generate retargetability rankings that are highly consistent with human labels. To further validate our model, we show the applications of image retargetability in retargeting method selection, retargeting method assessment and generating a photo collage.
Fan Tang, Weiming Dong, Yiping Meng, Chongyang Ma, Fuzhang Wu, Tong-Yee Lee
IEEE Trans. Multim.1
2019 Selective clustering for representative paintings selection
Yingying Deng, Fan Tang, Weiming Dong, Fuzhang Wu, Oliver Deussen, Changsheng Xu
Multim. Tools Appl.2
2019 Joint face alignment and segmentation via deep multi-task learning
Fan Tang, Weiming Dong, Feiyue Huang, Xiaopeng Zhang 0001
Multim. Tools Appl.2
2018 Photo Squarization by Deep Multi-Operator Retargeting
abstract
Squared forms of photos are widely used in social media as album covers or thumbnails of image streams. In this study, we realize photo squarization by modeling Retargeting Visual Perception Issues, which reflect human perception preference toward image ratargeting. General image retargeting techniques deal with three common issues, namely, salient content, object shape, and scene composition, to preserve the important information of original image. We propose a new way based on multi-operator techniques to investigate human behavior in balancing the three issues. We establish a new dataset and observe human behavior by inviting investigators to retarget images to square manually. We propose a data-driven approach composed of perception and distillation modules by using deep learning techniques to predict human perception preference. The perception part learns the relations among the three issues, and the distillation part transfers the learned relations to a simple but effective network. Our study contributes to deep learning literature by optimizing a network index and lightening its running burden. Experimental results show that photo squarization results generated by the proposed model are consistent with human visual perception results.
Fan Tang, Weiming Dong, Xiaopeng Zhang 0001, Oliver Deussen, Tong-Yee Lee
ACM Multimedia2
2018 Animated Construction of Chinese Brush Paintings
abstract
In this paper, we present a method for reconstructing the drawing process of Chinese brush paintings. We demonstrate the possibility of computing an artistically reasonable drawing order from a static brush painting that is consistent with the rules of art. We map the key principles of drawing composition to our computational framework, which first organizes the strokes in three stages and then optimizes stroke ordering with natural evolution strategies. Our system produces reasonable animated constructions of Chinese brush paintings with minimal or no user intervention. We test our algorithm on a range of input paintings with varying degrees of complexity and structure and then evaluate the results via a user study. We discuss the applications of the proposed system to painting instruction, painting animation, and image stylization, especially in the context of art teaching.
Fan Tang, Weiming Dong, Yiping Meng, Xing Mei, Feiyue Huang, Xiaopeng Zhang 0001, Oliver Deussen
IEEE Trans. Vis. Comput. Graph.1
2004 Pose Invariant, Robust Feature Extraction from Data with a Modified Scale Space Approach
abstract
Feature-based simultaneous localization and map building (SLAM) approaches require a robust method to extract position invariant landmarks from the surrounding environment. 2D laser range finders are currently one of the most common sensors used to obtain environmental information for mobile robot navigation due to their reliability, accuracy and low cost. However, the 2D laser scan data only give very limited information, making it difficult to extract meaningful features particularly in unstructured environments. The most important steps to extract features are segmentation and noise reduction. Scale space and adaptive smoothing are two common techniques within the vision community. They are used to remove high frequency noise and represent image data in multi-scale spaces. They allow for an easier segmentation of images and the extraction of features in the appropriate scale. In this paper, a modified adaptive smoothing algorithm is proposed and applied to laser range data within a modified scale space framework. This algorithm smoothes range data and segments it at the same time by translating a line model mask over the range data. Lines can be extracted from the segments by using a standard fitting algorithm.
Fan Tang, Martin David Adams, Javier Ibañez-Guzmán, W. Sardha Wijesoma
ICRA1