EDBT 2026 Demo / reviewers in the wild / expert
Zhanjie Zhang
dblp:327/9306
· DBLP profile ↗
33ranked-venue papers
8as first author
33since 2021 · last 2026
0000-0002-8966-1328ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 23 · 8 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 2 first-author · 20 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RelaCtrl: Relevance-Guided Efficient Control for Diffusion TransformersabstractThe Diffusion Transformer plays a pivotal role in advancing text-to-image and text-to-video generation, owing primarily to its inherent scalability. However, existing controlled diffusion transformer methods incur significant parameter and computational overheads and suffer from inefficient resource allocation due to their failure to account for the varying relevance of control information across different transformer layers. To address this, we propose the Relevance-Guided Efficient Controllable Generation framework, RelaCtrl, enabling efficient and resource-optimized integration of control signals into the Diffusion Transformer. First, we evaluate the relevance of each layer in the Diffusion Transformer to the control information by assessing the ControlNet Relevance Score, which measures the impact of skipping each control layer on both the quality of generation and the control effectiveness during inference. Based on the strength of the relevance, we then tailor the positioning, parameter scale, and modeling capacity of the control layers to reduce unnecessary parameters and redundant computations. Additionally, to further improve efficiency, we replace the self-attention and FFN in the commonly used copy block with the carefully designed Two-Dimensional Shuffle Mixer (TDSM), enabling efficient implementation of both the token mixer and channel mixer. Both qualitative and quantitative experimental results demonstrate that our approach achieves superior performance with only 15% of the parameters and computational complexity compared to PixArt-delta. Ke Cao 0001, Jing Wang 0021, Ao Ma 0005, Jiasong Feng, Xuanhua He, Run Ling, Haozhe Wang 0002, Hongjuan Pei, Yihua Shao, Zhanjie Zhang, Jie Zhang 0033 |
AAAI | 13 |
| 2026 | MoFu: Scale-Aware Modulation and Fourier Fusion for Multi-Subject Video GenerationabstractMulti-subject video generation aims to synthesize videos from textual prompts and multiple reference images, ensuring that each subject preserves natural scale and visual fidelity. However, current methods face two challenges: scale inconsistency, where variations in subject size lead to unnatural generation, and permutation sensitivity, where the order of reference inputs causes subject distortion. In this paper, we propose MoFu, a unified framework that tackles both challenges. For scale inconsistency, we introduce Scale-Aware Modulation (SMO), an LLM-guided module that extracts implicit scale cues from the prompt and modulates features to ensure consistent subject sizes. To address permutation sensitivity, we present a simple yet effective Fourier Fusion strategy that processes the frequency information of reference features via the Fast Fourier Transform to produce a unified representation. Besides, we design a Scale-Permutation Stability Loss to jointly encourage scale-consistent and permutation-invariant generation. To further evaluate these challenges, we establish a dedicated benchmark with controlled variations in subject scale and reference permutation. Extensive experiments demonstrate that MoFu significantly outperforms existing methods in preserving natural scale, subject fidelity, and overall visual quality. Run Ling, Ke Cao 0001, Ao Ma 0005, Runze He, Changwei Wang 0001, Rongtao Xu, Yihua Shao, Zhanjie Zhang, Guibing Guo, Jingjing Lv, Junjie Shen 0008, Ching Law, Xingwei Wang 0001 |
AAAI | 10 |
| 2026 | ENnRA: A Dual-Graph Architecture with Ego-Neighbor Alignment for Multimodal Recommendation Systems
Ruixiang Yu, Zhanjie Zhang, Yicheng Di, Yuan Liu 0021 |
ISCAS | 2 |
| 2026 | RAFed: Responsive Augmentation and Approximate Update Method for Federated Learning with Non-IID DataabstractFederated learning is a distributed collaborative training framework that enables multiple clients to share model updates and jointly train deep neural networks without exchanging raw data. Although extensive research has explored data augmentation techniques in federated settings, the naturally non-IID data distributions among clients render blind augmentation prone to severe degradation of the learned model. To solve this problem, we suggest Responsive Augmentation and Approximate Update Method for Fed erated Learning with Non-IID Data (RAFed), aimed at alleviating feature shift in client samples. We leverage a Responsive Augmentation Method to accumulate shared data augmentation policy knowledge through local learning, guiding the policy gradient to consider the impact of data augmentation on unseen local data, and employ an Approximate Update Mechanism to reduce communication costs and achieve efficient policy search. To improve the adaptability of data augmentation policies to local data distributions, we introduce a Dynamic Adaptive Method for searching personalized augmentation policies tailored to heterogeneous clients. Experiments on four popular datasets show that RAFed achieves superior test accuracy and lower communication costs compared to related baselines while providing privacy advantages. The code is available via https://github.com/anonymously123-stcak/RAFed. Yicheng Di, Zhanjie Zhang |
WWW | 2 |
| 2026 | Towards arbitrary-scale image super-resolution with prompting and diffusion prior
Xinyue Tu, Zhanjie Zhang, Wei Xing 0001, Lei Zhao 0011, Yuanxing Liu 0005 |
Knowl. Based Syst. | 3 |
| 2026 | IP-Controller: Decomposition and Optimization of Cross-Attention Maps for Accurate Subject-Driven Text-to-Image Diffusion GenerationabstractAlthough large pretrained stable diffusion (SD) models can generate high-quality images from prompts, they cannot generate images that are consistent with the fine-grained characteristics of a specific identityV∗(e.g., an anime character). Subject-driven generation focuses on exploring and leveraging the prior knowledge within a model to achieve the goals of ID and context preservation. There have been efforts, such as DreamBooth, to conduct subject-driven generation; however, they suffer from ID and context mistakes. An ID mistake means a feature loss ofV∗, and a context mistake means that the generated image does not align with the given prompt. To rectify these problems, in this paper, we propose masked fine-tuning for efficient feature learning ofV∗, then propose IP-Controller for decomposing and optimizing cross-attention maps ofV∗and prompt words other thanV∗. Specifically, we generate the cross-attention map using a vanilla input prompt and decompose it into an ID cross-attention map (matchingV∗) and a context cross-attention map (matching prompt words other thanV∗). Next, we generate fitter ID and context cross-attention maps on the basis of the input ID and context prompts, respectively. We optimize the ID and context cross-attention maps with the fitter ID and context cross-attention maps, respectively, so that the diffusion process pays fitter attention for specific contents. Experiments show that IP-Controller correctly integrates the core features ofV∗and the semantic context of the prompt words other thanV∗and generates high-quality images for the given prompt. Junsheng Luan, Zhanjie Zhang, Lei Zhao 0011, Wei Xing 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | SMFormer: Empowering Self-Supervised Stereo Matching via Foundation Models and Data AugmentationabstractRecent self-supervised stereo matching methods have made significant progress. They typically rely on the photometric consistency assumption, which presumes corresponding points across views share the same appearance. However, this assumption could be compromised by real-world disturbances, resulting in invalid supervisory signals and a significant accuracy gap compared to supervised methods. To address this issue, we propose SMFormer, a framework integrating more reliable self-supervision guided by the Vision Foundation Model (VFM) and data augmentation. We first incorporate the VFM with the Feature Pyramid Network (FPN), providing a discriminative and robust feature representation against disturbance in various scenarios. We then devise an effective data augmentation mechanism that ensures robustness to various transformations. The data augmentation mechanism explicitly enforces consistency between learned features and those influenced by illumination variations. Additionally, it regularizes the output consistency between disparity predictions of strong augmented samples and those generated from standard samples. Experiments on multiple mainstream benchmarks demonstrate that our SMFormer achieves state-of-the-art (SOTA) performance among self-supervised methods and even competes on par with supervised ones. Remarkably, in the challenging Booster benchmark, SMFormer even outperforms some SOTA supervised methods, such as CFNet. Yun Wang 0053, Zhengjie Yang, Jiahao Zheng 0001, Zhanjie Zhang, Dapeng Oliver Wu, Yulan Guo |
IEEE Trans. Image Process. | 4 |
| 2026 | Fast and Robust Deformable 3D Gaussian Splattingabstract3D Gaussian Splatting has demonstrated remarkable real-time rendering capabilities and superior visual quality in novel view synthesis for static scenes. Building upon these advantages, researchers have progressively extended 3D Gaussians to dynamic scene reconstruction. Deformation field-based methods have emerged as a promising approach among various techniques. These methods maintain 3D Gaussian attributes in a canonical field and employ the deformation field to transform this field across temporal sequences. Nevertheless, these approaches frequently encounter challenges such as suboptimal rendering speeds, significant dependence on initial point clouds, and vulnerability to local optima in dim scenes. To overcome these limitations, we present FRoG, an efficient and robust framework for high-quality dynamic scene reconstruction. FRoG integrates per-Gaussian embedding with a coarse-to-fine temporal embedding strategy, accelerating rendering through the early fusion of temporal embeddings. Moreover, to enhance robustness against sparse initializations, we introduce a novel depth- and error-guided sampling strategy. This strategy populates the canonical field with new 3D Gaussians at low-deviation initial positions, significantly reducing the optimization burden on the deformation field and improving detail reconstruction in both static and dynamic regions. Furthermore, by modulating opacity variations, we mitigate the local optima problem in dim scenes, improving color fidelity. Comprehensive experimental results validate that our method achieves accelerated rendering speeds while maintaining state-of-the-art visual quality. Han Jiao 0005, Jiakai Sun, Lei Zhao 0011, Zhanjie Zhang, Wei Xing 0001, Huaizhong Lin |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | DualNet: Robust Self-Supervised Stereo Matching with Pseudo-Label SupervisionabstractSelf-supervised stereo matching has drawn attention due to its ability to estimate disparity without needing ground-truth data. However, existing self-supervised stereo matching methods heavily rely on the photo-metric consistency assumption, which is vulnerable to natural disturbances, resulting in ambiguous supervision and inferior performance compared to the supervised ones. To relax the limitation of the photo-metric consistency assumption and even bypass this assumption, we propose a novel self-supervised framework named DualNet, which consists of two key steps: robust self-supervised teacher learning and pseudo-label supervised student training. Specifically, the teacher model is first trained in a self-supervised manner with a focus on feature-metric consistency and data augmentation consistency. Then, the output of the teacher model is geometrically constrained to obtain high-quality pseudo labels. Benefiting from these high-quality pseudo labels, the student model can outperform its teacher model by a large margin. With the two well-designed steps, the proposed framework DualNet ranks 1st among all self-supervised methods on multiple benchmarks, surprisingly even outperforming several supervised counterparts. Yun Wang 0053, Jiahao Zheng 0001, Chenghao Zhang 0003, Zhanjie Zhang, Kunhong Li 0001, Junjie Hu 0003 |
AAAI | 4 |
| 2025 | Complementary-Disentangled Neural Generalization: a Robust Framework for Stable Brain-Computer InterfacesabstractBrain-computer interfaces (BCIs) enable communication between the brain and the external environment, showing significant potential in restoration, rehabilitation and movement enhancement. However, neural drift causes BCI performance to degrade substantially over time, compromising their long-term reliability. A fundamental limitation of current methods is their failure to account for a key insight from neural preference theory that the magnitude of neural drift depends on specific motor parameters (e.g., velocity, direction, and speed), ultimately compromising performance. To overcome the limitation, we introduce a novel framework named ComplementaryDisentangled Neural Generalization (CDNG) inspired by neural preference theory. Specifically, we first conduct pre-experiments about the neural decoding preference, revealing that neural drifts differ across velocity, speed and direction. Then we adopt CDNG, which captures invariant neural representations through an ensemble of three specialized neural decoders after disentangling velocity into speed and direction. Extensive experiments on several datasets demonstrate that our method achieves state-of-the-art performance and significantly enhances cross-day generalization. Jiyu Wei, Dazhong Rong, Di Hong, Zhanjie Zhang, Xinyun Zhu, Qinming He, Yueming Wang 0001 |
BIBM | 4 |
| 2025 | Lay2Story: Extending Diffusion Transformers for Layout-Togglable Story Generation
Ao Ma 0005, Jiasong Feng, Ke Cao 0001, Jing Wang 0021, Yun Wang 0053, Quanwei Zhang, Zhanjie Zhang |
ICCV | 7 |
| 2025 | Learning Robust Stereo Matching in the Wild with Selective Mixture-of-ExpertsabstractRecently, learning-based stereo matching networks have advanced significantly. However, they often lack robustness and struggle to achieve impressive cross-domain performance due to domain shifts and imbalanced disparity distributions among diverse datasets. Leveraging Vision Foundation Models (VFMs) can intuitively enhance the model's robustness, but integrating such a model into stereo matching cost-effectively to fully realize their robustness remains a key challenge. To address this, we propose SMoEStereo, a novel framework that adapts VFMs for stereo matching through a tailored, scene-specific fusion of Low-Rank Adaptation (LoRA) and Mixture-of-Experts (MoE) modules. SMoEStereo introduces MoE-LoRA with adaptive ranks and MoE-Adapter with adaptive kernel sizes. The former dynamically selects optimal experts within MoE to adapt varying scenes across domains, while the latter injects inductive bias into frozen VFMs to improve geometric feature extraction. Importantly, to mitigate computational overhead, we further propose a lightweight decision network that selectively activates MoE modules based on input complexity, balancing efficiency with accuracy. Extensive experiments demonstrate that our method exhibits state-of-the-art cross-domain and joint generalization across multiple benchmarks without dataset-specific adaptation. The code is available at \textcolor{red}{https://github.com/cocowy1/SMoE-Stereo}. Yun Wang 0053, Longguang Wang, Chenghao Zhang 0003, Zhanjie Zhang, Ao Ma 0005, Chenyou Fan, Tin Lun Lam, Junjie Hu 0003 |
ICCV | 5 |
| 2025 | Unpaired Image Style Translations Using Mamba Adversarial Networks
Zhou Hong, Zhanjie Zhang, Juqin Wang, Yanzhao Shan, Jingwen Yu, Qingxia Chen |
ICIC (10) | 3 |
| 2025 | FancyVideo: Towards Dynamic and Consistent Video Generation via Cross-frame Textual GuidanceabstractSynthesizing motion-rich and temporally consistent videos remains a challenge in artificial intelligence, especially when dealing with extended durations. Existing text-to-video (T2V) models commonly employ spatial cross-attention for text control, equivalently guiding different frame generations without frame-specific textual guidance. Thus, the model's capacity to comprehend the temporal logic conveyed in prompts and generate videos with coherent motion is restricted. To tackle this limitation, we introduce FancyVideo, an innovative video generator that improves the existing text-control mechanism with the well-designed Cross-frame Textual Guidance Module (CTGM). Specifically, CTGM incorporates the Temporal Information Injector (TII) and Temporal Affinity Refiner (TAR) at the beginning and end of cross-attention, respectively, to achieve frame-specific textual guidance. Firstly, TII injects frame-specific information from latent features into text conditions, thereby obtaining cross-frame textual conditions. Then, TAR refines the correlation matrix between cross-frame textual conditions and latent features along the time dimension. Extensive experiments comprising both quantitative and qualitative evaluations demonstrate the effectiveness of FancyVideo. Our approach achieves state-of-the-art T2V generation results on the EvalCrafter benchmark and facilitates the synthesis of dynamic and consistent videos. Note that the T2V process of FancyVideo essentially involves a text-to-image step followed by T+I2V. This means it also supports the generation of videos from user images, i.e., the image-to-video (I2V) task. A significant number of experiments have shown that its performance is also outstanding. Jiasong Feng, Ao Ma 0005, Jing Wang 0021, Ke Cao 0001, Zhanjie Zhang |
IJCAI | 5 |
| 2025 | WISA: World simulator assistant for physics-aware text-to-video generationabstractRecent advances in text-to-video (T2V) generation, exemplified by models such as Sora and Kling, have demonstrated strong potential for constructing world simulators. However, existing T2V models still struggle to understand abstract physical principles and to generate videos that faithfully obey physical laws. This limitation stems primarily from the lack of explicit physical guidance, caused by a significant gap between high-level physical concepts and the generative capabilities of current models. To address this challenge, we propose the **W**orld **S**imulator **A**ssistant (**WISA**), a novel framework designed to systematically decompose and integrate physical principles into T2V models. Specifically, WISA decomposes physical knowledge into three hierarchical levels: textual physical descriptions, qualitative physical categories, and quantitative physical properties. It then incorporates several carefully designed modules—such as Mixture-of-Physical-Experts Attention (MoPA) and a Physical Classifier—to effectively encode these attributes and enhance the model’s adherence to physical laws during generation. In addition, most existing video datasets feature only weak or implicit representations of physical phenomena, limiting their utility for learning explicit physical principles. To bridge this gap, we present **WISA-80K**, a new dataset comprising 80,000 human-curated videos that depict 17 fundamental physical laws across three core domains of physics: dynamics, thermodynamics, and optics. Experimental results show that WISA substantially improves the alignment of T2V models (such as CogVideoX and Wan2.1) with real-world physical laws, achieving notable gains on the VideoPhy benchmark. Our data, code, and models are available in the [Project Page](https://wisav1.github.io/WISA/). Jing Wang 0021, Ao Ma 0005, Ke Cao 0001, Jiasong Feng, Zhanjie Zhang, Wanyuan Pang, Xiaodan Liang |
NeurIPS | 6 |
| 2025 | Personalized text-to-image generation with Large Language and Vision Assistant enhanced training
Junsheng Luan, Zhanjie Zhang, Wei Xing 0001, Lei Zhao 0011 |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | VectorSketcher: Learning to create a vector-based free-hand sketch
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | LGAST: Towards high-quality arbitrary style transfer with local-global style learning
Zhanjie Zhang, Ruichen Xia 0002, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011, Wei Xing 0001 |
Neurocomputing | 1 |
| 2025 | DyArtbank: Diverse artistic style transfer via pre-trained stable diffusion and dynamic style prompt Artbank
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Knowl. Based Syst. | 1 |
| 2025 | SPAST: Arbitrary style transfer with style priors via pre-trained large-scale model
Zhanjie Zhang, Quanwei Zhang, Junsheng Luan, Mengyuan Yang 0002, Yun Wang 0053, Lei Zhao 0011 |
Neural Networks | 1 |
| 2024 | ArtBank: Artistic Style Transfer with Pre-trained Diffusion Model and Implicit Style Prompt BankabstractArtistic style transfer aims to repaint the content image with the learned artistic style. Existing artistic style transfer methods can be divided into two categories: small model-based approaches and pre-trained large-scale model-based approaches. Small model-based approaches can preserve the content strucuture, but fail to produce highly realistic stylized images and introduce artifacts and disharmonious patterns; Pre-trained large-scale model-based approaches can generate highly realistic stylized images but struggle with preserving the content structure. To address the above issues, we propose ArtBank, a novel artistic style transfer framework, to generate highly realistic stylized images while preserving the content structure of the content images. Specifically, to sufficiently dig out the knowledge embedded in pre-trained large-scale models, an Implicit Style Prompt Bank (ISPB), a set of trainable parameter matrices, is designed to learn and store knowledge from the collection of artworks and behave as a visual prompt to guide pre-trained large-scale models to generate highly realistic stylized images while preserving content structure. Besides, to accelerate training the above ISPB, we propose a novel Spatial-Statistical-based self-Attention Module (SSAM). The qualitative and quantitative experiments demonstrate the superiority of our proposed method over state-of-the-art artistic style transfer methods. Code is available at https://github.com/Jamie-Cheung/ArtBank. Zhanjie Zhang, Quanwei Zhang, Wei Xing 0001, Lei Zhao 0011, Jiakai Sun, Zehua Lan, Junsheng Luan, Huaizhong Lin |
AAAI | 1 |
| 2024 | Rethinking Diffusion Model for Multi-Contrast MRI Super-ResolutionabstractRecently, diffusion models (DM) have been applied in magnetic resonance imaging (MRI) super-resolution (SR) reconstruction, exhibiting impressive performance, especially with regard to detailed reconstruction. However, the current DM-based SR reconstruction methods still face the following issues: (1) They require a large number of iterations to reconstruct the final image, which is inefficient and consumes a significant amount of computational re-sources. (2) The results reconstructed by these methods are often misaligned with the real high-resolution images, leading to remarkable distortion in the reconstructed MR images. To address the aforementioned issues, we propose an efficient diffusion model for multi-contrast MRI SR, named as DiffMSR. Specifically, we apply DM in a highly compact low-dimensional latent space to generate prior knowledge with high-frequency detail information. The highly compact latent space ensures that DM requires only a few simple iterations to produce accurate prior knowledge. In addition, we design the Prior-Guide Large Window Trans-former (PLWformer) as the decoder for DM, which can ex-tend the receptive field while fully utilizing the prior knowledge generated by DM to ensure that the reconstructed MR image remains undistorted. Extensive experiments on public and clinical datasets demonstrate that our DiffMSR11Code: https://github.com/GuangYuanKK/DiffMSR outperforms state-of-the-art methods. Chen Rao, Juncheng Mo, Zhanjie Zhang, Wei Xing 0001, Lei Zhao 0011 |
CVPR | 4 |
| 2024 | 3DGStream: On-the-Fly Training of 3D Gaussians for Efficient Streaming of Photo-Realistic Free-Viewpoint VideosabstractConstructing photo-realistic Free-Viewpoint Videos (FVVs) of dynamic scenes from multi-view videos remains a challenging endeavor. Despite the remarkable advance- ments achieved by current neural rendering techniques, these methods generally require complete video sequences for offline training and are not capable of real-time rendering. To address these constraints, we introduce 3DGStream, a method designed for efficient FVV streaming of real-world dynamic scenes. Our method achieves fast on-the-fly per- frame reconstruction within 12 seconds and real-time ren- dering at 200 FPS. Specifically, we utilize 3D Gaussians (3DGs) to represent the scene. Instead of the naï ve ap- proach of directly optimizing 3DGs per-frame, we employ a compact Neural Transformation Cache (NTC) to model the translations and rotations of 3DGs, markedly reducing the training time and storage required for each FVV frame. Furthermore, we propose an adaptive 3DG addition strat- egy to handle emerging objects in dynamic scenes. Exper- iments demonstrate that 3DGStream achieves competitive performance in terms of rendering speed, image quality, training time, and model storage when compared with state- of-the-art methods. Jiakai Sun, Han Jiao 0005, Zhanjie Zhang, Lei Zhao 0011, Wei Xing 0001 |
CVPR | 4 |
| 2024 | Towards Highly Realistic Artistic Style Transfer via Stable Diffusion with Step-aware and Layer-aware Prompt
Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin, Wei Xing 0001, Juncheng Mo, Shuaicheng Huang, Jinheng Xie, Junsheng Luan, Lei Zhao 0011, Dalong Zhang, Lixia Chen |
IJCAI | 1 |
| 2024 | Rethink arbitrary style transfer with transformer and contrastive learning
Zhanjie Zhang, Jiakai Sun, Lei Zhao 0011, Quanwei Zhang, Zehua Lan, Haolin Yin, Huaizhong Lin, Zhiwen Zuo |
Comput. Vis. Image Underst. | 1 |
| 2023 | Generative Image Inpainting with Segmentation Confusion Adversarial Training and Contrastive LearningabstractThis paper presents a new adversarial training framework for image inpainting with segmentation confusion adversarial training (SCAT) and contrastive learning. SCAT plays an adversarial game between an inpainting generator and a segmentation network, which provides pixel-level local training signals and can adapt to images with free-form holes. By combining SCAT with standard global adversarial training, the new adversarial training framework exhibits the following three advantages simultaneously: (1) the global consistency of the repaired image, (2) the local fine texture details of the repaired image, and (3) the flexibility of handling images with free-form holes. Moreover, we propose the textural and semantic contrastive learning losses to stabilize and improve our inpainting model's training by exploiting the feature representation space of the discriminator, in which the inpainting images are pulled closer to the ground truth images but pushed farther from the corrupted images. The proposed contrastive losses better guide the repaired images to move from the corrupted image data points to the real image data points in the feature representation space, resulting in more realistic completed images. We conduct extensive experiments on two benchmark datasets, demonstrating our model's effectiveness and superiority both qualitatively and quantitatively. Zhiwen Zuo, Lei Zhao 0011, Ailin Li, Zhizhong Wang, Zhanjie Zhang, Jiafu Chen, Wei Xing 0001, Dongming Lu |
AAAI | 5 |
| 2023 | Rethinking Multi-Contrast MRI Super-Resolution: Rectangle-Window Cross-Attention Transformer and Arbitrary-Scale UpsamplingabstractRecently, several methods have explored the potential of multi-contrast magnetic resonance imaging (MRI) super-resolution (SR) and obtain results superior to single-contrast SR methods. However, existing approaches still have two shortcomings: (1) They can only address fixed integer upsampling scales, such as 2×, 3×, and 4×, which require training and storing the corresponding model separately for each upsampling scale in clinic. (2) They lack direct interaction among different windows as they adopt the square window (e.g., 8×8) transformer network architecture, which results in inadequate modelling of longer-range dependencies. Moreover, the relationship between reference images and target images is not fully mined. To address these issues, we develop a novel network for multi-contrast MRI arbitrary-scale SR, dubbed as McASSR. Specifically, we design a rectangle-window cross-attention transformer to establish longer-range dependencies in MR images without increasing computational complexity and fully use reference information. Besides, we propose the reference-aware implicit attention as an upsampling module, achieving arbitrary-scale super-resolution via implicit neural representation, further fusing supplementary information of the reference image. Extensive and comprehensive experiments on both public and clinical datasets show that our McASSR yields superior performance over SOTA methods, demonstrating its great potential to be applied in clinical practice. Code will be available at https://github.com/GuangYuanKK/McASSR. Lei Zhao 0011, Jiakai Sun, Zehua Lan, Zhanjie Zhang, Jiafu Chen, Huaizhong Lin, Wei Xing 0001 |
ICCV | 5 |
| 2023 | TeSTNeRF: Text-Driven 3D Style Transfer via Cross-Modal LearningabstractText-driven 3D style transfer aims at stylizing a scene according to the text and generating arbitrary novel views with consistency. Simply combining image/video style transfer methods and novel view synthesis methods results in flickering when changing viewpoints, while existing 3D style transfer methods learn styles from images instead of texts. To address this problem, we for the first time design an efficient text-driven model for 3D style transfer, named TeSTNeRF, to stylize the scene using texts via cross-modal learning: we leverage an advanced text encoder to embed the texts in order to control 3D style transfer and align the input text and output stylized images in latent space. Furthermore, to obtain better visual results, we introduce style supervision, learning feature statistics from style images and utilizing 2D stylization results to rectify abrupt color spill. Extensive experiments demonstrate that TeSTNeRF significantly outperforms existing methods and provides a new way to guide 3D style transfer. Jiafu Chen, Boyan Ji, Zhanjie Zhang, Tianyi Chu, Zhiwen Zuo, Lei Zhao 0011, Wei Xing 0001, Dongming Lu |
IJCAI | 3 |
| 2023 | VGOS: Voxel Grid Optimization for View Synthesis from Sparse InputsabstractNeural Radiance Fields (NeRF) has shown great success in novel view synthesis due to its state-of-the-art quality and flexibility. However, NeRF requires dense input views (tens to hundreds) and a long training time (hours to days) for a single scene to generate high-fidelity images. Although using the voxel grids to represent the radiance field can significantly accelerate the optimization process, we observe that for sparse inputs, the voxel grids are more prone to overfitting to the training views and will have holes and floaters, which leads to artifacts. In this paper, we propose VGOS, an approach for fast (3-5 minutes) radiance field reconstruction from sparse inputs (3-10 views) to address these issues. To improve the performance of voxel-based radiance field in sparse input scenarios, we propose two methods: (a) We introduce an incremental voxel training strategy, which prevents overfitting by suppressing the optimization of peripheral voxels in the early stage of reconstruction. (b) We use several regularization techniques to smooth the voxels, which avoids degenerate solutions. Experiments demonstrate that VGOS achieves state-of-the-art performance for sparse inputs with super-fast convergence. Code will be available at https://github.com/SJoJoK/VGOS. Jiakai Sun, Zhanjie Zhang, Jiafu Chen, Boyan Ji, Lei Zhao 0011, Wei Xing 0001 |
IJCAI | 2 |
| 2023 | Self-Reference Image Super-Resolution via Pre-trained Diffusion Large Model and Window Adjustable TransformerabstractCurrently, reference-based super-resolution (RefSR) techniques leverage high-resolution (HR) reference images to provide useful content and texture information for low-resolution (LR) images during the super-resolution (SR) process. Nevertheless, it is time-consuming, laborious, and even impossible in some cases to find high-quality reference images. To tackle this problem, we propose a brand-new self-reference image super-resolution approach using a pre-trained diffusion large model and a window adjustable transformer, termed DWTrans. Our proposed method does not require explicitly inputting manually acquired reference images during training and inference. Specifically, we feed the degraded LR images into a pre-trained stable diffusion large model to automatically generate corresponding high-quality self-reference (SRef) images that provide valuable high-frequency details for the LR images in the process of SR. To extract valuable high-frequency information in SRef images, we design a window adjustable transformer with both non-adjustable window layer (NWL) and adjustable window layer (AWL). The NWL learns local features from LR images using a dense window, while the AWL acquires global features from the SRef images using a random sparse window. Furthermore, to fully utilize the high-frequency features in the SRef image, we introduce the adaptive deformable fusion module to adaptively fuse the features of the LR and SRef images. Experimental results validate that our proposed DWTrans outperforms state-of-the-art methods on various benchmark datasets both quantitatively and visually. Wei Xing 0001, Lei Zhao 0011, Zehua Lan, Jiakai Sun, Zhanjie Zhang, Quanwei Zhang, Huaizhong Lin |
ACM Multimedia | 6 |
| 2023 | DuDoINet: Dual-Domain Implicit Network for Multi-Modality MR Image Arbitrary-scale Super-ResolutionabstractCompared to single-modality magnetic resonance (MR) image super-resolution (SR) methods, multi-modality MR image methods can utilize high-resolution reference modality (e.g., T1 modality) to provide valuable complementary information for low-resolution target modality (e.g., T2 modality) in SR reconstruction, which can further improve the quality of the SR images. Although they have achieved impressive results, these methods still suffer from the following drawbacks: (1) They can only handle fixed integer upsampling factors, such as 2X, 3X, and 4X, and require training and storing corresponding models for each upsampling factor, which is infeasible in clinical practice; (2) They only perform feature extraction and reconstruction in the image domain. However, the aliasing artifacts produced in the image domain are structural and non-local. Therefore, using only the image domain cannot effectively reconstruct high-quality aliasing-free SR images. To address these issues, we develop a brand-new Dual-Domain Implicit Network (DuDoINet) for multi-modality MR image arbitrary-scale SR. Specifically, we propose a dual-domain learning scheme for multi-modality MR image SR, which allows the network to sufficiently exploit the frequency and image domain information in MR images. In addition, we design implicit attention to achieve arbitrary-scale upsampling of MR images, which utilizes a continuously differentiable function that generates pixel values from pixel coordinates. Furthermore, we designed a deformable cross-modality attention mechanism that can adaptively transfer high-frequency details from the T1 to the T2 modality, better integrating valuable complementary information from the T1 modality. Extensive and comprehensive experiments on healthy subjects and patient datasets demonstrate that our DuDoINet outperforms SOTA methods, demonstrating its great potential for clinical practice. Wei Xing 0001, Lei Zhao 0011, Zehua Lan, Zhanjie Zhang, Jiakai Sun, Haolin Yin, Huaizhong Lin |
ACM Multimedia | 5 |
| 2023 | Caster: Cartoon style transfer via dynamic cartoon style casting
Zhanjie Zhang, Jiakai Sun, Jiafu Chen, Lei Zhao 0011, Boyan Ji, Zehua Lan, Wei Xing 0001, Duanqing Xu |
Neurocomputing | 1 |
| 2022 | AesUST: Towards Aesthetic-Enhanced Universal Style TransferabstractRecent studies have shown remarkable success in universal style transfer which transfers arbitrary visual styles to content images. However, existing approaches suffer from the aesthetic-unrealistic problem that introduces disharmonious patterns and evident artifacts, making the results easy to spot from real paintings. To address this limitation, we propose AesUST, a novel Aesthetic-enhanced Universal Style Transfer approach that can generate aesthetically more realistic and pleasing results for arbitrary styles. Specifically, our approach introduces an aesthetic discriminator to learn the universal human-delightful aesthetic features from a large corpus of artist-created paintings. Then, the aesthetic features are incorporated to enhance the style transfer process via a novel Aesthetic-aware Style-Attention (AesSA) module. Such an AesSA module enables our AesUST to efficiently and flexibly integrate the style patterns according to the global aesthetic channel distribution of the style image and the local semantic spatial distribution of the content image. Moreover, we also develop a new two-stage transfer training strategy with two aesthetic regularizations to train our model more effectively, further improving stylization performance. Extensive experiments and user studies demonstrate that our approach synthesizes aesthetically more harmonious and realistic results than state of the art, greatly narrowing the disparity with real artist-created paintings. Our code is available at https://github.com/EndyWon/AesUST. Zhizhong Wang, Zhanjie Zhang, Lei Zhao 0011, Zhiwen Zuo, Ailin Li, Wei Xing 0001, Dongming Lu |
ACM Multimedia | 2 |