Yuechen Zhang

dblp:298/8473 · DBLP profile ↗
← Back
27ranked-venue papers
7as first author
27since 2021 · last 2026
0009-0000-9112-0216ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 5 first-author · 15 since 2021Graphics, computer vision, multimedia, augmented reality and games · 14 · 3 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 5 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 FWMamba-UNet: Frequency-Wavelet Enhanced Mamba UNet for Medical Image Segmentation
Zehan Zhao, Hongyan Yin, Yuechen Zhang
ICIC (12)6
2026 Extract Then Compile: Reliable Neuro-Symbolic Planning for Large Language Models
Yuechen Zhang, Hongyan Yin
ICIC (22)3
2026 Mini-Gemini: Mining the Potential of Multi-Modality Vision Language Models
abstract
In this work, we introduce Mini-Gemini, a simple and effective framework enhancing multi-modality Vision Language Models (VLMs). Despite the advancements in VLMs facilitating basic visual dialog and reasoning, a performance gap persists compared to advanced models like GPT-4 and Gemini. We propose a novel approach to narrow the gap by mining the potential of VLMs for better performance across various cross-modal tasks. It tackles the following questions: (1) How can high-resolution visual tokens improve image understanding without lengthening the token sequence? (2) How to improve reasoning and generation abilities of VLM with high-quality data? (3) How to close the gap between open-source VLMs and proprietary models on reasoning-driven generation? In particular, to enhance visual tokens, we propose to utilize an additional visual encoder for high-resolution refinement without increasing the visual token count. We further construct a high-quality dataset that promotes precise image comprehension and reasoning-based generation, expanding the operational scope of current VLMs. In general, Mini-Gemini further mines the potential of VLMs and empowers current frameworks with image understanding, reasoning, and generation simultaneously. The proposed model supports a series of dense and MoE Large Language Models (LLMs) from 2B to 34B, which achieve leading performance in several zero-shot benchmarks and even surpasses the developed private models. It is demonstrated to attain 80.6% accuracy on the MMB benchmark (+5.4 vs Gemini Pro) and 74.1% on TextVQA (+4.6 vs LLaVA-NeXT), achieving leading performance in several zero-shot benchmarks and even surpasses the developed private models. Furthermore, Mini-Gemini is proven to improve consistently with stronger LLM, visual encoder, and data in experiments.
Yuechen Zhang, Chengyao Wang, Zhisheng Zhong, Ruihang Chu, Shaoteng Liu, Jiaya Jia
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 ResMaster: Mastering High-Resolution Image Generation via Structural and Fine-Grained Guidance
abstract
Diffusion models excel at producing high-quality images; however, scaling to higher resolutions, such as 4K, often results in structural distortions, and repetitive patterns. To this end, we introduce ResMaster, a novel, training-free method that empowers resolution-limited diffusion models to generate high-quality images beyond resolution restrictions. Specifically, ResMaster leverages a low-resolution reference image created by a pre-trained diffusion model to provide structural and fine-grained guidance for crafting high-resolution images on a patch-by-patch basis. To ensure a coherent structure, ResMaster meticulously aligns the low-frequency components of high-resolution patches with the low-resolution reference at each denoising step. For fine-grained guidance, tailored image prompts based on the low-resolution reference and enriched textual prompts produced by a vision-language model are incorporated. This approach could significantly mitigate local pattern distortions and improve detail refinement. Extensive experiments validate that ResMaster sets a new benchmark for high-resolution image generation.
Shuwei Shi, Yuechen Zhang, Jingwen He, Biao Gong, Yinqiang Zheng
AAAI3
2025 DreamOmni: Unified Image Generation and Editing
abstract
projectpagepCurrently, the success of large language models (LLMs) illustrates that a unified multitasking approach can significantly enhance model usability, stream-line deployment, and foster synergistic benefits across different tasks. However, in computer vision, while text-to-image (T2I) models have significantly improved generation quality through scaling up, their framework design did not initially consider how to unify with downstream tasks, such as various types of editing. To address this, we introduce DreamOmni, a unified model for image generation and editing. We begin by analyzing existing frameworks and the requirements of downstream tasks, proposing a unified framework that integrates both T2I models and various editing tasks. Furthermore, another key challenge is the efficient creation of high-quality editing data, particularly for instruction-based and drag-based editing. To this end, we develop a synthetic data pipeline using sticker-like elements to synthesize accurate, high-quality datasets efficiently, which enables editing data scaling up for unified model training. For training, DreamOmni jointly trains T2I generation and downstream tasks. T2I training enhances the model’s understanding of specific concepts and improves generation quality, while editing training helps the model grasp the nuances of the editing task. This collaboration significantly boosts editing performance. Extensive experiments confirm the effectiveness of DreamOmni. The code and model will be released.
Bin Xia 0014, Yuechen Zhang, Jingyao Li 0001, Chengyao Wang, Bei Yu 0001, Jiaya Jia
CVPR2
2025 MagicMirror: ID-Preserved Video Generation in Video Diffusion Transformers
abstract
We present MagicMirror, a framework for generating identity-preserved videos with cinematic-level quality and dynamic motion. While recent advances in video diffusion models have shown impressive capabilities in text-to-video generation, maintaining consistent identity while producing natural motion remains challenging. Previous methods either require person-specific fine-tuning or struggle to balance identity preservation with motion diversity. Built upon Video Diffusion Transformers, our method introduces three key components: (1) a dual-branch facial feature extractor that captures both identity and structural features, (2) a lightweight cross-modal adapter with Conditioned Adaptive Normalization for efficient identity integration, and (3) a two-stage training strategy combining synthetic identity pairs with video data. Extensive experiments demonstrate that MagicMirror effectively balances identity consistency with natural motion, outperforming existing methods across multiple metrics while requiring minimal parameters added. The code and model will be made publicly available.
Yuechen Zhang, Yaoyang Liu, Bin Xia 0014, Bohao Peng, Zexin Yan, Eric Lo 0001, Jiaya Jia
ICCV1
2025 LYRA: An Efficient and Speech-Centric Framework for Omni-Cognition
abstract
As Multi-modal Large Language Models (MLLMs) evolve, expanding beyond single-domain capabilities is essential to meet the demands for more versatile and efficient AI. However, previous omni-models have insufficiently explored speech, neglecting its integration with multi-modality. We introduce Lyra, an efficient MLLM that enhances multimodal abilities, including advanced long-speech comprehension, sound understanding, cross-modality efficiency, and seamless speech interaction. To achieve efficiency and speech-centric capabilities, Lyra employs three strategies: (1) leveraging existing open-source large models and a proposed multi-modality LoRA to reduce training costs and data requirements; (2) using a latent multi-modality regularizer and extractor to strengthen the relationship between speech and other modalities, thereby enhancing model performance; and (3) constructing a high-quality, extensive dataset that includes 1.5M multi-modal (language, vision, audio) data samples and 12K long speech samples, enabling Lyra to handle complex long speech inputs and achieve more robust omni-cognition. Compared to other omni-methods, Lyra achieves state-of-the-art performance on various vision-language, vision-speech, and speech-language benchmarks, while also using fewer computational resources and less training data.
Zhisheng Zhong, Chengyao Wang, Yuqi Liu 0003, Senqiao Yang, Longxiang Tang, Yuechen Zhang, Jingyao Li 0001, Tianyuan Qu, Yukang Chen, Shaozuo Yu, Sitong Wu, Eric Lo 0001, Shu Liu 0005, Jiaya Jia
ICCV6
2025 Training-Free Efficient Video Generation via Dynamic Token Carving
abstract
Despite the remarkable generation quality of video Diffusion Transformer (DiT) models, their practical deployment is severely hindered by extensive computational requirements. This inefficiency stems from two key challenges: the quadratic complexity of self-attention with respect to token length and the multi-step nature of diffusion models. To address these limitations, we present Jenga, a novel inference pipeline that combines dynamic attention carving with progressive resolution generation. Our approach leverages two key insights: (1) early denoising steps do not require high-resolution latents, and (2) later steps do not require dense attention. Jenga introduces a block-wise attention mechanism that dynamically selects relevant token interactions using 3D space-filling curves, alongside a progressive resolution strategy that gradually increases latent resolution during generation. Experimental results demonstrate that Jenga achieves substantial speedups across multiple state-of-the-art video diffusion models while maintaining comparable generation quality (8.83$\times$ speedup with 0.01\% performance drop on VBench). As a plug-and-play solution, Jenga enables practical, high-quality video generation on modern hardware by reducing inference time from minutes to seconds---without requiring model retraining.
Yuechen Zhang, Jinbo Xing, Bin Xia 0014, Shaoteng Liu, Bohao Peng, Xin Tao 0001, Pengfei Wan 0001, Eric Lo 0001, Jiaya Jia
NeurIPS1
2025 MambaGuard: A CLIP-Mamba Approach for OOD Generated Image Detection
Xinchang Wang, Yuechen Zhang, Wenyao Qiu, Chunyang Cheng
PRCV (4)2
2025 Make-Your-Video: Customized Video Generation Using Textual and Structural Guidance
abstract
Creating a vivid video from the event or scenario in our imagination is a truly fascinating experience. Recent advancements in text-to-video synthesis have unveiled the potential to achieve this with prompts only. While text is convenient in conveying the overall scene context, it may be insufficient to control precisely. In this paper, we explore customized video generation by utilizing text as context description and motion structure (e.g., frame-wise depth) as concrete guidance. Our method, dubbed Make-Your-Video, involves joint-conditional video generation using a Latent Diffusion Model that is pre-trained for still image synthesis and then promoted for video generation with the introduction of temporal modules. This two-stage learning scheme not only reduces the computing resources required, but also improves the performance by transferring the rich concepts available in image datasets solely into video generation. Moreover, we use a simple yet effective causal attention mask strategy to enable longer video synthesis, which mitigates the potential quality degradation effectively. Experimental results show the superiority of our method over existing baselines, particularly in terms of temporal coherence and fidelity to users' guidance. In addition, our model enables several intriguing applications that demonstrate potential for practical usage.
Jinbo Xing, Menghan Xia, Yuechen Zhang, Yong Zhang 0034, Yingqing He, Hanyuan Liu, Haoxin Chen, Xiaodong Cun, Xintao Wang 0002, Ying Shan, Tien-Tsin Wong
IEEE Trans. Vis. Comput. Graph.4
2024 Progressively Knowledge Distillation via Re-parameterizing Diffusion Reverse Process
abstract
Knowledge distillation aims at transferring knowledge from the teacher model to the student one by aligning their distributions. Feature-level distillation often uses L2 distance or its variants as the loss function, based on the assumption that outputs follow normal distributions. This poses a significant challenge when distribution gaps are substantial since this loss function ignores the variance term. To address the problem, we propose to decompose the transfer objective into small parts and optimize it progressively. This process is inspired by diffusion models from which the noise distribution is mapped to the target distribution step by step. However, directly employing diffusion models is impractical in the distillation scenario due to its heavy reverse process. To overcome this challenge, we adopt the structural re-parameterization technique to generate multiple student features to approximate the teacher features sequentially. The multiple student features are combined linearly in inference time without extra cost. We present extensive experiments performed on various transfer scenarios, such as CNN-to-CNN and Transformer-to-CNN, that validate the effectiveness of our approach.
Xufeng Yao, Fanbin Lu, Yuechen Zhang, Xinyun Zhang 0001, Wenqian Zhao 0002, Bei Yu 0001
AAAI3
2024 MF-YOLO: Multimodal Fusion for Remote Sensing Object Detection Based on YOLOv5s
abstract
Remote sensing object detection has flourished at a rapid rate of development. In remote sensing images (RSI), most algorithms perform well in detecting medium and large objects but perform poorly on small objects. This is because small objects occupy fewer pixels in the entire image, and the detector is more likely to be biased towards detecting large objects. In addition, the huge background in the image also introduces noise, causing many detection failures. Most approaches solve the problem of small object detection by increasing the depth of the neural network, but the results are not satisfactory. To solve the above problems, we propose a new method, MF-YOLO, that utilizes fused infrared (IR) and red-green-blue (RGB) images for remote sensing object detection. Through multimodal fusion (MF), we can obtain more positive information, thereby improving detection accuracy. In addition, we introduced the Bi-Level Routing Attention (BRA) module to improve the YOLOv5s model structure, and proposed a new loss function to enhance the model’s learning discrimination ability. Experimental results show that the accuracy of MF-YOLO on the VEDAI and NWPU VHR-10 datasets is 76.62% and 91.63% respectively (in terms of [email protected]). Compared with the existing state-of-the-art, the detection accuracy is significantly improved. In addition, our model has fewer weight parameters than YOLOv5s, allowing it to achieve relatively high near-real-time detection speed.
Xiaotong Kong, Yuechen Zhang
CSCWD4
2024 FSD-YOLO: An Improved Method for Steel Surface Defect Detection Based on YOLOv5
abstract
Steel surface defect detection is very important in industrial quality inspection, and there have been many studies on Steel surface defect detection in recent years. Existing steel surface defect detection algorithms suffer from low detection accuracy and high model complexity. To improve the above problems, this paper proposes an optimized target detection algorithm based on YOLOv5. Firstly, the original spatial pyramid pooling (SPP) module is replaced by the SPPFCSPC module to better capture the target and scene information at different scales, and to improve the sensory field and feature expression ability of the model. Secondly, we propose the C2F-Faster module based on FasterNet and C2F module to replace the C3 module of the backbone network, which perfectly integrates the ideas of PConv and ELAN and ensures the detection accuracy while effectively reducing the model size. Finally, the Dynamic Head detection head is fused in the head of the model to improve the detection capability of the target detection head using the separated attention mechanism. We conducted experimental validation on the widely used NEU-DET dataset. The experimental results show that the improved model improves the original YOLOv5 model over the original YOLOv5 model by 3.6% in mAP@50, 9.9% and 2.8% in AP and AR, respectively, and the amount of model parameters is reduced by 13.2%. The improved model also outperforms SSD, FasterRCNN, YOLOv5, YOLOv6, YOLOv7, YOLOv8 and other mainstream target detection models, which better meets the requirements of actual industrial production on the accuracy and speed of steel surface defects detection model.
Yuechen Zhang, Xiaotong Kong
CSCWD1
2024 Video-P2P: Video Editing with Cross-Attention Control
abstract
Video-P2P is the first framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there are currently no large-scale video generation models publicly available. Video-P2P addresses this limitation by adapting an image generation diffusion model to complete various video editing tasks. Specifically, we propose to first tune a Text-to-Set (T2S) model to complete an approximate inversion and then optimize a shared unconditional embedding to achieve accurate video inversion with a small memory cost. We further prove that it is crucial for consistent video editing. For attention control, we introduce a novel decoupled-guidance strategy, which uses different guidance strategies for the source and target prompts. The optimized unconditional embedding for the source prompt improves reconstruction ability, while an initialized unconditional embedding for the target prompt enhances editability. Incorporating the attention maps of these two branches enables detailed editing. These technical designs enable various text-driven editing applications, including word swap, prompt refinement, and attention re-weighting. Video-P2P works well on real-world videos for generating new characters while optimally preserving their original poses and scenes. It significantly outperforms previous approaches.
Shaoteng Liu, Yuechen Zhang, Wenbo Li 0002, Zhe Lin 0001, Jiaya Jia
CVPR2
2024 Prompt Highlighter: Interactive Control for Multi-Modal LLMs
abstract
This study targets a critical aspect of multi-modal LLMs' (LLMs&VLMs) inference: explicit controllable text generation. Multi-modal LLMs empower multi-modality understanding with the capability of semantic generation yet bring less explainability and heavier reliance on prompt contents due to their autoregressive generative nature. While manipulating prompt formats could improve outputs, designing specific and precise prompts per task can be challenging and ineffective. To tackle this issue, we introduce a novel inference method, Prompt Highlighter, which enables users to highlight specific prompt spans to interactively control the focus during generation. Motivated by the classifier-free diffusion guidance, we form regular and unconditional context pairs based on highlighted tokens, demonstrating that the autoregressive generation in models can be guided in a classifier-free way. Notably, we find that, during inference, guiding the models with highlighted tokens through the attention weights leads to more desired outputs. Our approach is compatible with current LLMs and VLMs, achieving impressive customized generation results without training. Experiments confirm its effectiveness in focusing on input contexts and generating reliable content. Without tuning on LLaVA-v1.5, our method secured 70.7 in the MMBench test and 1552.5 in MME-perception.
Yuechen Zhang, Shengju Qian, Bohao Peng, Shu Liu 0005, Jiaya Jia
CVPR1
2024 Objective Evaluation of VR Sickness and Analysis of Its Relationship with VR Presence
Cheng Han 0002, Yuechen Zhang, Yongqing Cai
ICIC (10)4
2024 HSD-YOLO: A Lightweight and Accurate Method for PCB Defect Detection
abstract
PCB defect detection is a typical small objects detection task, as with other objects, there is a small object size, the detection process is susceptible to the problem of background interference. In practical industrial production, it is difficult for existing object detection models to realize the balance between accuracy and real-time performance. Therefore, we propose a new and lightweight object detection model. And we named it Hsd-YOLO. Firstly, HGNetv2, the backbone of the new paradigm RT-DETR for object detection, is chosen as the backbone of our model, which makes the model more lightweight and reduces the number of parameters and computation while guaranteeing the accuracy of the model’s detection. Secondly, we use the more lightweight convolutional GSConv, which is introduced into the neck to make the model balance between accuracy and speed. Finally, a unified dynamic head framework, DyHead (Dynamic Head), is introduced to make the model improve the representation of the object detection head without increasing the computational overhead. We perform comparison experiments as well as ablation experiments on a publicly available PCB defect datasets to fully illustrate the effectiveness of ours model. We conduct ablation experiments on the current state-of-the-art single-stage model YOLOv8, and our model improves AP, AR, [email protected] and [email protected] by 3.8%, 0.2%, 0.6% and 2.2%, respectively, and reduces the number of parameters in 385312 while ensuring accuracy. Comparing with major object detection models, our model performs the best in accuracy.
Xiaotong Kong, Yuechen Zhang
IJCNN5
2024 RFSD-YOLO: An Enhanced X-Ray Object Detection Model for Prohibited Items
abstract
X-ray image detection is essential for ensuring public safety, but traditional methods rely heavily on human analysis and are relatively inefficient. To address this issue, this paper proposes a new object detection model called RFSD-YOLO, which is based on the YOLOv8 model and is specifically designed for detecting prohibited items in X-ray images. The model adopts the RFCAConv structure instead of the conventional convolutional operation. This allows for independent parameterisation among the convolutional kernels, enhancing the model's ability to capture and express potential prohibited item features. The neck section of the model includes the GSConv and VoVgscspdesigns, which aim to balance complexity and parameter size while maintaining detection performance and reducing computational burden. Additionally, we have introduced a dynamic head structure, DyHead, which replaces the traditional detection head design and improves detection accuracy without adding computational cost. The experimental results demonstrate that our enhanced model surpasses the current state-of-the-art object detection models in detecting prohibited items. Additionally, we introduce a simplified version of the RFSD-YOLOnano model to cater to resource-constrained environments. This streamlined model improves AP, AR, and mAP by 4.4%, 14.6%, and 3.4%, respectively, compared to YOLOv8n. This series of innovations not only validates the superiority of our model, but also provides new solutions in the field of automated detection of prohibited items for X-ray security screening.
Xiaotong Kong, Yuechen Zhang
SMC5
2024 RTS-DETR: Efficient Real-Time DETR for Small Object Detection
abstract
In recent years, object detection models DEtection TRansformer (DETR) series based on Transformer architecture have played a huge role in various fields. However, the DETR series models are not satisfactory in small object detection. Mainly due to the huge amount of calculation of DETR, a lot of feature information will be lost in the feature fusion stage and the low tolerance of small objects to Intersection over Union (IoU). In order to solve the above problems, we propose a near real-time detection model RTS-DETR. In this paper, we revisit Real-Time DEtection TRansformer (RT-DETR), which effectively handles multi-scale features by decoupling intra-scale interaction and cross-scale fusion, but this will lose a lot of positive local information. To this end, we have improved the efficient hybrid encoder. We propose a new positional encoding method that enables the hybrid encoder to more accurately convert the input feature sequence into a high-dimensional representation, and propose a new feature fusion module to enhance the model's ability to capture local features. Furthermore, in order to improve the tolerance of small objects to IoU, we combine Normalized Wasserstein Distance (NWD) with Shape-IoU for the optimization model. This method more accurately takes into account the shape and size of objects, thereby improving detection accuracy. Our model achieves an accuracy of 38.8% (in terms of [email protected]) on the widely used VisDrone dataset, which improves the accuracy by 2.5% compared to RT-DETR with ResNet-18 as the backbone network.
Xiaotong Kong, Yuechen Zhang
SMC5
2024 LDD-YOLO: An Improved Lightweight Detection Method for Steel Surface Defects Based on YOLOv8
abstract
Steel is an indispensable raw material in the industrial field, steel surface defects seriously affect the quality of steel, in recent years a lot of research has been carried out on the detection of steel surface defects. Existing steel defect detection methods are unable to fully mine the underlying feature information of the target image and do not achieve a dynamic balance between accuracy and speed. To address the above problems, this paper proposes an optimised target detection algorithm based on YOLOv8. First, we proposed the DMCA module, which combines the ideas of deformable convolution and multi-channel self-attention mechanism. We developed a strengthen self-attention module to enhance the process of deformable convolutional generation of offsets, so that the model can better adapt to the complex shapes of different defective targets and extract features at a deeper level. Secondly, using the idea of LKA (Large Kernel Attention), we propose the LF-MSPP lightweight module with long-range dependence and adaptive capability to capture the tele-relationships with small computational cost and parameters, improved the problem of missing defective feature information. Finally, we replaced the head of the original YOLOv8 with a Dynamic Head and used the split attention mechanism to improve the head detection capabilities while ensuring lightweight. We conduct extensive experiments on the widely used Northeastern University steel defect dataset NEU-DET. Experimental results show that the improved model improves mAP@50, mAP@50-95, AP and AR indicators by 2.4%, 2.0%, 6.1%, and 3.4% respectively compared with the original YOLOv8 model, and the number of model parameters is reduced by 11.2%. The improved model is also better than mainstream defect target detection models such as SSD, Retinanet, FasterRCNN, YOLOv5, YOLOv6, YOLOv7, YOLOv8, etc, and can better meet the accuracy and speed requirements of actual industrial production for steel surface defect detection models.
Yuechen Zhang, Xiaotong Kong
SMC1
2023 CodeTalker: Speech-Driven 3D Facial Animation with Discrete Motion Prior
abstract
Speech-driven 3D facial animation has been widely studied, yet there is still a gap to achieving realism and vividness due to the highly ill-posed nature and scarcity of audio-visual data. Existing works typically formulate the cross-modal mapping into a regression task, which suffers from the regression-to-mean problem leading to over-smoothed facial motions. In this paper, we propose to cast speech-driven facial animation as a code query task in a finite proxy space of the learned codebook, which effectively promotes the vividness of the generated motions by reducing the cross-modal mapping uncertainty. The codebook is learned by self-reconstruction over real facial motions and thus embedded with realistic facial motion priors. Over the discrete motion space, a temporal autoregressive model is employed to sequentially synthesize facial motions from the input speech signal, which guarantees lip-sync as well as plausible facial expressions. We demonstrate that our approach outperforms current state-of-the-art methods both qualitatively and quantitatively. Also, a user study further justifies our superiority in perceptual quality. Code and video demo are available at https://doubiiu.github.io/projects/codetalker.
Jinbo Xing, Menghan Xia, Yuechen Zhang, Xiaodong Cun, Jue Wang 0001, Tien-Tsin Wong
CVPR3
2023 Ref-NPR: Reference-Based Non-Photorealistic Radiance Fields for Controllable Scene Stylization
abstract
Current 3D scene stylization methods transfer textures and colors as styles using arbitrary style references, lacking meaningful semantic correspondences. We introduce Reference-Based Non-Photorealistic Radiance Fields (Ref-NP R) to address this limitation. This controllable method stylizes a 3D scene using radiance fields with a single stylized 2D view as a reference. We propose a ray registration process based on the stylized reference view to obtain pseudo-ray supervision in novel views. Then we exploit semantic correspondences in content images to fill occluded regions with perceptually similar styles, resulting in non-photorealistic and continuous novel view sequences. Our experimental results demonstrate that Ref-NPR out-performs existing scene and video stylization methods regarding visual quality and semantic correspondence. The code and data are publicly available on the project page at https://ref-npr.github.io.
Yuechen Zhang, Zexin He, Jinbo Xing, Xufeng Yao, Jiaya Jia
CVPR1
2023 BEW-YOLO: An Improved Method for PCB Defect Detection Based on YOLOv7
abstract
The PCB defect object size is small and the detection process is susceptible to background interference, usually have the problem of missed and false detection. In order to solve the above problems, an improved method based on YOLOv7 is proposed in this paper. Firstly, the bi-level routing attention (BRA) has been added to the header of the original YOLOv7 model to capture global dependencies and ensure the accuracy of small object detection and localization. Secondly, the explicit visual center (EVC) block is introduced before the fusion of mid-level features and high-level features to capture the global remote dependencies of top-level features, extract the features in local corner regions, achieve a comprehensive feature representation. Finally, the loss function is improved by changing the loss function of the original model to Wise-IoU, which allows our model to focus more on ordinary-quality anchor boxes and improve the performance of the detector while decreasing the competitiveness of anchor boxes for high-quality samples and reducing the influence of low-quality samples on the detection results. The experimental results show that the improved model improves 5% in AP and 1.9% in both AR and mAP over the original YOLOv7 model. Meanwhile, comparison experiments are conducted on the data-enhanced PCB dataset, which proves the superiority of our model over other state-of-the-art models.
Yuechen Zhang, Xiaotong Kong
ICPADS3
2023 Real-World Image Variation by Aligning Diffusion Inversion Chain
abstract
Recent diffusion model advancements have enabled high-fidelity images to be generated using text prompts. However, a domain gap exists between generated images and real-world images, which poses a challenge in generating high-quality variations of real-world images. Our investigation uncovers that this domain gap originates from a latents' distribution gap in different diffusion processes. To address this issue, we propose a novel inference pipeline called Real-world Image Variation by ALignment (RIVAL) that utilizes diffusion models to generate image variations from a single image exemplar. Our pipeline enhances the generation quality of image variations by aligning the image generation process to the source image's inversion chain. Specifically, we demonstrate that step-wise latent distribution alignment is essential for generating high-quality variations. To attain this, we design a cross-image self-attention injection for feature interaction and a step-wise distribution normalization to align the latent features. Incorporating these alignment processes into a diffusion model allows RIVAL to generate high-quality image variations without further parameter optimization. Our experimental results demonstrate that our proposed approach outperforms existing methods concerning semantic similarity and perceptual quality. This generalized inference pipeline can be easily applied to other diffusion-based generation tasks, such as image-conditioned text-to-image generation and stylization. Project page: https://rival-diff.github.io
Yuechen Zhang, Jinbo Xing, Eric Lo 0001, Jiaya Jia
NeurIPS1
2022 High Quality Segmentation for Ultra High-resolution Images
abstract
To segment 4K or 6K ultra high-resolution images needs extra computation consideration in image segmentation. Common strategies, such as downsampling, patch cropping, and cascade model, cannot address well the balance issue between accuracy and computation cost. Motivated by the fact that humans distinguish among objects continuously from coarse to precise levels, we propose the Continuous Refinement Model (CRM) for the ultra high-resolution segmentation refinement task. CRM continuously aligns the feature map with the refinement target and aggregates features to reconstruct these image details. Besides, our CRM shows its significant generalization ability to fill the resolution gap between low-resolution training images and ultra high-resolution testing ones. We present quantitative performance evaluation and visualization to show that our proposed method is fast and effective on image segmentation refinement. Code is available at https://github.com/dvlab-research/Entity/tree/main/CRM.
Tiancheng Shen, Yuechen Zhang, Lu Qi 0001, Jason Kuen, Xingyu Xie, Jianlong Wu, Zhe Lin 0001, Jiaya Jia
CVPR2
2022 PCL: Proxy-based Contrastive Learning for Domain Generalization
abstract
Domain generalization refers to the problem of training a model from a collection of different source domains that can directly generalize to the unseen target domains. A promising solution is contrastive learning, which attempts to learn domain-invariant representations by exploiting rich semantic relations among sample-to-sample pairs from different domains. A simple approach is to pull positive sample pairs from different domains closer while pushing other negative pairs further apart. In this paper, we find that directly applying contrastive-based methods (e.g., supervised contrastive learning) are not effective in domain generalization. We argue that aligning positive sample-to-sample pairs tends to hinder the model generalization due to the significant distribution gaps between different domains. To address this issue, we propose a novel proxy-based contrastive learning method, which replaces the original sample-to-sample relations with proxy-to-sample relations, significantly alleviating the positive alignment issue. Experiments on the four standard benchmarks demonstrate the effectiveness of the proposed method. Furthermore, we also consider a more complex scenario where no ImageNet pre-trained models are provided. Our method consistently shows better performance.
Xufeng Yao, Xinyun Zhang 0001, Yuechen Zhang, Qi Sun 0002, Ran Chen 0001, Ruiyu Li, Bei Yu 0001
CVPR4
2021 Flow-aware synthesis: A generic motion model for video frame interpolation
abstract
A popular and challenging task in video research, frame interpolation aims to increase the frame rate of video. Most existing methods employ a fixed motion model, e.g., linear, quadratic, or cubic, to estimate the intermediate warping field. However, such fixed motion models cannot well represent the complicated non-linear motions in the real world or rendered animations. Instead, we present an adaptive flow prediction module to better approximate the complex motions in video. Furthermore, interpolating just one intermediate frame between consecutive input frames may be insufficient for complicated non-linear motions. To enable multi-frame interpolation, we introduce the time as a control variable when interpolating frames between original ones in our generic adaptive flow prediction module. Qualitative and quantitative experimental results show that our method can produce high-quality results and outperforms the existing state-of-the-art methods on popular public datasets.
Jinbo Xing, Wenbo Hu 0002, Yuechen Zhang, Tien-Tsin Wong
Comput. Vis. Media3