EDBT 2026 Demo / reviewers in the wild / expert
Zongxin Yang
dblp:249/5456
· DBLP profile ↗
55ranked-venue papers
9as first author
52since 2021 · last 2026
0000-0001-8783-8313ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 38 · 9 first-author · 35 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 4 first-author · 35 since 2021Computer networks · 4 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Efficient Training of Large Vision Models via Advanced Automated Progressive LearningabstractThe rapid advancements in Large Vision Models (LVMs), such as Vision Transformers (ViTs), diffusion models, and visual autoregressive models, have led to an increasing demand for computational resources, resulting in substantial financial and environmental costs. This growing challenge highlights the necessity of developing efficient training methods for LVMs. Progressive learning, a training strategy in which model capacity gradually increases during training, has shown promise in addressing these challenges. In this paper, we take a practical step toward the efficient training of LVMs by automating progressive learning. We focus first on the pre-training of LVMs, using ViTs as a case study. We propose AutoProg-One, an automated progressive learning scheme featuring momentum growth (MoGrow) and the one-shot growth schedule search. Additionally, we extend our approach beyond pre-training to address the transfer learning and fine-tuning of LVMs. We also expand the scope of AutoProg to encompass a wider range of LVMs, including diffusion models and visual autoregressive model. First, we introduce AutoProg-Zero, by enhancing the AutoProg framework with a novel zero-shot automated progressive learning method, eliminating the need for one-shot supernet training. Second, we introduce a novel Unique Stage Identifier (SID) scheme to bridge the gap during network growth. These innovations, integrated with the core principles of AutoProg, offer a comprehensive solution for efficient training across various LVM scenarios. Extensive experiments show that AutoProg accelerates ViT pre-training by up to 1.85 ×on ImageNet and accelerates the fine-tuning of diffusion models, and visual autoregressive model by up to 2.86 × and 1.89 ×, with comparable or even better performance. This work provides a robust and scalable approach to efficient training of LVMs, with potential applications in a wide range of vision tasks. Sihao Lin, Zongxin Yang, Junwei Liang 0001, Xiaodan Liang, Xiaojun Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | SELongVLM: Empowering Long Video Language Models With Self-Corrective Clip SelectionabstractRecent advances in multimodal large language models (MLLMs) have enabled impressive progress in visual-language reasoning, yet long-video understanding remains a formidable challenge due to the need for coherent reasoning over ultra-long spatiotemporal dependencies. Existing methods struggle with the vast candidate space for relevant information in long videos, often failing to distinguish meaningful events from redundant content. We identify two critical and previously under-explored issues: absolute redundancy, where static visual content inflates token counts without adding narrative value, and relative redundancy, where task-irrelevant segments introduce noise that impairs reasoning. Compounding these issues is the weak spatiotemporal modeling in current MLLMs, which limits their ability to capture complex event dynamics. To address these multifaceted challenges, we introduce SELongVLM, a dynamically lenient-to-stringent selection long video language model. SELongVLM integrates two coordinated branches: a Residual Token Pruner (RTP) that removes repetitive background tokens via inter-frame residual modeling thus mitigating absolute redundancy while preserving motion cues, and a Semantic-aware Self-Correction Selector (SCSelector) that progressively refines query-relevant clip selection without frame-level annotations to reduce relative redundancy, guided by a stringent-to-lenient self-correcting mechanism during optimization. To ensure causal continuity and bolster spatiotemporal reasoning across disjoint clips, the framework further incorporates an action-aware operation for intra-clip dynamics and a temporal memory for cross-clip context, enabling robust spatiotemporal inference on long videos. Extensive experiments across eight benchmarks demonstrate that SELongVLM markedly outperforms existing models on both general and specialized long-video tasks. Specifically, it achieves 65.5% on VideoMME and 69.8% on MLVU for general benchmarks, and delivers strong performance on four specialized benchmarks - for example, 39.2% on TOMATO for fine-grained temporal reasoning and 69.2% on EventBench for event-level understanding. Kecheng Zhang, Zongxin Yang, Mingfei Han 0002, Yunzhi Zhuge, Haihong Hao, Zhihui Li 0001, Xiaojun Chang |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Toward General-Purpose Video Reconstruction Through Synergy of Grid-Splicing Diffusion and Large Language ModelsabstractVarious forms of degradation, including noise, blur, and adverse weather conditions (e.g., rain, snow, and fog), significantly compromise video quality and system reliability across critical domains ranging from surveillance and medical imaging to entertainment. Previous research mainly focuses on network models tailored to specific degradation types, while recent unified frameworks and foundation models still face critical challenges in temporal consistency, automated degradation recognition, and detail preservation. Despite recent advances in foundation models, current approaches rely heavily on predefined degradation labels and remain focused on image-level operations, limiting their generalization to real-world scenarios and struggling with preserving fine-grained details. To address these challenges, we propose Grid Splicing Diffusion Model (GSDiff), a general framework for video reconstruction that leverages a novel grid splicing execution alongside instruction-tuned Large Language Model (LLM). GSDiff introduces three key innovative modules: (1) a LLM-driven degradation recognition module that enables automatic and fine-grained restoration guidance through zero-shot degradation analysis, (2) a Grid Splicing Module that organizes multiple frames into a unified grid structure to facilitate spatiotemporal feature processing, and (3) a Detail Preservation Module integrated with a Tail Refine Network to enhance fine-grained details during diffusion and post-processing. Extensive experiments demonstrate that GSDiff delivers state-of-the-art performance across a wide range of reconstruction tasks, including deraining, desnowing, denoising, and deblurring, propelling advancements in medical diagnostics and smart city applications. Sen Yang 0006, Jinxi Xiang, Jieqiong Zhao, Zongxin Yang, Junhan Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Photorealistic Text-to-3D Avatar Generation with Constraints for Decoupled Geometry and AppearanceabstractThis article presents SEEAvatar, a novel approach to generate photorealistic 3D avatars from text descriptions. Despite the fact that recent text-to-3D avatar generation methods have shown promising results, their joint representation and optimization of geometry and appearance often yield coarse results and limit practical applications. Our method introduces novel constraints for decoupled geometry and appearance. First, we constrain geometric optimization using a template avatar, which evolves periodically to enable flexible shape generation while maintaining decent human shape. The detailed geometry features in faces and hands are also preserved from static human priors. Second, we leverage diffusion models to guide a physically based rendering pipeline for texture generation, incorporating a lightness constraint on albedo textures to suppress incorrect lighting effects. Experimental results demonstrate that our method significantly outperforms existing methods in both global and local geometry quality as well as appearance fidelity. The high-quality meshes and textures produced by our approach are directly compatible with traditional graphics pipelines, enabling immediate practical applications. Project page at: https://yoxu515.github.io/SEEAvatar/ . Yuanyou Xu, Zongxin Yang, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2025 | Few-Shot Incremental Learning via Foreground Aggregation and Knowledge Transfer for Audio-Visual Semantic SegmentationabstractAudio-Visual Semantic Segmentation (AVSS) has gained significant attention in the multi-modal domain, aiming to segment video objects that produce specific sounds in the corresponding audio. Despite notable progress, existing methods still struggle to handle new classes not included in the original training set. To this end, we introduce Few-Shot Incremental Learning (FSIL) to the AVSS task, which seeks to seamlessly integrate new classes with limited incremental samples while preserving the knowledge of old classes. Two challenges arise in this new setting: (1) To reduce labeling costs, old classes within the incremental samples are treated as background, similar to silent objects. Training the model directly with background annotations may worsen the loss of distinctive knowledge about old classes, such as their outlines and sounds. (2) Most existing models adopt early cross-modal fusion with a single-tower design, incorporating more characteristics into class representations, which impedes knowledge transfer between classes based on similarity. To address these issues, we propose a Few-shot Incremental learning framework via class-centric foregrouNd aggreGation and dual-tower knowlEdge tRansfer (FINGER) for the AVSS task, which comprises two targeted modules: (1) The class-centric foreground aggregation gathers class-specific features for each foreground class while disregarding background features. The background class is excluded during training and inferred from the foreground predictions. (2) The dual-tower knowledge transfer postpones cross-modal fusion to separately conduct knowledge transfer for each modality. Extensive experiments validate the effectiveness of the FINGER model, significantly surpassing state-of-the-art methods. Jingqiao Xiu, Mengze Li 0001, Zongxin Yang, Wei Ji 0008, Yifang Yin, Roger Zimmermann |
AAAI | 3 |
| 2025 | The Devil is in Temporal Token: High Quality Video Reasoning SegmentationabstractExisting methods for Video Reasoning Segmentation rely heavily on a single special token to represent the object in the keyframe or the entire video, inadequately capturing spatial complexity and inter-frame motion. To overcome these challenges, we propose VRS-HQ, an end-to-end video reasoning segmentation approach that leverages Multimodal Large Language Models (MLLMs) to inject rich spatiotemporal features into hierarchical tokens. Our key innovations include a Temporal Dynamic Aggregation (TDA) and a Token-driven Keyframe Selection (TKS). Specifically, we design frame-leveland temporal-leveltokens that utilize MLLM’s autoregressive learning to effectively capture both local and global information. Subsequently, we apply a similarity-based weighted fusion and frame selection strategy, then utilize SAM2 to perform keyframe segmentation and propagation. To enhance keyframe localization accuracy, the TKS filters keyframes based on SAM2’s occlusion scores during inference. VRSHQ achieves state-of-the-art performance on ReVOS, surpassing VISA by 5.9%/12.5%/9.1% in ${\mathcal{J}}{{ \& }}{\mathcal{F}}$ scores across the three subsets. These results highlight the strong temporal reasoning and segmentation capabilities of our method. Code and model weights are available at VRS-HQ. Sitong Gong, Yunzhi Zhuge, Lu Zhang 0053, Zongxin Yang, Huchuan Lu |
CVPR | 4 |
| 2025 | SKDream: Controllable Multi-view and 3D Generation with Arbitrary SkeletonsabstractControllable generation has achieved substantial progress in both 2D and 3D domains, yet current conditional generation methods still face limitations in describing detailed shape structures. Skeletons can effectively represent and describe object anatomy and pose. Unfortunately, past studies are often limited to human skeletons. In this work, we generalize skeletal conditioned generation to arbitrary structures. First, we design a reliable mesh skeletonization pipeline to generate a large-scale mesh-skeleton paired dataset. Based on the dataset, a multi-view and 3D generation pipeline is built. We propose to represent 3D skeletons by Coordinate Color Encoding as 2D conditional images. A Skeletal Correlation Module is designed to extract global skeletal features for condition injection. After multi-view images are generated, 3D assets can be obtained by incorporating a large reconstruction model, followed by a UV texture refinement stage. As a result, our method achieves instant generation of multi-view and 3D contents that are aligned with given skeletons. The proposed techniques largely improve the object-skeleton alignment and generation quality. Project page at https://skdream3d.github.io/. Yuanyou Xu, Zongxin Yang, Yi Yang 0001 |
CVPR | 2 |
| 2025 | DreamRenderer: Taming Multi-Instance Attribute Control in Large-Scale Text-to-Image Models
Dewei Zhou, Zongxin Yang, Yi Yang 0001 |
ICCV | 3 |
| 2025 | Streaming Video Understanding and Multi-round Interaction with Memory-enhanced KnowledgeabstractRecent advances in Large Language Models (LLMs) have enabled the development of Video-LLMs, advancing multimodal learning by bridging video data with language tasks. However, current video understanding models struggle with processing long video sequences, supporting multi-turn dialogues, and adapting to real-world dynamic scenarios. To address these issues, we propose StreamChat, a training-free framework for streaming video reasoning and conversational interaction. StreamChat leverages a novel hierarchical memory system to efficiently process and compress video features over extended sequences, enabling real-time, multi-turn dialogue. Our framework incorporates a parallel system scheduling strategy that enhances processing speed and reduces latency, ensuring robust performance in real-world applications. Furthermore, we introduce StreamBench, a versatile benchmark that evaluates streaming video understanding across diverse media types and interactive scenarios, including multi-turn interactions and complex reasoning tasks. Extensive evaluations on StreamBench and other public benchmarks demonstrate that StreamChat significantly outperforms existing
state-of-the-art models in terms of accuracy and response times, confirming its effectiveness for streaming video understanding. Code is available at StreamChat. Haomiao Xiong, Zongxin Yang, Jiazuo Yu 0001, Yunzhi Zhuge, Lu Zhang 0053, Jiawen Zhu 0003, Huchuan Lu |
ICLR | 2 |
| 2025 | 3DIS: Depth-Driven Decoupled Image Synthesis for Universal Multi-Instance GenerationabstractThe increasing demand for controllable outputs in text-to-image generation has spurred advancements in multi-instance generation (MIG), allowing users to define both instance layouts and attributes. However, unlike image-conditional generation methods such as ControlNet, MIG techniques have not been widely adopted in state-of-the-art models like SD2 and SDXL, primarily due to the challenge of building robust renderers that simultaneously handle instance positioning and attribute rendering. In this paper, we introduce Depth-Driven Decoupled Image Synthesis (3DIS), a novel framework that decouples the MIG process into two stages: (i) generating a coarse scene depth map for accurate instance positioning and scene composition, and (ii) rendering fine-grained attributes using pre-trained ControlNet on any foundational model, without additional training. Our 3DIS framework integrates a custom adapter into LDM3D for precise depth-based layouts and employs a finetuning-free method for enhanced instance-level attribute rendering. Extensive experiments on COCO-Position and COCO-MIG benchmarks demonstrate that 3DIS significantly outperforms existing methods in both layout precision and attribute rendering. Notably, 3DIS offers seamless compatibility with diverse foundational models, providing a robust, adaptable solution for advanced multi-instance generation. The code is available at: https://github.com/limuloo/3DIS. Dewei Zhou, Ji Xie, Zongxin Yang, Yi Yang 0001 |
ICLR | 3 |
| 2025 | Origin Identification for Text-Guided Image-to-Image Diffusion ModelsabstractText-guided image-to-image diffusion models excel in translating images based on textual prompts, allowing for precise and creative visual modifications. However, such a powerful technique can be misused for spreading misinformation, infringing on copyrights, and evading content tracing. This motivates us to introduce the task of origin IDentification for text-guided Image-to-image Diffusion models (ID$\mathbf{^2}$), aiming to retrieve the original image of a given translated query. A straightforward solution to ID$^2$ involves training a specialized deep embedding model to extract and compare features from both query and reference images. However, due to visual discrepancy across generations produced by different diffusion models, this similarity-based approach fails when training on images from one model and testing on those from another, limiting its effectiveness in real-world applications. To solve this challenge of the proposed ID$^2$ task, we contribute the first dataset and a theoretically guaranteed method, both emphasizing generalizability. The curated dataset, OriPID, contains abundant Origins and guided Prompts, which can be used to train and test potential IDentification models across various diffusion models. In the method section, we first prove the existence of a linear transformation that minimizes the distance between the pre-trained Variational Autoencoder embeddings of generated samples and their origins. Subsequently, it is demonstrated that such a simple linear transformation can be generalized across different diffusion models. Experimental results show that the proposed method achieves satisfying generalization performance, significantly surpassing similarity-based methods (+31.6% mAP), even those with generalization designs. The project is available at https://id2icml.github.io. Yifan Sun 0003, Zongxin Yang, Zhentao Tan, Zhengdong Hu, Yi Yang 0001 |
ICML | 3 |
| 2025 | X-Field: A Physically Informed Representation for 3D X-ray ReconstructionabstractX-ray imaging is indispensable in medical diagnostics, yet its use is tightly regulated due to radiation exposure. Recent research borrows representations from the 3D reconstruction area to complete two tasks with reduced radiation dose: X-ray Novel View Synthesis (NVS) and Computed Tomography (CT) reconstruction.
However, these representations fail to fully capture the penetration and attenuation properties of X-ray imaging as they originate from visible light imaging.
In this paper, we introduce X-Field, a 3D representation informed in the physics of X-ray imaging.
First, we employ homogeneous 3D ellipsoids with distinct attenuation coefficients to accurately model diverse materials within internal structures. Second, we introduce an efficient path-partitioning algorithm that resolves the intricate intersection of ellipsoids to compute cumulative attenuation along an X-ray path.
We further propose a hybrid progressive initialization to refine the geometric accuracy of X-Field and incorporate material-based optimization to enhance model fitting along material boundaries.
Experiments show that X-Field achieves superior visual fidelity on both real-world human organ and synthetic object datasets, outperforming state-of-the-art methods in X-ray NVS and CT Reconstruction. Our code is
available on the project page: https://github.com/Brack-Wang/X-Field. Jiachen Tao, Junyi Wu 0002, Haoxuan Wang 0002, Bin Duan 0004, Kai Wang 0036, Zongxin Yang, Yan Yan 0002 |
NeurIPS | 7 |
| 2025 | Enabling Instructional Image Editing with In-Context Generation in Large Scale Diffusion TransformerabstractInstruction-based image editing enables precise modifications via natural language prompts, but existing methods face a precision-efficiency tradeoff: fine-tuning demands massive datasets (>10M) and computational resources, while training-free approaches suffer from weak instruction comprehension.
We address this by proposing \textbf{ICEdit}, which leverages the inherent comprehension and generation abilities of large-scale Diffusion Transformers (DiTs) through three key innovations: (1) An in-context editing paradigm without architectural modifications; (2) Minimal parameter-efficient fine-tuning for quality improvement; (3) Early Filter Inference-Time Scaling, which uses VLMs to select high-quality noise samples for efficiency.
Experiments show that ICEdit achieves state-of-the-art editing performance with only 0.1\% of the training data and 1\% trainable parameters compared to previous methods. Our approach establishes a new paradigm for balancing precision and efficiency in instructional image editing. Zechuan Zhang, Ji Xie, Yu Lu 0019, Zongxin Yang, Yi Yang 0001 |
NeurIPS | 4 |
| 2025 | Prompt-based multimodal representation learning for drug repurposingabstractDrug repurposing significantly reduces development costs and shortens research cycles, making it a critical strategy in drug discovery. An emerging class of drug repurposing approaches applies deep learning to structural data. However, these methods often depend on static representations of molecular and protein structures, which may not fully capture the dynamic character of compound-protein interactions. To address these challenges and enhance the accuracy of compound-protein interaction predictions, we introduce an innovative prompt-based multimodal representation learning framework that dynamically encodes task-specific contextual information for drug repurposing. Specifically, the framework includes a dynamic prompt generation module that adaptively creates receptor-specific prompts and a prompt calibration module for effective multimodal feature integration and optimization. When applied to identifying FDA-approved drug candidates targeting G-protein-coupled receptors, our method achieved a 7.4% improvement in mean absolute error compared with state-of-the-art methods, with up to a 25.1% improvement for specific target-of-interest. By demonstrating potential in repurposing non-opioid treatments without the risk of addiction for safe pain management, our method has the capacity to advance drug discovery and meet a wide range of therapeutic needs. Kaicheng U, Dhruv Rana, Sophia Meixuan Zhang, Sen Yang 0006, Zongxin Yang, Hongping Tang, Junhan Zhao |
Briefings Bioinform. | 9 |
| 2025 | MIGC++: Advanced Multi-Instance Generation Controller for Image SynthesisabstractWe introduce the Multi-Instance Generation (MIG) task, which focuses on generating multiple instances within a single image, each accurately placed at predefined positions with attributes such as category, color, and shape, strictly following user specifications. MIG faces three main challenges: avoiding attribute leakage between instances, supporting diverse instance descriptions, and maintaining consistency in iterative generation. To address attribute leakage, we propose the Multi-Instance Generation Controller (MIGC). MIGC generates multiple instances through a divide-and-conquer strategy, breaking down multi-instance shading into single-instance tasks with singular attributes, later integrated. To provide more types of instance descriptions, we developed MIGC++. MIGC++ allows attribute control through text & images and position control through boxes & masks. Lastly, we introduced the Consistent-MIG algorithm to enhance the iterative MIG ability of MIGC and MIGC++. This algorithm ensures consistency in unmodified regions during the addition, deletion, or modification of instances, and preserves the identity of instances when their attributes are changed. We introduce the COCO-MIG and Multimodal-MIG benchmarks to evaluate these methods. Extensive experiments on these benchmarks, along with the COCO-Position benchmark and DrawBench, demonstrate that our methods substantially outperform existing techniques, maintaining precise control over aspects including position, attribute, and quantity. Project page: https://github.com/limuloo/MIGC. Dewei Zhou, Fan Ma, Zongxin Yang, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Test-Time Adaptation for Real-World Video Adverse Weather Restoration With Meta Batch NormalizationabstractAdverse weather conditions like rain, fog snow reduce visibility and degrade image quality, challenging the reliability of outdoor vision systems. Previous research mainly focuses on network models tailored to specific adverse weather conditions, limiting their effectiveness in addressing diverse weather scenarios in video processing. Recent research focuses on unified models for weather removal, significantly improving video quality in adverse conditions. However, the performance of these methods notably deteriorates in real environments due to the domain gap between synthetic and actual environments. In this paper, we present a meta-learning framework featuring a self-supervised learning (SSL) branch, aimed at boosting adaptability. In particular, we employ a two-stage training process. Initially, Joint training is implemented to establish a comprehensive model for weather reconstruction. Following this, Meta-BN training is applied to fine-tune the affine coefficients of the Batch Normalization (BN) layers, thus enabling the model to quickly adjust to different weather scenarios and maintain its efficacy in reconstruction. Moreover, an SSL-driven update strategy bolsters this targeted optimization, facilitating Test-time Weather Adaptation (TT-WA) and ensuring effective generalization to unfamiliar weather conditions. Experimental results across multiple benchmark datasets demonstrate that TT-WA consistently achieves state-of-the-art (SOTA) performance in both qualitative and quantitative evaluations under a variety of weather conditions, including rain, haze, and snow, outperforming existing methods. More critically, our approach exhibits robust adaptive reconstruction capabilities when applied to unseen real-world videos, further underscoring its effectiveness in generalizing to diverse and complex weather scenarios. Zongxin Yang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Exploiting EfficientSAM and Temporal Coherence for Audio-Visual Segmentation
Kun Li 0008, Zongxin Yang |
IEEE Trans. Multim. | 3 |
| 2025 | GD-NeRF: Generative Detail Compensation for One-shot Generalizable Neural Radiance FieldsabstractIn this article, we focus on the one-shot novel view synthesis task which targets synthesizing photo-realistic novel views given only one reference image per scene. Previous One-shot Generalizable Neural Radiance Field (OG-NeRF) methods solve this task in a finetuning-free manner, yet suffer from the blurry issue due to the encoder-only architecture that highly relies on the limited reference image. On the other hand, recent diffusion-based image-to-3D methods show vivid plausible results via distilling pre-trained 2D diffusion models, yet require tedious per-scene optimization. Targeting these issues, we propose GD-NeRF, a generative detail compensation framework that is both capable of producing vivid plausible details and is finetuning-free. Following a coarse-to-fine strategy, it is mainly composed of a One-stage Parallel Pipeline (OPP) and a Diffusion-based 3D-consistent Enhancer (Diff3DE). At the coarse stage, OPP first efficiently integrates the GAN model into the existing OG-NeRF pipeline for injecting primary in-distribution details. Then, at the fine stage, Diff3DE further leverages the pre-trained diffusion models to complement rich out-distribution details while maintaining decent 3D consistency. Extensive experiments on both the synthetic and real-world datasets show that GD-NeRF noticeably improves the vivid details while eliminating the need for per-scene finetuning. Xiao Pan 0001, Zongxin Yang, Shuai Bai, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Controllable 3D Face Generation with Conditional Style Code DiffusionabstractGenerating photorealistic 3D faces from given conditions is a challenging task. Existing methods often rely on time-consuming one-by-one optimization approaches, which are not efficient for modeling the same distribution content, e.g., faces. Additionally, an ideal controllable 3D face generation model should consider both facial attributes and expressions. Thus we propose a novel approach called TEx-Face(TExt & Expression-to-Face) that addresses these challenges by dividing the task into three components, i.e., 3D GAN Inversion, Conditional Style Code Diffusion, and 3D Face Decoding. For 3D GAN inversion, we introduce two methods, which aim to enhance the representation of style codes and alleviate 3D inconsistencies. Furthermore, we design a style code denoiser to incorporate multiple conditions into the style code and propose a data augmentation strategy to address the issue of insufficient paired visual-language data. Extensive experiments conducted on FFHQ, CelebA-HQ, and CelebA-Dialog demonstrate the promising performance of our TEx-Face in achieving the efficient and controllable generation of photorealistic 3D faces. The code will be publicly available. Xiaolong Shen, Chang Zhou 0005, Zongxin Yang |
AAAI | 4 |
| 2024 | SIFU: Side-view Conditioned Implicit Function for Real-world Usable Clothed Human ReconstructionabstractCreating high-quality 3D models of clothed humans from single images for real-world applications is crucial. De-spite recent advancements, accurately reconstructing hu-mans in complex poses or with loose clothing from in-the-wild images, along with predicting textures for unseen areas, remains a significant challenge. A key limitation of previous methods is their insufficient prior guidance in transitioning from 2D to 3D and in texture prediction. In response, we introduce SIFU (Side-view Conditioned Implicit Function for Real-world Usable Clothed Human Reconstruction), a novel approach combining a Side-view Decoupling Transformer with a 3D Consistent Texture Re-finement pipeline. SIFU employs a cross-attention mech-anism within the transformer, using SMPL-X normals as queries to effectively decouple side-view features in the process of mapping 2D features to 3D. This method not only improves the precision of the 3D models but also their ro-bustness, especially when SMPL-X estimates are not per-fect. Our texture refinement process leverages text-to-image diffusion-based prior to generate realistic and consistent textures for invisible views. Through extensive experiments, SIFU surpasses SOTA methods in both geometry and texture reconstruction, showcasing enhanced robustness in com-plex scenarios and achieving an unprecedented Chamfer and P2S measurement. Our approach extends to practi-cal applications such as 3D printing and scene building, demonstrating its broad utility in real-world scenarios. Zechuan Zhang, Zongxin Yang, Yi Yang 0001 |
CVPR | 2 |
| 2024 | HeadStudio: Text to Animatable Head Avatars with 3D Gaussian Splatting
Zhenglin Zhou, Fan Ma, Hehe Fan, Zongxin Yang, Yi Yang 0001 |
ECCV (32) | 4 |
| 2024 | DoraemonGPT: Toward Understanding Dynamic Scenes with Large Language Models (Exemplified as A Video Agent)abstractRecent LLM-driven visual agents mainly focus on solving image-based tasks, which limits their ability to understand dynamic scenes, making it far from real-life applications like guiding students in laboratory experiments and identifying their mistakes. Hence, this paper explores DoraemonGPT, a comprehensive and conceptually elegant system driven by LLMs to understand dynamic scenes. Considering the video modality better reflects the ever-changing nature of real-world scenarios, we exemplify DoraemonGPT as a video agent. Given a video with a question/task, DoraemonGPT begins by converting the input video into a symbolic memory that stores task-related attributes. This structured representation allows for spatial-temporal querying and reasoning by well-designed sub-task tools, resulting in concise intermediate results. Recognizing that LLMs have limited internal knowledge when it comes to specialized domains (e.g., analyzing the scientific principles underlying experiments), we incorporate plug-and-play tools to assess external knowledge and address tasks across different domains. Moreover, a novel LLM-driven planner based on Monte Carlo Tree Search is introduced to explore the large planning space for scheduling various tools. The planner iteratively finds feasible solutions by backpropagating the result’s reward, and multiple solutions can be summarized into an improved final answer. We extensively evaluate DoraemonGPT’s effectiveness on three benchmarks and several in-the-wild scenarios. Project page: https://z-x-yang.github.io/doraemon-gpt. Zongxin Yang, Guikun Chen, Wenguan Wang, Yi Yang 0001 |
ICML | 1 |
| 2024 | DRIP: Unleashing Diffusion Priors for Joint Foreground and Alpha Prediction in Image MattingabstractRecovering the foreground color and opacity/alpha matte from a single image (i.e., image matting) is a challenging and ill-posed problem where data priors play a critical role in achieving precise results. Traditional methods generally predict the alpha matte and then extract the foreground through post-processing, often failing to produce high-fidelity foreground color. This failure stems from the models' difficulty in learning robust color predictions from limited matting datasets. To address this, we explore the potential of leveraging vision priors embedded in pre-trained latent diffusion models (LDM) for estimating foreground RGBA values in challenging scenarios and rare objects. We introduce Drip, a novel approach for image matting that harnesses the rich prior knowledge of LDM models. Our method incorporates a switcher and a cross-domain attention mechanism to extend the original LDM for joint prediction of the foreground color and opacity. This setup facilitates mutual information exchange and ensures high consistency across both modalities. To mitigate the inherent reconstruction errors of the LDM's VAE decoder, we propose a latent transparency decoder to align the RGBA prediction with the input image, thereby reducing discrepancies. Comprehensive experimental results demonstrate that our approach achieves state-of-the-art performance in foreground and alpha predictions and shows remarkable generalizability across various benchmarks. Zongxin Yang, Ruijie Quan, Yi Yang 0001 |
NeurIPS | 2 |
| 2024 | Scalable Video Object Segmentation With Identification MechanismabstractThis paper delves into the challenges of achieving scalable and effective multi-object modeling for semi-supervised Video Object Segmentation (VOS). Previous VOS methods decode features with a single positive object, limiting the learning of multi-object representation as they must match and segment each target separately under multi-object scenarios. Additionally, earlier techniques catered to specific application objectives and lacked the flexibility to fulfill different speed-accuracy requirements. To address these problems, we present two innovative approaches, Associating Objects with Transformers (AOT) and Associating Objects with Scalable Transformers (AOST). In pursuing effective multi-object modeling, AOT introduces the IDentification (ID) mechanism to allocate each object a unique identity. This approach enables the network to model the associations among all objects simultaneously, thus facilitating the tracking and segmentation of objects in a single network pass. To address the challenge of inflexible deployment, AOST further integrates scalable long short-term transformers that incorporate scalable supervision and layer-wise ID-based attention. This enables online architecture scalability in VOS for the first time and overcomes ID embeddings' representation limitations. Given the absence of a benchmark for VOS involving densely multi-object annotations, we propose a challenging Video Object Segmentation in the Wild (VOSW) benchmark to validate our approaches. We evaluated various AOT and AOST variants using extensive experiments across VOSW and five commonly used VOS benchmarks, including YouTube-VOS 2018 & 2019 Val, DAVIS-2017 Val & Test, and DAVIS-2016. Our approaches surpass the state-of-the-art competitors and display exceptional efficiency and scalability consistently across all six benchmarks. Moreover, we notably achieved the$\mathbf {1^{st}}$position in the 3 rd Large-scale Video Object Segmentation Challenge. Project page:https://github.com/yoxu515/aot-benchmark. Zongxin Yang, Jiaxu Miao, Yunchao Wei, Wenguan Wang, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2024 | IDPro: Flexible Interactive Video Object Segmentation by ID-Queried Concurrent PropagationabstractInteractive Video Object Segmentation (iVOS) is inherently demanding, requiring real-time interaction between humans and computers. Enhancing user experience involves considerations such as user input habits, segmentation quality, running time, and memory consumption. However, existing methods compromise user experience by employing a single input mode and exhibiting slow running speeds. Specifically, these approaches restrict user interaction to a single frame, limiting the expression of user intent. To overcome these limitations and better align with user habits, we introduce a framework that facilitates flexible input modes by ID-queried concurrent propagation (IDPro). In particular, we have devised the Across-Frame Interaction Module (AFI), allowing users to freely annotate various objects across multiple frames. The AFI module transfers scribble information across interactive frames, generating multi-frame masks. Additionally, we leverage an id-queried mechanism to process multiple objects. To achieve more efficient propagation and a lightweight model, we propose a truncated re-propagation strategy, replacing the previous multi-round fusion module, which employs an across-round memory that stores crucial interaction information. Our SwinB-IDPro attains a new state-of-the-art performance on DAVIS 2017 (89.6%,${\mathcal {J}}\& {\mathcal {F}}\text{@}60$). Furthermore, our R50-IDPro exhibits over${3 \times }$faster performance than the leading competitor in challenging multi-object scenarios. Tao Jiang 0042, Zongxin Yang, Yi Yang 0001, Yueting Zhuang, Jun Xiao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | MuscleParseNet: A Novel Framework for Parsing Muscles of Drosophila Larva in Light-Sheet Fluorescence Microscopy ImagesabstractAccurately parsing (i.e., segmenting and recognizing) muscles of freely-moving animals such as Drosophila larva in light-sheet fluorescence microscopy images is necessary to study the relationship between muscle activity and animal motions. However, this task is challenging due to the large inter-class similarity and intra-class variance of muscles, as well as the in-homogeneous intensity and blurred boundaries of neighboring muscles. Existing semantic and instance segmentation methods cannot effectively overcome these challenges, resulting in poor segmentation and unreliable classification. In this work, we propose a novel framework named MuscleParseNet that explicitly utilizes sequential and spatial contexts to address these challenges. MuscleParseNet contains a deformable muscle candidate detector (D-CMD) to detect candidate muscles, and a sequential and spatial context-based fine muscle parser (SS-FMP) to refine the candidates. D-CMD boosts Mask RCNN with deformable convolutions to capture shape variations for more accurate muscle segmentation. Moreover, SS-FMP re-classifies the detected candidates by establishing a global spatial context to explicitly reflect spatial relative location, then optimizes the classification using the sequential associations of candidates in adjacent frames, which significantly improves muscle recognition accuracy. Experiments on the synchronized muscle-motion dataset of nearly freely-moving larvae show that MuscleParseNet produces promising results, outperforming state-of-the-art semantic and instance segmentation methods. Zhiying Song, Jinrun Zhou, Zongxin Yang, Yi Yang 0001, Zhefeng Gong, Nenggan Zheng |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | Show Me a Video: A Large-Scale Narrated Video Dataset for Coherent Story IllustrationabstractIllustrating a multi-sentence story with visual content is a significant challenge in multimedia research. While previous works have focused on sequential story-to-visual representations at the image level or representing a single sentence with a video clip, illustrating a long multi-sentence story with coherent videos remains an under-explored area. In this paper, we propose the task of video-based story illustration that focuses on the goal of visually illustrating a story with retrieved video clips. To support this task, we first create a large-scale dataset of coherent video stories in each sample, consisting of 85K narrative stories with 60 pairs of consistent clips and texts. We then propose the Story Context-Enhanced Model, which leverages local and global contextual information within the story, inspired by sequence modeling in language understanding. Through comprehensive quantitative experiments, we demonstrate the effectiveness of our baseline model. In addition, qualitative results and detailed user studies reveal that our method can retrieve coherent video sequences from stories. The dataset and code will be made publicly athttps://nfy-dot.github.io/CVSV-dataset/. Yu Lu 0019, Feiyue Ni, Linchao Zhu, Zongxin Yang, Ruihua Song, Lele Cheng, Yi Yang 0001 |
IEEE Trans. Multim. | 6 |
| 2024 | Noise-Tolerant Hybrid Prototypical Learning with Noisy Web DataabstractWe focus on the challenging problem of learning an unbiased classifier from a large number of potentially relevant but noisily labeled web images given only a few clean labeled images. This problem is particularly practical because it reduces the expensive annotation costs by utilizing freely accessible web images with noisy labels. Typically, prototypes are representative images or features used to classify or identify other images. However, in the few clean and many noisy scenarios, the class prototype can be severely biased due to the presence of irrelevant noisy images. The resulting prototypes are less compact and discriminative, as previous methods do not take into account the diverse range of images in the noisy web image collections. On the other hand, the relation modeling between noisy and clean images is not learned for the class prototype generation in an end-to-end manner, which results in a suboptimal class prototype. In this article, we introduce a similarity maximization loss named SimNoiPro. Our SimNoiPro first generates noise-tolerant hybrid prototypes composed of clean and noise-tolerant prototypes and then pulls them closer to each other. Our approach considers the diversity of noisy images by explicit division and overcomes the optimization discrepancy issue. This enables better relation modeling between clean and noisy images and helps extract judicious information from the noisy image set. The evaluation results on two extended few-shot classification benchmarks confirm that our SimNoiPro outperforms prior methods in measuring image relations and cleaning noisy data. Chao Liang 0002, Linchao Zhu, Zongxin Yang, Wei Chen 0001, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | High Fidelity Makeup via 2D and 3D Identity Preservation NetabstractIn this article, we address the challenging makeup transfer task, aiming to transfer makeup from a reference image to a source image while preserving facial geometry and background consistency. Existing deep neural network-based methods have shown promising results in aligning facial parts and transferring makeup textures. However, they often neglect the facial geometry of the source image, leading to two adverse effects: (1) alterations in geometrically relevant facial features, causing face flattening and loss of personality, and (2) difficulties in maintaining background consistency, as networks cannot clearly determine the face-background boundary. To jointly tackle these issues, we propose the High Fidelity Makeup via two-dimensional (2D) and 3D Identity Preservation Network (IP23-Net), to the best of our knowledge, a novel framework that leverages facial geometry information to generate more realistic results. Our method comprises a 3D Shape Identity Encoder, which extracts identity and 3D shape features. We incorporate a 3D face reconstruction model to ensure the three-dimensional effect of face makeup, thereby preserving the characters’ depth and natural appearance. To preserve background consistency, our Background Correction Decoder automatically predicts an adaptive mask for the source image, distinguishing the foreground and background. In addition to popular benchmarks, we introduce a new large-scale High Resolution Synthetic Makeup Dataset containing 335,230 diverse high-resolution face images to evaluate our method’s generalization ability. Experiments demonstrate that IP23-Net achieves high-fidelity makeup transfer while effectively preserving background consistency. The code will be made publicly available. Zhedong Zheng, Zongxin Yang, Yi Yang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | ProD: Prompting-to-disentangle Domain Knowledge for Cross-domain Few-shot Image ClassificationabstractThis paper considers few-shot image classification under the cross-domain scenario, where the train-to-test domain gap compromises classification accuracy. To mitigate the domain gap, we propose a prompting-to-disentangle (ProD) method through a novel exploration with the prompting mechanism. ProD adopts the popular multi-domain training scheme and extracts the backbone feature with a standard Convolutional Neural Network. Based on these two common practices, the key point of ProD is using the prompting mechanism in the transformer to disentangle the domain-general (DG) and domain-specific (DS) knowledge from the backbone feature. Specifically, ProD concatenates a DG and a DS prompt to the backbone feature and feeds them into a lightweight transformer. The DG prompt is learnable and shared by all the training domains, while the DS prompt is generated from the domain-of-interest on the fly. As a result, the transformer outputs DG and DS features in parallel with the two prompts, yielding the disentangling effect. We show that: 1) Simply sharing a single DG prompt for all the training domains already improves generalization towards the novel test domain. 2) The cross-domain generalization can be further reinforced by making the DG prompt neutral towards the training domains. 3) When inference, the DS prompt is generated from the support samples and can capture test domain knowledge through the prompting mechanism. Combining all three benefits, ProD significantly improves cross-domain few-shot classification. For instance, on CUB, ProD improves the 5-way 5-shot ac-curacy from 73.56% (baseline) to 79.19%, setting a new state of the art. Yifan Sun 0003, Zongxin Yang, Yi Yang 0001 |
CVPR | 3 |
| 2023 | FedSeg: Class-Heterogeneous Federated Learning for Semantic SegmentationabstractFederated Learning (FL) is a distributed learning paradigm that collaboratively learns a global model by multiple clients with data privacy-preserving. Although many FL algorithms have been proposed for classification tasks, few works focus on more challenging semantic segmentation tasks, especially in the class-heterogeneous FL situation. Compared with classification, the issues from heterogeneous FL for semantic segmentation are more severe: (1) Due to the non-IID distribution, different clients may contain inconsistent foreground-background classes, resulting in divergent local updates. (2) Class-heterogeneity for complex dense prediction tasks makes the local optimum of clients farther from the global optimum. In this work, we propose FedSeg, a basic federated learning approach for class-heterogeneous semantic segmentation. We first propose a simple but strong modified cross-entropy loss to correct the local optimization and address the foreground-background inconsistency problem. Based on it, we introduce pixel-level contrastive learning to enforce local pixel embeddings belonging to the global semantic space. Extensive experiments on four semantic segmentation benchmarks (Cityscapes, CamVID, PascalVOC and ADE20k) demonstrate the effectiveness of our FedSeg. We hope this work will attract more attention from the FL community to the challenging semantic segmentation federated learning. Jiaxu Miao, Zongxin Yang, Leilei Fan, Yi Yang 0001 |
CVPR | 2 |
| 2023 | Global-to-Local Modeling for Video-Based 3D Human Pose and Shape EstimationabstractVideo-based 3D human pose and shape estimations are evaluated by intra-frame accuracy and inter-frame smoothness. Although these two metrics are responsible for different ranges of temporal consistency, existing state-of-the-art methods treat them as a unified problem and use monotonous modeling structures (e.g., RNN or attention-based block) to design their networks. However, using a single kind of modeling structure is difficult to balance the learning of short-term and long-term temporal correlations, and may bias the network to one of them, leading to undesirable predictions like global location shift, temporal inconsistency, and insufficient local details. To solve these problems, we propose to structurally decouple the modeling of long-term and short-term correlations in an end-to-end framework, Global-to-Local Transformer (GLoT). First, a global transformer is introduced with a Masked Pose and Shape Estimation strategy for long-term modeling. The strategy stimulates the global transformer to learn more inter-frame correlations by randomly masking the features of several frames. Second, a local transformer is responsible for exploiting local details on the human mesh and interacting with the global transformer by leveraging cross-attention. Moreover, a Hierarchical Spatial Correlation Regressor is further introduced to refine intra-frame estimations by decoupled global-local representation and implicit kinematic constraints. Our GLoT surpasses previous state-of-the-art methods with the lowest model parameters on popular benchmarks, i.e., 3DPW, MPI-INF-3DHP, and Human3.6M. Codes are available at https://github.com/sxl142/GLoT. Xiaolong Shen, Zongxin Yang, Chang Zhou 0005, Yi Yang 0001 |
CVPR | 2 |
| 2023 | Shuffled Autoregression for Motion InterpolationabstractThis work aims to provide a deep-learning solution for the motion interpolation task. Previous studies solve it with geometric weight functions. Some other works propose neural networks for different problem settings with consecutive pose sequences as input. However, motion interpolation is a more complex problem that takes isolated poses (e.g., only one start pose and one end pose) as input. When applied to motion interpolation, these deep learning methods have limited performance since they do not leverage the flexible dependencies between interpolation frames as the original geometric formulas do. To realize this interpolation characteristic, we propose a novel framework, referred to as Shuffled AutoRegression, which expands the autoregression to generate in arbitrary (shuffled) order and models any inter-frame dependencies as a directed acyclic graph. We further propose an approach to constructing a particular kind of dependency graph, with three stages assembled into an end-to-end spatial-temporal motion Transformer. Experimental results on one of the current largest datasets show that our model generates vivid and coherent motions from only one start frame to one end frame and outperforms competing methods by a large margin. The proposed model is also extensible to multiple keyframes’ motion interpolation tasks and other areas’ interpolation. Shuo Huang 0005, Jia Jia 0001, Zongxin Yang, Wei Wang 0010, Haozhe Wu, Yi Yang 0001, Junliang Xing |
ICASSP | 3 |
| 2023 | Efficient Emotional Adaptation for Audio-Driven Talking-Head GenerationabstractAudio-driven talking-head synthesis is a popular research topic for virtual human-related applications. However, the inflexibility and inefficiency of existing methods, which necessitate expensive end-to-end training to transfer emotions from guidance videos to talking-head predictions, are significant limitations. In this work, we propose the Emotional Adaptation for Audio-driven Talking-head (EAT) method, which transforms emotion-agnostic talking-head models into emotion-controllable ones in a cost-effective and efficient manner through parameter-efficient adaptations. Our approach utilizes a pretrained emotion-agnostic talking-head transformer and introduces three lightweight adaptations (the Deep Emotional Prompts, Emotional Deformation Network, and Emotional Adaptation Module) from different perspectives to enable precise and realistic emotion controls. Our experiments demonstrate that our approach achieves state-of-the-art performance on widely-used benchmarks, including LRW and MEAD. Additionally, our parameter-efficient adaptations exhibit remarkable generalization ability, even in scenarios where emotional training videos are scarce or nonexistent. Project website: https://yuangan.github.io/eat/ Yuan Gan, Zongxin Yang, Xihang Yue, Lingyun Sun, Yi Yang 0001 |
ICCV | 2 |
| 2023 | JOTR: 3D Joint Contrastive Learning with Transformers for Occluded Human Mesh RecoveryabstractIn this study, we focus on the problem of 3D human mesh recovery from a single image under obscured conditions. Most state-of-the-art methods aim to improve 2D alignment technologies, such as spatial averaging and 2D joint sampling. However, they tend to neglect the crucial aspect of 3D alignment by improving 3D representations. Furthermore, recent methods struggle to separate the target human from occlusion or background in crowded scenes as they optimize the 3D space of target human with 3D joint coordinates as local supervision. To address these issues, a desirable method would involve a framework for fusing 2D and 3D features and a strategy for optimizing the 3D space globally. Therefore, this paper presents 3D JOint contrastive learning with TRansformers (JOTR) framework for handling occluded 3D human mesh recovery. Our method includes an encoder-decoder transformer architecture to fuse 2D and 3D representations for achieving 2D&3D aligned results in a coarse-to-fine manner and a novel 3D joint contrastive learning approach for adding explicitly global supervision for the 3D feature space. The contrastive learning approach includes two contrastive losses: joint-to-joint contrast for enhancing the similarity of semantically similar voxels (i.e., human joints), and joint-to-non-joint contrast for ensuring discrimination from others (e.g., occlusions and background). Qualitative and quantitative analyses demonstrate that our method outperforms state-of-the-art competitors on both occlusion-specific and standard benchmarks, significantly improving the reconstruction of occluded humans. Code is available at https://github.com/xljh0520/JOTR. Jiahao Li 0005, Zongxin Yang, Chang Zhou 0005, Yi Yang 0001 |
ICCV | 2 |
| 2023 | TransHuman: A Transformer-based Human Representation for Generalizable Neural Human RenderingabstractIn this paper, we focus on the task of generalizable neural human rendering which trains conditional Neural Radiance Fields (NeRF) from multi-view videos of different characters. To handle the dynamic human motion, previous methods have primarily used a SparseConvNet (SPC)-based human representation to process the painted SMPL. However, such SPC-based representation i) optimizes under the volatile observation space which leads to the pose-misalignment between training and inference stages, and ii) lacks the global relationships among human parts that is critical for handling the incomplete painted SMPL. Tackling these issues, we present a brand-new framework named TransHuman, which learns the painted SMPL under the canonical space and captures the global relationships between human parts with transformers. Specifically, TransHuman is mainly composed of Transformer-based Human Encoding (TransHE), Deformable Partial Radiance Fields (DPaRF), and Fine-grained Detail Integration (FDI). TransHE first processes the painted SMPL under the canonical space via transformers for capturing the global relationships between human parts. Then, DPaRF binds each output token with a deformable radiance field for encoding the query point under the observation space. Finally, the FDI is employed to further integrate fine-grained information from reference images. Extensive experiments on ZJU-MoCap and H36M show that our TransHuman achieves a significantly new state-of-the-art performance with high efficiency. Project page: https://pansanity666.github.io/TransHuman/ Xiao Pan 0001, Zongxin Yang, Chang Zhou 0005, Yi Yang 0001 |
ICCV | 2 |
| 2023 | Integrating Boxes and Masks: A Multi-Object Framework for Unified Visual Tracking and SegmentationabstractTracking any given object(s) spatially and temporally is a common purpose in Visual Object Tracking (VOT) and Video Object Segmentation (VOS). Joint tracking and segmentation have been attempted in some studies but they often lack full compatibility of both box and mask in initialization and prediction, and mainly focus on single-object scenarios. To address these limitations, this paper proposes a Multi-object Mask-box Integrated framework for unified Tracking and Segmentation, dubbed MITS. Firstly, the unified identification module is proposed to support both box and mask reference for initialization, where detailed object information is inferred from boxes or directly retained from masks. Additionally, a novel pinpoint box predictor is proposed for accurate multi-object box prediction, facilitating target-oriented representation learning. All target objects are processed simultaneously from encoding to propagation and decoding, as a unified pipeline for VOT and VOS. Experimental results show MITS achieves state-of-the-art performance on both VOT and VOS benchmarks. Notably, MITS surpasses the best prior VOT competitor by around 6% on the GOT-10k test set, and significantly improves the performance of box initialization on VOS benchmarks. The code is available at https://github.com/yoxu515/MITS. Yuanyou Xu, Zongxin Yang, Yi Yang 0001 |
ICCV | 2 |
| 2023 | Decompose to Generalize: Species-Generalized Animal Pose Estimation
Guangrui Li 0005, Yifan Sun 0003, Zongxin Yang, Yi Yang 0001 |
ICLR | 3 |
| 2023 | Video Object Segmentation in Panoptic Wild ScenesabstractIn this paper, we introduce semi-supervised video object segmentation (VOS) to panoptic wild scenes and present a large-scale benchmark as well as a baseline method for it. Previous benchmarks for VOS with sparse annotations are not sufficient to train or evaluate a model that needs to process all possible objects in real-world scenarios. Our new benchmark (VIPOSeg) contains exhaustive object annotations and covers various real-world object categories which are carefully divided into subsets of thing/stuff and seen/unseen classes for comprehensive evaluation. Considering the challenges in panoptic VOS, we propose a strong baseline method named panoptic object association with transformers (PAOT), which associates multiple objects by panoptic identification in a pyramid architecture on multiple scales. Experimental results show that VIPOSeg can not only boost the performance of VOS models by panoptic training but also evaluate them comprehensively in panoptic scenes. Previous methods for classic VOS still need to improve in performance and efficiency when dealing with panoptic scenes, while our PAOT achieves SOTA performance with good efficiency on VIPOSeg and previous VOS benchmarks. PAOT also ranks 1st in the VOT2022 challenge. Our dataset and code are available at https://github.com/yoxu515/VIPOSeg-Benchmark. Yuanyou Xu, Zongxin Yang, Yi Yang 0001 |
IJCAI | 2 |
| 2023 | Pyramid Diffusion Models for Low-light Image EnhancementabstractRecovering noise-covered details from low-light images is challenging, and the results given by previous methods leave room for improvement. Recent diffusion models show realistic and detailed image generation through a sequence of denoising refinements and motivate us to introduce them to low-light image enhancement for recovering realistic details. However, we found two problems when doing this, i.e., 1) diffusion models keep constant resolution in one reverse process, which limits the speed; 2) diffusion models sometimes result in global degradation (e.g., RGB shift). To address the above problems, this paper proposes a Pyramid Diffusion model (PyDiff) for low-light image enhancement. PyDiff uses a novel pyramid diffusion method to perform sampling in a pyramid resolution style (i.e., progressively increasing resolution in one reverse process). Pyramid diffusion makes PyDiff much faster than vanilla diffusion models and introduces no performance degradation. Furthermore, PyDiff uses a global corrector to alleviate the global degradation that may occur in the reverse process, significantly improving the performance and making the training of diffusion models easier with little additional computational consumption. Extensive experiments on popular benchmarks show that PyDiff achieves superior performance and efficiency. Moreover, PyDiff can generalize well to unseen noise and illumination distributions. Code and supplementary materials are available at https://github.com/limuloo/PyDIff.git. Dewei Zhou, Zongxin Yang, Yi Yang 0001 |
IJCAI | 2 |
| 2023 | AvatarFusion: Zero-shot Generation of Clothing-Decoupled 3D Avatars Using 2D DiffusionabstractLarge-scale pre-trained vision-language models allow for the zero-shot text-based generation of 3D avatars. The previous state-of-the-art method utilized CLIP to supervise neural implicit models that reconstructed a human body mesh. However, this approach has two limitations. Firstly, the lack of avatar-specific models can cause facial distortion and unrealistic clothing in the generated avatars. Secondly, CLIP only provides optimization direction for the overall appearance, resulting in less impressive results. To address these limitations, we propose AvatarFusion, the first framework to use a latent diffusion model to provide pixel-level guidance for generating human-realistic avatars while simultaneously segmenting clothing from the avatar's body. AvatarFusion includes the first clothing-decoupled neural implicit avatar model that employs a novel Dual Volume Rendering strategy to render the decoupled skin and clothing sub-models in one space. We also introduce a novel optimization method, called Pixel-Semantics Difference-Sampling (PS-DS), which semantically separates the generation of body and clothes, and generates a variety of clothing styles. Moreover, we establish the first benchmark for zero-shot text-to-avatar generation. Our experimental results demonstrate that our framework outperforms previous approaches, with significant improvements observed in all metrics. Additionally, since our model is clothing-decoupled, we can exchange the clothes of avatars. Code are available on our project page https://hansenhuang0823.github.io/AvatarFusion. Shuo Huang 0005, Zongxin Yang, Liangting Li, Yi Yang 0001, Jia Jia 0001 |
ACM Multimedia | 2 |
| 2023 | CATR: Combinatorial-Dependence Audio-Queried Transformer for Audio-Visual Video SegmentationabstractAudio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects within image frames and ensure the maps faithfully adheres to the given audio, such as identifying and segmenting a singing person in a video. However, existing methods exhibit two limitations: 1) they address video temporal features and audio-visual interactive features separately, disregarding the inherent spatial-temporal dependence of combined audio and video, and 2) they inadequately introduce audio constraints and object-level information during the decoding stage, resulting in segmentation outcomes that fail to comply with audio directives. To tackle these issues, we propose a decoupled audio-video transformer that combines audio and video features from their respective temporal and spatial dimensions, capturing their combined dependence. To optimize memory consumption, we design a block, which, when stacked, enables capturing audio-visual fine-grained combinatorial-dependence in a memory-efficient manner. Additionally, we introduce audio-constrained queries during the decoding phase. These queries contain rich object-level information, ensuring the decoded mask adheres to the sounds. Experimental results confirm our approach's effectiveness, with our framework achieving a new SOTA performance on all three datasets using two backbones. The code is available at https://github.com/aspirinone/CATR.github.io. Zongxin Yang, Lei Chen 0082, Yi Yang 0001, Jun Xiao 0001 |
ACM Multimedia | 2 |
| 2023 | Global-correlated 3D-decoupling Transformer for Clothed Avatar ReconstructionabstractReconstructing 3D clothed human avatars from single images is a challenging task, especially when encountering complex poses and loose clothing. Current methods exhibit limitations in performance, largely attributable to their dependence on insufficient 2D image features and inconsistent query methods. Owing to this, we present the Global-correlated 3D-decoupling Transformer for clothed Avatar reconstruction (GTA), a novel transformer-based architecture that reconstructs clothed human avatars from monocular images. Our approach leverages transformer architectures by utilizing a Vision Transformer model as an encoder for capturing global-correlated image features. Subsequently, our innovative 3D-decoupling decoder employs cross-attention to decouple tri-plane features, using learnable embeddings as queries for cross-plane generation. To effectively enhance feature fusion with the tri-plane 3D feature and human body prior, we propose a hybrid prior fusion strategy combining spatial and prior-enhanced queries, leveraging the benefits of spatial localization and human body prior knowledge. Comprehensive experiments on CAPE and THuman2.0 datasets illustrate that our method outperforms state-of-the-art approaches in both geometry and texture reconstruction, exhibiting high robustness to challenging poses and loose clothing, and producing higher-resolution textures. Codes are available at https://github.com/River-Zhang/GTA. Zechuan Zhang, Zongxin Yang, Ling Chen 0001, Yi Yang 0001 |
NeurIPS | 3 |
| 2023 | Collaborative Content-Dependent Modeling: A Return to the Roots of Salient Object DetectionabstractSalient object detection (SOD) aims to identify the most visually distinctive object(s) from each given image. Most recent progresses focus on either adding elaborative connections among different convolution blocks or introducing boundary-aware supervision to help achieve better segmentation, which is actually moving away from the essence of SOD, i.e., distinctiveness/salience. This paper goes back to the roots of SOD and investigates the principles of how to identify distinctive object(s) in a more effective and efficient way. Intuitively, the salience of one object should largely depend on its global context within the input image. Based on this, we devise a clean yet effective architecture for SOD, named Collaborative Content-Dependent Networks (CCD-Net). In detail, we propose a collaborative content-dependent head whose parameters are conditioned on the input image's global context information. Within the content-dependent head, a hand-crafted multi-scale (HMS) module and a self-induced (SI) module are carefully designed to collaboratively generate content-aware convolution kernels for prediction. Benefited from the content-dependent head, CCD-Net is capable of leveraging global context to detect distinctive object(s) while keeping a simple encoder-decoder design. Extensive experimental results demonstrate that our CCD-Net achieves state-of-the-art results on various benchmarks. Our architecture is simple and intuitive compared to previous solutions, resulting in competitive characteristics with respect to model complexity, operating efficiency, and segmentation accuracy. Siyu Jiao, Vidit Goel, Shant Navasardyan, Zongxin Yang, Levon Khachatryan, Yi Yang 0001, Yunchao Wei, Yao Zhao 0001, Humphrey Shi |
IEEE Trans. Image Process. | 4 |
| 2023 | Co-Learning Meets Stitch-Up for Noisy Multi-Label Visual RecognitionabstractIn real-world scenarios, collected and annotated data often exhibit the characteristics of multiple classes and long-tailed distribution. Additionally, label noise is inevitable in large-scale annotations and hinders the applications of learning-based models. Although many deep learning based methods have been proposed for handling long-tailed multi-label recognition or label noise respectively, learning with noisy labels in long-tailed multi-label visual data has not been well-studied because of the complexity of long-tailed distribution entangled with multi-label correlation. To tackle such a critical yet thorny problem, this paper focuses on reducing noise based on some inherent properties of multi-label classification and long-tailed learning under noisy cases. In detail, we propose a Stitch-Up augmentation to synthesize a cleaner sample, which directly reduces multi-label noise by stitching up multiple noisy training samples. Equipped with Stitch-Up, a Heterogeneous Co-Learning framework is further designed to leverage the inconsistency between long-tailed and balanced distributions, yielding cleaner labels for more robust representation learning with noisy long-tailed data. To validate our method, we build two challenging benchmarks, named VOC-MLT-Noise and COCO-MLT-Noise, respectively. Extensive experiments are conducted to demonstrate the effectiveness of our proposed method. Compared to a variety of baselines, our method achieves superior results. Chao Liang 0002, Zongxin Yang, Linchao Zhu, Yi Yang 0001 |
IEEE Trans. Image Process. | 2 |
| 2022 | H2FA R-CNN: Holistic and Hierarchical Feature Alignment for Cross-domain Weakly Supervised Object DetectionabstractCross-domain weakly supervised object detection (CD-WSOD) aims to adapt the detection model to a novel target domain with easily acquired image-level annotations. How to align the source and target domains is critical to the CDWSOD accuracy. Existing methods usually focus on partial detection components for domain alignment. In contrast, this paper considers that all the detection components are important and proposes a Holistic and Hier-archical Feature Alignment (H2FA) R-CNN. H2FA R-CNN enforces two image-level alignments for the backbone features, as well as two instance-level alignments for the RPN and detection head. This coarse-to-fine aligning hierarchy is in pace with the detection pipeline, i.e., processing the image-level feature and the instance-level features from bottom to top. Importantly, we devise a novel hybrid supervision method for learning two instance-level align-ments. It enables the RPN and detection head to simultane-ously receive weak/full supervision from the target/source domains. Combining all these feature alignments, H2 FA R-CNN effectively mitigates the gap between the source and target domains. Experimental results show that H2 FA R-CNN significantly improves cross-domain object detection accuracy and sets new state of the art on popular benchmarks. Code and pre-trained models are available at https://github.com/XuYunqiu/H2FA_R-CNN. Yunqiu Xu, Yifan Sun 0003, Zongxin Yang, Jiaxu Miao, Yi Yang 0001 |
CVPR | 3 |
| 2022 | Instance as Identity: A Generic Online Paradigm for Video Instance Segmentation
Feng Zhu 0005, Zongxin Yang, Xin Yu 0002, Yi Yang 0001, Yunchao Wei |
ECCV (29) | 2 |
| 2022 | In-N-Out Generative Learning for Dense Unsupervised Video SegmentationabstractIn this paper, we focus on unsupervised learning for Video Object Segmentation (VOS) which learns visual correspondence (i.e., the similarity between pixel-level features) from unlabeled videos. Previous methods are mainly based on the contrastive learning paradigm, which optimize either in image level or pixel level. Image-level optimization (e.g., the spatially pooled feature of ResNet) learns robust high-level semantics but is sub-optimal since the pixel-level features are optimized implicitly. By contrast, pixel-level optimization is more explicit, however, it is sensitive to the visual quality of training data and is not robust to object deformation. To complementarily perform these two levels of optimization in a unified framework, we propose the In-aNd-Out (INO) generative learning from a purely generative perspective with the help of naturally designed class tokens and patch tokens in Vision Transformer (ViT). Specifically, for image-level optimization, we force the out-view imagination from local to global views on class tokens, which helps capture high-level semantics, and we name it as out-generative learning. As to pixel-level optimization, we perform in-view masked image modeling on patch tokens, which recovers the corrupted parts of an image via inferring its fine-grained structure, and we term it as in-generative learning. To discover the temporal information better, we additionally force the inter-frame consistency from both feature and affinity matrix levels. Extensive experiments on DAVIS-2017 val and YouTube-VOS 2018 val show that our INO outperforms previous state-of-the-art methods by significant margins. Xiao Pan 0001, Peike Li, Zongxin Yang, Huiling Zhou, Chang Zhou 0005, Hongxia Yang, Jingren Zhou 0001, Yi Yang 0001 |
ACM Multimedia | 3 |
| 2022 | Decoupling Features in Hierarchical Propagation for Video Object SegmentationabstractThis paper focuses on developing a more effective method of hierarchical propagation for semi-supervised Video Object Segmentation (VOS). Based on vision transformers, the recently-developed Associating Objects with Transformers (AOT) approach introduces hierarchical propagation into VOS and has shown promising results. The hierarchical propagation can gradually propagate information from past frames to the current frame and transfer the current frame feature from object-agnostic to object-specific. However, the increase of object-specific information will inevitably lead to the loss of object-agnostic visual information in deep propagation layers. To solve such a problem and further facilitate the learning of visual embeddings, this paper proposes a Decoupling Features in Hierarchical Propagation (DeAOT) approach. Firstly, DeAOT decouples the hierarchical propagation of object-agnostic and object-specific embeddings by handling them in two independent branches. Secondly, to compensate for the additional computation from dual-branch propagation, we propose an efficient module for constructing hierarchical propagation, i.e., Gated Propagation Module, which is carefully designed with single-head attention. Extensive experiments show that DeAOT significantly outperforms AOT in both accuracy and efficiency. On YouTube-VOS, DeAOT can achieve 86.0% at 22.4fps and 82.0% at 53.4fps. Without test-time augmentations, we achieve new state-of-the-art performance on four benchmarks, i.e., YouTube-VOS (86.2%), DAVIS 2017 (86.2%), DAVIS 2016 (92.9%), and VOT 2020 (0.622 EAO). Project page: https://github.com/z-x-yang/AOT. Zongxin Yang, Yi Yang 0001 |
NeurIPS | 1 |
| 2022 | Collaborative Video Object Segmentation by Multi-Scale Foreground-Background IntegrationabstractThis paper investigates the principles of embedding learning to tackle the challenging semi-supervised video object segmentation. Unlike previous practices that focus on exploring the embedding learning of foreground object (s), we consider background should be equally treated. Thus, we propose a Collaborative video object segmentation by Foreground-Background Integration (CFBI) approach. CFBI separates the feature embedding into the foreground object region and its corresponding background region, implicitly promoting them to be more contrastive and improving the segmentation results accordingly. Moreover, CFBI performs both pixel-level matching processes and instance-level attention mechanisms between the reference and the predicted sequence, making CFBI robust to various object scales. Based on CFBI, we introduce a multi-scale matching structure and propose an Atrous Matching strategy, resulting in a more robust and efficient framework, CFBI+. We conduct extensive experiments on two popular benchmarks, i.e., DAVIS and YouTube-VOS. Without applying any simulated data for pre-training, our CFBI+ achieves the performance ( J& F) of 82.9 and 82.8 percent, outperforming all the other state-of-the-art methods. Code: https://github.com/z-x-yang/CFBI. Zongxin Yang, Yunchao Wei, Yi Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | DSC-PoseNet: Learning 6DoF Object Pose Estimation via Dual-Scale ConsistencyabstractCompared to 2D object bounding-box labeling, it is very difficult for humans to annotate 3D object poses, especially when depth images of scenes are unavailable. This paper investigates whether we can estimate the object poses effectively when only RGB images and 2D object annotations are given. To this end, we present a two-step pose estimation framework to attain 6DoF object poses from 2D object bounding-boxes. In the first step, the framework learns to segment objects from real and synthetic data in a weakly-supervised fashion, and the segmentation masks will act as a prior for pose estimation. In the second step, we design a dual-scale pose estimation network, namely DSC-PoseNet, to predict object poses by employing a differential renderer. To be specific, our DSC-PoseNet firstly predicts object poses in the original image scale by comparing the segmentation masks and the rendered visible object masks. Then, we resize object regions to a fixed scale to estimate poses once again. In this fashion, we eliminate large scale variations and focus on rotation estimation, thus facilitating pose estimation. Moreover, we exploit the initial pose estimation to generate pseudo ground-truth to train our DSC-PoseNet in a self-supervised manner. The estimation results in these two scales are ensembled as our final pose estimation. Extensive experiments on widely-used benchmarks demonstrate that our method outperforms state-of-the-art models trained on synthetic data by a large margin and even is on par with several fully-supervised methods. Zongxin Yang, Xin Yu 0002, Yi Yang 0001 |
CVPR | 1 |
| 2021 | Associating Objects with Transformers for Video Object SegmentationabstractThis paper investigates how to realize better and more efficient embedding learning to tackle the semi-supervised video object segmentation under challenging multi-object scenarios. The state-of-the-art methods learn to decode features with a single positive object and thus have to match and segment each target separately under multi-object scenarios, consuming multiple times computing resources. To solve the problem, we propose an Associating Objects with Transformers (AOT) approach to match and decode multiple objects uniformly. In detail, AOT employs an identification mechanism to associate multiple targets into the same high-dimensional embedding space. Thus, we can simultaneously process multiple objects' matching and segmentation decoding as efficiently as processing a single object. For sufficiently modeling multi-object association, a Long Short-Term Transformer is designed for constructing hierarchical matching and propagation. We conduct extensive experiments on both multi-object and single-object benchmarks to examine AOT variant networks with different complexities. Particularly, our R50-AOT-L outperforms all the state-of-the-art competitors on three popular benchmarks, i.e., YouTube-VOS (84.1% J&F), DAVIS 2017 (84.9%), and DAVIS 2016 (91.1%), while keeping more than 3X faster multi-object run-time. Meanwhile, our AOT-T can maintain real-time multi-object speed on the above benchmarks. Based on AOT, we ranked 1st in the 3rd Large-scale VOS Challenge. Zongxin Yang, Yunchao Wei, Yi Yang 0001 |
NeurIPS | 1 |
| 2020 | Gated Channel Transformation for Visual RecognitionabstractIn this work, we propose a generally applicable transformation unit for visual recognition with deep convolutional neural networks. This transformation explicitly models channel relationships with explainable control variables. These variables determine the neuron behaviors of competition or cooperation, and they are jointly optimized with the convolutional weight towards more accurate recognition. In Squeeze-and-Excitation (SE) Networks, the channel relationships are implicitly learned by fully connected layers, and the SE block is integrated at the block-level. We instead introduce a channel normalization layer to reduce the number of parameters and computational complexity. This lightweight layer incorporates a simple l2 normalization, enabling our transformation unit applicable to operator-level without much increase of additional parameters. Extensive experiments demonstrate the effectiveness of our unit with clear margins on many vision tasks, i.e., image classification on ImageNet, object detection and instance segmentation on COCO, video classification on Kinetics. Zongxin Yang, Linchao Zhu, Yu Wu 0011, Yi Yang 0001 |
CVPR | 1 |
| 2020 | Collaborative Video Object Segmentation by Foreground-Background Integration
Zongxin Yang, Yunchao Wei, Yi Yang 0001 |
ECCV (5) | 1 |
| 2019 | Very Long Natural Scenery Image Prediction by OutpaintingabstractComparing to image inpainting, image outpainting receives less attention due to two challenges in it. The first challenge is how to keep the spatial and content consistency between generated images and original input. The second challenge is how to maintain high quality in generated results, especially for multi-step generations in which generated regions are spatially far away from the initial input. To solve the two problems, we devise some innovative modules, named Skip Horizontal Connection and Recurrent Content Transfer, and integrate them into our designed encoder-decoder structure. By this design, our network can generate highly realistic outpainting prediction effectively and efficiently. Other than that, our method can generate new images with very long sizes while keeping the same style and semantic content as the given input. To test the effectiveness of the proposed architecture, we collect a new scenery dataset with diverse, complicated natural scenes. The experimental results on this dataset have demonstrated the efficacy of our proposed network. Zongxin Yang, Jian Dong 0011, Ping Liu 0004, Yi Yang 0001, Shuicheng Yan |
ICCV | 1 |