Tian Ye 0001

dblp:230/9616 · DBLP profile ↗
← Back
44ranked-venue papers
6as first author
44since 2021 · last 2026
0000-0002-8255-2997ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 36 · 6 first-author · 36 since 2021Artificial intelligence and machine learning · 28 · 5 first-author · 28 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021
YearPublicationVenuePosition
2026 Improved and Accelerated Text-to-Image Generation With Collect, Reflect, and Refine
abstract
Recently, enhancing the generative capability of text-to-image (T2I) models has become a promising direction in both academia and industry. Prior studies often focused on either improving generative quality or reducing inference latency, but typically failed to improve both quality and speed simultaneously. Moreover, existing inference-enhancement methods do not achieve significant improvements simultaneously across both diffusion models (DMs) and autoregressive models (ARMs). In this paper, we introduce a general tuning-based inference-enhancement framework, named CoRe$^{2}$2, which is the first to simultaneously achieve significant generative quality and reduced inference overhead across DMs and ARMs, to the best of our knowledge. CoRe$^{2}$2 comprises three stages: Collect, Reflect, and Refine. During the Collect stage, classifier-free guidance (CFG) trajectories are collected and subsequently used in the Reflect stage to train a weak model capable of reflecting the "easy-to-learn" content. Finally, during the Refine stage, CoRe$^{2}$2 can utilize the trained weak model to achieve speedup and performance gain in inference. Specifically, in the early sampling steps, CoRe$^{2}$2 employs weak-to-strong guidance to refine the "difficult-to-learn" and realistic content, thereby improving generative quality. In the later sampling steps, CoRe$^{2}$2 can use the weak model to generate "easy-to-learn" content instead of CFG, dramatically reducing inference time. Experimental outcomes substantiates CoRe$^{2}$2 achieve significant performance improvements on HPD v2, Pick-of-Pic, Drawbench, GenEval, and T2I-Compbench across SDXL, SD3.5, FLUX and LlamaGen. Notably, for SD3.5, CoRe$^{2}$2 can be seamlessly integrated with the state-of-the-art inference-enhancement algorithm Z-Sampling, outperforming it even with less time.
Shitong Shao, Zikai Zhou, Dian Xie, Yuetong Fang, Tian Ye 0001, Lichen Bai, Bo Han 0003, Zeke Xie
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 MovieChat+: Question-Aware Sparse Memory for Long Video Question Answering
abstract
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific vision tasks. Yet, existing methods either employ complex spatial-temporal modules or rely heavily on additional perception models to extract temporal features for video understanding, performing well only on short videos. For long videos, the computational complexity and memory costs associated with long-term temporal connections are significantly increased, posing additional challenges. Leveraging the hierarchical memory structure of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination, we propose MovieChat within a training-free memory consolidation mechanism to overcome these challenges, which transfers dense frames from short-term memory into sparse tokens in long-term memory by temporally merging adjacent frames. We lift pre-trained large multi-modal models for understanding long videos without additional trainable modules, employing a zero-shot approach. Additionally, in our new version, MovieChat+, we design an enhanced training-free vision-question matching-based memory consolidation mechanism to better anchor predictions to relevant visual content. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1 K benchmark with 1 K long video, 2 K temporal grounding labels, and 14 K manual annotations.
Enxin Song, Wenhao Chai, Tian Ye 0001, Jenq-Neng Hwang, Xi Li 0001, Gaoang Wang
IEEE Trans. Pattern Anal. Mach. Intell.3
2026 Standing on the Giants: Informative Messenger Prompts With Self-Adapter for Image Restoration
abstract
Despite the recent advances in the application of diffusion models to the realm of image restoration, their inherent stochastic nature can often lead to inaccuracies in reconstructing spatial structures and fine details. This paper introduces an innovative paradigm that harnesses the rich knowledge encapsulated in existing high-level pre-trained models to guide the diffusion process, offering a flexible and potent approach to restoration tasks. However, there are significant challenges in leveraging a pre-trained model for image restoration, including a data gap between natural clean images and degradation images, and paradigm differences that cause insufficient intermediate knowledge. To tackle these issues, we introduce informative Messenger prompts and a Self-adapter for the pre-trained model, whose appropriate information acts as explicit constraints for diffusion, enabling reliable result generation. Specifically, ourMeSa-IRsuccessfully adapts to feature exploration for degraded samples, via disseminating distinctive information from degraded instances. Furthermore, it bolsters knowledge representation through the bidirectional interchange of hierarchical information facilitated by the innovative use of messenger prompts. Experimental results demonstrate the state-of-the-art performance of our framework on five tasks in terms of perceptual and distortion metrics. We will release codes at https://github.com/Ephemeral182/MeSa-IR.
Sixiang Chen, Tian Ye 0001, Yulun Zhang 0001, Haoyu Chen 0003, Zhaohu Xing, Fugee Tsung, Lei Zhu 0003
IEEE Trans. Circuits Syst. Video Technol.2
2026 SegMamba-V2: Long-Range Sequential Modeling Mamba for General 3-D Medical Image Segmentation
abstract
The Transformer architecture has demonstrated remarkable results in 3D medical image segmentation due to its capability of modeling global relationships. However, it poses a significant computational burden when processing high-dimensional medical images. Mamba, as a State Space Model (SSM), has recently emerged as a notable approach for modeling long-range dependencies in sequential data. Although a substantial amount of Mamba-based research has focused on natural language and 2D image processing, few studies explore the capability of Mamba on 3D medical images. In this paper, we propose SegMamba-V2, a novel 3D medical image segmentation model, to effectively capture long-range dependencies within whole-volume features at each scale. To achieve this goal, we first devise a hierarchical scale downsampling strategy to enhance the receptive field and mitigate information loss during downsampling. Furthermore, we design a novel tri-orientated spatial Mamba block that extends the global dependency modeling process from one plane to three orthogonal planes to improve feature representation capability. Moreover, we collect and annotate a large-scale dataset (named CRC-2000) with fine-grained categories to facilitate benchmarking evaluation in 3D colorectal cancer (CRC) segmentation. We evaluate the effectiveness of our SegMamba-V2 on CRC-2000 and three other large-scale 3D medical image segmentation datasets, covering various modalities, organs, and segmentation targets. Experimental results demonstrate that our Segmamba-V2 outperforms state-of-the-art methods by a significant margin, which indicates the universality and effectiveness of the proposed model on 3D medical image segmentation tasks. The code for SegMamba-V2 is publicly available at: https://github.com/ge-xing/SegMamba-V2.
Zhaohu Xing, Tian Ye 0001, Du Cai, Baowen Gai, Xiao-Jian Wu, Feng Gao 0023, Lei Zhu 0003
IEEE Trans. Medical Imaging2
2026 Temporal Prompt Learning With Depth Memory for Video Mirror Detection
abstract
Mirror detection in dynamic scenes plays a crucial role in ensuring safety for various applications, such as drone tracking and robot navigation. However, current mirror detection models often fail in areas with mirrors that have a similar visual and color appearance to their surrounding objects. They also struggle to generalize well in complex cases, primarily due to limited annotated datasets. In this work, we propose a novel temporal prompt learning network with depth memory (TPD-Net) to address these critical challenges. Our approach includes several key components. First, we introduce a Temporal Prompt Generator (TPG) to learn temporal prompt features. Then, we devise Multi-layer Depth-aware Adaptor (MDA) modules to progressively adapt prompt features from the TPG, thereby learning mirror-related features by embedding temporal depth information as guidance. Moreover, we further refine these mirror-related features by constructing a depth memory and a Depth Memory Read module to read the temporal depths stored in the memory, boosting video mirror detection. Experimental results on a benchmark dataset show that our TPD-Net significantly outperforms 22 state-of-the-art methods in video mirror detection tasks. Our code, models, and results are publicly available athttps://github.com/ge-xing/TPDNet.
Zhaohu Xing, Tian Ye 0001, Xin Yang 0011, Sixiang Chen, Huazhu Fu, Yan Nei Law, Lei Zhu 0003
IEEE Trans. Multim.2
2025 PromptHaze: Prompting Real-world Dehazing via Depth Anything Model
abstract
Real-world image dehazing remains a challenging task due to the diverse nature of haze degradation and the lack of large-scale paired datasets. Existing methods based on hand-crafted priors or generative priors struggle to recover accurate backgrounds and fine details from dense haze regions. In this work, we propose a novel paradigm, PromptHaze, for real-world image dehazing via the depth prompt from the Depth Anything model. By employing a prompt-by-prompt strategy, our method iteratively updates the depth prompt and progressively restores the background through a dehazing network with controllable dehazing strength. Extensive experiments on widely-used real-world dehazing benchmarks demonstrate the superiority of PromptHaze in recovering authentic backgrounds and fine details from various haze scenes, outperforming state-of-the-art methods across multiple quality metrics.
Tian Ye 0001, Sixiang Chen, Haoyu Chen 0003, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003
AAAI1
2025 Residual Diffusion Deblurring Model for Single Image Defocus Deblurring
abstract
Defocus deblurring is a challenging task due to the spatially varying nature of defocus blur with multiple plausible solutions of a single given image. However, most existing methods falter when faced with extensive and variable defocus blur, either ignoring it or relying on additional loss functions to enhance perceptual quality. This often results in unrealistic reconstructions and compromised generalizability. In this paper, we propose a novel Residual Diffusion Deblurring Model framework for single image defocus deblurring. Our approach integrates a pre-trained defocus map estimator and a lightweight pre-deblur module with a learnable receptive field, providing crucial posterior information to effectively address large-scale and varying shaped defocus blur. In addition, a carefully-design denoising network enables the generation of diverse reconstructions from a single input. This approach not only significantly improves the perceptual quality of defocus deblurring outputs through multi-step residual learning, but also offers a more efficient inference strategy. Experimental results demonstrate that our method achieves competitive performance on real-world defocus deblurring image datasets across both perceptual and distortion evaluation metrics.
Haoxuan Feng, Haohui Zhou, Tian Ye 0001, Sixiang Chen, Lei Zhu 0003
AAAI3
2025 AGLLDiff: Guiding Diffusion Models Towards Unsupervised Training-free Real-world Low-light Image Enhancement
abstract
Existing low-light image enhancement (LIE) methods have achieved noteworthy success in solving synthetic distortions, yet they often fall short in practical applications. The limitations arise from two inherent challenges in real-world LIE: 1) the collection of distorted/clean image pairs is often impractical and sometimes even unavailable, and 2) accurately modeling complex degradations presents a non-trivial problem. To overcome them, we propose the Attribute Guidance Diffusion framework (AGLLDiff), a training-free method for effective real-world LIE. Instead of specifically defining the degradation process, AGLLDiff shifts the paradigm and models the desired attributes, such as image exposure, structure and color of normal-light images. These attributes are readily available and impose no assumptions about the degradation process, which guides the diffusion sampling process to a reliable high-quality solution space. Extensive experiments demonstrate that our approach outperforms the current leading unsupervised LIE methods across benchmarks in terms of distortion-based and perceptual-based metrics, and it performs well even in sophisticated wild degradation.
Yunlong Lin, Tian Ye 0001, Sixiang Chen, Zhenqi Fu, Yingying Wang 0005, Wenhao Chai, Zhaohu Xing, Wenxue Li 0003, Lei Zhu 0003, Xinghao Ding
AAAI2
2025 DPLUT: Unsupervised Low-light Image Enhancement with Lookup Tables and Diffusion Priors
abstract
Low-light image enhancement (LIE) aims at precisely and efficiently recovering an image degraded in poor illumination environments. Recent advanced LIE techniques are using deep neural networks, which require lots of low-normal light image pairs, network parameters, and computational resources. As a result, their practicality is limited. In this work, we devise a novel unsupervised LIE framework based on diffusion priors and lookup tables (DPLUT) to achieve efficient low-light image recovery. The proposed approach comprises two critical components: a light adjustment lookup table (LLUT) and a noise suppression lookup table (NLUT). LLUT is optimized with a set of unsupervised losses. It aims at predicting pixel-wise curve parameters for the dynamic range adjustment of a specific image. NLUT is designed to remove the amplified noise after the light brightens. As diffusion models are sensitive to noise, diffusion priors are introduced to achieve high-performance noise suppression. Extensive experiments demonstrate that our approach outperforms state-of-the-art methods in terms of visual quality and efficiency.
Yunlong Lin, Zhenqi Fu, Kairun Wen, Tian Ye 0001, Sixiang Chen, Ge Meng, Yingying Wang 0005, Chui Kong, Yue Huang 0001, Xiaotong Tu, Xinghao Ding
AAAI4
2025 Dehaze-RetinexGAN: Real-World Image Dehazing via Retinex-based Generative Adversarial Network
abstract
Deep learning based dehazing networks trained on paired synthetic data have shown impressive performance, but they struggle with significant degradation in generalization ability on real-world hazy scenes. In this paper, we propose Dehaze-RetinexGAN, a lightweight Retinex-based Generative Adversarial Network for real-world image Dehazing using unpaired data. Our Dehaze-RetinexGAN consists of two stages: self-supervised pre-training and weakly-supervised fine-tuning. During the pre-training, we reduce the image dehazing task to an illumination-reflectance decomposition task based on the duality correlation between Retinex and dehazing. Specifically, a decomposition network named DecomNet is constructed to obtain an illumination and a reflectance, simultaneously. Moreover, a self-supervised learning strategy is developed to construct the connection between the preliminary dehazed result and the input hazy image, which constrains the solution space of DecomNet and accelerates training, leading to a more realistic dehazed result. In the fine-tuning stage, we develop a dual DTCWT-based attention module and embed it into the U-Net architecture to further improve the quality of preliminary result in the frequency domain. In addition, the adversarial learning is employed to constrain the relevance between the clean image and the final dehazed result in a weakly supervised manner, which can promote more natural performance. Extensive experiments on several real-world datasets demonstrate that our proposed framework performs favorably over state-of-the-art dehazing methods in visual quality and quantitative evaluation.
Tian Ye 0001, Yun Liu 0002
AAAI3
2025 POSTA: A Go-to Framework for Customized Artistic Poster Generation
abstract
Poster design is a critical medium for visual communication. Prior work has explored automatic poster design using deep learning techniques, but these approaches lack text accuracy, user customization, and aesthetic appeal, limiting their applicability in artistic domains such as movies and exhibitions, where both clear content delivery and visual impact are essential. To address these limitations, we present POSTA: a modular framework powered by diffusion models and multimodal large language models (MLLMs) for customized artistic poster generation. The framework consists of three modules. Background Diffusion creates a themed background based on user input. Design MLLM then generates layout and typography elements that align with and complement the background style. Finally, to enhance the poster’s aesthetic appeal, ArtText Diffusion applies additional stylization to key text elements. The final result is a visually cohesive and appealing poster, with a fully modular process that allows for complete customization. To train our models, we develop the PosterArt dataset, comprising high-quality artistic posters annotated with layout, typography, and pixel-level stylized text segmentation. Our comprehensive experimental analysis demonstrates POSTA’s exceptional controllability and design diversity, outperforming existing models in both text accuracy and aesthetic quality.
Haoyu Chen 0003, Wenbo Li 0002, Tian Ye 0001, Songhua Liu, Ying-Cong Chen, Lei Zhu 0003, Xinchao Wang
CVPR5
2025 SnowMaster: Comprehensive Real-world Image Desnowing via MLLM with Multi-Model Feedback Optimization
abstract
Snowfall presents significant challenges for visual data processing, necessitating specialized desnowing algorithms. However, existing models often fail to generalize effectively due to their heavy reliance on synthetic datasets. Furthermore, current real-world snowfall datasets are limited in scale and lack dedicated evaluation metrics designed specifically for snowfall degradation, thus hindering the effective integration of real snowy images into model training to reduce domain gaps. To address these challenges, we first introduce RealSnow10K, a large-scale, high-quality dataset consisting of over 10,000 annotated real-world snowy images. In addition, we curate a preference dataset comprising 36,000 expert-ranked image pairs, enabling the adaptation of multimodal large language models (MLLMs) to better perceive snowy image quality through our innovative Multi-Model Preference Optimization (MMPO). Finally, we propose the SnowMaster, which employs MMPO-enhanced MLLM to perform accurate snowy image evaluation and pseudo-label filtering for semi-supervised training. Experiments demonstrate that SnowMaster delivers superior desnowing performance under real-world conditions.
Jianyu Lai, Sixiang Chen, Yunlong Lin, Tian Ye 0001, Yun Liu 0002, Song Fei, Zhaohu Xing, Weiming Wang 0002, Lei Zhu 0003
CVPR4
2025 Detect Any Mirrors: Boosting Learning Reliability on Large-Scale Unlabeled Data with an Iterative Data Engine
abstract
Mirror detection is a challenging task because a mirror’s visual appearance varies depending on the reflected content. Due to limited annotated data, current methods failed to generalize well for detecting diverse mirror scenes. Semi-supervised learning with large-scale unlabeled data can improve generalization capabilities on mirror detection, but these methods often suffer from unreliable pseudo-labels due to distribution differences between labeled and unlabeled data, therefore affecting the learning process. To address this issue, we first collect a large-scale dataset of approximately 0.4 million mirror-related images from the internet, significantly expanding the data scale for mirror detection. To effectively exploit this unlabeled dataset, we propose the first semi-supervised framework (namely an iterative data engine) consisting of four steps: (1) mirror detection model training, (2) pseudo label prediction, (3) dual guidance scoring, and (4) selection of highly reliable pseudo labels. In each iteration of the data engine, we employ a geometric accuracy scoring approach to assess pseudo labels based on multiple segmentation metrics, and design a multi-modal agent-driven semantic scoring approach to enhance the semantic perception of pseudo labels. These two scoring approaches can effectively improve the reliability of pseudo labels by selecting unlabeled samples with higher scores. Our method demonstrates promising performance across three mirror detection tasks and exhibits strong generalization on unseen examples. Our code will be publicly available at https://github.com/ge-xing/DAM.
Zhaohu Xing, Hongqiu Wang, Tian Ye 0001, Sixiang Chen, Wenxue Li 0003, Guang Liu 0006, Lei Zhu 0003
CVPR5
2025 Genhaze: Pioneering Controllable One-Step Realistic Haze Generation for Real-World Dehazing
abstract
Real-world image dehazing is crucial for enhancing visual quality in computer vision applications. However, existing physics-based haze generation paradigms struggle to model the complexities of real-world haze and lack controllability, limiting the performance of existing baselines on real-world images. In this paper, we introduce GenHaze, a pioneering haze generation framework that enables the one-step generation of high-quality, reference-controllable hazy images. GenHaze leverages the pre-trained latent diffusion model (LDM) with a carefully designed clean-to-haze generation protocol to produce realistic hazy images. Additionally, by leveraging its fast, controllable generation of paired highquality hazy images, we illustrate that existing dehazing baselines can be unleashed in a simple and efficient manner. Extensive experiments indicate that GenHaze achieves visually convincing and quantitatively superior hazy images. It also significantly improves multiple existing dehazing models across 7 non-reference metrics with minimal fine-tuning epochs. Our work demonstrates that LDM possesses the potential to generate realistic degradations, providing an effective alternative to prior generation pipelines.
Sixiang Chen, Tian Ye 0001, Yunlong Lin, Yeying Jin, Haoyu Chen 0003, Jianyu Lai, Song Fei, Zhaohu Xing, Fugee Tsung, Lei Zhu 0003
ICCV2
2025 GlassWizard: Harvesting Diffusion Priors for Glass Surface Detection
Wenxue Li 0003, Tian Ye 0001, Xinyu Xiong, Jinbin Bai, Wenxuan Song, Zhaohu Xing, Lie Ju, Guanbin Li, Lei Zhu 0003
ICCV2
2025 Bringing RNNs Back to Efficient Open-Ended Video Understanding
Weili Xu, Enxin Song, Wenhao Chai, Xuexiang Wen, Tian Ye 0001, Gaoang Wang
ICCV5
2025 Meissonic: Revitalizing Masked Generative Transformers for Efficient High-Resolution Text-to-Image Synthesis
abstract
We present Meissonic, which elevates non-autoregressive text-to-image Masked Image Modeling (MIM) to a level comparable with state-of-the-art diffusion models like SDXL. By incorporating a comprehensive suite of architectural innovations, advanced positional encoding strategies, and optimized sampling conditions, Meissonic substantially improves MIM's performance and efficiency. Additionally, we leverage high-quality training data, integrate micro-conditions informed by human preference scores, and employ feature compression layers to further enhance image fidelity and resolution. Our model not only matches but often exceeds the performance of existing methods in generating high-quality, high-resolution images. Extensive experiments validate Meissonic’s capabilities, demonstrating its potential as a new standard in text-to-image synthesis.
Jinbin Bai, Tian Ye 0001, Wei Chow, Enxin Song, Xiangtai Li, Zhen Dong 0003, Lei Zhu 0003, Shuicheng Yan
ICLR2
2025 Farther Than Mirror: Explore Pattern-Compensated Depth of Mirror with Temporal Changes for Video Mirror Detection
Zhaohu Xing, Tian Ye 0001, Sixiang Chen, Guang Liu 0006, Lei Zhu 0003
ACM Multimedia3
2025 Triplane-Smoothed Video Dehazing with CLIP-Enhanced Generalization
Haoyu Chen 0003, Tian Ye 0001, Lei Zhu 0003
Int. J. Comput. Vis.3
2024 VQCNIR: Clearer Night Image Restoration with Vector-Quantized Codebook
abstract
Night photography often struggles with challenges like low light and blurring, stemming from dark environments and prolonged exposures. Current methods either disregard priors and directly fitting end-to-end networks, leading to inconsistent illumination, or rely on unreliable handcrafted priors to constrain the network, thereby bringing the greater error to the final result. We believe in the strength of data-driven high-quality priors and strive to offer a reliable and consistent prior, circumventing the restrictions of manual priors. In this paper, we propose Clearer Night Image Restoration with Vector-Quantized Codebook (VQCNIR) to achieve remarkable and consistent restoration outcomes on real-world and synthetic benchmarks. To ensure the faithful restoration of details and illumination, we propose the incorporation of two essential modules: the Adaptive Illumination Enhancement Module (AIEM) and the Deformable Bi-directional Cross-Attention (DBCA) module. The AIEM leverages the inter-channel correlation of features to dynamically maintain illumination consistency between degraded features and high-quality codebook features. Meanwhile, the DBCA module effectively integrates texture and structural information through bi-directional cross-attention and deformable convolution, resulting in enhanced fine-grained detail and structural fidelity across parallel decoders. Extensive experiments validate the remarkable benefits of VQCNIR in enhancing image quality under low-light conditions, showcasing its state-of-the-art performance on both synthetic and real-world datasets. The code is available at https://github.com/AlexZou14/VQCNIR.
Wenbin Zou, Hongxia Gao, Tian Ye 0001, Liang Chen 0026, Weipeng Yang 0002, Shasha Huang, Sixiang Chen
AAAI3
2024 Learning Diffusion Texture Priors for Image Restoration
abstract
Diffusion Models have shown remarkable performance in image generation tasks, which are capable of generating diverse and realistic image content. When adopting diffusion models for image restoration, the crucial challenge lies in how to preserve high-level image fidelity in the random-ness diffusion process and generate accurate background structures and realistic texture details. In this paper, we propose a general framework and develop a Diffusion Texture Prior Model (DTPM) for image restoration tasks. DTPM explicitly models high-quality texture details through the diffusion process, rather than global contextual content. In phase one of the training stage, we pretrain DTPM on approximately 55K high-quality image samples, after which we freeze most of its parameters. In phase two, we insert conditional guidance adapters into DTPM and equip it with an initial predictor, thereby facilitating its rapid adaptation to downstream image restoration tasks. Our DTPM could mitigate the randomness of traditional diffusion models by utilizing encapsulated rich and diverse texture knowledge and background structural information provided by the initial predictor during the sampling process.
Tian Ye 0001, Sixiang Chen, Wenhao Chai, Zhaohu Xing, Harry Qin, Lei Zhu 0003
CVPR1
2024 MovieChat: From Dense Token to Sparse Memory for Long Video Understanding
abstract
Recently, integrating video foundation models and large language models to build a video understanding system can overcome the limitations of specific pre-defined vision tasks. Yet, existing systems can only handle videos with very few frames. For long videos, the computation complexity, memory cost, and long-term temporal connection impose additional challenges. Taking advantage of the Atkinson-Shiffrin memory model, with tokens in Transformers being employed as the carriers of memory in combination with our specially designed memory mechanism, we propose the MovieChat to overcome these challenges. MovieChat achieves state-of-the-art performance in long video understanding, along with the released MovieChat-1K benchmark with 1K long video and 14K manual annotations for validation of the effectiveness of our method. The code, models and data can be found in https://reself.github.io/MovieChat.
Enxin Song, Wenhao Chai, Guanhong Wang, Haoyang Zhou, Feiyang Wu, Haozhe Chi, Xun Guo 0002, Tian Ye 0001, Yanting Zhang 0001, Yan Lu 0001, Jenq-Neng Hwang, Gaoang Wang
CVPR9
2024 Teaching Tailored to Talent: Adverse Weather Restoration via Prompt Pool and Depth-Anything Constraint
Sixiang Chen, Tian Ye 0001, Kai Zhang 0008, Zhaohu Xing, Yunlong Lin, Lei Zhu 0003
ECCV (9)2
2024 See and Think: Embodied Agent in Virtual Environment
Zhonghan Zhao, Wenhao Chai, Boyi Li 0002, Shengyu Hao, Shidong Cao, Tian Ye 0001, Gaoang Wang
ECCV (8)7
2024 Integrating View Conditions for Image Synthesis
Jinbin Bai, Zhen Dong 0003, Aosong Feng, Tian Ye 0001, Kaicheng Zhou
IJCAI5
2024 Cross-conditioned Diffusion Model for Medical Image to Image Translation
Zhaohu Xing, Sicheng Yang 0001, Sixiang Chen, Tian Ye 0001, Harry Qin, Lei Zhu 0003
MICCAI (7)4
2024 SegMamba: Long-Range Sequential Modeling Mamba for 3D Medical Image Segmentation
Zhaohu Xing, Tian Ye 0001, Guang Liu 0006, Lei Zhu 0003
MICCAI (8)2
2024 Timeline and Boundary Guided Diffusion Network for Video Shadow Detection
abstract
Video Shadow Detection (VSD) aims to detect the shadow masks with frame sequence. Existing works suffer from inefficient temporal learning. Moreover, few works address the VSD problem by considering the characteristic (i.e., boundary) of shadow. Motivated by this, we propose a Timeline and Boundary Guided Diffusion (TBGDiff) network for VSD where we take account of the past-future temporal guidance and boundary information jointly. In detail, we design a Dual Scale Aggregation (DSA) module for better temporal understanding by rethinking the affinity of the long-term and short-term frames for the clipped video. Next, we introduce Shadow Boundary Aware Attention (SBAA) to utilize the edge contexts for capturing the characteristics of shadows. Moreover, we are the first to introduce the Diffusion model for VSD in which we explore a Space-Time Encoded Embedding (STEE) to inject the temporal guidance for Diffusion to conduct shadow detection. Benefiting from these designs, our model can not only capture the temporal information but also the shadow property. Extensive experiments show that the performance of our approach overtakes the state-of-the-art methods, verifying the effectiveness of our components. We release the codes, weights, and results at \url{https://github.com/haipengzhou856/TBGDiff}.
Haipeng Zhou, Hongqiu Wang, Tian Ye 0001, Zhaohu Xing, Jun Ma 0008, Ping Li 0016, Qiong Wang 0001, Lei Zhu 0003
ACM Multimedia3
2024 RestoreAgent: Autonomous Image Restoration Agent via Multimodal Large Language Models
abstract
Natural images captured by mobile devices often suffer from multiple types of degradation, such as noise, blur, and low light. Traditional image restoration methods require manual selection of specific tasks, algorithms, and execution sequences, which is time-consuming and may yield suboptimal results. All-in-one models, though capable of handling multiple tasks, typically support only a limited range and often produce overly smooth, low-fidelity outcomes due to their broad data distribution fitting. To address these challenges, we first define a new pipeline for restoring images with multiple degradations, and then introduce RestoreAgent, an intelligent image restoration system leveraging multimodal large language models. RestoreAgent autonomously assesses the type and extent of degradation in input images and performs restoration through (1) determining the appropriate restoration tasks, (2) optimizing the task sequence, (3) selecting the most suitable models, and (4) executing the restoration. Experimental results demonstrate the superior performance of RestoreAgent in handling complex degradation, surpassing human experts. Furthermore, the system’s modular design facilitates the fast integration of new tasks and models.
Haoyu Chen 0003, Wenbo Li 0002, Jinjin Gu, Sixiang Chen, Tian Ye 0001, Renjing Pei, Kaiwen Zhou 0001, Fenglong Song, Lei Zhu 0003
NeurIPS6
2024 MPM: A Unified 2D-3D Human Pose Representation via Masked Pose Modeling
Zhenyu Zhang 0030, Wenhao Chai, Zhongyu Jiang, Tian Ye 0001, Mingli Song, Jenq-Neng Hwang, Gaoang Wang
PRCV (11)4
2024 Degradation-adaptive neural network for jointly single image dehazing and desnowing
Erkang Chen, Sixiang Chen, Tian Ye 0001, Yun Liu 0002
Frontiers Comput. Sci.3
2023 Five A+ Network: You Only Need 9K Parameters for Underwater Image Enhancement
Jingxia Jiang, Tian Ye 0001, Sixiang Chen, Erkang Chen, Yun Liu 0002, Jinbin Bai, Wenhao Chai
BMVC2
2023 MSP-Former: Multi-Scale Projection Transformer for Single Image Desnowing
abstract
Snow removal causes challenges due to its characteristic of complex degradations. To this end, targeted treatment of multi-scale snow degradations is critical for the network to learn effective snow removal. In order to handle the diverse scenes, we propose a multi-scale projection transformer (MSP-Former), which understands and covers a variety of snow degradation features in a multi-path manner, and integrates comprehensive scene context information for clean reconstruction via self-attention operation. For the local details of various snow degradations, the local capture module is introduced in parallel to assist in the rebuilding of a clean image. Such design achieves the SOTA performance on three desnowing benchmark datasets while costing the low parameters and computational complexity, providing a guarantee of practicality.
Sixiang Chen, Tian Ye 0001, Yun Liu 0002, Taodong Liao, Jingxia Jiang, Erkang Chen
ICASSP2
2023 DEHRFormer: Real-Time Transformer for Depth Estimation and Haze Removal from Varicolored Haze Scenes
abstract
Varicolored haze caused by chromatic casts poses haze removal and depth estimation challenges. Recent learning-based depth estimation methods are mainly targeted at dehazing first and estimating depth subsequently from haze-free scenes. This way, the inner connections between colored haze and scene depth are lost. In this paper, we propose a real-time transformer for simultaneous single image Depth Estimation and Haze Removal (DEHRFormer). DEHRFormer consists of a single encoder and two task-specific decoders. The transformer decoders with learnable queries are designed to decode coupling features from the task-agnostic encoder and project them into clean image and depth map, respectively. In addition, we introduce a novel learning paradigm that utilizes contrastive learning and domain consistency learning to tackle weak-generalization problem for real-world dehazing, while predicting the same depth map from the same scene with varicolored haze. Experiments demonstrate that DEHRFormer achieves significant performance improvement across diverse varicolored haze scenes over previous depth estimation networks and dehazing approaches.
Sixiang Chen, Tian Ye 0001, Yun Liu 0002, Jingxia Jiang, Erkang Chen
ICASSP2
2023 Sparse Sampling Transformer with Uncertainty-Driven Ranking for Unified Removal of Raindrops and Rain Streaks
abstract
In the real world, image degradations caused by rain often exhibit a combination of rain streaks and raindrops, thereby increasing the challenges of recovering the underlying clean image. Note that the rain streaks and raindrops have diverse shapes, sizes, and locations in the captured image, and thus modeling the correlation relationship between irregular degradations caused by rain artifacts is a necessary prerequisite for image deraining. This paper aims to present an efficient and flexible mechanism to learn and model degradation relationships in a global view, thereby achieving a unified removal of intricate rain scenes. To do so, we propose a Sparse Sampling Transformer based on Uncertainty-Driven Ranking, dubbed UDR-S2Former. Compared to previous methods, our UDR-S2Former has three merits. First, it can adaptively sample relevant image degradation information to model underlying degradation relationships. Second, explicit application of the uncertainty-driven ranking strategy can facilitate the network to attend to degradation features and understand the reconstruction process. Finally, experimental results show that our UDR-S2Former clearly outperforms state-of-the-art methods for all benchmarks.
Sixiang Chen, Tian Ye 0001, Jinbin Bai, Erkang Chen, Lei Zhu 0003
ICCV2
2023 Adverse Weather Removal with Codebook Priors
abstract
Despite recent advancements in unified adverse weather removal methods, there remains a significant challenge of achieving realistic fine-grained texture and reliable background reconstruction to mitigate serious distortions.Inspired by recent advancements in codebook and vector quantization (VQ) techniques, we present a novel Adverse Weather Removal network with Codebook Priors (AWRCP) to address the problem of unified adverse weather removal. AWRCP leverages high-quality codebook priors derived from undistorted images to recover vivid texture details and faithful background structures. However, simply utilizing high-quality features from the codebook does not guarantee good results in terms of fine-grained details and structural fidelity. Therefore, we develop a deformable cross-attention with sparse sampling mechanism for flexible perform feature interaction between degraded features and high-quality features from codebook priors. In order to effectively incorporate high-quality texture features while maintaining the realism of the details generated by codebook priors, we propose a hierarchical texture warping head that gradually fuses hierarchical codebook prior features into high-resolution features at final restoring stage.With the utilization of the VQ codebook as a feature dictionary of high quality and the proposed designs, AWRCP can largely improve the restored quality of texture details, achieving the state-of-the-art performance across multiple adverse weather removal benchmark.
Tian Ye 0001, Sixiang Chen, Jinbin Bai, Chenghao Xue, Jingxia Jiang, Junjie Yin, Erkang Chen, Yun Liu 0002
ICCV1
2023 RSFDM-Net: Real-Time Spatial and Frequency Domains Modulation Network for Underwater Image Enhancement
abstract
Underwater images typically experience mixed degradations of brightness and structure caused by the absorption and scattering of light by suspended particles. To address this issue, we propose a Real-time Spatial and Frequency Domains Modulation Network (RSFDM-Net) for the efficient enhancement of colors and details in underwater images. Specifically, our proposed conditional network is designed with Adaptive Fourier Gating Mechanism (AFGM) and Multiscale Convolutional Attention Module (MCAM) to generate vectors carrying low-frequency background information and high-frequency detail features, which effectively promote the network to model global background information and local texture details. To more precisely correct the color cast and low saturation of the image, we introduce a Three-branch Feature Extraction (TFE) block in the primary net that processes images pixel by pixel to integrate the color information extended by the same channel (R, G, or B). This block consists of three small branches, each of which has its own weights. Extensive experiments demonstrate that our network significantly outperforms over state-of-the-art methods in both visual quality and quantitative metrics.
Jingxia Jiang, Jinbin Bai, Yun Liu 0002, Junjie Yin, Sixiang Chen, Tian Ye 0001, Erkang Chen
ICIP6
2023 CPLFormer: Cross-scale Prototype Learning Transformer for Image Snow Removal
abstract
Removing snow from a single image poses a significant challenge within the image restoration domain, as snowfall's effects are in various scales and forms. Existing methods have tried to tackle this issue by using multi-scale approaches, but their reliance on targeted design for handling each single-scale feature has resulted in unsatisfactory performance. This is primarily due to a lack of cross-scale knowledge, making it difficult to effectively handle degradations. To this end, we propose a novel approach, CPLFormer, which uses snow prototypes to own comprehensive clean scene understanding through learning from cross-scale features, outperforming convolutional network and vanilla transformer-based solutions. CPLFormer has several advantages: firstly, learnable snow prototypes learn global context information from multiple scales to uncover hidden clean cues; secondly, prototypes can propagate cross-scale information to each patch through cross-attention to assist with clean patch reconstruction; thirdly, CPLFormer surpasses advanced state-of-the-art desnowing networks and the prevalent universal image restoration transformers on six synthetic and real-world benchmark tests.
Sixiang Chen, Tian Ye 0001, Yun Liu 0002, Jinbin Bai, Haoyu Chen 0003, Yunlong Lin, Erkang Chen
ACM Multimedia2
2023 Uncertainty-Driven Dynamic Degradation Perceiving and Background Modeling for Efficient Single Image Desnowing
abstract
Single-image snow removal aims to restore clean images from heterogeneous and irregular snow degradations. Recent methods utilize neural networks to remove various degradations directly. However, these approaches suffer from the limited ability to flexibly perceive complicated snow degradation patterns and insufficient representation of background structure information. To further improve the performance and generalization ability of snow removal, this paper aims to develop a novel and efficient paradigm from the perspective of degradation perceiving and background modeling.
Sixiang Chen, Tian Ye 0001, Chenghao Xue, Haoyu Chen 0003, Yun Liu 0002, Erkang Chen, Lei Zhu 0003
ACM Multimedia2
2023 NightHazeFormer: Single Nighttime Haze Removal Using Prior Query Transformer
abstract
Nighttime image dehazing is a challenging task due to the presence of multiple types of adverse degrading effects including glow, haze, blur, noise, color distortion, and so on. However, most previous studies mainly focus on daytime image dehazing or partial degradations presented in nighttime hazy scenes, which may lead to unsatisfactory restoration results. In this paper, we propose an end-to-end transformer-based framework for nighttime haze removal, called NightHazeFormer. Our proposed approach consists of two stages: supervised pre-training and semi-supervised fine-tuning. During the pre-training stage, we introduce two powerful priors into the transformer decoder to generate the non-learnable prior queries, which guide the model to extract specific degradations. For the fine-tuning, we combine the generated pseudo ground truths with input real-world nighttime hazy images as paired images and feed into the synthetic domain to fine-tune the pre-trained model. This semi-supervised fine-tuning paradigm helps improve the generalization to real domain. In addition, we also propose a large-scale synthetic dataset called UNREAL-NH, to simulate the real-world nighttime haze scenarios comprehensively. Extensive experiments on several synthetic and real-world datasets demonstrate the superiority of our NightHazeFormer over state-of-the-art nighttime haze removal methods in terms of both visually and quantitatively.
Yun Liu 0002, Zhongsheng Yan, Sixiang Chen, Tian Ye 0001, Wenqi Ren, Erkang Chen
ACM Multimedia4
2023 Sequential Affinity Learning for Video Restoration
abstract
Video restoration networks aim to restore high-quality frame sequences from degraded ones. However, traditional video restoration methods heavily rely on temporal modeling operators or optical flow estimation, which limits their versatility. The aim of this work is to present a novel approach for video restoration that eliminates inefficient temporal modeling operators and pixel-level feature alignment in the network architecture. The proposed method, Sequential Affinity Learning Network (SALN), is designed based on an affinity mechanism that establishes direct correspondences between the Query frame, degraded sequence, and restored frames in latent space. This unique perspective allows for more accurate and effective restoration of video content without relying on temporal modeling operators or optical flow estimation techniques. Moreover, we enhanced the design of the channel-wise self-attention block to improve the decoder's performance for video restoration. Our method outperformed previous state-of-the-art methods by a significant margin in several classic video tasks, including video deraining, video dehazing, and video waterdrop removal, demonstrating excellent efficiency. As a novel network that differs significantly from previous video restoration methods, SALN aims to provide innovative ideas and directions for video restoration. Our contributions include proposing a novel affinity-based approach for video restoration, enhancing the design of the channel-wise self-attention block, and achieving state-of-the-art performance on several classic video tasks.
Tian Ye 0001, Sixiang Chen, Yun Liu 0002, Wenhao Chai, Jinbin Bai, Wenbin Zou, Yunchen Zhang, Mingchao Jiang, Erkang Chen, Chenghao Xue
ACM Multimedia1
2022 Towards Real-Time High-Definition Image Snow Removal: Efficient Pyramid Network with Asymmetrical Encoder-Decoder Architecture
Tian Ye 0001, Sixiang Chen, Yun Liu 0002, Yi Ye, Jinbin Bai, Erkang Chen
ACCV (3)1
2022 Perceiving and Modeling Density for Image Dehazing
Tian Ye 0001, Yunchen Zhang, Mingchao Jiang, Liang Chen 0026, Yun Liu 0002, Sixiang Chen, Erkang Chen
ECCV (19)1
2022 Single nighttime image dehazing based on unified variational decomposition model and multi-scale contrast enhancement
Yun Liu 0002, Zhongsheng Yan, Tian Ye 0001, Aimin Wu, Yuche Li
Eng. Appl. Artif. Intell.3