EDBT 2026 Demo / reviewers in the wild / expert
Zhentao Tan
dblp:211/5776
· DBLP profile ↗
38ranked-venue papers
12as first author
34since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 25 · 6 first-author · 21 since 2021Artificial intelligence and machine learning · 23 · 8 first-author · 23 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Flora: Effortless Context Construction to Arbitrary Length and ScaleabstractEffectively handling long contexts is challenging for Large Language Models (LLMs) due to the rarity of long texts, high computational demands, and substantial forgetting of short-context abilities. Recent approaches have attempted to construct long contexts for instruction tuning, but these methods often require LLMs or human interventions, which are both costly and limited in length and diversity. Also, the drop in short-context performances of present long-context LLMs remains significant. In this paper, we introduce Flora, an effortless (human/LLM-free) long-context construction strategy. Flora can markedly enhance the long-context performance of LLMs by arbitrarily assembling short instructions based on categories and instructing LLMs to generate responses based on long-context meta-instructions. This enables Flora to produce contexts of arbitrary length and scale with rich diversity, while only slightly compromising short-context performance. Experiments on Llama3-8B-Instruct and QwQ-32B show that LLMs enhanced by Flora excel in three long-context benchmarks while maintaining strong performances in short-context tasks. Zhentao Tan, Xiaofan Bo, Qi Chu 0001, Jieping Ye |
AAAI | 2 |
| 2026 | MagicPaint: Operate Anything for Image Inpainting with Diffusion ModelabstractRecent diffusion-based models have significantly improved inpainting quality. However, existing methods struggle with multi-task inpainting due to conflicting optimization objectives, and current datasets are typically limited to task-specific scenarios, hindering joint training. To address these challenges, we propose MagicPaint, a unified diffusion-based inpainting model that supports object addition, removal, and unconditional inpainting across both text and image modalities. MagicPaint semantically decouples operation types and target content by learnable tokens in MMToken Module, effectively reconciling conflicting optimization objectives and enabling robust multi-task, multi-modal inpainting. Besides, a novel inpainting paradigm named MagicMask, encodes operating intent directly into the mask and applies a mask loss for spatially precise supervision. In addition, existing inpainting datasets are insufficient for multi-task and multi-modal scenarios, limiting the capability of inpainting models. Thus, we further introduce a new dataset comprising 2.1M image tuples. It is dedicatedly designed to support diverse inpainting scenarios and significantly improves upon existing datasets, particularly in object removal. Through efforts from both model and data perspectives, MagicPaint enables users to operate anything—add, remove or inpaint content which is specified through either text or image modalities in a seamless and unified manner. Extensive experiments demonstrate that MagicPaint achieves state-of-the-art performance across three key tasks (i.e., text-guided addition, image-guided addition, and object removal) and produces outputs with superior visual consistency and contextual fidelity compared to existing methods. Qinhong Yang, Dongdong Chen 0001, Qi Chu 0001, Qiankun Liu 0001, Zhentao Tan, Xulin Li, Huamin Feng, Nenghai Yu |
AAAI | 6 |
| 2026 | AnyPattern: Towards In-context Image Copy Detection
Yifan Sun 0003, Zhentao Tan, Yi Yang 0001 |
Int. J. Comput. Vis. | 3 |
| 2026 | Neural Wave Propagation for Surgical Video Action Recognition: A New Dataset and BaselineabstractAccurate and efficient recognition of surgical actions in videos is critical for advancing AI-driven surgical robotics. However, current surgical video action recognition (SVAR) datasets suffer from limitations such as small scale, low resolution, inconsistent annotations, and insufficient action coverage. Most latest video recognition models are trained on large-scale common datasets and underperform in SVAR due to architectures that suppress high-frequency visual details (crucial for recognizing surgical tools and motions) and lack a strong spatial inductive bias, requiring extensive training data for good convergence. This is particularly challenging in the surgical domain, where data access is limited. Therefore, a new baseline is required. To address these issues, we introduce LapSurg-230K, an SVAR dataset of 7,569 high-resolution laparoscopic surgical video clips with 230,246 frames, well-annotated for 11 key actions across 9 surgery types. It supports both full and progressive data volume evaluation settings. We further propose WaveR, an attention-free baseline based on physical wave propagation. WaveR embeds an innate physical inductive bias: each video patch acts as a wave source that propagates waves toward action-critical regions (e.g., instrument tips), adaptively aggregating spatial-temporal context while preserving high-frequency surgical cues. This mechanism eliminates dependency on massive training data. Experiments demonstrate WaveR's robustness under extreme data scarcity ( $\leq 30\%$ training samples), achieving state-of-the-art accuracy on both surgical video action recognition and phase recognition tasks. The complete dataset, licensed under CC-BY 4.0, is available at https://doi.org/10.6084/m9.figshare.32237319. Our code is available at https://github.com/yezizi1022/WaveR_TIP. Zhentao Tan, Ru Zhou, Le Lu 0001 |
IEEE Trans. Image Process. | 3 |
| 2025 | Training-free Open-Vocabulary Semantic Segmentation via Diverse Prototype Construction and Sub-region MatchingabstractOpen-vocabulary semantic segmentation (OVSS) aims to segment images of arbitrary categories specified by class labels. While previous approaches relied on extensive image-text pairs or dense semantic annotations, recent training-free methods attempted to overcome these limitations by constructing semantic prototypes in the construction stage and image-to-image matching (i.e., prototype matching) during testing. However, these methods often struggle to effectively capture the visual characteristics of categories and fail to utilize local features during prototype matching. To deal with these problems, we propose a novel training-free framework for OVSS that constructs diverse prototypes and performs fine-grained sub-region matching. Specifically, our method leverages Large Language Models (LLMs) to guide support image generation by descriptions of different attributes of categories and employs coarse-fine clustering to obtain diverse and robust part-level prototypes in the construction stage. During testing, we propose a sub-region matching method, which assigns part-level prototypes to sub-regions utilizing optimal transport, to fully utilize local image features among part-level prototypes. Extensive experiments demonstrate the effectiveness of our method and show that our method achieves state-of-the-art performance, outperforming previous methods across five datasets. Xuanpu Zhao, Dianmo Sheng, Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
AAAI | 3 |
| 2025 | UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype DiscoveryabstractRecent advancements in in-context segmentation generalists have demonstrated significant success in performing various image segmentation tasks using a limited number of labeled example images. However, real-world applications present challenges due to the variability of support examples, which often exhibit quality issues resulting from various sources and inaccurate labeling. How to extract more robust representations from these examples has always been one of the goals of in-context visual learning. In response, we propose UNICL-SAM, to better model the example distribution and extract robust representations to help in-context segmentation. We incorporate an uncertainty probabilistic module to quantify each example’s reliability during both the training and testing phases. Utilizing this uncertainty estimation, we introduce an uncertainty-guided graph augmentation and feature refinement strategy, aimed at mitigating the impact of high-uncertainty regions to enhance the learning of robust representations. Subsequently, we construct prototypes for each example by aggregating part information, thereby creating reliable in-context instruction that effectively represents fine-grained local semantics. This approach serves as a valuable complement to traditional global pooling features. Experimental results demonstrate the effectiveness of the proposed framework, underscoring its potential for real-world applications. Dianmo Sheng, Dongdong Chen 0001, Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Wenbin Tu, Shengwei Xu, Nenghai Yu |
CVPR | 3 |
| 2025 | SweetTok: Semantic-Aware Spatial-Temporal Tokenizer for Compact Video Discretization
Zhentao Tan, Ben Xue, Jian Jia, Wencai Ye, Shaoyun Shi, Mingjie Sun, Wenjin Wu, Quan Chen 0006, Peng Jiang 0002 |
ICCV | 1 |
| 2025 | Origin Identification for Text-Guided Image-to-Image Diffusion ModelsabstractText-guided image-to-image diffusion models excel in translating images based on textual prompts, allowing for precise and creative visual modifications. However, such a powerful technique can be misused for spreading misinformation, infringing on copyrights, and evading content tracing. This motivates us to introduce the task of origin IDentification for text-guided Image-to-image Diffusion models (ID$\mathbf{^2}$), aiming to retrieve the original image of a given translated query. A straightforward solution to ID$^2$ involves training a specialized deep embedding model to extract and compare features from both query and reference images. However, due to visual discrepancy across generations produced by different diffusion models, this similarity-based approach fails when training on images from one model and testing on those from another, limiting its effectiveness in real-world applications. To solve this challenge of the proposed ID$^2$ task, we contribute the first dataset and a theoretically guaranteed method, both emphasizing generalizability. The curated dataset, OriPID, contains abundant Origins and guided Prompts, which can be used to train and test potential IDentification models across various diffusion models. In the method section, we first prove the existence of a linear transformation that minimizes the distance between the pre-trained Variational Autoencoder embeddings of generated samples and their origins. Subsequently, it is demonstrated that such a simple linear transformation can be generalized across different diffusion models. Experimental results show that the proposed method achieves satisfying generalization performance, significantly surpassing similarity-based methods (+31.6% mAP), even those with generalization designs. The project is available at https://id2icml.github.io. Yifan Sun 0003, Zongxin Yang, Zhentao Tan, Zhengdong Hu, Yi Yang 0001 |
ICML | 4 |
| 2025 | Mixture-of-Noises Enhanced Forgery-Aware Predictor for Multi-Face Manipulation Detection and LocalizationabstractWith the advancement of face manipulation technology, forgery images in multi-face scenarios are gradually becoming a more complex and realistic challenge. Despite this, detection and localization methods for such multi-face manipulations remain underdeveloped. Traditional manipulation localization methods either indirectly derive detection results from localization masks, resulting in limited detection performance, or employ a naive two-branch structure to simultaneously obtain detection and localization results, which cannot effectively benefit the localization capability due to limited interaction between the two tasks. This paper proposes a new framework, namely MoNFAP, specifically tailored for multi-face manipulation detection and localization. The MoNFAP primarily introduces two novel modules: the Forgery-aware Unified Predictor (FUP) Module and the Mixture-of-Noises Module (MNM). The proposed FUP integrates detection and localization tasks using a token learning strategy and multiple forgery-aware transformers, which facilitates the use of classification information to enhance localization capability. Furthermore, to mitigate the interference from general semantic object information, we propose the MNM that leverages multiple noise extractors based on the mixture of experts concept. This allows the MNM to learn semantic-agnostic forgery features from general RGB features, further boosting the performance of our proposed framework. Finally, we establish a comprehensive benchmark for multi-face detection and localization, and the proposed MoNFAP achieves significant performance. The code is available: https://github.com/miaoct/MoNFAP. Changtao Miao, Qi Chu 0001, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Honggang Hu, Nenghai Yu |
ACM Multimedia | 4 |
| 2025 | Bootstrapping Audio-Visual Video Segmentation by Strengthening Audio CuesabstractHow to effectively interact audio with vision has garnered considerable interest within the multi-modality research field. Recently, a novel audio-visual video segmentation (AVS) task has been proposed, aiming to segment the sounding objects in video frames under the guidance of audio cues. However, most existing AVS methods are hindered by a modality imbalance where the visual features tend to dominate those of the audio modality, due to a unidirectional and insufficient integration of audio cues. This imbalance skews the feature representation towards the visual aspect, impeding the learning of joint audio-visual representations and potentially causing segmentation inaccuracies. To address this issue, we propose AVSAC. Our approach features a Bidirectional Audio-Visual Decoder (BAVD) with integrated bidirectional bridges, enhancing audio cues and fostering continuous interplay between audio and visual modalities. This bidirectional interaction narrows the modality imbalance, facilitating more effective learning of integrated audio-visual representations. Additionally, we present a strategy for audio-visual frame-wise synchrony as fine-grained guidance of BAVD. This strategy enhances the share of auditory components in visual features, contributing to a more balanced audio-visual representation learning. Extensive experiments show that our method has state-of-the-art performance on several AVS public benchmarks. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu, Le Lu 0001, Jieping Ye |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Zig-RiR: Zigzag RWKV-in-RWKV for Efficient Medical Image SegmentationabstractMedical image segmentation has made significant strides with the development of basic models. Specifically, models that combine CNNs with transformers can successfully extract both local and global features. However, these models inherit the transformer's quadratic computational complexity, limiting their efficiency. Inspired by the recent Receptance Weighted Key Value (RWKV) model, which achieves linear complexity for long-distance modeling, we explore its potential for medical image segmentation. While directly applying vision-RWKV yields suboptimal results due to insufficient local feature exploration and disrupted spatial continuity, we propose a novel nested structure, Zigzag RWKV-in-RWKV (Zig-RiR), to address these issues. It consists of Outer and Inner RWKV blocks to adeptly capture both global and local features without disrupting spatial continuity. We treat local patches as "visual sentences" and use the Outer Zig-RWKV to explore global information. Then, we decompose each sentence into sub-patches ("visual words") and use the Inner Zig-RWKV to further explore local information among words, at negligible computational cost. We also introduce a Zigzag-WKV attention mechanism to ensure spatial continuity during token scanning. By aggregating visual word and sentence features, our Zig-RiR can effectively explore both global and local information while preserving spatial continuity. Experiments on four medical image segmentation datasets of both 2D and 3D modalities demonstrate the superior accuracy and efficiency of our method, outperforming the state-of-the-art method 14.4 times in speed and reducing GPU memory usage by 89.5% when testing on ${1024} \times {1024}$ high-resolution medical images. Our code is available at https://github.com/txchen-USTC/Zig-RiR. Zhentao Tan, Qi Chu 0001, Nenghai Yu, Le Lu 0001 |
IEEE Trans. Medical Imaging | 3 |
| 2025 | Multi-spectral Class Center Network for Face Manipulation LocalizationabstractAs Deepfake content proliferates online, advancing face manipulation forensics has become crucial. To combat this emerging threat, previous methods mainly focus on studying how to distinguish authentic and manipulated face images. Although impressive, image-level classification lacks explainability and is limited to specific application scenarios, spurring recent research on pixel-level prediction for face manipulation forensics. However, existing forgery localization methods suffer from exploring frequency-based forgery traces in the localization network. In this paper, we observe that multi-frequency spectrum information is effective for identifying tampered regions. To this end, a novel Multi-spectral Class Center Network (MSCCNet) is proposed for face manipulation localization. Specifically, we design a Multi-spectral Class Center (MSCC) module to learn more generalizable and multi-frequency features. Based on the features of different frequency bands, the MSCC module collects multi-spectral class centers and computes pixel-to-class relations. Applying multi-spectral class-level representations suppresses the semantic information of the visual concepts which is insensitive to manipulated regions of forgery images. Furthermore, we propose a Multi-level Features Aggregation (MFA) module to employ more low-level forgery artifacts and structural textures. Meanwhile, we conduct a comprehensive localization benchmark based on pixel-level FF++ and Dolos datasets. Experimental results quantitatively and qualitatively demonstrate the effectiveness and superiority of the proposed MSCCNet. We expect this work to inspire more studies on pixel-level face manipulation localization. The codes are available. Changtao Miao, Qi Chu 0001, Zhentao Tan, Zhenchao Jin, Wanyi Zhuang, Bin Liu 0016, Honggang Hu, Nenghai Yu |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2024 | TCI-Former: Thermal Conduction-Inspired Transformer for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) is critical to national security and has been extensively applied in military areas. ISTD aims to segment small target pixels from background. Most ISTD networks focus on designing feature extraction blocks or feature fusion modules, but rarely describe the ISTD process from the feature map evolution perspective. In the ISTD process, the network attention gradually shifts towards target areas. We abstract this process as the directional movement of feature map pixels to target areas through convolution, pooling and interactions with surrounding pixels, which can be analogous to the movement of thermal particles constrained by surrounding variables and particles. In light of this analogy, we propose Thermal Conduction-Inspired Transformer (TCI-Former) based on the theoretical principles of thermal conduction. According to thermal conduction differential equation in heat dynamics, we derive the pixel movement differential equation (PMDE) in the image domain and further develop two modules: Thermal Conduction-Inspired Attention (TCIA) and Thermal Conduction Boundary Module (TCBM). TCIA incorporates finite difference method with PMDE to reach a numerical approximation so that target body features can be extracted. To further remove errors in boundary areas, TCBM is designed and supervised by boundary masks to refine target body features with fine boundary details. Experiments on IRSTD-1k and NUAA-SIRST demonstrate the superiority of our method. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
AAAI | 2 |
| 2024 | Towards More Unified In-Context Visual UnderstandingabstractThe rapid advancement of large language models (LLMs) has accelerated the emergence of in-context learning (ICL) as a cutting-edge approach in the natural language processing domain. Recently, ICL has been employed in visual understanding tasks, such as semantic segmentation and image captioning, yielding promising results. However, existing visual ICL framework can not enable producing content across multiple modalities, whicd limits their potential usage scenarios. To address this issue, we present a new ICLframeworkfor visual understanding with multi-modal output enabled. First, we quantize and embed both text and visual prompt into a unified representational space, structured as interleaved in-context sequences. Then a decoder-only sparse transformer architecture is employed to perform generative modeling on them, facilitating in-context learning. Thanks to this design, the model is capable of handling in-context vision understanding tasks with multimodal output in a unified pipeline. Experimental re-sults demonstrate that our model achieves competitive performance compared with specialized models and previous ICL baselines. Overall, our research takes a further step toward unified multimodal in-context learning. Dianmo Sheng, Dongdong Chen 0001, Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Jianmin Bao, Bin Liu 0016, Shengwei Xu, Nenghai Yu |
CVPR | 3 |
| 2024 | SimAC: A Simple Anti-Customization Method for Protecting Face Privacy Against Text-to-Image Synthesis of Diffusion ModelsabstractDespite the success of diffusion-based customization methods on visual content creation, increasing concerns have been raised about such techniques from both privacy and political perspectives. To tackle this issue, several anti-customization methods have been proposed in very recent months, predominantly grounded in adversarial attacks. Unfortunately, most of these methods adopt straight-forward designs, such as end-to-end optimization with afocus on adversarially maximizing the original training loss, thereby neglecting nuanced internal properties intrinsic to the diffusion model, and even leading to ineffective optimization in some diffusion time steps. In this paper, we strive to bridge this gap by undertaking a comprehensive exploration of these inherent properties, to boost the performance of current anti-customization approaches. Two aspects of properties are investigated: 1) We examine the relationship between time step selection and the model's perception in the frequency domain of images and find that lower time steps can give much more contributions to adversarial noises. This inspires us to propose an adaptive greedy search for optimal time steps that seamlessly integrates with existing anti-customization methods. 2) We scrutinize the roles of features at different layers during denoising and devise a sophisticated feature-based optimization framework for anti-customization. Experiments on facial benchmarks demonstrate that our approach significantly increases identity disruption, thereby protecting user privacy and copyright. Our code is available at:: http//github.com/somuchtome/SimAC. Zhentao Tan, Tianyi Wei, Qidong Huang |
CVPR | 2 |
| 2024 | Boosting Vanilla Lightweight Vision Transformers via Re-parameterizationabstractLarge-scale Vision Transformers have achieved promising performance on downstream tasks through feature pre-training. However, the performance of vanilla lightweight Vision Transformers (ViTs) is still far from satisfactory compared to that of recent lightweight CNNs or hybrid networks. In this paper, we aim to unlock the potential of vanilla lightweight ViTs by exploring the adaptation of the widely-used re-parameterization technology to ViTs for improving learning ability during training without increasing the inference cost. The main challenge comes from the fact that CNNs perfectly complement with re-parameterization over convolution and batch normalization, while vanilla Transformer architectures are mainly comprised of linear and layer normalization layers. We propose to incorporate the nonlinear ensemble into linear layers by expanding the depth of the linear layers with batch normalization and fusing multiple linear features with hierarchical representation ability through a pyramid structure. We also discover and solve a new transformer-specific distribution rectification problem caused by multi-branch re-parameterization. Finally, we propose our Two-Dimensional Re-parameterized Linear module (TDRL) for ViTs. Under the popular self-supervised pre-training and supervised fine-tuning strategy, our TDRL can be used in these two stages to enhance both generic and task-specific representation. Experiments demonstrate that our proposed method not only boosts the performance of vanilla Vit-Tiny on various vision tasks to new state-of-the-art (SOTA) but also shows promising generality ability on other networks. Code will be available. Zhentao Tan, Qi Chu 0001, Le Lu 0001, Nenghai Yu, Jieping Ye |
ICLR | 1 |
| 2024 | Learning Solution-Aware Transformers for Efficiently Solving Quadratic Assignment ProblemabstractRecently various optimization problems, such as Mixed Integer Linear Programming Problems (MILPs), have undergone comprehensive investigation, leveraging the capabilities of machine learning. This work focuses on learning-based solutions for efficiently solving the Quadratic Assignment Problem (QAPs), which stands as a formidable challenge in combinatorial optimization. While many instances of simpler problems admit fully polynomial-time approximate solution (FPTAS), QAP is shown to be strongly NPhard. Even finding a FPTAS for QAP is difficult, in the sense that the existence of a FPTAS implies P = NP. Current research on QAPs suffer from limited scale and computational inefficiency. To attack the aforementioned issues, we here propose the first solution of its kind for QAP in the learn-to-improve category. This work encodes facility and location nodes separately, instead of forming computationally intensive association graphs prevalent in current approaches. This design choice enables scalability to larger problem sizes. Furthermore, a Solution AWare Transformer (SAWT) architecture integrates the incumbent solution matrix with the attention score to effectively capture higher-order information of the QAPs. Our model’s effectiveness is validated through extensive experiments on self-generated QAP instances of varying sizes and the QAPLIB benchmark. Zhentao Tan, Yadong Mu |
ICML | 1 |
| 2024 | Image Copy Detection for Diffusion ModelsabstractImages produced by diffusion models are increasingly popular in digital artwork and visual marketing. However, such generated images might replicate content from existing ones and pose the challenge of content originality. Existing Image Copy Detection (ICD) models, though accurate in detecting hand-crafted replicas, overlook the challenge from diffusion models. This motivates us to introduce ICDiff, the first ICD specialized for diffusion models. To this end, we construct a Diffusion-Replication (D-Rep) dataset and correspondingly propose a novel deep embedding method. D-Rep uses a state-of-the-art diffusion model (Stable Diffusion V1.5) to generate 40, 000 image-replica pairs, which are manually annotated into 6 replication levels ranging from 0 (no replication) to 5 (total replication). Our method, PDF-Embedding, transforms the replication level of each image-replica pair into a probability density function (PDF) as the supervision signal. The intuition is that the probability of neighboring replication levels should be continuous and smooth. Experimental results show that PDF-Embedding surpasses protocol-driven methods and non-PDF choices on the D-Rep test set. Moreover, by utilizing PDF-Embedding, we find that the replication ratios of well-known diffusion models against an open-source gallery range from 10% to 20%. The project is publicly available at https://icdiff.github.io/. Yifan Sun 0003, Zhentao Tan, Yi Yang 0001 |
NeurIPS | 3 |
| 2024 | Vision transformers are active learners for image copy detectionabstractImage Copy Detection (ICD) is developed to identify and track duplicated or manipulated images. The majority of existing methods rely on Convolutional Neural Networks (CNNs) and are trained using unsupervised learning techniques, which leads to subpar performance. We discover that by carefully designing the training process, Vision Transformer (ViT) backbones yield superior results. Specifically, directly training a ViT for ICD often leads to overfitting on the training images, which in turn results in poor generalization to unseen (test) images. Consequently, we initially train a CNN (such as ResNet-50), and during the ViT training, the distances between the features of CNN and ViT are regularized. We also incorporate an active learning method to further enhance performance. Notably, due to the visual discrepancy between auto-generated transformations and those used in the query set, we incorporate a small number (approximately 0.5% of unlabeled training images) of manually produced and labeled positive pairs. Training models on these pairs results in a significant performance boost though with little cost. Experimental findings demonstrate the effectiveness of our approach, and our method achieves state-of-the-art performance. Our code is available at: https://github.com/WangWenhao0716/ViT4ICD. Zhentao Tan, Caifeng Shan |
Neurocomputing | 1 |
| 2024 | Transformer Based Pluralistic Image Completion With Reduced Information LossabstractTransformer based methods have achieved great success in image inpainting recently. However, we find that these solutions regard each pixel as a token, thus suffering from an information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration. 2) They quantize 2563RGB values to a small number (such as 512) of quantized color values. The indices of quantized pixels are used as tokens for the inputs and prediction targets of the transformer. To mitigate these issues, we propose a new transformer based framework called “PUT”. Specifically, to avoid input downsampling while maintaining computation efficiency, we design a patch-based auto-encoder P-VQVAE. The encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by input quantization, an Un-quantized Transformer is applied. It directly takes features from the P-VQVAE encoder as input without any quantization and only regards the quantized tokens as prediction targets.Furthermore, to make the inpainting process more controllable, we introduce semantic and structural conditions as extra guidance. Extensive experiments show that our method greatly outperforms existing transformer based methods on image fidelity and achieves much higher diversity and better fidelity than state-of-the-art pluralistic inpainting methods on complex large-scale datasets (e.g., ImageNet). Codes are available athttps://github.com/liuqk3/PUT. Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Ying Fu 0001, Qi Chu 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Hierarchical reinforcement learning for chip-macro placement in integrated circuit
Zhentao Tan, Yadong Mu |
Pattern Recognit. Lett. | 1 |
| 2024 | Feature Preservation and Shape Cues Assist Infrared Small Target DetectionabstractInfrared small target detection (ISTD) aims to segment small target pixels from infrared images and has extensive applications in many fields. Despite multiple progress, challenges remain as present methods still easily suffer from missed detection. Also, present methods are not sensitive enough to irregular target shapes. We argue that the main reason is that some informative small target features get lost during the aggressive downsampling in the encoder without effective recovery. In this article, we propose a new network with a dual-branch encoder-decoder structure for ISTD to address the two challenges. Specifically, to better preserve small target body features for more accurate target locations, we propose to maintain a relatively high resolution of feature maps in one encoder branch. For the other encoder branch, we gradually enlarge feature channels while shrinking resolutions and devise Perona-Malik diffusion (PMD) blocks to preserve shape cues inspired by the shape-preserving effect of PMD in denoising. The encoded high-resolution target body features and high-channel shape cues actually complement each other, so we design channel-resolution interact modules (CRIMs) to combine them. In the decoder, we propose orthogonal central difference fusion (OCDF) that relies on mining contrast differences to further refine shape-aware ISTD quality. Experiments on NUAA-SIRST and IRSTD-1k prove the superiority of our method. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | MiM-ISTD: Mamba-in-Mamba for Efficient Infrared Small-Target DetectionabstractRecently, infrared small-target detection (ISTD) has made significant progress, thanks to the development of basic models. Specifically, the models combining CNNs with Transformers can successfully extract both local and global features. However, the disadvantage of the Transformer is also inherited, that is, the quadratic computational complexity to sequence length. Inspired by the recent basic model with linear complexity for long-distance modeling, Mamba, we explore the potential of this state-space model (SSM) for ISTD tasks in terms of effectiveness and efficiency in the article. However, directly applying Mamba achieves suboptimal performances due to the insufficient harnessing of local features, which are imperative for detecting small targets. Instead, we tailor a nested structure, Mamba-in-Mamba (MiM-ISTD), for efficient ISTD. It consists of Outer and Inner Mamba blocks to adeptly capture both global and local features. Specifically, we treat the local patches as “visual sentences” and use the Outer Mamba to explore the global information. We then decompose each visual sentence into subpatches as “visual words” and use the Inner Mamba to further explore the local information among words in the visual sentence with negligible computational costs. By aggregating the visual word and visual sentence features, our MiM-ISTD can effectively explore both global and local information. Experiments on NUAA-SIRST and IRSTD-1k show the superior accuracy and efficiency of our method. Specifically, MiM-ISTD is$8\times $faster than the SOTA method and reduces GPU memory usage by 62.2% when testing on$2048 \times 2048$images, overcoming the computation and memory constraints on high-resolution infrared images. Zhentao Tan, Qi Chu 0001, Bin Liu 0016, Nenghai Yu, Jieping Ye |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Exploring the Application of Large-Scale Pre-Trained Models on Adverse Weather RemovalabstractImage restoration under adverse weather conditions (e.g., rain, snow, and haze) is a fundamental computer vision problem that has important implications for various downstream applications. Distinct from early methods that are specially designed for specific types of weather, recent works tend to simultaneously remove various adverse weather effects based on either spatial feature representation learning or semantic information embedding. Inspired by various successful applications incorporating large-scale pre-trained models (e.g., CLIP), in this paper, we explore their potential benefits for leveraging large-scale pre-trained models in this task based on both spatial feature representation learning and semantic information embedding aspects: 1) spatial feature representation learning, we design a Spatially Adaptive Residual (SAR) encoder to adaptively extract degraded areas. To facilitate training of this model, we propose a Soft Residual Distillation (CLIP-SRD) strategy to transfer spatial knowledge from CLIP between clean and adverse weather images; 2) semantic information embedding, we propose a CLIP Weather Prior (CWP) embedding module to enable the network to adaptively respond to different weather conditions. This module integrates the sample-specific weather priors extracted by the CLIP image encoder with the distribution-specific information (as learned by a set of parameters) and embeds these elements using a cross-attention mechanism. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art performance under various and severe adverse weather conditions. The code will be made available. Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Le Lu 0001, Jieping Ye, Nenghai Yu |
IEEE Trans. Image Process. | 1 |
| 2023 | BAUENet: Boundary-Aware Uncertainty Enhanced Network for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) is indispensable in remote sensing and military surveillance. Existing ISTD methods can discover regularly-shaped and clear objects well, but tend to overlook the tough-to-detect ones, such as targets with irregular shapes or blurry boundaries, causing inaccurate segmentation and missed detection. Considering that boundary areas assemble rich uncertainty information, we propose the Boundary-Aware Uncertainty Enhanced Network (BAUENet), where Uncertainty Enhanced Context Refinement (UECR) and Adaptive Feature Fusion Modules (AFFM) are devised to address this problem. Specifically, UECR extracts spatial contexts and refines them with uncertain area maps derived from backbone intermediate outputs, so as to distinguish boundary areas from other regions. AFFM adaptively aggregates cross-level features via balancing low-level details and high-level semantics for finer boundary preservation in both channel and spatial dimensions during up-sampling feature fusion. Experiments on several public datasets demonstrate the effectiveness of the proposed method, especially for irregular shape and blurry boundary cases. Qi Chu 0001, Zhentao Tan, Bin Liu 0016, Nenghai Yu |
ICASSP | 3 |
| 2023 | Video Action Segmentation via Contextually Refined Temporal KeypointsabstractVideo action segmentation involves categorizing each frame or short snippet of an untrimmed video into predefined action categories. Despite notable advancements in recent years, a considerable number of current approaches still rely on frame-wise segmentation that tends to render fragmentary results. To address it, we present an innovative approach for video action segmentation, centered around contextually refined temporal keypoints. Initially, our method identifies a set of sparse, over-complete temporal keypoints through non-local visual cues, with each keypoint representing a potential action segment candidate. Subsequent enhancements to these initial keypoints are achieved through iterative refining and re-assembling operations. Driven by the notion that optimal temporal keypoints should collectively resemble the true ground-truth structurally, we introduce a module that conducts graph matching between the keypoint-derived graph and the reference graph constructed from accurate annotations. This module effectively learns structural features used to further refine the initial keypoints. Moreover, a set of predefined rules is applied to re-assemble all temporal keypoints. The unfiltered temporal keypoints, resulting from these operations, are harnessed to generate the final action segments. We extensively evaluate our method across three video benchmarks: 50salads, GTEA, and Breakfast. Our proposed approach consistently demonstrates substantial improvements over existing methods, establishing its superiority in video action segmentation. It achieves F 1@50 scores (one of the key performance metrics for this task) of 79.5%, 83.4%, and 60.5%, respectively, v.s. previous state-of-the-art 78.5%, 79.8% and 57.4%. Borui Jiang, Zhentao Tan, Yadong Mu |
ICCV | 3 |
| 2023 | ABMNet: Coupling Transformer with CNN Based on Adams-Bashforth-Moulton Method for Infrared Small Target DetectionabstractInfrared small target detection (ISTD) aims at segmenting the small targets from infrared images, which has wide applications in military surveillance. Present methods are mainly based on CNN and focus on modelling locality while ignoring global dependencies, which are indispensable because the local areas similar to small targets always spread over most of the background, causing heavy target ambiguity. Recently, RKformer [1] has combined local features with global dependencies and further introduced Runge-Kutta method, a one-step Ordinary Differential Equation (ODE) solver, to ISTD and performed well. However, the method simply fuses features from original transformer and residual blocks by naive concatenation, causing insufficient feature interaction. Also, it inevitably brings effective information loss, which greatly impairs ambiguous target features. To address above problems and target ambiguity, we introduce Adams-Bashforth-Moulton method and propose ABMNet, which has (1) multi-step memory and self-rectification mechanisms, guaranteeing more sufficient information usage and more accurate detection, (2) and achieves more sufficient interaction of both local and global information. Experiments on MDFA and IRSTD-1k demonstrate the superiority of our method. Qi Chu 0001, Zhentao Tan, Bin Liu 0016, Nenghai Yu |
ICME | 3 |
| 2023 | Semantic Probability Distribution Modeling for Diverse Semantic Image SynthesisabstractSemantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level or even instance-level multimodal results, still remains a challenge. In this article, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at both semantics and instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. To further extend the user interaction function of the proposed method, we also introduce sketches into the network. In addition, specially designed generator modules, Progressive Growing Module and Multi-Scale Refinement Module, can be used as a general module to improve the performance of complex scene generation. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Codes are available at https://github.com/tzt101/INADE.git. Zhentao Tan, Qi Chu 0001, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qiankun Liu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Reduce Information Loss in Transformers for Pluralistic Image InpaintingabstractTransformers have achieved great success in pluralistic image inpainting recently. However, we find existing transformer based solutions regard each pixel as a token, thus suffer from information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration, incurring information loss and extra misalignment for the boundaries of masked regions. 2) They quantize 2563RGB pixels to a small number (such as 512) of quantized pixels. The indices of quantized pixels are used as tokens for the inputs and prediction targets of transformer. Although an extra CNN network is used to upsample and refine the low-resolution results, it is difficult to retrieve the lost information back. To keep input information as much as possible, we propose a new transformer based framework “PUT”. Specifically, to avoid input downsampling while maintaining the computation efficiency, we design a patch-based auto-encoder P-VQVAE, where the encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by quantization, an Un-Quantized Transformer (UQ-Transformer) is applied, which directly takes the features from P-VQVAE encoder as input without quantization and regards the quantized tokens only as prediction targets. Extensive experiments show that PUT greatly outperforms state-of-the-art methods on image fidelity, especially for large masked regions and complex large-scale datasets. Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Qi Chu 0001, Xiyang Dai, Yinpeng Chen, Mengchen Liu, Lu Yuan 0001, Nenghai Yu |
CVPR | 2 |
| 2022 | HairCLIP: Design Your Hair by Text and Reference ImageabstractHair editing is an interesting and challenging problem in computer vision and graphics. Many existing methods require well-drawn sketches or masks as conditional inputs for editing, however these interactions are neither straight-forward nor efficient. In order to free users from the tedious interaction process, this paper proposes a new hair editing interaction mode, which enables manipulating hair attributes individually or jointly based on the texts or reference images provided by users. For this purpose, we encode the image and text conditions in a shared embedding space and propose a unified hair editing framework by leveraging the powerful image text representation capability of the Contrastive Language-Image Pre-Training (CLIP) model. With the carefully designed network structures and loss functions, our framework can perform high-quality hair editing in a disentangled manner. Extensive experiments demonstrate the superiority of our approach in terms of manipulation accuracy, visual realism of editing results, and irrelevant attribute preservation. Tianyi Wei, Dongdong Chen 0001, Wenbo Zhou 0004, Jing Liao 0001, Zhentao Tan, Lu Yuan 0001, Weiming Zhang 0001, Nenghai Yu |
CVPR | 5 |
| 2022 | UIA-ViT: Unsupervised Inconsistency-Aware Method Based on Vision Transformer for Face Forgery Detection
Wanyi Zhuang, Qi Chu 0001, Zhentao Tan, Qiankun Liu 0001, Changtao Miao, Zixiang Luo, Nenghai Yu |
ECCV (5) | 3 |
| 2022 | Efficient Semantic Image Synthesis via Class-Adaptive NormalizationabstractSpatially-adaptive normalization (SPADE) is remarkably successful recently in conditional semantic image synthesis in T. Park et al. 2019 which modulates the normalized activation with spatially-varying transformations learned from semantic layouts, to prevent the semantic information from being washed away. Despite its impressive performance, a more thorough understanding of the advantages inside the box is still highly demanded to help reduce the significant computation and parameter overhead introduced by this novel structure. In this paper, from a return-on-investment point of view, we conduct an in-depth analysis of the effectiveness of this spatially-adaptive normalization and observe that its modulation parameters benefit more from semantic-awareness rather than spatial-adaptiveness, especially for high-resolution input masks. Inspired by this observation, we propose class-adaptive normalization (CLADE), a lightweight but equally-effective variant that is only adaptive to semantic class. In order to further improve spatial-adaptiveness, we introduce intra-class positional map encoding calculated from semantic layouts to modulate the normalization parameters of CLADE and propose a truly spatially-adaptive variant of CLADE, namely CLADE-ICPE. Through extensive experiments on multiple challenging datasets, we demonstrate that the proposed CLADE can be generalized to different SPADE-based methods while achieving comparable generation quality compared to SPADE, but it is much more efficient with fewer extra parameters and lower computational cost. The code and pretrained models are available at https://github.com/tzt101/CLADE.git. Zhentao Tan, Dongdong Chen 0001, Qi Chu 0001, Menglei Chai, Jing Liao 0001, Mingming He, Lu Yuan 0001, Gang Hua 0001, Nenghai Yu |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Diverse Semantic Image Synthesis via Probability Distribution ModelingabstractSemantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level multimodal results, still remains a challenge. In this paper, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at semantic or even instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Code will be available at https://github.com/tzt101/INADE.git Zhentao Tan, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qi Chu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu |
CVPR | 1 |
| 2021 | Real Time Video Object Segmentation in Compressed DomainabstractMany of the recent methods for semi-supervised video object segmentation are still far from being applicable for real time applications due to their slow inference speed. Therefore, we explore a propagation based segmentation method in compressed domain to accelerate inference speed in this paper. In particular, we only extract the features of I-frames by traditional deep convolutional neural network and produce the features of P-frames through information flow propagation. In the process of feature propagation, we propose two effective components to enhance the representation ability of simply warped features in terms of appearance and location. Specifically, we propose a residual supplement module to supplement appearance information which is lost in direct warping and a spatial attention module that can mine extra spatial saliency to provide the location information of the specified object. Besides, we propose a metric based decoder module which consists of a feature match module and a multi-level refinement module to transform information from semantic representation to shape segmentation mask. Extensive experiments on several video datasets demonstrate that the proposed method can achieve comparable accuracy while much faster inference speed when compared to the state-of-the-art algorithms. Zhentao Tan, Bin Liu 0016, Qi Chu 0001, Hangshi Zhong, Weihai Li, Nenghai Yu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | MichiGAN: multi-input-conditioned hair image generation for portrait editingabstractDespite the recent success of face image generation with GANs, conditional hair editing remains challenging due to the under-explored complexity of its geometry and appearance. In this paper, we present MichiGAN (Multi-Input-Conditioned Hair Image GAN), a novel conditional image generation method for interactive portrait hair manipulation. To provide user control over every major hair visual factor, we explicitly disentangle hair into four orthogonal attributes, including shape, structure, appearance, and background. For each of them, we design a corresponding condition module to represent, process, and convert user inputs, and modulate the image generation pipeline in ways that respect the natures of different visual attributes. All these condition modules are integrated with the backbone generator to form the final end-to-end network, which allows fully-conditioned hair generation from multiple user inputs. Upon it, we also build an interactive portrait hair editing system that enables straightforward manipulation of hair by projecting intuitive and high-level user inputs such as painted masks, guiding strokes, or reference photos to well-defined condition representations. Through extensive experiments and evaluations, we demonstrate the superiority of our method regarding both result quality and user controllability. Zhentao Tan, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qi Chu 0001, Lu Yuan 0001, Sergey Tulyakov, Nenghai Yu |
ACM Trans. Graph. | 1 |
| 2019 | Enhanced Video Segmentation with Object Tracking
Zheran Hong, Zhentao Tan, Qiankun Liu 0001, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 3 |
| 2019 | PPML: Metric Learning with Prior Probability for Video Object SegmentationabstractVideo object segmentation plays an important role in computer vision and has attracted much attention. Although many recent works have removed the fine-tuning process in pursuit of fast inference speed, while achieving high segmentation accuracy, they are still far from being real-time. In this paper, we regard this task as a feature matching problem and propose a prior probability based metric learning (PPML) method for faster inference speed and higher segmentation accuracy. The proposed method consists of two ingredients: a novel template space updating strategy that improves the efficiency of segmentation by avoiding the explosion of data in template space, and a novel feature matching method which applies more potential probability information through integrating the prior of the first frame and the predicted score of previous frames. Experimental results on DAVIS datasets demonstrate that the proposed method reaches the state-of-the-art competitive performance and is more efficient in time consumption. Hangshi Zhong, Zhentao Tan, Bin Liu 0016, Weihai Li, Nenghai Yu |
VCIP | 2 |
| 2017 | PPEDNet: Pyramid Pooling Encoder-Decoder Network for Real-Time Semantic Segmentation
Zhentao Tan, Bin Liu 0016, Nenghai Yu |
ICIG (1) | 1 |