Tao Zhang 0042

dblp:15/4777-42 · DBLP profile ↗
← Back
37ranked-venue papers
11as first author
33since 2021 · last 2026
0000-0002-7358-0603ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 26 · 10 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 1 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Real Noise Decoupling for Hyperspectral Image Denoising
abstract
Hyperspectral image (HSI) denoising is a crucial step in enhancing the quality of HSIs. Noise modeling methods can fit noise distributions to generate synthetic HSIs to train denoising networks. However, the noise in captured HSIs is usually complex and difficult to model accurately, which significantly limits the effectiveness of these approaches. In this paper, we propose a multi-stage noise-decoupling framework that decomposes complex noise into explicitly modeled and implicitly modeled components. This decoupling reduces the complexity of noise and enhances the learnability of HSI denoising methods when applied to real paired data. Specifically, for explicitly modeled noise, we utilize an existing noise model to generate paired data for pre-training a denoising network, equipping it with prior knowledge to handle the explicitly modeled noise effectively. For implicitly modeled noise, we introduce a high-frequency wavelet guided network. Leveraging the prior knowledge from the pre-trained module, this network adaptively extracts high-frequency features to target and remove the implicitly modeled noise from real paired HSIs. Furthermore, to effectively eliminate all noise components and mitigate error accumulation across stages, a multi-stage learning strategy, comprising separate pre-training and joint fine-tuning, is employed to optimize the entire framework. Extensive experiments on public and our captured datasets demonstrate that our proposed framework outperforms state-of-the-art methods, effectively handling complex real-world noise and significantly enhancing HSI quality.
Yingkai Zhang, Tao Zhang 0042, Ying Fu 0001
AAAI2
2026 Secrecy Energy Efficiency Maximization for RIS-Assisted UAV-MEC Networks: A Deep Reinforcement Learning Based Approach
Shutian Li, Jinguo Zhang, Tao Zhang 0042, Nan Qi 0001
WCNC4
2025 Point Cloud Mamba: Point Cloud Learning via State Space Model
abstract
Recently, state space models have exhibited strong global modeling capabilities and linear computational complexity in contrast to transformers. This research focuses on applying such architecture to more efficiently and effectively model point cloud data globally with linear computational complexity. In particular, for the first time, we demonstrate that Mamba-based point cloud methods can outperform previous methods based on transformer or multi-layer perceptrons (MLPs). To enable Mamba to process 3-D point cloud data more effectively, we propose a novel Consistent Traverse Serialization method to convert point clouds into 1-D point sequences while ensuring that neighboring points in the sequence are also spatially adjacent. Consistent Traverse Serialization yields six variants by permuting the order of x, y, and z coordinates, and the synergistic use of these variants aids Mamba in comprehensively observing point cloud data. Furthermore, to assist Mamba in handling point sequences with different orders more effectively, we introduce point prompts to inform Mamba of the sequence’s arrangement rules. Finally, we propose positional encoding based on spatial coordinate mapping to inject positional information into point cloud sequences more effectively. Point Cloud Mamba surpasses the state-of-the-art (SOTA) point-based method PointNeXt and achieves new SOTA performance on the ScanObjectNN, ModelNet40, ShapeNetPart, and S3DIS datasets. It is worth mentioning that when using a more powerful local feature extraction module, our PCM achieves 79.6 mIoU on S3DIS, significantly surpassing the previous SOTA models, DeLA and PTv3, by 5.5 mIoU and 4.9 mIoU, respectively.
Tao Zhang 0042, Haobo Yuan, Lu Qi 0001, Jiangning Zhang, Qianyu Zhou 0001, Shunping Ji, Shuicheng Yan, Xiangtai Li
AAAI1
2025 Noise Calibration and Spatial-Frequency Interactive Network for STEM Image Enhancement
abstract
Scanning Transmission Electron Microscopy (STEM) enables the observation of atomic arrangements at sub-angstrom resolution, allowing for atomically resolved analysis of the physical and chemical properties of materials. However, due to the effects of noise, electron beam damage, sample thickness, etc, obtaining satisfactory atomic-level images is often challenging. Enhancing STEM images can reveal clearer structural details of materials. Nonetheless, existing STEM image enhancement methods usually overlook unique features in the frequency domain, and existing datasets lack realism and generality. To resolve these issues, in this paper, we develop noise calibration, data synthesis, and enhancement methods for STEM images. We first present a STEM noise calibration method, which is used to synthesize more realistic STEM images. The parameters of background noise, scan noise, and pointwise noise are obtained by statistical analysis and fitting of real STEM images containing atoms. Then we use these parameters to develop a more general dataset that considers both regular and random atomic arrangements and includes both HAADF and BF mode images. Finally, we design a spatial-frequency interactive network for STEM image enhancement, which can explore the information in the frequency domain formed by the periodicity of atomic arrangement. Experimental results show that our data is closer to real STEM images and achieves better enhancement performances together with our network. Code will be available at https://github.com/HeasonLee/SFIN.
Hesong Li, Ziqi Wu, Ruiwen Shao, Tao Zhang 0042, Ying Fu 0001
CVPR4
2025 Are They the Same? Exploring Visual Correspondence Shortcomings of Multimodal LLMs
abstract
Recent advancements in multimodal large language models (MLLM) have shown a strong ability in visual perception, reasoning abilities, and vision-language understanding. However, the visual matching ability of MLLMs is rarely studied, despite finding the visual correspondence of objects is essential in computer vision. Our research reveals that the matching capabilities in recent MLLMs still exhibit systematic shortcomings, even with current strong MLLMs models, GPT-4o. In particular, we construct a Multimodal Visual Matching (MMVM) benchmark to fairly benchmark over 30 different MLLMs. The MMVM benchmark is built from 15 open-source datasets and Internet videos with manual annotation. We categorize the data samples of MMVM benchmark into eight aspects based on the required cues and capabilities to more comprehensively evaluate and analyze current MLLMs. In addition, we have designed an automatic annotation pipeline to generate the MMVM SFT dataset, including 220K visual matching data with reasoning annotation. To our knowledge, this is the first visual corresponding dataset and benchmark for the MLLM community. Finally, we present CoLVA, a novel contrastive MLLM with two novel technical designs: fine-grained vision expert with object-level contrastive learning and instruction augmentation strategy. The former learns instance discriminative tokens, while the latter further improves instruction following ability. CoLVA-InternVL2-4B achieves an overall accuracy (OA) of 49.80\% on the MMVM benchmark, surpassing GPT-4o and the best open-source MLLM, Qwen2VL-72B, by 7.15\% and 11.72\% OA, respectively. These results demonstrate the effectiveness of our MMVM SFT dataset and our novel technical designs. Code, benchmark, dataset, and models will be released.
Yikang Zhou, Tao Zhang 0042, Shilin Xu 0001, Shihao Chen, Qianyu Zhou 0001, Yunhai Tong, Shunping Ji, Jiangning Zhang, Lu Qi 0001, Xiangtai Li
ICCV2
2025 On Path to Multimodal Generalist: General-Level and General-Bench
abstract
The Multimodal Large Language Model (MLLM) is currently experiencing rapid growth, driven by the advanced capabilities of language-based LLMs. Unlike their specialist predecessors, existing MLLMs are evolving towards a Multimodal Generalist paradigm. Initially limited to understanding multiple modalities, these models have advanced to not only comprehend but also generate across modalities. Their capabilities have expanded from coarse-grained to fine-grained multimodal understanding and from supporting singular modalities to accommodating a wide array of or even arbitrary modalities. To assess the capabilities of various MLLMs, a diverse array of benchmark test sets has been proposed. This leads to a critical question: Can we simply assume that higher performance across tasks indicates a stronger MLLM capability, bringing us closer to human-level AI? We argue that the answer is not as straightforward as it seems. In this project, we introduce an evaluation framework to delineate the capabilities and behaviors of current multimodal generalists. This framework, named General-Level, establishes 5-scale levels of MLLM performance and generality, offering a methodology to compare MLLMs and gauge the progress of existing systems towards more robust multimodal generalists and, ultimately, towards AGI (Artificial General Intelligence). Central to our framework is the use of Synergy as the evaluative criterion, categorizing capabilities based on whether MLLMs preserve synergy across comprehension and generation, as well as across multimodal interactions. To evaluate the comprehensive abilities of various generalists, we present a massive multimodal benchmark, General-Bench, which encompasses a broader spectrum of skills, modalities, formats, and capabilities, including over 700 tasks and 325,800 instances. The evaluation results that involve over 100 existing state-of-the-art MLLMs uncover the capability rankings of generalists, highlighting the challenges in reaching genuine AI. We expect this project to pave the way for future research on next-generation multimodal foundation models, providing a robust infrastructure to accelerate the realization of AGI. Project Page: https://generalist.top/, Leaderboard: https://generalist.top/leaderboard/, Benchmark: https://huggingface.co/General-Level/.
Hao Fei 0001, Yuan Zhou 0016, Juncheng Li 0006, Xiangtai Li, Qingshan Xu 0001, Bobo Li 0001, Shengqiong Wu, Yaoting Wang, Junbao Zhou, Jiahao Meng, Liangtao Shi, Minghe Gao, Daoan Zhang, Zhiqi Ge, Siliang Tang, Kaihang Pan, Yaobo Ye, Haobo Yuan, Tao Zhang 0042, Weiming Wu, Tianjie Ju, Zixiang Meng, Shilin Xu 0001, Liyu Jia, Meng Luo 0010, Jiebo Luo 0001, Tat-Seng Chua, Shuicheng Yan, Hanwang Zhang
ICML21
2025 LL-Gaussian: Low-Light Scene Reconstruction and Enhancement via Gaussian Splatting for Novel View Synthesis
Fenggen Yu, Huiyao Xu, Tao Zhang 0042, Changqing Zou
ACM Multimedia4
2025 Latent Diffusion Enhanced Rectangle Transformer for Hyperspectral Image Restoration
abstract
The restoration of hyperspectral image (HSI) plays a pivotal role in subsequent hyperspectral image applications. Despite the remarkable capabilities of deep learning, current HSI restoration methods face challenges in effectively exploring the spatial non-local self-similarity and spectral low-rank property inherently embedded with HSIs. This paper addresses these challenges by introducing a latent diffusion enhanced rectangle Transformer for HSI restoration, tackling the non-local spatial similarity and HSI-specific latent diffusion low-rank property. In order to effectively capture non-local spatial similarity, we propose the multi-shape spatial rectangle self-attention module in both horizontal and vertical directions, enabling the model to utilize informative spatial regions for HSI restoration. Meanwhile, we propose a spectral latent diffusion enhancement module that generates the image-specific latent dictionary based on the content of HSI for low-rank vector extraction and representation. This module utilizes a diffusion model to generatively obtain representations of global low-rank vectors, thereby aligning more closely with the desired HSI. A series of comprehensive experiments were carried out on four common hyperspectral image restoration tasks, including HSI denoising, HSI super-resolution, HSI reconstruction, and HSI inpainting. The results of these experiments highlight the effectiveness of our proposed method, as demonstrated by improvements in both objective metrics and subjective visual quality.
Miaoyu Li, Ying Fu 0001, Tao Zhang 0042, Ji Liu 0003, Dejing Dou, Chenggang Yan 0001, Yulun Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 DVIS++: Improved Decoupled Framework for Universal Video Segmentation
abstract
We present the Decoupled VIdeo Segmentation (DVIS) framework, a novel approach for the challenging task of universal video segmentation, including video instance segmentation (VIS), video semantic segmentation (VSS), and video panoptic segmentation (VPS). Unlike previous methods that model video segmentation in an end-to-end manner, our approach decouples video segmentation into three cascaded sub-tasks: segmentation, tracking, and refinement. This decoupling design allows for simpler and more effective modeling of the spatio-temporal representations of objects, especially in complex scenes and long videos. Accordingly, we introduce two novel components: the referring tracker and the temporal refiner. These components track objects frame by frame and model spatio-temporal representations based on pre-aligned features. To improve the tracking capability of DVIS, we propose a denoising training strategy and introduce contrastive learning, resulting in a more robust framework named DVIS++. The proposed decoupled framework efficiently handles universal and open-vocabulary object representations, allowing DVIS++ to conduct universal and open-vocabulary video segmentation. We conduct extensive experiments on six mainstream benchmarks, including the VIS, VSS, and VPS datasets. Using a unified architecture, DVIS++ significantly outperforms state-of-the-art specialized methods on these benchmarks in closed- and open-vocabulary settings.
Tao Zhang 0042, Xingye Tian, Yikang Zhou, Shunping Ji, Xuebo Wang, Xin Tao 0001, Yuan Zhang 0020, Pengfei Wan 0001, Zhongyuan Wang 0006, Yu Wu 0011
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Calibration-Free Raw Image Denoising via Fine-Grained Noise Estimation
abstract
Image denoising has progressed significantly due to the development of effective deep denoisers. To improve the performance in real-world scenarios, recent trends prefer to formulate superior noise models to generate realistic training data, or estimate noise levels to steer non-blind denoisers. In this paper, we bridge both strategies by presenting an innovative noise estimation and realistic noise synthesis pipeline. Specifically, we integrates a fine-grained statistical noise model and contrastive learning strategy, with a unique data augmentation to enhance learning ability. Then, we use this model to estimate noise parameters on evaluation dataset, which are subsequently used to craft camera-specific noise distribution and synthesize realistic noise. One distinguishing feature of our methodology is its adaptability: our pre-trained model can directly estimate unknown cameras, making it possible to unfamiliar sensor noise modeling using only testing images, without calibration frames or paired training data. Another highlight is our attempt in estimating parameters for fine-grained noise models, which extends the applicability to even more challenging low-light conditions. Through empirical testing, our calibration-free pipeline demonstrates effectiveness in both normal and low-light scenarios, further solidifying its utility in real-world noise synthesis and denoising tasks.
Yunhao Zou, Ying Fu 0001, Yulun Zhang 0001, Tao Zhang 0042, Chenggang Yan 0001, Radu Timofte
IEEE Trans. Pattern Anal. Mach. Intell.4
2025 P2PFormerV2: Improving Primitive-Based Regular Building Contour Extraction Methods via Contour Feature Enhancement
abstract
Deep learning methods have been widely used to map building contours automatically in high-resolution remote sensing images over recent years. However, most deep learning-based methods still require complex post-processing to generate regular contours. P2PFormer is an advanced method that can directly obtain the positions and sequences of general geometric primitives like points, lines, and angles (corners) of a building instance without complex post-processing. Nevertheless, due to the inherent limitations of ROI-Align, P2PFormer faces significant challenges in primitive detection. ROI-Align utilizes a uniform sampling approach for feature extraction, it inevitably introduces a high proportion of invalid sampling points into the extracted features. Resulting in missed and false detections of building primitives, ultimately affecting the accuracy of building contour extraction. To address these issues, we propose P2PFormerV2, which introduces a contour feature enhancer to improve P2PFormer. The contour feature enhancer increases the proportion of valid feature sampling points and enhances the extraction of contour-aware features, significantly improving the accuracy of primitive segmentation. This enhancer comprises three key components: the sparse feature extractor, the dense feature extractor, and the feature fusion module. The sparse feature extractor optimizes the sampling strategy to increase the proportion of valid feature sampling points; the dense feature extractor generates rich contour features and provides additional supervision signals; the feature fusion module integrates the outputs of the first two components to further enhance the extraction of contour features. Experimental results demonstrate that P2PFormerV2 achieves average precisions (AP) of 74.7%, 79.6%, and 64.2% in the WHU, CrowdAI, and WHU-Mix datasets, respectively, significantly outperforming the original P2PFormer and other existing advanced methods. Our findings about the shortcomings of ROI-Align and the importance of improving the effective feature extraction provides insights for future building extraction research.
Wenling Yu, Tao Zhang 0042, Shunping Ji, Bo Liu 0068, Hua Liu 0002, Jianya Gong
IEEE Trans. Geosci. Remote. Sens.2
2025 Supervise-Assisted Self-Supervised Deep-Learning Method for Hyperspectral Image Restoration
abstract
Hyperspectral image (HSI) restoration is a challenging research area, covering a variety of inverse problems. Previous works have shown the great success of deep learning in HSI restoration. However, facing the problem of distribution gaps between training HSIs and target HSI, those data-driven methods falter in delivering satisfactory outcomes for the target HSIs. In addition, the degradation process of HSIs is usually disturbed by noise, which is not well taken into account in existing restoration methods. The existence of noise further exacerbates the dissimilarities within the data, rendering it challenging to attain desirable results without an appropriate learning approach. To track these issues, in this article, we propose a supervise-assisted self-supervised deep-learning method to restore noisy degraded HSIs. Initially, we facilitate the restoration network to acquire a generalized prior through supervised learning from extensive training datasets. Then, the self-supervised learning stage is employed and utilizes the specific prior of the target HSI. Particularly, to restore clean HSIs during the self-supervised learning stage from noisy degraded HSIs, we introduce a noise-adaptive loss function that leverages inner statistics of noisy degraded HSIs for restoration. The proposed noise-adaptive loss consists of Stein's unbiased risk estimator (SURE) and total variation (TV) regularizer and fine-tunes the network with the presence of noise. We demonstrate through experiments on different HSI tasks, including denoising, compressive sensing, super-resolution, and inpainting, that our method outperforms state-of-the-art methods on benchmarks under quantitative metrics and visual quality. The code is available at https://github.com/ying-fu/SSDL-HSI.
Miaoyu Li, Ying Fu 0001, Tao Zhang 0042, Guanghui Wen
IEEE Trans. Neural Networks Learn. Syst.3
2024 Improving Video Segmentation via Dynamic Anchor Queries
Yikang Zhou, Tao Zhang 0042, Shunping Ji, Shuicheng Yan, Xiangtai Li
ECCV (50)2
2024 A Cross-modal Fusion Method for Multispectral Small Ship Detection
abstract
The fusion module of RGB and infrared (IR) remote sensing images is the key of multispectral ship detection. Existing works have shown that the cross-attention-based feature fusion can achieve good performance by extracting the complementary information of RGB and IR modalities. However, the existing commonly used cross-attention mechanisms introduce lots of redundancy parameters and mainly focus on global feature interaction of multispectral images, ignoring local detail information that is also important for small ship detection. In this paper, we propose a novel multispectral ship detection approach named LoGFusion. In LoGFusion, we design the cross stage partial module with partial convolution (CSPMPC) to reduce feature redundancy and utilize the local cross-modal fusion module (LoCFM) and global cross-modal fusion module (GCFM) to capture both local and global cross-modal features. Furthermore, we introduce a Multispectral Small Ship Dataset (MSSD) containing over 5k ship targets for small target detection. Experiments on MSSD validate the effectiveness of our method in terms of small ship detection in multispectral images.
Yang Liu 0119, Yu Liu 0005, Xueqian Wang 0002, Linping Zhang, Zhizhuo Jiang, Yaowen Li, Chenggang Yan 0001, Ying Fu 0001, Tao Zhang 0042
FUSION9
2024 MuChin: A Chinese Colloquial Description Benchmark for Evaluating Language Models in the Field of Music
Tao Zhang 0042, Jinyang Luo, Yan Liu 0004, Ming Xi
IJCAI3
2024 OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding
abstract
Current universal segmentation methods demonstrate strong capabilities in pixel-level image and video understanding. However, they lack reasoning abilities and cannot be controlled via text instructions. In contrast, large vision-language multimodal models exhibit powerful vision-based conversation and reasoning capabilities but lack pixel-level understanding and have difficulty accepting visual prompts for flexible user interaction. This paper proposes OMG-LLaVA, a new and elegant framework combining powerful pixel-level vision understanding with reasoning abilities. It can accept various visual and text prompts for flexible user interaction. Specifically, we use a universal segmentation method as the visual encoder, integrating image information, perception priors, and visual prompts into visual tokens provided to the LLM. The LLM is responsible for understanding the user's text instructions and providing text responses and pixel-level segmentation results based on the visual information. We propose perception prior embedding to better integrate perception priors with image features. OMG-LLaVA achieves image-level, object-level, and pixel-level reasoning and understanding in a single model, matching or surpassing the performance of specialized methods on multiple benchmarks. Rather than using LLM to connect each specialist, our work aims at end-to-end training on one encoder, one decoder, and one LLM. The code and model have been released for further research.
Tao Zhang 0042, Xiangtai Li, Hao Fei 0001, Haobo Yuan, Shengqiong Wu, Shunping Ji, Chen Change Loy, Shuicheng Yan
NeurIPS1
2024 SAM-RSIS: Progressively Adapting SAM With Box Prompting to Remote Sensing Image Instance Segmentation
abstract
The recent segment anything model (SAM) trained on massive close-range images has demonstrated impressive performance on general segmentation or specific segmentation tasks with manual prompts. However, the significant domain shift problem between remote sensing and close-range images should be tackled before introducing the pretrained SAM to remote sensing instance segmentation (RSIS). To address this and unlock the potential of SAM in RSIS, this article proposes a novel framework called SAM for remote sensing instance segmentation (SAM-RSIS), which overcomes the problems in a few recent works that only adapt a part of SAM to remote sensing. SAM-RSIS fine-tunes the vision transformer (ViT) backbone and mask decoder of SAM progressively on remote sensing data and uses automatic box prompting to eliminate the need for manual prompting. SAM-RSIS consists of an object detection stage and a mask generation stage. In object detection, we introduce an adapter to adapt knowledge embedded in the pretrained ViT backbone to remote sensing images and then build an object detector. In mask generation, using the detected bounding boxes as prompts, along with two learnable mask output tokens, and the two-layer high-resolution features from the adapter, we fine-tune the mask decoder of SAM to produce high-quality masks. Experimental results on the WHU, WHU-Mix, and NWPU datasets for binary and multiclass RSIS demonstrate the effectiveness and robustness of the proposed method, surpassing various derivative methods of SAM and achieving performance comparable to and even better than the specific state-of-the-art instance segmentation methods.
Muying Luo, Tao Zhang 0042, Shiqing Wei, Shunping Ji
IEEE Trans. Geosci. Remote. Sens.2
2024 P2PFormer: A Primitive-to-Polygon Method for Regular Building Contour Extraction From Remote Sensing Images
abstract
Extracting building contours from remote sensing imagery is a significant challenge due to buildings’ complex and diverse shapes, occlusions, and noise. Existing methods often struggle with irregular contours, rounded corners, and redundancy points, necessitating extensive postprocessing to produce regular polygonal building contours. To address these challenges, we introduce a novel, streamlined pipeline that generates regular building contours without postprocessing. Our approach begins with the segmentation of generic geometric primitives (which can include vertices, lines, and corners), followed by the prediction of their sequence. This allows for the direct construction of regular building contours by sequentially connecting the segmented primitives. Building on this pipeline, we developed primitive-to-polygon using transformer (P2PFormer), which uses a transformer-based architecture to segment geometric primitives and predict their order. To enhance the segmentation of primitives, we introduce a unique representation called group queries. This representation comprises a set of queries and a singular query position, which improve the focus on multiple midpoints of primitives and their efficient linkage. Furthermore, we propose an innovative implicit update strategy for the query position embedding aimed at sharpening the focus of queries on the correct positions and, consequently, enhancing the quality of primitive segmentation. Our experiments demonstrate that P2PFormer achieves new state-of-the-art (SOTA) performance on the WHU, CrowdAI, and WHU-Mix datasets, surpassing the previous SOTA PolyWorld by a margin of 2.7 AP and 6.5 AP75 on the largest CrowdAI dataset. We intend to make the code and trained weights publicly available to promote their use and facilitate further research.
Tao Zhang 0042, Shiqing Wei, Yikang Zhou, Muying Luo, Wenling Yu, Shunping Ji
IEEE Trans. Geosci. Remote. Sens.1
2023 DVIS: Decoupled Video Instance Segmentation Framework
abstract
Video instance segmentation (VIS) is a critical task with diverse applications, including autonomous driving and video editing. Existing methods often underperform on complex and long videos in real world, primarily due to two factors. Firstly, offline methods are limited by the tightly-coupled modeling paradigm, which treats all frames equally and disregards the interdependencies between adjacent frames. Consequently, this leads to the introduction of excessive noise during long-term temporal alignment. Secondly, online methods suffer from inadequate utilization of temporal information. To tackle these challenges, we propose a decoupling strategy for VIS by dividing it into three independent sub-tasks: segmentation, tracking, and refinement. The efficacy of the decoupling strategy relies on two crucial elements: 1) attaining precise long-term alignment outcomes via frame-by-frame association during tracking, and 2) the effective utilization of temporal information predicated on the aforementioned accurate alignment outcomes during refinement. We introduce a novel referring tracker and temporal refiner to construct the Decoupled VIS framework (DVIS). DVIS achieves new SOTA performance in both VIS and VPS, surpassing the current SOTA methods by 7.3 AP and 9.6 VPQ on the OVIS and VIPSeg datasets, which are the most challenging and realistic benchmarks. Moreover, thanks to the decoupling strategy, the referring tracker and temporal refiner are super light-weight (only 1.69% of the segmenter FLOPs), allowing for efficient training and inference on a single GPU with 11G memory. The code is available at https://github.com/zhang-tao-whu/DVIS.
Tao Zhang 0042, Xingye Tian, Yu Wu 0011, Shunping Ji, Xuebo Wang, Yuan Zhang 0020, Pengfei Wan 0001
ICCV1
2023 Deep external and internal learning for noisy compressive sensing
Tao Zhang 0042, Ying Fu 0001, Debing Zhang, Chun Hu
Neurocomputing1
2023 Blind Super-Resolution of Single Remotely Sensed Hyperspectral Image
abstract
Hyperspectral image (HSI) super-resolution has recently advanced with significant progress by utilizing the powerful representation capabilities of deep neural networks. These approaches, however, inevitably rely on a sizable amount of training data which can be difficult to acquire for remotely sensed HSIs. In many cases, these methods are designed and tailored for only one or a few specific super-resolution scenarios, making them inflexible for handling images with different unknown degradations. In this paper, we introduce a two-step framework for blind remotely sensed HSI super-resolution, where the degradation is unknown. Specifically, in the first step, we propose to leverage the abundant remotely sensed color images to address the data insufficiency for remotely sensed HSI super-resolution. It is achieved by exploring the spatial knowledge from remotely sensed color images with a super-resolution network for a predefined degradation, which is then transferred to HSIs via band-by-band super-resolution. Direct use of the results from the transferred super-resolution network is suboptimal as it neglects the spectral correlations of different bands and the gap between predefined degradation and the real one. To make further refinements, we present an unsupervised scheme that simultaneously refines the super-resolved HSI and the unknown degradation by a non-negative matrix factorization network and a learnable degradation prior. To validate the effectiveness of our method, we conducted extensive experiments on a variety of remotely sensed HSI datasets. The results demonstrate that our method could generalize on various unknown degradations with superior performance against the state-of-the-art methods.
Zhiyuan Liang, Shuai Wang 0049, Tao Zhang 0042, Ying Fu 0001
IEEE Trans. Geosci. Remote. Sens.3
2023 From Image Transfer to Object Transfer: Cross-Domain Instance Segmentation Based on Center Point Feature Alignment
abstract
Remote sensing images can have significant appearance differences due to various factors such as atmospheric conditions, sensor types, seasons, and capture times. Therefore, when applying a pre-trained instance segmentation deep learning model to newly accessed remote sensing images, the model’s performance tends to decrease significantly. Current mainstream image-based or feature-based domain adaptation methods are not designed specifically for the cross-domain instance segmentation problem. These methods attempt to align the whole images, which may not be optimal for instance segmentation tasks. To address this issue, we propose a cross-domain instance segmentation method based on object-level alignment. Instead of aligning the entire images from both datasets, we only align the features of each object instance, particularly the representative center point features. Our approach mainly consists of an improved contour-based instance segmentation model for object-based domain adaptation, an object-pasting enhancement technique based on Fourier domain adaptation (FDA) that effectively reduces the gap between the source and target domains of the object instances, and a self-training strategy that dynamically generates pseudo-labels for iterative model training. Our experiments on cross-domain building instance segmentation demonstrate that the proposed method achieves a 9.5 intersection over union (IoU) improvement over the current best method. Additionally, experiments on a cross-domain close-range dataset involving transfer between simulated and real street images show that our method significantly outperforms the current best method by 6.5 mean average precision (mAP). These results on remote sensing and close-range datasets validate the universality of our approach.
Shunping Ji, Tao Zhang 0042
IEEE Trans. Geosci. Remote. Sens.3
2023 Low-Light Raw Video Denoising With a High-Quality Realistic Motion Dataset
abstract
Recently, supervised deep-learning methods have shown their effectiveness on raw video denoising in low-light. However, existing training datasets have specific drawbacks, e.g., inaccurate noise modeling in synthetic datasets, simple motion created by hand or fixed motion, and limited-quality ground truth caused by the beam splitter in real captured datasets. These defects significantly decline the performance of network when tackling real low-light video sequences, where noise distribution and motion patterns are extremely complex. In this paper, we collect a raw video denoising dataset in low-light with complex motion and high-quality ground truth, overcoming the drawbacks of previous datasets. Specifically, we capture 210 paired videos, each containing short/long exposure pairs of real video frames with dynamic objects and diverse scenes displayed on a high-end monitor. Besides, since spatial self-similarity has been extensively utilized in image tasks, harnessing this property for network design is more crucial for video denoising as temporal redundancy. To effectively exploit the intrinsic temporal-spatial self-similarity of complex motion in real videos, we propose a new Transformer-based network, which can effectively combine the locality of convolution with the long-range modeling ability of 3D temporal-spatial self-attention. Extensive experiments verify the value of our dataset and the effectiveness of our method on various metrics.
Ying Fu 0001, Tao Zhang 0042, Jun Zhang 0007
IEEE Trans. Multim.3
2022 Deep Spatial Adaptive Network for Real Image Demosaicing
abstract
Demosaicing is the crucial step in the image processing pipeline and is a highly ill-posed inverse problem. Recently, various deep learning based demosaicing methods have achieved promising performance, but they often design the same nonlinear mapping function for different spatial location and are not well consider the difference of mosaic pattern for each color. In this paper, we propose a deep spatial adaptive network (SANet) for real image demosaicing, which can adaptively learn the nonlinear mapping function for different locations. The weights of spatial adaptive convolution layer are generated by the pattern information in the receptive filed. Besides, we collect a paired real demosaicing dataset to train and evaluate the deep network, which can make the learned demosaicing network more practical in the real world. The experimental results show that our SANet outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality in both noiseless and noisy cases.
Tao Zhang 0042, Ying Fu 0001, Cheng Li 0009
AAAI1
2022 E2EC: An End-to-End Contour-based Method for High-Quality High-Speed Instance Segmentation
abstract
Contour-based instance segmentation methods have developed rapidly recently but feature rough and hand-crafted front-end contour initialization, which restricts the model performance, and an empirical and fixed backend predicted-label vertex pairing, which contributes to the learning difficulty. In this paper, we introduce a novel contour-based method, named E2EC, for high-quality instance segmentation. Firstly, E2EC applies a novel learnable contour initialization architecture instead of hand-crafted contour initialization. This consists of a contour initialization module for constructing more explicit learning goals and a global contour deformation module for taking advantage of all of the vertices' features better. Secondly, we propose a novel label sampling scheme, named multi-direction alignment, to reduce the learning difficulty. Thirdly, to improve the quality of the boundary details, we dynamically match the most appropriate predicted-ground truth vertex pairs and propose the corresponding loss function named dynamic matching loss. The experiments showed that E2EC can achieve a state-of-the-art performance on the KITTI INStance (KINS) dataset, the Semantic Boundaries Dataset (SBD), the Cityscapes and the COCO dataset. E2EC is also efficient for use in real-time applications, with an inference speed of 36 fps for$512\times 512$images on an NVIDIA A6000 GPU. Code will be released at https://github.com/zhang-tao-whu/e2ec.
Tao Zhang 0042, Shiqing Wei, Shunping Ji
CVPR1
2022 Guided Hyperspectral Image Denoising with Realistic Data
Tao Zhang 0042, Ying Fu 0001, Jun Zhang 0007
Int. J. Comput. Vis.1
2022 A large-scale hyperspectral dataset for flower classification
Yongrong Zheng, Tao Zhang 0042, Ying Fu 0001
Knowl. Based Syst.2
2022 Coded Hyperspectral Image Reconstruction Using Deep External and Internal Learning
abstract
To solve the low spatial and/or temporal resolution problem which the conventional hyperspectral cameras often suffer from, coded hyperspectral imaging systems have attracted more attention recently. Recovering a hyperspectral image (HSI) from its corresponding coded image is an ill-posed inverse problem, and learning accurate prior of HSI is essential to solve this inverse problem. In this paper, we present an effective convolutional neural network (CNN) based method for coded HSI reconstruction, which learns the deep prior from the external dataset as well as the internal information of input coded image with spatial-spectral constraint. Specifically, we first develop a CNN-based channel attention reconstruction network to effectively exploit the spatial-spectral correlation of the HSI. Then, the reconstruction network is learned by leveraging an arbitrary external hyperspectral dataset to exploit the general spatial-spectral correlation under adversarial loss. Finally, we customize the network by internal learning with spatial-spectral constraint and total variation regularization for each coded image, which can make use of the internal imaging model to learn specific prior for current desirable image and effectively avoids overfitting. Experimental results using both synthetic data and real images show that our method outperforms the state-of-the-art methods on several popular coded hyperspectral imaging systems under both comprehensive quantitative metrics and perceptive quality.
Ying Fu 0001, Tao Zhang 0042, Lizhi Wang 0001, Hua Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 Joint Camera Spectral Response Selection and Hyperspectral Image Recovery
abstract
Hyperspectral image (HSI) recovery from a single RGB image has attracted much attention, whose performance has recently been shown to be sensitive to the camera spectral response (CSR). In this paper, we present an efficient convolutional neural network (CNN) based method, which can jointly select the optimal CSR from a candidate dataset and learn a mapping to recover HSI from a single RGB image captured with this algorithmically selected camera under multi-chip or single-chip setups. Given a specific CSR, we first present a HSI recovery network, which accounts for the underlying characteristics of the HSI, including spectral nonlinear mapping and spatial similarity. Later, we append a CSR selection layer onto the recovery network, and the optimal CSR under both multi-chip and single-chip setups can thus be automatically determined from the network weights under the nonnegative sparse constraint. Experimental results on three hyperspectral datasets and two camera spectral response datasets demonstrate that our HSI recovery network outperforms state-of-the-art methods in terms of both quantitative metrics and perceptive quality, and the selection layer always returns a CSR consistent to the best one determined by exhaustive search. Finally, we show that our method can also perform well in the real capture system, and collect a hyperspectral flower dataset to evaluate the effect from HSI recovery on classification problem.
Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2022 A Concentric Loop Convolutional Neural Network for Manual Delineation-Level Building Boundary Segmentation From Remote-Sensing Images
abstract
To date, accurate building footprint delineation in the surveying, mapping, and geographic information system (GIS) communities has been dependent on human labor. In this article, to address this issue, we propose a concentric loop convolutional neural network (CLP-CNN) method for the automatic segmentation of building boundaries from remote-sensing images. The proposed method consists of three components: 1) a boundary detector to extract coarse polygonal boundaries of individual regions of interest; 2) a concentric loop-shaped convolutional network with bidirectional pairing loss to fine-tune the vertices of the polygons; and 3) a refinement block, which removes redundant vertices and regularizes the boundaries to polygons at the manual delineation level. We also demonstrate that the proposed CLP-CNN method is applicable to other generic objects in natural images. Experiments on two building datasets confirmed that more than 77%/67% of the building polygons predicted by the proposed method are on par with the manual delineation level, representing a significant saving in the labor cost of manual annotation. In generic object boundary delineation tests performed on the Semantic Boundaries Dataset (SBD), the proposed method outperformed the most recent state-of-the-art methods by at least 3.1% in average precision (AP). Furthermore, compared with other vertex matching methods, the learning process of the proposed method converges faster. The source code will be available athttp://gpcv.whu.edu.cn/data.
Shiqing Wei, Tao Zhang 0042, Shunping Ji
IEEE Trans. Geosci. Remote. Sens.2
2021 Hyperspectral Image Denoising with Realistic Data
abstract
The hyperspectral image (HSI) denoising has been widely utilized to improve HSI qualities. Recently, learning-based HSI denoising methods have shown their effectiveness, but most of them are based on synthetic dataset and lack the generalization capability on real testing HSI. Moreover, there is still no public paired real HSI denoising dataset to learn HSI denoising network and quantitatively evaluate HSI methods. In this paper, we mainly focus on how to produce realistic dataset for learning and evaluating HSI denoising network. On the one hand, we collect a paired real HSI denoising dataset, which consists of short-exposure noisy HSIs and the corresponding long-exposure clean HSIs. On the other hand, we propose an accurate HSI noise model which matches the distribution of real data well and can be employed to synthesize realistic dataset. On the basis of the noise model, we present an approach to calibrate the noise parameters of the given hyperspectral camera. The extensive experimental results show that a network learned with only synthetic data generated by our noise model performs as well as it is learned with paired real data. Our code and data are available at: https://github.com/ColinTaoZhang/HSIDwRD.
Tao Zhang 0042, Ying Fu 0001, Cheng Li 0009
ICCV1
2021 M3D-VTON: A Monocular-to-3D Virtual Try-On Network
abstract
Virtual 3D try-on can provide an intuitive and realistic view for online shopping and has a huge potential commercial value. However, existing 3D virtual try-on methods mainly rely on annotated 3D human shapes and garment templates, which hinders their applications in practical scenarios. 2D virtual try-on approaches provide a faster alternative to manipulate clothed humans, but lack the rich and realistic 3D representation. In this paper, we propose a novel Monocular-to-3D Virtual Try-On Network (M3D-VTON) that builds on the merits of both 2D and 3D approaches. By integrating 2D information efficiently and learning a mapping that lifts the 2D representation to 3D, we make the first attempt to reconstruct a 3D try-on mesh only taking the target clothing and a person image as inputs. The proposed M3D-VTON includes three modules: 1) The Monocular Prediction Module (MPM) that estimates an initial full-body depth map and accomplishes 2D clothes-person alignment through a novel two-stage warping procedure; 2) The Depth Refinement Module (DRM) that refines the initial body depth to produce more detailed pleat and face characteristics; 3) The Texture Fusion Module (TFM) that fuses the warped clothing with the non-target body part to refine the results. We also construct a high-quality synthesized Monocular-to-3D virtual try-on dataset, in which each person image is associated with a front and a back depth map. Extensive experiments demonstrate that the proposed M3D-VTON can manipulate and reconstruct the 3D human body wearing the given clothing with compelling details and is more efficient than other 3D approaches.1
Fuwei Zhao, Zhenyu Xie, Michael Kampffmeyer, Haoye Dong, Songfang Han, Tianxiang Zheng 0001, Tao Zhang 0042, Xiaodan Liang
ICCV7
2021 Residual scale attention network for arbitrary scale image super-resolution
Ying Fu 0001, Tao Zhang 0042, Yonggang Lin
Neurocomputing3
2019 Hyperspectral Image Super-Resolution With Optimized RGB Guidance
abstract
To overcome the limitations of existing hyperspectral cameras on spatial/temporal resolution, fusing a low resolution hyperspectral image (HSI) with a high resolution RGB (or multispectral) image into a high resolution HSI has been prevalent. Previous methods for this fusion task usually employ hand-crafted priors to model the underlying structure of the latent high resolution HSI, and the effect of the camera spectral response (CSR) of the RGB camera on super-resolution accuracy has rarely been investigated. In this paper, we first present a simple and efficient convolutional neural network (CNN) based method for HSI super-resolution in an unsupervised way, without any prior training. Later, we append a CSR optimization layer onto the HSI super-resolution network, either to automatically select the best CSR in a given CSR dataset, or to design the optimal CSR under some physical restrictions. Experimental results show our method outperforms the state-of-the-arts, and the CSR optimization can further boost the accuracy of HSI super-resolution.
Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001
CVPR2
2019 Hyperspectral Image Reconstruction Using Deep External and Internal Learning
abstract
To solve the low spatial and/or temporal resolution problem which the conventional hypelrspectral cameras often suffer from, coded snapshot hyperspectral imaging systems have attracted more attention recently. Recovering a hyperspectral image (HSI) from its corresponding coded image is an ill-posed inverse problem, and learning accurate prior of HSI is essential to solve this inverse problem. In this paper, we present an effective convolutional neural network (CNN) based method for coded HSI reconstruction, which learns the deep prior from the external dataset as well as the internal information of input coded image with spatial-spectral constraint. Our method can effectively exploit spatial-spectral correlation and sufficiently represent the variety nature of HSIs. Experimental results show our method outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality.
Tao Zhang 0042, Ying Fu 0001, Lizhi Wang 0001, Hua Huang 0001
ICCV1
2019 HyperReconNet: Joint Coded Aperture Optimization and Image Reconstruction for Compressive Hyperspectral Imaging
abstract
Coded aperture snapshot spectral imaging (CASSI) system encodes the 3D hyperspectral image (HSI) within a single 2D compressive image and then reconstructs the underlying HSI by employing an inverse optimization algorithm, which equips with the distinct advantage of snapshot but usually results in low reconstruction accuracy. To improve the accuracy, existing methods attempt to design either alternative coded apertures or advanced reconstruction methods, but cannot connect these two aspects via a unified framework, which limits the accuracy improvement. In this paper, we propose a convolution neural network (CNN) based endto- end method to boost the accuracy by jointly optimizing the coded aperture and the reconstruction method. On the one hand, based on the nature of CASSI forward model, we design a repeated pattern for the coded aperture, whose entities are learned by acting as the network weights. On the other hand, we conduct the reconstruction through simultaneously exploiting intrinsic properties within HSI - the extensive correlations across the spatial and the spectral dimensions. By leveraging the power of deep learning, the coded aperture design and the image reconstruction are connected and optimized via a unified framework. Experimental results show that our method outperforms the state-of-the-art methods under both comprehensive quantitative metrics and perceptive quality.
Lizhi Wang 0001, Tao Zhang 0042, Ying Fu 0001, Hua Huang 0001
IEEE Trans. Image Process.2
2018 Joint Camera Spectral Sensitivity Selection and Hyperspectral Image Recovery
Ying Fu 0001, Tao Zhang 0042, Yinqiang Zheng, Debing Zhang, Hua Huang 0001
ECCV (3)2