Feng Zhao 0004

dblp:181/2734-4 · DBLP profile ↗
← Back
135ranked-venue papers
5as first author
123since 2021 · last 2026
0000-0001-6767-8105ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 90 · 2 first-author · 87 since 2021Graphics, computer vision, multimedia, augmented reality and games · 89 · 3 first-author · 83 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 6 since 2021
YearPublicationVenuePosition
2026 Token Painter: Training-Free Text-Guided Image Inpainting via Mask Autoregressive Models
abstract
Text-guided image inpainting aims to inpaint masked image regions based on a textual prompt while preserving the background. Although diffusion-based methods have become dominant, their property of modeling the entire image in latent space makes it challenging for the results to align well with prompt details and maintain a consistent background. To address these issues, we explore Mask AutoRegressive (MAR) models for this task. MAR naturally supports image inpainting by generating latent tokens corresponding to mask regions, enabling better local controllability without altering the background. However, directly applying MAR to this task makes the inpainting content either ignore the prompts or be disharmonious with the background context. Through analysis of the attention maps from the inpainting images, we identify the impact of background tokens on text tokens during the MAR generation, and leverage this to designToken Painter, a training-free text-guided image inpainting method based on MAR. Our approach introduces two key components: (1) Dual-Stream Encoder Information Fusion (DEIF), which fuses the semantic and context information from text and background in frequency domain to produce novel guidance tokens, allowing MAR to generate text-faithful inpainting content while keeping harmonious with background context. (2) Adaptive Decoder Attention Score Enhancing (ADAE), which adaptively enhances attention scores on guidance tokens and inpainting tokens to further enhance the alignment of prompt details and the content visual quality. Extensive experiments demonstrate that our training-free method outperforms prior state-of-the-art methods across almost all metrics.
Longtao Jiang, Jie Huang 0017, Mingfei Han 0002, Yongqiang Yu, Feng Zhao 0004, Xiaojun Chang, Zhihui Li 0001
AAAI6
2026 Towards Efficient and Robust Manipulation via Multi-Frame Vision-Language-Action Modeling
abstract
Recent vision-language-action (VLA) models built on pretrained vision-language models (VLMs) have demonstrated strong performance in robotic manipulation. However, these models remain constrained by the single-frame image paradigm and fail to fully leverage the temporal information offered by multi-frame histories, as directly feeding multiple frames into VLM backbones incurs substantial computational overhead and inference latency. We propose CronusVLA, a unified framework that extends single-frame VLA models to the multi-frame paradigm. CronusVLA follows a two-stage process: (1) Single-frame pretraining on large-scale embodied datasets with autoregressive prediction of action tokens, establishing an effective embodied vision-language foundation; (2) Multi-frame post-training, which adapts the prediction of the vision-language backbone from discrete tokens to learnable features, and aggregates historical information via feature chunking. CronusVLA effectively addresses the existing challenges of multi-frame modeling while enhancing performance. To evaluate the robustness under temporal and spatial disturbances, we introduce SimplerEnv-OR, a novel benchmark featuring 24 types of observational disturbances and 120 severity levels. Experiments across three embodiments in simulated and real-world environments demonstrate that CronusVLA achieves leading performance and superior robustness, with a 70.9% success rate on SimplerEnv, a 26.8% improvement over OpenVLA on LIBERO, and the highest robustness score on SimplerEnv-OR, showing the promise of efficient multi-frame adaptation for real-world VLA deployment.
Hao Li 0069, Shuai Yang 0001, Xiaoda Yang, Dahua Lin, Feng Zhao 0004, Jiangmiao Pang
AAAI10
2026 UniCorn: Towards Self-Improving Unified Multimodal Models through Self-Generated Supervision
abstract
Zhen Fang, Ruiyan Han, XinYu Sun, Yuchen Ma, Ziheng Wang, Yu Zeng, Zehui Chen, Lin Chen, Wenxuan Huang, Wei-Jie Xu, Yi Cao, Feng Zhao. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Ruiyan Han, XinYu Sun, Lin Chen 0026, Wenxuan Huang 0001, Wei-Jie Xu, Yi Cao 0006, Feng Zhao 0004
ACL (1)12
2026 Tracing the Roots: A Multi-Agent Framework for Uncovering Data Lineage in Post-Training LLMs
abstract
Yu Li, Xiaoran Shang, Qizhi Pei, Yun Zhu, Xin Gao, Honglin Lin, Zhanping Zhong, Zhuoshi Pan, Zheng Liu, Xiaoyang Wang, Conghui He, Dahua Lin, Feng Zhao, Lijun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yu Li 0006, Xiaoran Shang, Qizhi Pei, Yun Zhu 0007, Xin Gao 0001, Honglin Lin, Zhanping Zhong, Zhuoshi Pan, Xiaoyang Wang 0007, Conghui He, Dahua Lin, Feng Zhao 0004, Lijun Wu 0003
ACL (1)13
2026 Beyond Accuracy: Unveiling Inefficiency Patterns in Tool-Integrated Reasoning
abstract
In real-world Tool-Integrated Reasoning (TIR) scenarios, where LLMs interleave reasoning with external tool calls, a major source of inefficiency is that the toolcalls create pauses between LLM requests and cause KV-Cache eviction, forcing recomputation.Also, the long, unfiltered response returned by external tools inflates the KV-Cache, so each decode step spends more time loading the growing cache and thus becomes steadily slower as context length increases.However, existing efficiency metrics like token counts and toolcall counts fail to capture the real model inference latency.To address this, we introduce PTE (Prefill Token Equivalents), a hardwareaware TIR-efficiency metric that unifies internal reasoning and external tool-use costs while explicitly accounting for non-reusable KV-Cache and long-tool-response scenarios.Validation in a high-concurrency industrial setting indicates that PTE aligns significantly better with wall-clock latency than standard token counts, while maintaining consistent efficiency rankings across diverse hardware profiles.We conduct extensive experiments across five TIR benchmarks, quantify their PTE costs, and identify four inefficiency patterns that appear in TIR.We also discover that trajectories with higher PTE costs tend to have lower reasoning correctness, indicating that simply using more tools does not improve the quality of the answer.The code is available at https://github.com/sqs-ustc/ tool-reasoning-framework-PTE.
Qisheng Su, Shiting Huang, Feng Zhao 0004
ACL (1)6
2026 Breaking Block Boundaries: Anchor-based History-stable Decoding for Diffusion Large Language Models
abstract
Shun Zou, Yong Wang, Zehui Chen, Lin Chen, Chongyang Tao, Feng Zhao, Xiangxiang Chu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shun Zou, Chongyang Tao, Feng Zhao 0004, Xiangxiang Chu
ACL (1)6
2026 MFP-DETR: A Multi-scale Frequency-Aware and Prototype-Guided Transformer Detection Model for Power Inspection Defects
Feng Zhao 0004, Zilei Wang
ICIC (10)5
2026 Breaking prejudice: Empowering curve-based exposure corrector with denoising
Naishan Zheng, Jie Huang 0017, Feng Zhao 0004
Neurocomputing4
2025 VFM-Adapter: Adapting Visual Foundation Models for Dense Prediction with Dynamic Hybrid Operation Mapping
abstract
Although pre-trained large vision foundation models (VFM) yield superior results on various downstream tasks, full fine-tuning is often impractical due to its high computational cost and storage requirements. Recent advancements in parameter-efficient fine-tuning (PEFT) of VFM for image classification show significant promise. However, the application of PEFT techniques to dense prediction tasks remains largely unexplored. Our analysis of existing methods reveals that the underlying premise of utilizing low-rank parameter matrices, despite their efficacy in specific applications, may not be adequately suitable for dense prediction tasks. To this end, we propose a novel PEFT learning approach tailored for dense prediction tasks, namely VFM-Adapter. Specifically, the VFM-Adapter introduces a hybrid operation mapping technique that seamlessly integrates local information with global modeling to the adapter module. It capitalizes on the distinct inductive biases inherent in different operations. Additionally, we dynamically generate parameters for the VFM-Adapter, enabling flexibility of feature extraction given specific inputs. To validate the efficacy of VFM-Adapter, we conduct extensive experiments across object detection, semantic segmentation, and instance segmentation tasks. Results on multiple benchmarks consistently demonstrate the superiority of our method over previous approaches. Notably, with only three percent of the trainable parameters of the SAM-Base backbone, our approach achieves competitive or even superior performance compared to full fine-tuning. The code will be available.
Hongzhi Gao, Lin Chen 0019, Jiaming Liu 0003, Feng Zhao 0004
AAAI7
2025 Horizon-GS: Unified 3D Gaussian Splatting for Large-Scale Aerial-to-Ground Scenes
abstract
Seamless integration of both aerial and street view images remains a significant challenge in neural scene reconstruction and rendering. Existing methods predominantly focus on single domain, limiting their applications in immersive environments, which demand extensive free view exploration with large view changes both horizontally and vertically. We introduce Horizon-Gs, a novel approach built upon Gaussian Splatting techniques, tackles the unified reconstruction and rendering for aerial and street views. Our method addresses the key challenges of combining these perspectives with a new training strategy, overcoming viewpoint discrepancies to generate high-fidelity scenes. We also curate a high-quality aerial-to-ground views dataset encompassing both synthetic and real-world scene to advance further research. Experiments across diverse urban scene datasets confirm the effectiveness of our method.
Lihan Jiang, Kerui Ren, Mulin Yu, Linning Xu, Junting Dong, Tao Lu 0005, Feng Zhao 0004, Dahua Lin, Bo Dai 0002
CVPR7
2025 FreePCA: Integrating Consistency Information across Long-short Frames in Training-free Long Video Generation via Principal Component Analysis
abstract
Long video generation involves generating extended videos using models trained on short videos, suffering from distribution shifts due to varying frame counts. It necessitates the use of local information from the original short frames to enhance visual and motion quality, and global information from the entire long frames to ensure appearance consistency. Existing training-free methods struggle to effectively integrate the benefits of both, as appearance and motion in videos are closely coupled, leading to motion inconsistency and visual quality. In this paper, we reveal that global and local information can be precisely decoupled into consistent appearance and motion intensity information by applying Principal Component Analysis (PCA), allowing for refined complementary integration of global consistency and local quality. With this insight, we propose FreePCA, a training-free long video generation paradigm based on PCA that simultaneously achieves high consistency and quality. Concretely, we decouple consistent appearance and motion intensity features by measuring cosine similarity in the principal component space. Critically, we progressively integrate these features to preserve original quality and ensure smooth transitions, while further enhancing consistency by reusing the mean statistics of the initial noise. Experiments demonstrate that FreePCA can be applied to various video diffusion models without requiring training, leading to substantial improvements. Code is available at https://github.com/JosephTiTan/FreePCA.
Jiangtong Tan, Hu Yu 0001, Jie Huang 0017, Jie Xiao 0002, Feng Zhao 0004
CVPR5
2025 Navigating Image Restoration with VAR's Distribution Alignment Prior
abstract
Generative models trained on extensive high-quality datasets effectively capture the structural and statistical properties of clean images, rendering them powerful priors for transforming degraded features into clean ones in image restoration. VAR, a novel image generative paradigm, surpasses diffusion models in generation quality by applying a next-scale prediction approach. It progressively captures both global structures and fine-grained details through the autoregressive process, consistent with the multi-scale restoration principle widely acknowledged in the restoration community. Furthermore, we observe that during the image reconstruction process utilizing VAR, scale predictions automatically modulate the input, facilitating the alignment of representations at subsequent scales with the distribution of clean images. To harness VAR’s adaptive distribution alignment capability in image restoration tasks, we formulate the multi-scale latent representations within VAR as the restoration prior, thus advancing our delicately designed VarFormer framework. The strategic application of these priors enables our VarFormer to achieve remarkable generalization on unseen tasks while also reducing training computational costs. Extensive experiments underscores that our VarFormer outperforms existing multitask image restoration methods across various restoration tasks. The code is available at https://github.com/siywang541/Varformer.
Naishan Zheng, Jie Huang 0017, Feng Zhao 0004
CVPR4
2025 Adaptive Dropout: Unleashing Dropout across Layers for Generalizable Image Super-Resolution
abstract
Blind Super-Resolution (blind SR) aims to enhance the model’s generalization ability with unknown degradation, yet it still encounters severe overfitting issues. Some previous methods inspired by dropout, which enhances generalization by regularizing features, have shown promising results in blind SR. Nevertheless, these methods focus solely on regularizing features before the final layer and overlook the need for generalization in features at intermediate layers. Without explicit regularization of features at intermediate layers, the blind SR network struggles to obtain well-generalized feature representations. However, the key challenge is that directly applying dropout to intermediate layers leads to a significant performance drop, which we attribute to the inconsistency in training-testing and across layers it introduced. Therefore, we propose Adaptive Dropout, a new regularization method for blind SR models, which mitigates the inconsistency and facilitates application across intermediate layers of networks. Specifically, for training-testing inconsistency, we re-design the form of dropout and integrate the features before and after dropout adaptively. For inconsistency in generalization requirements across different layers, we innovatively design an adaptive training strategy to strengthen feature propagation by layer-wise annealing. Experimental results show that our method outperforms all past regularization methods on both synthetic and real-world benchmark datasets, also highly effective in other image restoration tasks. Code is available at https://github.com/xuhang07/Adpative-Dropout.
Hang Xu 0004, Jie Huang 0017, Jiangtong Tan, Zhen Zou, Feng Zhao 0004
CVPR6
2025 SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Models
abstract
The emergence of Vision Language Models (VLMs) has brought unprecedented advances in understanding multi-modal information. The combination of textual and visual semantics in VLMs is highly complex and diverse, making the safety alignment of these models challenging. Furthermore, due to the limited study on the safety alignment of VLMs, there is a lack of large-scale, high-quality datasets. To address these limitations, we propose a Safety Preference Alignment dataset for Vision Language Models named SPA-VL. In terms of breadth, SPA-VL covers 6 harmfulness domains, 13 categories, and 53 subcategories, and contains 100,788 samples of the quadruple (question, image, chosen response, rejected response). In terms of depth, the responses are collected from 12 open-source (e.g., QwenVL) and closed-source (e.g., Gemini) VLMs to ensure diversity. The construction of preference data is fully automated, and the experimental results indicate that models trained with alignment techniques on the SPA-VL dataset exhibit substantial improvements in harmlessness and helpfulness while maintaining core capabilities. SPA-VL, as a large-scale, high-quality, and diverse dataset, represents a significant milestone in ensuring that VLMs achieve both harmlessness and helpfulness.
Yongting Zhang, Lu Chen 0001, Guodong Zheng, Yifeng Gao 0002, Jinlan Fu, Zhenfei Yin, Senjie Jin, Yu Qiao 0001, Xuanjing Huang 0001, Feng Zhao 0004, Tao Gui
CVPR11
2025 SkySense-O: Towards Open-World Remote Sensing Interpretation with Vision-Centric Visual-Language Modeling
abstract
Open-world interpretation aims to accurately localize and recognize all objects within images by vision-language models (VLMs). While substantial progress has been made in this task for natural images, the advancements for remote sensing (RS) images still remain limited, primarily due to these two challenges. 1) Existing RS semantic categories are limited, particularly for pixel-level interpretation datasets. 2) Distinguishing among diverse RS spatial regions solely by language space is challenging due to the dense and intricate spatial distribution in open-world RS imagery. To address the first issue, we develop a fine-grained RS interpretation dataset, Sky-SA, which contains 183,375 high-quality local image-text pairs with full-pixel manual annotations, covering 1,763 category labels, exhibiting richer semantics and higher density than previous datasets. Afterwards, to solve the second issue, we introduce the vision-centric principle for vision-language modeling. Specifically, in the pre-training stage, the visual self-supervised paradigm is incorporated into image-text alignment, reducing the degradation of general visual representation capabilities of existing paradigms. Then, we construct a visual-relevance knowledge graph across open-category texts and further develop a novel vision-centric image-text contrastive loss for fine-tuning with text prompts. This new model, denoted as SkySense-O, demonstrates impressive zero-shot capabilities on a thorough evaluation encompassing 14 datasets over 4 tasks, from recognizing to reasoning and classification to localization. Specifically, it outperforms the latest models such as SegEarthOV, GeoRSCLIP, and VHM by a large margin, i.e., 11.95%, 8.04% and 3.55% on average respectively. The code is publicly available to facilitate further research at https://github.com/zqcrafts/SkySense-O.
Qi Zhu 0010, Jiangwei Lao, Deyi Ji, Lixiang Ru, Jian Wang 0108, Jingdong Chen, Ming Yang 0007, Dong Liu 0002, Feng Zhao 0004
CVPR12
2025 CRITICTOOL: Evaluating Self-Critique Capabilities of Large Language Models in Tool-Calling Error Scenarios
abstract
Pretrained language models (LMs) are prone to arithmetic errors.Existing work showed limited success in probing numeric values from models' representations, indicating that these errors can be attributed to the inherent unreliability of distributionally learned embeddings in representing exact quantities.However, we observe that previous probing methods are inadequate for the emergent structure of learned number embeddings with sinusoidal patterns.In response, we propose a novel probing technique that decodes numeric values from input embeddings with near-perfect accuracy across a range of open-source LMs.This proves that after the sole pre-training, LMs represent numbers with remarkable precision.Finally, we find that the embeddings' precision, judged by our probe's accuracy, explains a large portion of LM's errors in elementary arithmetic, and show that aligning the embeddings with the pattern our probes discover can mitigate these errors.
Shiting Huang, Junjie Ye 0005, Lin Chen 0019, Feng Zhao 0004
EMNLP9
2025 ViDoRAG: Visual Document Retrieval-Augmented Generation via Dynamic Iterative Reasoning Agents
abstract
Understanding information from visually rich documents remains a significant challenge for traditional Retrieval-Augmented Generation (RAG) methods.Existing benchmarks predominantly focus on image-based question answering (QA), overlooking the fundamental challenges of efficient retrieval, comprehension, and reasoning within dense visual documents.To bridge this gap, we introduce ViDoSeek, a novel dataset designed to evaluate RAG performance on visually rich documents requiring complex reasoning.Based on it, we identify key limitations in current RAG approaches: (i) purely visual retrieval methods struggle to effectively integrate both textual and visual features, and (ii) previous approaches often allocate insufficient reasoning tokens, limiting their effectiveness.To address these challenges, we propose ViDoRAG, a novel multi-agent RAG framework tailored for complex reasoning across visual documents.ViDoRAG employs a Gaussian Mixture Model (GMM)-based hybrid strategy to effectively handle multimodal retrieval.To further elicit the model's reasoning capabilities, we introduce an iterative agent workflow incorporating exploration, summarization, and reflection, providing a framework for investigating test-time scaling in RAG domains.Extensive experiments on ViDoSeek validate the effectiveness and generalization of our approach.Notably, ViDoRAG outperforms existing methods by over 10% on the competitive benchmark.The code is available at https: //github.com/Alibaba-NLP/ViDoRAG.
Qiuchen Wang, Ruixue Ding, Weiqi Wu, Pengjun Xie, Feng Zhao 0004
EMNLP7
2025 Enhancing Large Vision-Language Models with Ultra-Detailed Image Caption Generation
abstract
High-quality image captions are essential for improving modality alignment and visual understanding in Large Vision-Language Models (LVLMs).However, the scarcity of ultradetailed image caption data limits further advancements.This paper presents a systematic pipeline for generating high-quality, ultradetailed image captions, encompassing both pre-processing and post-processing stages.In the pre-processing stage, we classify and deduplicate images, extract visual information using expert tools, and leverage GPT-4o with structured prompts to generate initial captions.To enhance comprehensiveness, we introduce an expansion strategy based on Large Language Models (LLMs), defining eight descriptive dimensions to refine and extend captions, which serve as seed data for training a proprietary captioner model.In the post-processing stage, we incorporate human error-correction annotations and an active learning-inspired approach to refine low-quality samples.Using high-quality corrected data, we apply Direct Preference Optimization (DPO) and develop a critic-rewrite pipeline, training a sentence-level critic model to mitigate hallucinations.Experimental results demonstrate that our ultra-detailed captions significantly enhance LVLMs' perception and cognitive abilities across multiple vision-language benchmarks.The code and dataset are available at https://github.com/yuzeng0-0/UltraCaption.
Yukun Qi, Xikun Bao, Lin Chen 0026, Shiting Huang, Feng Zhao 0004
EMNLP9
2025 FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise Alignment
abstract
Domain Adaptation(DA) for dense prediction tasks is an important topic, which enhances the dense prediction model's performance when tested on its unseen domain. Recently, with the development of Diffusion-based Dense Prediction (DDP) models, the exploration of DA designs tailored to this framework is worth exploring, since the diffusion model is effective in modeling the distribution transformation that comprises domain information. In this work, we propose a training-free mechanism for DDP frameworks, endowing them with DA capabilities. Our motivation arises from the observation that the exposure bias (e.g., noise statistics bias) in diffusion brings domain shift, and different domains in conditions of DDP models can also be effectively captured by the noise prediction statistics. Based on this, we propose a training-free Domain Noise Alignment (DNA) approach, which alleviates the variations of noise statistics to domain changes during the diffusion sampling process, thereby achieving domain adaptation. Specifically, when the source domain is available, we directly adopt the DNA method to achieve domain adaptation by aligning the noise statistics of the target domain with those of the source domain. For the more challenging source-free DA, inspired by the observation that regions closer to the source domain exhibit higher confidence meeting variations of sampling noise, we utilize the statistics from the high-confidence regions progressively to guide the noise statistic adjustment during the sampling process. Notably, our method demonstrates the effectiveness of enhancing the DA capability of DDP models across four common dense prediction tasks. Code is available at \href{https://github.com/xuhang07/FreeDNA}{https://github.com/xuhang07/FreeDNA}.
Hang Xu 0004, Jie Huang 0017, Linjiang Huang, Dong Liu 0002, Yidi Liu, Feng Zhao 0004
ICCV6
2025 Gaussian Variation Field Diffusion for High-Fidelity Video-to-4D Synthesis
abstract
In this paper, we present a novel framework for video-to-4D generation that creates high-quality dynamic 3D content from single video inputs. Direct 4D diffusion modeling is extremely challenging due to costly data construction and the high-dimensional nature of jointly representing 3D shape, appearance, and motion. We address these challenges by introducing a Direct 4DMesh-to-GS Variation Field VAE that directly encodes canonical Gaussian Splats (GS) and their temporal variations from 3D animation data without per-instance fitting, and compresses high-dimensional animations into a compact latent space. Building upon this efficient representation, we train a Gaussian Variation Field diffusion model with temporal-aware Diffusion Transformer conditioned on input videos and canonical GS. Trained on carefully-curated animatable 3D objects from the Objaverse dataset, our model demonstrates superior generation quality compared to existing methods. It also exhibits remarkable generalization to in-the-wild video inputs despite being trained exclusively on synthetic data, paving the way for generating high-quality animated 3D content. Project page: https://gvfdiffusion.github.io/.
Bowen Zhang 0002, Sicheng Xu, Chuxin Wang, Jiaolong Yang, Feng Zhao 0004, Dong Chen 0003, Baining Guo
ICCV5
2025 MindSearch: Mimicking Human Minds Elicits Deep AI Searcher
abstract
Information seeking and integration is a complex cognitive task that consumes enormous time and effort. Inspired by the remarkable progress of Large Language Models, recent works attempt to solve this task by combining LLMs and search engines. However, these methods still obtain unsatisfying performance due to three challenges: (1) complex requests often cannot be accurately and completely retrieved by the search engine once (2) corresponding information to be integrated is spread over multiple web pages along with massive noise, and (3) a large number of web pages with long contents may quickly exceed the maximum context length of LLMs. Inspired by the cognitive process when humans solve these problems, we introduce MindSearch to mimic the human minds in web information seeking and integration, which can be instantiated by a simple yet effective LLM-based multi-agent framework. The WebPlanner models the human mind of multi-step information seeking as a dynamic graph construction process: it decomposes the user query into atomic sub-questions as nodes in the graph and progressively extends the graph based on the search result from WebSearcher. Tasked with each sub-question, WebSearcher performs hierarchical information retrieval with search engines and collects valuable information for WebPlanner. The multi-agent design of MindSearch enables the whole framework to seek and integrate information parallelly from larger-scale (e.g., more than 300) web pages in 3 minutes, which is worth 3 hours of human effort. MindSearch demonstrates significant improvement in the response quality in terms of depth and breadth, on both close-set and open-set QA problems. Besides, responses from MindSearch based on InternLM2.5-7B are preferable by humans to ChatGPT-Web and Perplexity.ai applications, which implies that MindSearch can already deliver a competitive solution to the proprietary AI search engine.
Kuikun Liu, Qiuchen Wang, Jiangning Liu, Kai Chen 0026, Feng Zhao 0004
ICLR7
2025 PseDet: Revisiting the Power of Pseudo Label in Incremental Object Detection
abstract
Incremental Objection Detection (IOD) facilitates the expansion of the usage scope of object detectors without forgetting previously acquired knowledge. Current approaches mostly adopt response-level knowledge distillation to overcome forgetting issues, by conducting implicit memory replay from the teacher model on new training data. However, this indirect learning paradigm does not fully leverage the knowledge generated by the teacher model. In this paper, we dive deeper into the mechanism of pseudo-labeling in incremental object detection by investigating three critical problems: (a) the upper bound quality of the pseudo labels is greatly limited by the previous model, (b) fixed score thresholds for label filtering, without considering the distribution across categories, and (c) the confidence score generated by the model does not well reflect the quality of the localization. Based on these observations, we propose a simple yet effective pseudo-labeling continual object detection framework, namely PseDet. Specifically, we introduce the spatio-temporal enhancement module to alleviate the negative effects when learning noisy data from the previous model. Considering the score distribution divergence across different classes, we propose the Categorical Adaptive Label Selector with a simple mathematical prior and fast K-Means pre-computation to dynamically determine the class-wise filtering threshold. In order to align the label score with the localization quality of the pseudo labels, we project the score through non-linear mapping to calibrate the distribution and integrate it into the new-step supervision. Extensive experiments on the competitive COCO benchmarks demonstrate the effectiveness and generalization of PseDet. Notably, it achieves 43.5+/41.2+ mAP under the 1/4-step incremental settings, achieving new state-of-the-art performance.
Qiuchen Wang, Chenhongyi Yang, Jiaming Liu 0003, Zhenyu Li 0007, Feng Zhao 0004
ICLR6
2025 AB-Cache: Training-Free Acceleration of Diffusion Models via Adams-Bashforth Cached Feature Reuse
abstract
Diffusion models have demonstrated remarkable success in generative tasks, yet their iterative denoising process results in slow inference, limiting their practicality. While existing acceleration methods exploit the well-known U-shaped similarity pattern between adjacent steps through caching mechanisms, they lack theoretical foundation and rely on simplistic computation reuse, often leading to performance degradation. In this work, we provide a theoretical understanding by analyzing the denoising process through the second-order Adams-Bashforth method, revealing a linear relationship between the outputs of consecutive steps. This analysis explains why the outputs of adjacent steps exhibit a U-shaped pattern. Furthermore, extending Adams-Bashforth method to higher order, we propose a novel caching-based acceleration approach for diffusion models, instead of directly reusing cached results, with a truncation error bound of only (O(hk) where h is the step size. Extensive validation across diverse image and video diffusion models (including HunyuanVideo and FLUX.1-dev) with various schedulers demonstrates our method's effectiveness in achieving nearly 3× speedup while maintaining original performance levels, offering a practical real-time solution without compromising generation quality.
Zichao Yu 0002, Zhen Zou, Guojiang Shao, Shengze Xu, Jie Huang 0017, Feng Zhao 0004, Xiaodong Cun, Wenyi Zhang 0001
ACM Multimedia7
2025 Towards Better & Faster Autoregressive Image Generation: From the Perspective of Entropy
abstract
In this work, we first revisit the sampling issues in current autoregressive (AR) image generation models and identify that image tokens, unlike text tokens, exhibit lower information density and non-uniform spatial distribution. Accordingly, we present an entropy-informed decoding strategy that facilitates higher autoregressive generation quality with faster synthesis speed. Specifically, the proposed method introduces two main innovations: 1) dynamic temperature control guided by spatial entropy of token distributions, enhancing the balance between content diversity, alignment accuracy, and structural coherence in both mask-based and scale-wise models, without extra computational overhead, and 2) entropy-aware acceptance rules in speculative decoding, achieving near-lossless generation at about 85% of the inference cost of conventional acceleration methods. Extensive experiments across multiple benchmarks using diverse AR image generation models demonstrate the effectiveness and generalizability of our approach in enhancing both generation quality and sampling speed.
Feng Zhao 0004, Pengyang Ling, Haibo Qiu, Zhixiang Wei, Hu Yu 0001, Jie Huang 0017, Zhixiong Zeng, Lin Ma 0002
NeurIPS2
2025 VRAG-RL: Empower Vision-Perception-Based RAG for Visually Rich Information Understanding via Iterative Reasoning with Reinforcement Learning
abstract
Effectively retrieving, reasoning and understanding visually rich information remains a challenge for traditional Retrieval-Augmented Generation (RAG) methods. On the one hand, traditional text-based methods cannot handle visual-related information. On the other hand, current vision-based RAG approaches are often limited by fixed pipelines and frequently struggle to reason effectively due to the insufficient activation of the fundamental capabilities of models. As reinforcement learning (RL) has been proven to be beneficial for model reasoning, we introduce VRAG-RL, a novel RL framework tailored for complex reasoning across visually rich information. With this framework, VLMs interact with search engines, autonomously sampling single-turn or multi-turn reasoning trajectories with the help of visual perception tokens and undergoing continual optimization based on these samples. Our approach highlights key limitations of RL in RAG domains: (i) Prior Multi-modal RAG approaches tend to merely incorporate images into the context, leading to insufficient reasoning token allocation and neglecting visual-specific perception; and (ii) When models interact with search engines, their queries often fail to retrieve relevant information due to the inability to articulate requirements, thereby leading to suboptimal performance. To address these challenges, we define an action space tailored for visually rich inputs, with actions including cropping and scaling, allowing the model to gather information from a coarse-to-fine perspective. Furthermore, to bridge the gap between users' original inquiries and the retriever, we employ a simple yet effective reward that integrates query rewriting and retrieval performance with a model-based reward. Our VRAG-RL optimizes VLMs for RAG tasks using specially designed RL strategies, aligning the model with real-world applications. Extensive experiments on diverse and challenging benchmarks show that our VRAG-RL outperforms existing methods by 20\% (Qwen2.5-VL-7B) and 30\% (Qwen2.5-VL-3B), demonstrating the effectiveness of our approach. The code is available at https://github.com/Alibaba-NLP/VRAG.
Qiuchen Wang, Ruixue Ding, Pengjun Xie, Fei Huang 0002, Feng Zhao 0004
NeurIPS9
2025 VideoMAR: Autoregressive Video Generation with Continuous Tokens
abstract
Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored. Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored. In this paper, we propose \textbf{VideoMAR}, a concise and efficient decoder-only autoregressive image-to-video model with continuous tokens, composing temporal frame-by-frame and spatial masked generation. We first identify temporal causality and spatial bi-directionality as the first principle of video AR models, and propose the next-frame diffusion loss for the integration of mask and video generation. Besides, the huge cost and difficulty of long sequence autoregressive modeling is a basic but crucial issue. To this end, we propose the temporal short-to-long curriculum learning and spatial progressive resolution training, and employ progressive temperature strategy at inference time to mitigate the accumulation error. Furthermore, VideoMAR replicates several unique capacities of language models to video generation. It inherently bears high efficiency due to simultaneous temporal-wise KV cache and spatial-wise parallel generation, and presents the capacity of spatial and temporal extrapolation via 3D rotary embeddings. On the VBench-I2V benchmark, VideoMAR surpasses the previous state-of-the-art (Cosmos I2V) while requiring significantly fewer parameters ($9.3\%$), training data ($0.5\%$), and GPU resources ($0.2\%$).
Hu Yu 0001, Biao Gong, Hangjie Yuan, Weilong Chai, Jingdong Chen, Kecheng Zheng, Feng Zhao 0004
NeurIPS8
2025 Cleanness-navigated-contamination network: A unified framework for recovering regional degradation
Qianhao Yu, Naishan Zheng, Jie Huang 0017, Feng Zhao 0004
Comput. Vis. Image Underst.4
2025 Context Sensitive Network for weakly-supervised fine-grained temporal action localization
Cerui Dong, Qinying Liu, Zilei Wang, Yixin Zhang 0007, Feng Zhao 0004
Neural Networks5
2025 Structural and Statistical Texture Knowledge Distillation and Learning for Segmentation
abstract
Low-level texture feature/knowledge is also of vital importance for characterizing the local structural pattern and global statistical properties, such as boundary, smoothness, regularity, and color contrast, which may not be well addressed by high-level deep features. In this paper, we aim to re-emphasize the low-level texture information in deep networks for semantic segmentation and related knowledge distillation tasks. To this end, we take full advantage of both structural and statistical texture knowledge and propose a novel Structural and Statistical Texture Knowledge Distillation (SSTKD) framework for semantic segmentation. Specifically, Contourlet Decomposition Module (CDM) is introduced to decompose the low-level features with iterative Laplacian pyramid and directional filter bank to mine the structural texture knowledge, and Texture Intensity Equalization Module (TIEM) is designed to extract and enhance the statistical texture knowledge with the corresponding Quantization Congruence Loss (QDL). Moreover, we propose the Co-occurrence TIEM (C-TIEM) and generic segmentation frameworks, namely STLNet++ and U-SSNet, to enable existing segmentation networks to harvest the structural and statistical texture information more effectively. Extensive experimental results on three segmentation tasks demonstrate the effectiveness of the proposed methods and their state-of-the-art performance on seven popular benchmark datasets, respectively.
Deyi Ji, Feng Zhao 0004, Hongtao Lu 0001, Feng Wu 0005, Jieping Ye
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 AnySplat: Feed-forward 3D Gaussian Splatting from Unconstrained Views
abstract
We introduce AnySplat, a feed-forward network for novel-view synthesis from uncalibrated image collections. In contrast to traditional neural-rendering pipelines that demand known camera poses and per-scene optimization, or recent feed-forward methods that buckle under the computational weight of dense views—our model predicts everything in one shot. A single forward pass yields a set of 3D Gaussian primitives encoding both scene geometry and appearance, and the corresponding camera intrinsics and extrinsics for each input image. This unified design scales effortlessly to casually captured, multi-view datasets without any pose annotations. In extensive zero-shot evaluations, AnySplat matches the quality of pose-aware baselines in both sparse- and dense-view scenarios while surpassing existing pose-free approaches. Moreover, it greatly reduces rendering latency compared to optimization-based neural fields, bringing real-time novel-view synthesis within reach for unconstrained capture settings. Project page: https://city-super.github.io/anysplat/.
Lihan Jiang, Yucheng Mao, Linning Xu, Tao Lu 0005, Kerui Ren, Yichen Jin, Xudong Xu, Mulin Yu, Jiangmiao Pang, Feng Zhao 0004, Dahua Lin, Bo Dai 0002
ACM Trans. Graph.10
2024 Leveraging Imagery Data with Spatial Point Prior for Weakly Semi-supervised 3D Object Detection
abstract
Training high-accuracy 3D detectors necessitates massive labeled 3D annotations with 7 degree-of-freedom, which is laborious and time-consuming. Therefore, the form of point annotations is proposed to offer significant prospects for practical applications in 3D detection, which is not only more accessible and less expensive but also provides strong spatial information for object localization. In this paper, we empirically discover that it is non-trivial to merely adapt Point-DETR to its 3D form, encountering two main bottlenecks: 1) it fails to encode strong 3D prior into the model, and 2) it generates low-quality pseudo labels in distant regions due to the extreme sparsity of LiDAR points. To overcome these challenges, we introduce Point-DETR3D, a teacher-student framework for weakly semi-supervised 3D detection, designed to fully capitalize on point-wise supervision within a constrained instance-wise annotation budget. Different from Point-DETR which encodes 3D positional information solely through a point encoder, we propose an explicit positional query initialization strategy to enhance the positional prior. Considering the low quality of pseudo labels at distant regions produced by the teacher model, we enhance the detector's perception by incorporating dense imagery data through a novel Cross-Modal Deformable RoI Fusion (D-RoI). Moreover, an innovative point-guided self-supervised learning technique is proposed to allow for fully exploiting point priors, even in student models. Extensive experiments on representative nuScenes dataset demonstrate our Point-DETR3D obtains significant improvements compared to previous works. Notably, with only 5% of labeled data, Point-DETR3D achieves over 90% performance of its fully supervised counterpart.
Hongzhi Gao, Lin Chen 0019, Jiaming Liu 0003, Shanghang Zhang, Feng Zhao 0004
AAAI7
2024 T-Eval: Evaluating the Tool Utilization Capability of Large Language Models Step by Step
abstract
Zehui Chen, Weihua Du, Wenwei Zhang, Kuikun Liu, Jiangning Liu, Miao Zheng, Jingming Zhuo, Songyang Zhang, Dahua Lin, Kai Chen, Feng Zhao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Weihua Du, Kuikun Liu, Jiangning Liu, Miao Zheng, Jingming Zhuo, Songyang Zhang 0001, Dahua Lin, Kai Chen 0026, Feng Zhao 0004
ACL (1)11
2024 PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
abstract
Zaibin Zhang, Yongting Zhang, Lijun Li, Hongzhi Gao, Lijun Wang, Huchuan Lu, Feng Zhao, Yu Qiao, Jing Shao. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Zaibin Zhang, Yongting Zhang, Hongzhi Gao, Yu Qiao 0001, Lijun Wang 0001, Huchuan Lu, Feng Zhao 0004
ACL (1)9
2024 Bimodality Approach to Mental Disorder Diagnosis via Dual-pathway Fusion Perceptual Network
abstract
Neuroimaging technologies such as functional MRI (fMRI) are used for studying the neural mechanisms underlying mental disorders. Deep learning models offer promising solutions for analyzing neuroimaging data. However, singlemodality methods suffer from performance bottle-necks, while existing multi-modality methods fail to fully utilize the complimentary information between different modality data. We proposed a dual-pathway fusion perceptual (DP-FP) network that leverages multi-stage attentional perception, effectively distinguishing psychiatric disorder patients from healthy controls. Specifically, DP-FP consists of two pathways to receive time course and functional connection inputs respectively. Self-attentive modules were introduced into the network structure, which not only enhance the capability of fine-grained feature extraction, but also improve the interpretability of model. Furthermore, a coarse perception module and a fusion strategy were proposed, which achieved deep integration of feature representations from the two pathways and optimized the complementarity of information. We validated our proposed DP-FP network using the public schizophrenia dataset (COBRE) and the public autism dataset (ABIDE). On COBRE, the accuracy for predicting schizophrenia was found to be 87.6%. On ABIDE, the accuracy for predicting autism was found to be 71.1%. Group-discriminative brain regions were attributed and visualized through gradient-based interpretation. DP-FP focused more on the paracentral lobe and near the angular gyrus in the diagnostic task of schizophrenia, and the frontal lobe, temporal lobe and angular gyrus in the di-agnostic task of autism. DP-FP achieved state-of-the-art accuracy on benchmark datasets for both schizophrenia and autism. DP-FP also offers high interpretability, providing novel insights into exploring pathological mechanisms.
Ruipeng Xu, Shuoqiu Gan, Feng Zhao 0004, Zhentao Zuo, Tiangang Zhou
BIBM3
2024 Rectifying Shortcut Learning through Cellular Differentiation in Deep Learning Neurons
Hongjing Niu, Hanting Li, Guoping Wu, Bin Li 0025, Feng Zhao 0004
BMVC5
2024 Frequency Decomposition to Tap the Potential of Single Domain for Generalization
Hongjing Niu, Qingyue Yang, Wei Zhang 0251, Bin Li 0025, Feng Zhao 0004
BMVC6
2024 Task-Related Feature Enhancement Network for Neuronal Morphology Classification
Chunli Sun, Feng Zhao 0004
BMVC2
2024 Revisiting Spatial-Frequency Information Integration from a Hierarchical Perspective for Panchromatic and Multi-Spectral Image Fusion
abstract
Pan-sharpening is a super-resolution problem that essentially relies on spectra fusion of panchromatic (PAN) images and low-resolution multi-spectral (LRMS) images. The previous methods have validated the effectiveness of information fusion in the Fourier space of the whole image. However, they haven't fully explored the Fourier relationships at different hierarchies between PAN and LRMS images. To this end, we propose a Hierarchical Frequency Integration Network (HFIN) to facilitate hierarchical Fourier information integration for pan-sharpening. Specifically, our network consists of two designs: information stratification and information integration. For information stratification, we hierarchically decompose PAN and LRMS information into spatial, global Fourier and local Fourier information, and fuse them independently. For information integration, the above hierarchical fused information is processed to further enhance their relationships and undergo comprehensive integration. Our method extend a new space for exploring the relationships of PAN and LRMS images, enhancing the integration of spatial-frequency information. Extensive experiments robustly validate the effectiveness of the proposed network, showcasing its superior performance compared to other state-of-the-art methods and generalization in real-world scenes and other fusion tasks as a general image fusion framework. Code is available at https://github.com/JosephTiTan/HFIN.
Jiangtong Tan, Jie Huang 0017, Naishan Zheng, Man Zhou 0003, Danfeng Hong, Feng Zhao 0004
CVPR7
2024 Empowering Resampling Operation for Ultra-High-Definition Image Enhancement with Model-Aware Guidance
abstract
Image enhancement algorithms have made remarkable advancements in recent years, but directly applying them to Ultra-high-definition (UHD) images presents intractable computational overheads. Therefore, previous straightforward solutions employ resampling techniques to reduce the resolution by adopting a “Downsampling-Enhancement-Upsampling” processing paradigm. However, this paradigm disentangles the resampling operators and inner enhancement algorithms, which results in the loss of information that is favored by the model, further leading to sub-optimal outcomes. In this paper, we propose a novel method of Learning Model-Aware Resampling (LMAR), which learns to customize resampling by extracting model-aware information from the UHD input image, under the guidance of model knowledge. Specifically, our method consists of two core designs, namely compensatory kernel estimation and steganographic resampling. At the first stage, we dynamically predict compensatory kernels tailored to the specific input and resampling scales. At the second stage, the image-wise compensatory information is derived with the compensatory kernels and embedded into the rescaled input images. This promotes the representation of the newly derived downscaled inputs to be more consistent with the full-resolution UHD inputs, as perceived by the model. Our LMAR enables model-aware and model-favored resampling while maintaining compatibility with existing resampling operators. Extensive experiments on multiple UHD image enhancement datasets and different backbones have shown consistent performance gains after correlating resizer and enhancer; e.g., up to 1.2dB PSNR gain for ×1.8 resampling scale on UHD-LOL4K. The code is available at https://github.com/YPatrickW/LMAR.
Jie Huang 0017, Bing Li 0024, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004
CVPR7
2024 Probing Synergistic High-Order Interaction in Infrared and Visible Image Fusion
abstract
Infrared and visible image fusion aims to generate a fused image by integrating and distinguishing complementary information from multiple sources. While the cross-attention mechanism with global spatial interactions appears promising, it only capture second-order spatial inter-actions, neglecting higher-order interactions in both spatial and channel dimensions. This limitation hampers the ex-ploitation of synergies between multi-modalities. To bridge this gap, we introduce a Synergistic High-order Interaction Paradigm (SHIP), designed to systematically investigate the spatial fine-grained and global statistics collaborations between infrared and visible images across two fundamental dimensions: 1) Spatial dimension: we construct spatial fine-grained interactions through element-wise multiplication, mathematically equivalent to global interactions, and then foster high-order formats by iteratively aggregating and evolving complementary information, enhancing both efficiency andflexibility; 2) Channel dimension: expanding on channel interactions with first-order statistics (mean), we devise high-order channel interactions to facilitate the discernment of inter-dependencies between source images based on global statistics. Harnessing high-order interactions significantly enhances our model's ability to exploit multi-modal synergies, leading to superior performance over state-of-the-art alternatives, as shown through comprehensive experiments across various benchmarks. Code is available at https://github.com/zheng980629/SHIP.
Naishan Zheng, Man Zhou 0003, Jie Huang 0017, Junming Hou, Haoying Li, Feng Zhao 0004
CVPR7
2024 ShareGPT4V: Improving Large Multi-modal Models with Better Captions
Lin Chen 0026, Jinsong Li 0001, Xiaoyi Dong, Pan Zhang 0001, Conghui He, Jiaqi Wang 0003, Feng Zhao 0004, Dahua Lin
ECCV (17)7
2024 Stable Preference: Redefining Training Paradigm of Human Preference Model for Text-to-Image Synthesis
Hanting Li, Hongjing Niu, Feng Zhao 0004
ECCV (28)3
2024 Idling Neurons, Appropriately Lenient Workload During Fine-Tuning Leads to Better Generalization
Hongjing Niu, Hanting Li, Bin Li 0025, Feng Zhao 0004
ECCV (53)4
2024 Stream Query Denoising for Vectorized HD-Map Construction
Fan Jia 0006, Weixin Mao, Yingfei Liu, Tiancai Wang, Chi Zhang 0026, Xiangyu Zhang 0005, Feng Zhao 0004
ECCV (19)10
2024 Unleashing the Potential of the Semantic Latent Space in Diffusion Models for Image Dehazing
Zizheng Yang, Hu Yu 0001, Bing Li 0024, Jie Huang 0017, Feng Zhao 0004
ECCV (44)6
2024 Unmasking Bias in Diffusion Model Training
Hu Yu 0001, Jie Huang 0017, Feng Zhao 0004
ECCV (66)5
2024 RodinHD: High-Fidelity 3D Avatar Generation with Diffusion Models
Bowen Zhang 0010, Yiji Cheng, Chunyu Wang 0001, Ting Zhang 0002, Jiaolong Yang, Yansong Tang, Feng Zhao 0004, Dong Chen 0003, Baining Guo
ECCV (14)7
2024 Changenet: Multi-Temporal Asymmetric Change Detection Dataset
abstract
Change Detection (CD) has been attracting extensive interests with the availability of bi-temporal datasets. However, due to the huge cost of multi-temporal images acquisition and labeling, existing change detection datasets are small in quantity, short in temporal, and low in practicability. Therefore, a large-scale practical-oriented dataset covering wide temporal phases is urgently needed to facilitate the community. To this end, the ChangeNet dataset is presented especially for multi-temporal change detection, along with the new task of "Asymmetric Change Detection". Specifically, ChangeNet consists of 31,000 multi-temporal images pairs, a wide range of complex scenes from 100 cities, and 6 pixel-level annotated categories, which is far superior to all the existing change detection datasets including LEVIR-CD, WHU Building CD, etc.. In addition, ChangeNet contains amounts of real-world perspective distortions in different temporal phases on the same areas, which is able to promote the practical application of change detection algorithms. The ChangeNet dataset is suitable for both binary change detection (BCD) and semantic change detection (SCD) tasks. Accordingly, we benchmark the ChangeNet dataset on six BCD methods and two SCD methods, and extensive experiments demonstrate its challenges and great significance. The dataset is available at https://github.com/jankyee/ChangeNet.
Deyi Ji, Mingyuan Tao, Hongtao Lu 0001, Feng Zhao 0004
ICASSP5
2024 CLIPER: A Unified Vision-Language Framework for In-the-Wild Facial Expression Recognition
abstract
As one of the most informative behaviors of humans, facial expressions are often compound and variable, which is manifested by the fact that different people may express the same expression in very different ways. However, most facial expression recognition (FER) methods still use one-hot or soft labels as the supervision, which lack sufficient semantic descriptions of facial expressions and are less interpretable. Recently, contrastive vision-language pre-training models (e.g., CLIP) use text as the supervision and have injected new vitality into various computer vision tasks, benefiting from the rich semantics in text. Therefore, we propose CLIPER, a unified framework for both static and dynamic facial Expression Recognition based on CLIP. Besides, we introduce multiple expression text descriptors (METD) to learn fine-grained expression representations and a two-stage training paradigm to reserve the interpretability of CLIP. We conduct extensive experiments on several popular FER benchmarks to demonstrates the effectiveness of CLIPER. The source code will be available at https://github.com/muse1998/CLIPER.
Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao 0004
ICME4
2024 Discrete Latent Perspective Learning for Segmentation and Detection
abstract
In this paper, we address the challenge of Perspective-Invariant Learning in machine learning and computer vision, which involves enabling a network to understand images from varying perspectives to achieve consistent semantic interpretation. While standard approaches rely on the labor-intensive collection of multi-view images or limited data augmentation techniques, we propose a novel framework, Discrete Latent Perspective Learning (DLPL), for latent multi-perspective fusion learning using conventional single-view images. DLPL comprises three main modules: Perspective Discrete Decomposition (PDD), Perspective Homography Transformation (PHT), and Perspective Invariant Attention (PIA), which work together to discretize visual features, transform perspectives, and fuse multi-perspective semantic information, respectively. DLPL is a universal perspective learning framework applicable to a variety of scenarios and vision tasks. Extensive experiments demonstrate that DLPL significantly enhances the network’s capacity to depict images across diverse scenarios (daily photos, UAV, auto-driving) and tasks (detection, segmentation).
Deyi Ji, Feng Zhao 0004, Lanyun Zhu, Wenwei Jin, Hongtao Lu 0001, Jieping Ye
ICML2
2024 Unsupervised Low-Light Image Enhancement via Spectral Consistency
Bing Li 0024, Naishan Zheng, Jie Huang 0017, Feng Zhao 0004
ICPR (22)5
2024 Unsupervised Low-Light Image Enhancement with Dual Contrastive Learning
Bing Li 0024, Jie Huang 0017, Feng Zhao 0004
ICPR (21)4
2024 PPTFormer: Pseudo Multi-Perspective Transformer for UAV Segmentation
Deyi Ji, Wenwei Jin, Hongtao Lu 0001, Feng Zhao 0004
IJCAI4
2024 Training Pansharpening Networks at Full Resolution Using Degenerate Invariance
abstract
Pansharpening is an important technique for remote sensing imaging systems to obtain high-resolution multispectral images. Existing deep learning-based methods mostly rely on using pseudo-groundtruth multi-spectral images for supervised learning. The whole training process only remains at the scale of reduced resolution, which means that the impact of the degradation process is ignored and high-quality images cannot be guaranteed at full resolution. To address the challenge, we propose a new unsupervised framework that does not rely on pseudo-groundtruth but uses the invariance of the degradation process to build a consistent loss function on the original scale for network training. Specifically, we first introduce the operator learning method to build an exact mapping function from multi-spectral to panchromatic images and decouple both spectral and texture features. Then, through joint training, operators and convolutional networks can learn the spatial degradation process and spectral degradation process at full resolution, respectively. By introducing them to build consistency constraints, we can train the pansharpening network at the original full resolution. Our approach can be applied to existing pansharpening methods, improving their usability on original data, which matches practical application requirements. The experimental results on different kinds of satellite datasets demonstrate that the proposed network outperforms state-of-the-art methods both visually and quantitatively. Our code is available at https://github.com/quycruin/Qvac.
Yichang Qu, Bing Li 0024, Jie Huang 0017, Feng Zhao 0004
ACM Multimedia4
2024 Image-free Pre-training for Low-Level Vision
abstract
The constrained data scale in low-level vision often induces the demon overfitting hazard for restoration networks, necessitating the adoption of the pre-training paradigm. Mirroring the success of the high-level pre-training approaches, recent methods in the low-level community aim to derive general visual representation from extensive data with synthesized degradation. In this paper, we propose a new perspective beyond the data-driven image pre-training paradigm for low-level vision, building upon the following examination. First, unlike the semantic extraction prevalent in high-level vision, low-level vision primarily focuses on the continuous and content-agnostic pixel-level regression, indicating that the diversified contents inherent in large-scale data are potentially unnecessary for low-level vision pre-training. Second, considering the low-level degradations are highly relevant to the frequency spectrum, we discern that the low-level pre-training paradigm can be implemented in the Fourier space with fostered degradation sensibility. Therefore, we develop an Image-free Pre-training (IFP) paradigm, a novel low-level pre-training approach with necessity of single randomly sampled Gaussian noise image, streamlining complicated data collection and synthesis procedure. The principle of the IFP involves reconstructing the original Gaussian noise from the randomly perturbed counterpart with partially masked spectrum band, facilitating the capability for robust spectrum representation extraction in response to the capricious downstream degradations. Extensive experiments demonstrate the significant improvements brought by IFP to various downstream tasks, such as 1.31 dB boost in low-light enhancement for Restormer, and improvements of 1.2 dB in deblurring, and 2.42 dB in deraining for Uformer. Code is publicly available at https://github.com/siywang541/IFP.
Jie Huang 0017, Feng Zhao 0004
ACM Multimedia4
2024 FreqMamba: Viewing Mamba from a Frequency Perspective for Image Deraining
abstract
Images corrupted by rain streaks often lose vital frequency information for perception, and image deraining aims to solve this problem, which relies on global and local degradation modeling. Recent studies have witnessed the effectiveness and efficiency of Mamba for perceiving global and local information based on its exploiting local correlation among patches, however, rarely attempts have been explored to extend it with frequency analysis for image deraining, limiting its ability to perceive global degradation that is relevant to frequency modeling (e.g. Fourier transform). In this paper, we propose FreqMamba, an effective and efficient paradigm that leverages the complementary between Mamba and frequency analysis for image deraining. The core of our method lies in extending Mamba with frequency analysis from two perspectives: extending it with frequency band for exploiting frequency correlation, and connecting it with Fourier transform for global degradation modeling. Specifically, FreqMamba introduces complementary triple interaction structures including spatial Mamba, frequency-band Mamba, and Fourier global modeling. Frequency-Band Mamba decomposes the image into sub-bands of different frequencies to allow 2D scanning from the frequency dimension. Furthermore, leveraging Mamba's unique data-dependent properties, we use rainy images at different scales to provide degradation priors to the network, thereby facilitating efficient training. Extensive experiments show that our method outperforms state-of-the-art methods both visually and quantitatively. Our code is available at: https://github.com/aSleepyTree/FreqMamba.
Zhen Zou, Hu Yu 0001, Jie Huang 0017, Feng Zhao 0004
ACM Multimedia4
2024 ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
abstract
We present the ShareGPT4Video series, aiming to facilitate the video understanding of large video-language models (LVLMs) and the video generation of text-to-video models (T2VMs) via dense and precise captions. The series comprises: 1) ShareGPT4Video, 40K GPT4V annotated dense captions of videos with various lengths and sources, developed through carefully designed data filtering and annotating strategy. 2) ShareCaptioner-Video, an efficient and capable captioning model for arbitrary videos, with 4.8M high-quality aesthetic videos annotated by it. 3) ShareGPT4Video-8B, a simple yet superb LVLM that reached SOTA performance on three advancing video benchmarks. To achieve this, taking aside the non-scalable costly human annotators, we find using GPT4V to caption video with a naive multi-frame or frame-concatenation input strategy leads to less detailed and sometimes temporal-confused results. We argue the challenge of designing a high-quality video captioning strategy lies in three aspects: 1) Inter-frame precise temporal change understanding. 2) Intra-frame detailed content description. 3) Frame-number scalability for arbitrary-length videos. To this end, we meticulously designed a differential video captioning strategy, which is stable, scalable, and efficient for generating captions for videos with arbitrary resolution, aspect ratios, and length. Based on it, we construct ShareGPT4Video, which contains 40K high-quality videos spanning a wide range of categories, and the resulting captions encompass rich world knowledge, object attributes, camera movements, and crucially, detailed and precise temporal descriptions of events. Based on ShareGPT4Video, we further develop ShareCaptioner-Video, a superior captioner capable of efficiently generating high-quality captions for arbitrary videos. We annotated 4.8M aesthetically appealing videos by it and verified their effectiveness on a 10-second text2video generation task. For video understanding, we verified the effectiveness of ShareGPT4Video on several current LVLM architectures and presented our superb new LVLM ShareGPT4Video-8B. All the models, strategies, and annotations will be open-sourced and we hope this project can serve as a pivotal resource for advancing both the LVLMs and T2VMs community.
Lin Chen 0016, Xilin Wei, Jinsong Li 0001, Xiaoyi Dong, Pan Zhang 0001, Yuhang Zang, Haodong Duan, Lin Bin, Zhenyu Tang 0004, Li Yuan 0007, Yu Qiao 0001, Dahua Lin, Feng Zhao 0004, Jiaqi Wang 0003
NeurIPS14
2024 Are We on the Right Way for Evaluating Large Vision-Language Models?
abstract
Large vision-language models (LVLMs) have recently achieved rapid progress, sparking numerous studies to evaluate their multi-modal capabilities. However, we dig into current evaluation works and identify two primary issues: 1) Visual content is unnecessary for many samples. The answers can be directly inferred from the questions and options, or the world knowledge embedded in LLMs. This phenomenon is prevalent across current benchmarks. For instance, GeminiPro achieves 42.7% on the MMMU benchmark without any visual input, and outperforms the random choice baseline across six benchmarks near 24% on average. 2) Unintentional data leakage exists in LLM and LVLM training. LLM and LVLM could still answer some visual-necessary questions without visual content, indicating the memorizing of these samples within large-scale training data. For example, Sphinx-X-MoE gets 43.6% on MMMU without accessing images, surpassing its LLM backbone with 17.9%. Both problems lead to misjudgments of actual multi-modal gains and potentially misguide the study of LVLM. To this end, we present MMStar, an elite vision-indispensable multi-modal benchmark comprising 1,500 samples meticulously selected by humans. MMStar benchmarks 6 core capabilities and 18 detailed axes, aiming to evaluate LVLMs' multi-modal capacities with carefully balanced and purified samples. These samples are first roughly selected from current benchmarks with an automated pipeline, human review is then involved to ensure each curated sample exhibits visual dependency, minimal data leakage, and requires advanced multi-modal capabilities. Moreover, two metrics are developed to measure data leakage and actual performance gain in multi-modal training. We evaluate 16 leading LVLMs on MMStar to assess their multi-modal capabilities, and on 7 benchmarks with the proposed metrics to investigate their data leakage and actual multi-modal gain.
Jinsong Li 0001, Xiaoyi Dong, Pan Zhang 0001, Yuhang Zang, Haodong Duan, Jiaqi Wang 0003, Yu Qiao 0001, Dahua Lin, Feng Zhao 0004
NeurIPS11
2024 GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative Modeling
abstract
We introduce a radiance representation that is both structured and fully explicit and thus greatly facilitates 3D generative modeling. Existing radiance representations either require an implicit feature decoder, which significantly degrades the modeling power of the representation, or are spatially unstructured, making them difficult to integrate with mainstream 3D diffusion methods. We derive GaussianCube by first using a novel densification-constrained Gaussian fitting algorithm, which yields high-accuracy fitting using a fixed number of free Gaussians, and then rearranging these Gaussians into a predefined voxel grid via Optimal Transport. Since GaussianCube is a structured grid representation, it allows us to use standard 3D U-Net as our backbone in diffusion modeling without elaborate designs. More importantly, the high-accuracy fitting of the Gaussians allows us to achieve a high-quality representation with orders of magnitude fewer parameters than previous structured representations for comparable quality, ranging from one to two orders of magnitude. The compactness of GaussianCube greatly eases the difficulty of 3D generative modeling. Extensive experiments conducted on unconditional and class-conditioned object generation, digital avatar creation, and text-to-3D synthesis all show that our model achieves state-of-the-art generation results both qualitatively and quantitatively, underscoring the potential of GaussianCube as a highly accurate and versatile radiance representation for 3D generative modeling.
Bowen Zhang 0010, Yiji Cheng, Jiaolong Yang, Chunyu Wang 0001, Feng Zhao 0004, Yansong Tang, Dong Chen 0003, Baining Guo
NeurIPS5
2024 Few-shot adaptation of GANs using self-supervised consistency regularization
Syed Muhammad Israr, Rehan Saeed, Feng Zhao 0004
Knowl. Based Syst.3
2024 Learning Spatio-Temporal Sharpness Map for Video Deblurring
abstract
Video deblurring is a challenging task because only input blurry sequences are available. To further constrain the optimization process, existing methods explore various additional information,e.g., events, depth and sharpness prior. However, they consume large computing costs or generate unpleasant visual results due to the insufficient exploitation of spatio-temporal information. In this work, we develop a novel spatio-temporal sharpness map learned by a prior-based generation network implicitly. The proposed generation network blends both spatial and temporal sharpness priors in a blurry sequence, while few extra parameters are added. We show that the proposed map has better spatial continuity and guidance for video deblurring than the previous method. Furthermore, different from the simply concatenation in the previous work, we allow the sharpness map to accommodate to more effective video deblurring via a dual-stream network. Specifically, the network is decomposed by two branches, namely inter-frame and intra-frame reconstructions. The inter-frame reconstruction obtains the sharp patches of cecutive frames from the sharpness map to restore textures well. Meanwhile, the other intra-frame branch is responsible for recovering structures of the latent frame, where a novel histogram statistical method is developed to quantify and count textures in the feature under the modulation of the sharpness map. Quantitative and qualitative experiments successfully validate the effectiveness of our proposed method.
Qi Zhu 0010, Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Feng Zhao 0004
IEEE Trans. Circuits Syst. Video Technol.6
2024 Graph-DETR4D: Spatio-Temporal Graph Modeling for Multi-View 3D Object Detection
abstract
Multi-View 3D object detection (MV3D) has made tremendous progress by leveraging multiple perspective features through surrounding cameras. Despite demonstrating promising prospects in various applications, accurately detecting objects through camera view in the 3D space is extremely difficult due to the ill-posed issue in monocular depth estimation. Recently, Graph-DETR3D presents a novel graph-based 3D-2D query paradigm in aggregating multi-view images for 3D object detection and achieves competitive performance. Although it enriches the query representations with 2D image features through a learnable 3D graph, it still suffers from limited depth and velocity estimation abilities due to the adoption of a single-frame input setting. To solve this problem, we introduce a unified spatial-temporal graph modeling framework to fully leverage the multi-view imagery cues under the multi-frame inputs setting. Thanks to the flexibility and sparsity of the dynamic graph architecture, we lift the original 3D graph into the 4D space with an effective attention mechanism to automatically perceive imagery information at both spatial and temporal levels. Moreover, considering the main latency bottleneck lies in the image backbone, we propose a novel dense-sparse distillation framework for multi-view 3D object detection, to reduce the computational budget while sacrificing no detection accuracy, making it more suitable for real-world deployment. To this end, we propose Graph-DETR4D, a faster and stronger multi-view 3D object detection framework, built on top of Graph-DETR3D. Extensive experiments on nuScenes and Waymo benchmarks demonstrate the effectiveness and efficiency of Graph-DETR4D. Notably, our best model achieves 62.0% NDS on nuScenes test leaderboard. Code is available at https://github.com/zehuichen123/Graph-DETR4D.
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Wu 0005, Feng Zhao 0004
IEEE Trans. Image Process.8
2024 DDOD: Dive Deeper into the Disentanglement of Object Detector
abstract
Compared to many other dense prediction tasks, object detection plays a fundamental role in visual perception and scene understanding. Dense object detection, aiming at localizing objects directly from the feature map, has drawn great attention due to its low cost and high efficiency. Though it has been developed for a long time, the training pipeline of dense object detectors is still compromised to lots of conjunctions. In this paper, we demonstrate the existence of three conjunctions lying in the current paradigm of one-stage detectors: 1) only samples assigned as positive in classification head are used to train the regression head; 2) classification and regression share the same input feature and computational fields defined by the parallel head architecture; and 3) samples distributed in different feature pyramid layers are treated equally when computing the loss. Based on this, we propose Disentangled Dense Object Detector (DDOD), a simple, direct, and efficient framework for 2D detection with strong performance. We derive two DDOD variants (i.e., DR-CNN, and DDETR) following the basic one-stage/two-stage and recently developed transformer-based pipelines. Specifically, we develop three effective disentanglement mechanisms and integrate them into the current state-of-the-art object detectors. Extensive experiments on MS COCO benchmark show that our approach obtains significant enhancements with negligible extra overhead on various detectors. Notably, our best model reaches 55.4 mAP on the COCOtest-devset, achieving new state-of-the-art performance on this competitive benchmark. Additionally, we validate our model on several challenging tasks including small object detection and crowded object detection. The experimental results further prove the superiority of disentanglement on these conjunctions. Code is available athttps://github.com/zehuichen123/DDOD.
Chenhongyi Yang, Feng Zhao 0004, Zhengjun Zha, Feng Wu 0001
IEEE Trans. Multim.4
2024 IRVR: A General Image Restoration Framework for Visual Recognition
abstract
Images corrupted with degradations often result in a performance drop in downstream image recognition models trained on clean images. Previous image restoration (IR) methods either restore the images without delicately considering the semantic recovery, or the training objectives cannot meet unseen recognition models, leading to poor and non-generalizable performance for various downstream recognition tasks. In this paper, we propose a general image restoration framework for visual recognition, IRVR, which is addressed for generalized and effective semantic recovery in image restoration for a range of high-level tasks. Concretely, for better generalization, we train the IR models with semantic recovery as the primary objective, and image regression as a regularization term, respectively, where the primary objective gradient is calibrated with the regularization gradient to ensure the generalization of IR to unseen recognition models. For effectiveness, we introduce an intrinsic semantic consistency constraint to match the semantic statistical distribution between restored and clean image pairs. Our IRVR is recognition-agnostic and orthogonal to IR, making it a plug-and-play component that can be incorporated into existing IR methods without adding any computation cost during inference. Extensive experiments demonstrate the effectiveness and generalization of our IRVR for improving the performance of IR in diverse downstream high-level tasks. The IRVR's ability to accurately recover intrinsic semantics in images is instrumental in high-level machine analysis, which ensures the integrity and authenticity of multimedia content.
Zizheng Yang, Jie Huang 0017, Man Zhou 0003, Naishan Zheng, Feng Zhao 0004
IEEE Trans. Multim.5
2024 Rethinking Pan-Sharpening in Closed-Loop Regularization
abstract
It is generally known that pan-sharpening is fundamentally a PAN-guided multispectral (MS) image super-resolution problem that involves learning the nonlinear mapping from low-resolution (LR) to high-resolution (HR) MS images. Since an infinite number of HR-MS images can be downsampled to produce the same corresponding LR-MS image, learning the mapping from LR-MS to HR-MS image is typically ill-posed and the space of the possible pan-sharpening functions can be extremely large, making it difficult to estimate the optimal mapping solution. To address the above issue, we propose a closed-loop scheme that learns the two opposite mapping including the pan-sharpening and its corresponding degradation process simultaneously to regularize the solution space in a single pipeline. More specifically, an invertible neural network (INN) is introduced to perform a bidirectional closed-loop: the forward operation for LR-MS pan-sharpening and the backward operation for learning the corresponding HR-MS image degradation process. In addition, given the vital importance of high-frequency textures for the Pan-sharpened MS images, we further strengthen the INN by designing a specified multiscale high-frequency texture extraction module. Extensive experimental results demonstrate that the proposed algorithm performs favorably against state-of-the-art methods qualitatively and quantitatively with fewer parameters. Ablation studies also verify the effectiveness of the closed-loop mechanism in pan-sharpening. The source code is made publicly available at https://github.com/manman1995/pan-sharpening-Team-zhouman/.
Man Zhou 0003, Jie Huang 0017, Danfeng Hong, Feng Zhao 0004, Chongyi Li, Jocelyn Chanussot
IEEE Trans. Neural Networks Learn. Syst.4
2023 Intensity-Aware Loss for Dynamic Facial Expression Recognition in the Wild
abstract
Compared with the image-based static facial expression recognition (SFER) task, the dynamic facial expression recognition (DFER) task based on video sequences is closer to the natural expression recognition scene. However, DFER is often more challenging. One of the main reasons is that video sequences often contain frames with different expression intensities, especially for the facial expressions in the real-world scenarios, while the images in SFER frequently present uniform and high expression intensities. Nevertheless, if the expressions with different intensities are treated equally, the features learned by the networks will have large intra-class and small inter-class differences, which are harmful to DFER. To tackle this problem, we propose the global convolution-attention block (GCA) to rescale the channels of the feature maps. In addition, we introduce the intensity-aware loss (IAL) in the training process to help the network distinguish the samples with relatively low expression intensities. Experiments on two in-the-wild dynamic facial expression datasets (i.e., DFEW and FERV39k) indicate that our method outperforms the state-of-the-art DFER approaches. The source code will be available at https://github.com/muse1998/IAL-for-Facial-Expression-Recognition.
Hanting Li, Hongjing Niu, Zhaoqing Zhu, Feng Zhao 0004
AAAI4
2023 Learning Semantic Degradation-Aware Guidance for Recognition-Driven Unsupervised Low-Light Image Enhancement
abstract
Low-light images suffer severe degradation of low lightness and noise corruption, causing unsatisfactory visual quality and visual recognition performance. To solve this problem while meeting the unavailability of paired datasets in wide-range scenarios, unsupervised low-light image enhancement (ULLIE) techniques have been developed. However, these methods are primarily guided to alleviate the degradation effect on visual quality rather than semantic levels, hence limiting their performance in visual recognition tasks. To this end, we propose to learn a Semantic Degradation-Aware Guidance (SDAG) that perceives the low-light degradation effect on semantic levels in a self-supervised manner, which is further utilized to guide the ULLIE methods. The proposed SDAG utilizes the low-light degradation factors as augmented signals to degrade the low-light images, and then capture their degradation effect on semantic levels. Specifically, our SDAG employs the subsequent pre-trained recognition model extractor to extract semantic representations, and then learns to self-reconstruct the enhanced low-light image and its augmented degraded images. By constraining the relative reconstruction effect between the original enhanced image and the augmented formats, our SDAG learns to be aware of the degradation effect on semantic levels in a relative comparison manner. Moreover, our SDAG is general and can be plugged into the training paradigm of the existing ULLIE methods. Extensive experiments demonstrate its effectiveness for improving the ULLIE approaches on the downstream recognition tasks while maintaining a competitive visual quality. Code will be available at https://github.com/zheng980629/SDAG.
Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Zizheng Yang, Qi Zhu 0010, Feng Zhao 0004
AAAI6
2023 Frequency-consistent Optimization for Image Enhancement Networks
Bing Li 0024, Naishan Zheng, Qi Zhu 0010, Jie Huang 0017, Feng Zhao 0004
BMVC5
2023 Learning Sample Relationship for Exposure Correction
abstract
Exposure correction task aims to correct the underexposure and its adverse overexposure images to the normal exposure in a single network. As well recognized, the optimization flow is the opposite. Despite great advancement, existing exposure correction methods are usually trained with a mini-batch of both underexposure and overexposure mixed samples and have not explored the relationship between them to solve the optimization inconsistency. In this paper, we introduce a new perspective to conjunct their optimization processes by correlating and constraining the relationship of correction procedure in a mini-batch. The core designs of our framework consist of two steps: 1) formulating the exposure relationship of samples across the batch dimension via a context-irrelevant pretext task. 2) delivering the above sample relationship design as the regularization term within the loss function to promote optimization consistency. The proposed sample relationship design as a general term can be easily integrated into existing exposure correction methods without any computational burden in inference time. Extensive experiments over multiple representative exposure correction benchmarks demonstrate consistent performance gains by introducing our sample relationship design.
Jie Huang 0017, Feng Zhao 0004, Man Zhou 0003, Jie Xiao 0002, Naishan Zheng, Zhiwei Xiong
CVPR2
2023 Ultra-High Resolution Segmentation with Ultra-Rich Context: A Novel Benchmark
abstract
With the increasing interest and rapid development of methods for Ultra-High Resolution (UHR) segmentation, a large-scale benchmark covering a wide range of scenes with full fine-grained dense annotations is urgently needed to facilitate the field. To this end, the URUR dataset is introduced, in the meaning of Ultra-High Resolution dataset with Ultra-Rich Context. As the name suggests, URUR contains amounts of images with high enough resolution (3,008 images of size 5,120 × 5,120), a wide range of complex scenes (from 63 cities), rich-enough context (1 million instances with 8 categories) and fine-grained annotations (about 80 billion manually annotated pixels), which is far superior to all the existing UHR datasets including DeepGlobe, Inria Aerial, UDD, etc.. Moreover, we also propose WSDNet, a more efficient and effective framework for UHR segmentation especially with ultra-rich context. Specifically, multi-level Discrete Wavelet Transform (DWT) is naturally integrated to release computation burden while preserve more spatial details, along with a Wavelet Smooth Loss (WSL) to reconstruct original structured context and texture with a smooth constrain. Experiments on several UHR datasets demonstrate its state-of-the-art performance. The dataset is available at https://github.com/jankyee/URUR.
Deyi Ji, Feng Zhao 0004, Hongtao Lu 0001, Mingyuan Tao, Jieping Ye
CVPR2
2023 Towards Domain Generalization for Multi-view 3D Object Detection in Bird-Eye-View
abstract
Multi-view 3D object detection (MV3D-Det) in Bird-Eye-View (BEV) has drawn extensive attention due to its low cost and high efficiency. Although new algorithms for camera-only 3D object detection have been continuously proposed, most of them may risk drastic performance degradation when the domain of input images differs from that of training. In this paper, we first analyze the causes of the domain gap for the MV3D-Det task. Based on the covariate shift assumption, we find that the gap mainly attributes to the feature distribution of BEV, which is determined by the quality of both depth estimation and 2D image's feature representation. To acquire a robust depth prediction, we propose to decouple the depth estimation from the intrinsic parameters of the camera (i.e. the focal length) through converting the prediction of metric depth to that of scale-invariant depth and perform dynamic perspective augmentation to increase the diversity of the extrinsic parameters (i.e. the camera poses) by utilizing homography. Moreover, we modify the focal length values to create multiple pseudo-domains and construct an adversarial training loss to encourage the feature representation to be more domain-agnostic. Without bells and whistles, our approach, namely DG-BEV, successfully alleviates the performance drop on the unseen target domain without impairing the accuracy of the source domain. Extensive experiments on Waymo, nuScenes, and Lyft, demonstrate the generalization and effectiveness of our approach.
Xinhai Zhao, Dameng Yu, Feng Zhao 0004
CVPR8
2023 Visual Recognition-Driven Image Restoration for Multiple Degradation with Intrinsic Semantics Recovery
abstract
Deep image recognition models suffer a significant performance drop when applied to low-quality images since they are trained on high-quality images. Although many studies have investigated to solve the issue through image restoration or domain adaptation, the former focuses on visual quality rather than recognition quality, while the latter requires semantic annotations for task-specific training. In this paper, to address more practical scenarios, we propose a Visual Recognition-Driven Image Restoration network for multiple degradation, dubbed VRD-IR, to recover high-quality images from various unknown corruption types from the perspective of visual recognition within one model. Concretely, we harmonize the semantic representations of diverse degraded images into a unified space in a dynamic manner, and then optimize them towards intrinsic semantics recovery. Moreover, a prior-ascribing optimization strategy is introduced to encourage VRD-IR to couple with various downstream recognition tasks better. Our VRD-IR is corruption- and recognition-agnostic, and can be inserted into various recognition tasks directly as an image enhancement module. Extensive experiments on multiple image distortions demonstrate that our VRD-IR surpasses existing image restoration methods and show superior performance on diverse high-level tasks, including classification, detection, and person re-identification.
Zizheng Yang, Jie Huang 0017, Man Zhou 0003, Hu Yu 0001, Feng Zhao 0004
CVPR7
2023 Ingredient-oriented Multi-Degradation Learning for Image Restoration
abstract
Learning to leverage the relationship among diverse image restoration tasks is quite beneficial for unraveling the intrinsicingredients behind the degradation. Recent years have witnessed the flourish of various All-in-one methods, which handle multiple image degradations within a single model. In practice, however, few attempts have been made to excavate task correlations in that exploring the underlying fundamentalingredients of various image degradations, resulting in poor scalability as more tasks are involved. In this paper, we propose a novel perspective to delve into the degradation via aningredients-oriented rather than previous task-oriented manner for scalable learning. Specifically, our method, named Ingredients-oriented Degradation Reformulation framework (IDR), consists of two stages, namely task-oriented knowledge collection and ingredients-oriented knowledge integration. In the first stage, we conduct ad hoc operations on different degradations according to the underlying physics principles, and establish the corresponding prior hubs for each type of degradation. While the second stage progressively reformulates the preceding task-oriented hubs into single ingredients-oriented hub via learnable Principal Component Analysis (PCA), and employs a dynamic routing mechanism for probabilistic unknown degradation removal. Extensive experiments on various image restoration tasks demonstrate the effectiveness and scalability of our method. More importantly, our IDR exhibits the favorable generalization ability to unknown downstream tasks.
Jie Huang 0017, Mingde Yao, Zizheng Yang, Hu Yu 0001, Man Zhou 0003, Feng Zhao 0004
CVPR7
2023 DETRDistill: A Universal Knowledge Distillation Framework for DETR-families
abstract
Transformer-based detectors (DETRs) are becoming popular for their simple framework, but the large model size and heavy time consumption hinder their deployment in the real world. While knowledge distillation (KD) can be an appealing technique to compress giant detectors into small ones for comparable detection performance and low inference cost. Since DETRs formulate object detection as a set prediction problem, existing KD methods designed for classic convolution-based detectors may not be directly applicable. In this paper, we propose DETRDistill, a novel knowledge distillation method dedicated to DETR-families. Specifically, we first design a Hungarian-matching logits distillation to encourage the student model to have the exact predictions as those of the teacher DETRs. Then, we propose a target-aware feature distillation to help the student model learn from the object-centric features of the teacher model. Finally, in order to improve the convergence rate of the student DETR, we introduce a query-prior assignment distillation to speed up the student model learning from well-trained queries and stable assignment of the teacher model. Extensive experimental results on the COCO dataset validate the effectiveness of our approach. Notably, DETRDistill consistently improves various DETRs by more than 2.0 mAP, even surpassing their teacher models.
Chenhongyi Yang, Feng Zhao 0004
ICCV6
2023 Learning with Noisy Data for Semi-Supervised 3D Object Detection
abstract
Pseudo-Labeling (PL) is a critical approach in semisupervised 3D object detection (SSOD). In PL, delicately selected pseudo-labels, generated by the teacher model, are provided for the student model to supervise the semisupervised detection framework. However, such a paradigm may introduce misclassified labels or loose localized box predictions, resulting in a sub-optimal solution of detection performance. In this paper, we take PL from a noisy learning perspective: instead of directly applying vanilla pseudo-labels, we design a noise-resistant instance supervision module for better generalization. Specifically, we soften the classification targets by considering both the quality of pseudo labels and the network learning ability, and convert the regression task into a probabilistic modeling problem. Besides, considering that self-supervised learning works in the absence of labels, we incorporate dense pixel-wise feature consistency constraints to eliminate the negative impact of noisy labels. To this end, we propose NoiseDet, a simple yet effective framework for semi-supervised 3D object detection. Extensive experiments on competitive ONCE and Waymo benchmarks demonstrate that our method outperforms current semisupervised approaches by a large margin. Notably, our NoiseDet achieves state-of-the-art performance under various dataset scales on ONCE dataset. For example, NoiseDet improves its NoiseyStudent baseline from 55.5 mAP to 58.0 mAP, and further reaches 60.2 mAP with enhanced pseudo-label generation. Code will be available at https://github.com/zehuichen123/NoiseDet.
Zhenyu Li 0007, Dengpan Fu, Feng Zhao 0004
ICCV5
2023 FrozenRecon: Pose-free 3D Scene Reconstruction with Frozen Depth Models
abstract
3D scene reconstruction is a long-standing vision task. Existing approaches can be categorized into geometry-based and learning-based methods. The former leverages multi-view geometry but may face catastrophic failures due to the reliance on accurate pixel correspondence across views, while the latter mitigates these issues by learning 2D or 3D representation directly. However, without a largescale video or 3D training data, it can hardly be generalized to diverse real-world scenarios due to the presence of tens of millions or even billions of optimization parameters in the deep network.Recently, robust monocular depth estimation models trained with large-scale datasets have been proven to possess weak 3D geometry prior, but they are insufficient for reconstruction due to the unknown camera parameters, the affine-invariant property, and inter-frame inconsistency. To address these issues, we propose a novel test-time optimization approach that can transfer the robustness of affine- invariant depth models such as LeReS to challenging diverse scenes while ensuring inter-frame consistency, with only dozens of parameters to optimize per video frame. Specifically, our approach involves freezing the pre-trained affine-invariant depth model’s depth predictions, rectifying them by optimizing the unknown scale-shift values with a geometric consistency alignment module, and employing the resulting scale-consistent depth maps to robustly obtain camera poses and achieve dense scene reconstruction, even in low-texture regions. Experiments show that our method achieves state-of-the-art cross-dataset reconstruction on five zero-shot testing datasets. Code is available at: https://aim-uofa.github.io/FrozenRecon/
Guangkai Xu, Wei Yin 0006, Hao Chen 0041, Chunhua Shen, Feng Zhao 0004
ICCV6
2023 Empowering Low-Light Image Enhancer through Customized Learnable Priors
abstract
Deep neural networks have achieved remarkable progress in enhancing low-light images by improving their brightness and eliminating noise. However, most existing methods construct end-to-end mapping networks heuristically, neglecting the intrinsic prior of image enhancement task and lacking transparency and interpretability. Although some unfolding solutions have been proposed to relieve these issues, they rely on proximal operator networks that deliver ambiguous and implicit priors. In this work, we propose a paradigm for low-light image enhancement that explores the potential of customized learnable priors to improve the transparency of the deep unfolding paradigm. Motivated by the powerful feature representation capability of Masked Autoencoder (MAE), we customize MAE-based illumination and noise priors and redevelop them from two perspectives: 1) structure flow: we train the MAE from a normal-light image to its illumination properties and then embed it into the proximal operator design of the unfolding architecture; and 2) optimization flow: we train MAE from a normal-light image to its gradient representation and then employ it as a regularization term to constrain noise in the model output. These designs improve the interpretability and representation capability of the model. Extensive experiments on multiple low-light image enhancement datasets demonstrate the superiority of our proposed paradigm over state-of-the-art methods. Code is available at https://github.com/zheng980629/CUE.
Naishan Zheng, Man Zhou 0003, Yanmeng Dong, Xiangyu Rui, Jie Huang 0017, Chongyi Li, Feng Zhao 0004
ICCV7
2023 Exploring Temporal Frequency Spectrum in Deep Video Deblurring
abstract
Video deblurring aims to restore the latent video frames from their blurred counterparts. Despite the remarkable progress, most promising video deblurring methods only investigate the temporal priors in the spatial domain and rarely explore their its potential in the frequency domain. In this paper, we revisit the blurred sequence in the Fourier space and figure out some intrinsic frequency-temporal priors that imply the temporal blur degradation can be accessibly decoupled in the potential frequency domain. Based on these priors, we propose a novel Fourier-based frequency-temporal video deblurring solution, where the core design accommodates the temporal spectrum to a popular video deblurring pipeline of feature extraction, alignment, aggregation, and optimization. Specifically, we design a Spectrum Prior-guided Alignment module by leveraging enlarged blur information in the potential spectrum to mitigate the blur effects on the alignment. Then, Temporal Energy prior-driven Aggregation is implemented to replenish the original local features by estimating the temporal spectrum energy as the global sharpness guidance. In addition, the customized frequency loss is devised to optimize the proposed method for decent spectral distribution. Extensive experiments demonstrate that our model performs favorably against other state-of-the-art methods, thus confirming the effectiveness of frequency-temporal prior modeling.
Qi Zhu 0010, Man Zhou 0003, Naishan Zheng, Chongyi Li, Jie Huang 0017, Feng Zhao 0004
ICCV6
2023 AFNet-M: Adaptive Fusion Network with Masks for 2D+3D Facial Expression Recognition
abstract
2D+3D facial expression recognition (FER) can effectively cope with illumination and pose changes by merging texture and robust depth information. Most deep learning-based approaches employ the simple fusion strategy that concatenates the multimodal features directly after fully-connected layers, without considering the different degrees of significance for each modality. Meanwhile, how to focus more on both 2D and 3D local features is still a great challenge. In this paper, we propose the adaptive fusion network with masks (AFNet-M) for 2D+3D FER. To enhance 2D and 3D local features, we take the masks annotating salient regions of the face as prior knowledge and design the mask attention module (MA) which can automatically learn two modulation vectors to scale the feature maps. We also introduce an adaptive fusion module (AF) at convolutional layers through the computed importance weights. Experimental results demonstrate that our AFNet-M achieves the state-of-the-art performance on BU-3DFE and Bosphorus datasets and requires fewer parameters in comparison with other models.
Mingzhe Sui, Hanting Li, Zhaoqing Zhu, Feng Zhao 0004
ICIP4
2023 BEVDistill: Cross-Modal BEV Distillation for Multi-View 3D Object Detection
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004
ICLR6
2023 SAFE: Simultaneous Alignment of Features and Predictions for Dense Object Detectors
abstract
Dense detectors have gained increasing attention due to their simplicity and efficiency. However, there are two inherent misalignments lying in the one-stage architecture that greatly inhibit the performance. One is the misalignment between features and predicted boxes, the other is the spatial prediction misalignment between classification and regression branches. In this work, we propose a Simultaneous Alignment of Features and prEdictions (SAFE) for dense object detectors to mitigate both issues. Regarding of the misalignment between features and corresponding boxes, we design an Adaptive Feature Alignment module that leverages the location information to align features in the classification branch. As for the inconsistency of spatial prediction, we introduce a Weighted Average of Predictions module to refine the predicted sub-optimal coordinates by autonomously selecting higher-quality predicted ones in the surrounding area. Extensive experiments validate the persistent improvement of our approach. Our best model SAFE-S achieves state-of-the-art performance (56.5 mAP) on MS-COCO.
Xuesong Guo, Feng Zhao 0004
ICME5
2023 Guided Patch-Grouping Wavelet Transformer with Spatial Congruence for Ultra-High Resolution Segmentation
abstract
Most existing ultra-high resolution (UHR) segmentation methods always struggle in the dilemma of balancing memory cost and local characterization accuracy, which are both taken into account in our proposed Guided Patch-Grouping Wavelet Transformer (GPWFormer) that achieves impressive performances. In this work, GPWFormer is a Transformer (T)-CNN (C) mutual leaning framework, where T takes the whole UHR image as input and harvests both local details and fine-grained long-range contextual dependencies, while C takes downsampled image as input for learning the category-wise deep context. For the sake of high inference speed and low computation complexity, T partitions the original UHR image into patches and groups them dynamically, then learns the low-level local details with the lightweight multi-head Wavelet Transformer (WFormer) network. Meanwhile, the fine-grained long-range contextual dependencies are also captured during this process, since patches that are far away in the spatial domain can also be assigned to the same group. In addition, masks produced by C are utilized to guide the patch grouping process, providing a heuristics decision. Moreover, the congruence constraints between the two branches are also exploited to maintain the spatial consistency among the patches. Overall, we stack the multi-stage process in a pyramid way. Experiments show that GPWFormer outperforms the existing methods with significant improvements on five benchmark datasets.
Deyi Ji, Feng Zhao 0004, Hongtao Lu 0001
IJCAI2
2023 Learning Non-Uniform-Sampling for Ultra-High-Definition Image Enhancement
abstract
Ultra-high-definition (UHD) image enhancement is a challenging problem that aims to effectively and efficiently recover clean UHD images. To maintain efficiency, the straightforward approach is to downsample and perform most computations on low-resolution images. However, previous studies typically rely on the uniform and content-agnostic downsampling method that equally treats various regions regardless of their complexities, thus limiting the detail reconstruction in UHD image enhancement. To alleviate this issue, we propose a novel spatial-variant and invertible non-uniform downsampler that adaptively adjusts the sampling rate according to the richness of details. It magnifies important regions to preserve more information (e.g., sparse sampling points for sky, dense sampling points for buildings). Therefore, we propose a novel Non-uniform-Sampling Enhancement Network (NSEN) consisting of two core designs: 1) content-guided downsampling that extracts texture representation to guide the sampler to perform content-aware downsampling for producing detail-preserved low-resolution images; 2) invertible pixel-alignment which remaps the forward sampling process in an iterative manner to eliminate the deformations caused by the non-uniform downsampling, thus producing detail-rich clean UHD images. To demonstrate the superiority of our proposed model, we conduct extensive experiments on various UHD enhancement tasks. The results show that the proposed NSEN yields better performance against other state-of-the-art methods both visually and quantitatively.
Qi Zhu 0010, Naishan Zheng, Jie Huang 0017, Man Zhou 0003, Feng Zhao 0004
ACM Multimedia6
2023 Transition-constant Normalization for Image Enhancement
abstract
Normalization techniques that capture image style by statistical representation have become a popular component in deep neural networks. Although image enhancement can be considered as a form of style transformation, there has been little exploration of how normalization affect the enhancement performance. To fully leverage the potential of normalization, we present a novel Transition-Constant Normalization (TCN) for various image enhancement tasks. Specifically, it consists of two streams of normalization operations arranged under an invertible constraint, along with a feature sub-sampling operation that satisfies the normalization constraint. TCN enjoys several merits, including being parameter-free, plug-and-play, and incurring no additional computational costs. We provide various formats to utilize TCN for image enhancement, including seamless integration with enhancement networks, incorporation into encoder-decoder architectures for downsampling, and implementation of efficient architectures. Through extensive experiments on multiple image enhancement tasks, like low-light enhancement, exposure correction, SDR2HDR translation, and image dehazing, our TCN consistently demonstrates performance improvements. Besides, it showcases extensive ability in other tasks including pan-sharpening and medical segmentation. The code is available at \textit{\textcolor{blue}{https://github.com/huangkevinj/TCNorm}}.
Jie Huang 0017, Man Zhou 0003, Mingde Yao, Chongyi Li, Zhiwei Xiong, Feng Zhao 0004
NeurIPS8
2023 Deep Fractional Fourier Transform
abstract
Existing deep learning-based computer vision methods usually operate in the spatial and frequency domains, which are two orthogonal \textbf{individual} perspectives for image processing. In this paper, we introduce a new spatial-frequency analysis tool, Fractional Fourier Transform (FRFT), to provide comprehensive \textbf{unified} spatial-frequency perspectives. The FRFT is a unified continuous spatial-frequency transform that simultaneously reflects an image's spatial and frequency representations, making it optimal for processing non-stationary image signals. We explore the properties of the FRFT for image processing and present a fast implementation of the 2D FRFT, which facilitates its widespread use. Based on these explorations, we introduce a simple yet effective operator, Multi-order FRactional Fourier Convolution (MFRFC), which exhibits the remarkable merits of processing images from more perspectives in the spatial-frequency plane. Our proposed MFRFC is a general and basic operator that can be easily integrated into various tasks for performance improvement. We experimentally evaluate the MFRFC on various computer vision tasks, including object detection, image classification, guided super-resolution, denoising, dehazing, deraining, and low-light enhancement. Our proposed MFRFC consistently outperforms baseline methods by significant margins across all tasks.
Hu Yu 0001, Jie Huang 0017, Lingzhi Li 0002, Man Zhou 0003, Feng Zhao 0004
NeurIPS5
2023 FouriDown: Factoring Down-Sampling into Shuffling and Superposing
abstract
Spatial down-sampling techniques, such as strided convolution, Gaussian, and Nearest down-sampling, are essential in deep neural networks. In this study, we revisit the working mechanism of the spatial down-sampling family and analyze the biased effects caused by the static weighting strategy employed in previous approaches. To overcome this limitation, we propose a novel down-sampling paradigm in the Fourier domain, abbreviated as FouriDown, which unifies existing down-sampling techniques. Drawing inspiration from the signal sampling theorem, we parameterize the non-parameter static weighting down-sampling operator as a learnable and context-adaptive operator within a unified Fourier function. Specifically, we organize the corresponding frequency positions of the 2D plane in a physically-closed manner within a single channel dimension. We then perform point-wise channel shuffling based on an indicator that determines whether a channel's signal frequency bin is susceptible to aliasing, ensuring the consistency of the weighting parameter learning. FouriDown, as a generic operator, comprises four key components: 2D discrete Fourier transform, context shuffling rules, Fourier weighting-adaptively superposing rules, and 2D inverse Fourier transform. These components can be easily integrated into existing image restoration networks. To demonstrate the efficacy of FouriDown, we conduct extensive experiments on image de-blurring and low-light image enhancement. The results consistently show that FouriDown can provide significant performance improvements. We will make the code publicly available to facilitate further exploration and application of FouriDown.
Qi Zhu 0010, Man Zhou 0003, Jie Huang 0017, Naishan Zheng, Hongzhi Gao, Chongyi Li, Feng Zhao 0004
NeurIPS8
2023 Counterfactual attention alignment for visible-infrared cross-modality person re-identification
Zongzhe Sun, Feng Zhao 0004
Pattern Recognit. Lett.2
2023 Whole-Body Control of an Autonomous Mobile Manipulator Using Model Predictive Control and Adaptive Fuzzy Technique
abstract
Whole-body control (WBC) has emerged as an important framework in manipulation for mobile manipulators. However, most existing WBC frameworks require known dynamics. Considering whole-body manipulation and optimization with unknown dynamics, this article presents the WBC of a nonholonomic mobile manipulator using model predictive control (MPC) and fuzzy logic system. First, by constructing a dynamics-based feedback linearized robotic multi-input-multi-output (MIMO) system, an MPC-based WBC strategy is proposed for mobile manipulator. Such a strategy can provide the optimal control inputs with the specified optimization index and constraints. Thereafter, a primal-dual neural network effectively addresses the constrained quadratic programming (QP) problem over a finite receding horizon brought by the MPC. Then, in order to convert the intermediate control signals into the optimal control torques that can be executed by actuators, an adaptive FLS is employed to approximate the unknown dynamics. The novel elements of the current design control approach refer to the dynamics-based feedback linearized robotic MIMO system and the combination of an MPC module with an adaptive fuzzy controller. Finally, the trajectory tracking experiments performed on a mobile dual-arm robot demonstrate the effectiveness of the proposed method.
Wang Yuan, Yong-Hua Liu, Chun-Yi Su, Feng Zhao 0004
IEEE Trans. Fuzzy Syst.4
2023 Deep Adaptive Pansharpening via Uncertainty-Aware Image Fusion
abstract
Pansharpening is a procedure that fuses high-resolution panchromatic (PAN) images and low-resolution multispectral (LMS) images to derive high-resolution multispectral (HMS) images. Despite its rapid development, most existing pansharpening techniques integrate the information of PAN and LMS invariantly in the spatial dimension, ignoring the uneven spatial dependence of restoring HMS with the aid of PAN information and resulting in ineffective fusion results. In this work, we propose an Uncertainty-aware Adaptive Pansharpening Network (UAPN) that integrates PAN information spatial-variantly to restore LMS information with an uncertainty mechanism. Specifically, we first estimate the epistemic and aleatoric uncertainties together, which model the spatial-variant distributions of restoring the LMS image to the HMS image. Then, we introduce Uncertainty-conditioned Adaptive Convolution (UAC) to adaptively integrate LMS and PAN information, where its parameters are spatially variable by conditioning on the uncertainty estimations. Furthermore, we propose a multi-stage uncertainty-driven loss function to explicitly force the network to concentrate on restoring challenging areas of the LMS image. Extensive experimental results demonstrate the superiority of our UAPN with fewer parameters and flops, outperforming other state-of-the-art methods both qualitatively and quantitatively on multiple satellite datasets. The code is available at https://github.com/keviner1/UAPN..
Jie Huang 0017, Man Zhou 0003, Danfeng Hong, Feng Zhao 0004
IEEE Trans. Geosci. Remote. Sens.5
2023 Modality-Aware Feature Integration for Pan-Sharpening
abstract
Pan-sharpening aims to super-solve low-spatial resolution multiple spectral (MS) images with the guidance of high-resolution (HR) texture-rich panchromatic (PAN) images. Recently, deep-learning-based pan-sharpening approaches have dominated this field and achieved remarkable advancement. However, most promising algorithms are devised in one-way mapping and have not fully explored the mutual dependencies between PAN and MS modalities, thus impacting the model performance. To address this issue, we propose a novel information compensation and integration network for pan-sharpening by effective cross-modality joint learning in this work. First, the cross-central difference convolution is employed to explicitly extract the texture details of the PAN images. Second, we implement the compensation process by imitating the classical back-projection (BP) technique where the extracted PAN textures are employed to guide the intrinsic information learning of MS images iteratively. Subsequently, we devise the hierarchical transformer to integrate the comprehensive relations of stage-iteration information from spatial and temporal contexts. Extensive experiments over multiple satellite datasets demonstrate the superiority of our method to the existing state-of-the-art methods. The source code is available athttps://github.com/manman1995/pansharpening.
Man Zhou 0003, Jie Huang 0017, Feng Zhao 0004, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.3
2023 Low-Light Stereo Image Enhancement
abstract
Stereo cameras are now commonly used in more and more devices. Nevertheless, visually unpleasant images captured under low-light conditions hinder their practical application. As an initial attempt at low-light stereo image enhancement, we propose a novel Dual-View Enhancement Network (DVENet) based on the Retinex theory, which consists of two stages. The first stage estimates an illumination map to obtain a coarse enhancement result, which boosts the correlation of two views, while the second stage recovers details by integrating the information from two views to achieve fine image quality improvement with the guidance of the illumination map. To fully utilize the dual-view correlation, we further design a wavelet-based view transfer module to efficiently carry out multi-scale detail recovery. Then, we design an illumination-aware attention fusion module to exploit the complementarity between the fused features from two views and the single-view features. Experiments on both synthetic and real-world stereo datasets demonstrate the superiority of our proposed method over existing solutions. The code and model are publicly available at:https://github.com/KevinJ-Huang/Stereo-Low-Light.
Jie Huang 0017, Xueyang Fu, Zeyu Xiao 0002, Feng Zhao 0004, Zhiwei Xiong
IEEE Trans. Multim.4
2023 Unsupervised Underexposed Image Enhancement via Self-Illuminated and Perceptual Guidance
abstract
Underexposed images inevitably suffer severe degradation due to light distortion and noise corruption. Motivated by the limited samples of paired datasets, several unsupervised enhancement methods have been developed. However, these techniques heavily rely on pre-defined fixed lightness and noise removal constraints. Correspondingly, they cannot match the image-specific lightness when performing enhancement and can only refine details in a non-perceptual way. In this paper, we propose an Unsupervised Underexposed Image Enhancement Network (U2IENet) with self-illuminated and perceptual guidance. Specifically, to adjust the illumination for matching the image-specific lightness adaptively, we utilize the bright area of the underexposed image as the self-illuminated guidance to constrain the training process and modulate the features. Meanwhile, we introduce the perceptual guidance as a constraint to remove the noise based on illumination distribution, thus refining the details perceptually. Experiments on both underexposed datasets and public low-light datasets demonstrate the superiority of the proposed approach with higher flexibility over state-of- the-art solutions. In addition, our U2IENet also provides a side function that enables users to adjust the lightness via interactive tuning of a single parameter.
Naishan Zheng, Jie Huang 0017, Feng Zhao 0004, Xueyang Fu, Feng Wu 0001
IEEE Trans. Multim.3
2022 DiffLoss: Unleashing Diffusion Model as Constraint for Training Image Restoration Network
Jiangtong Tan, Hu Yu 0001, Jie Huang 0017, Zizheng Yang, Feng Zhao 0004
ACCV (4)5
2022 Exposure Normalization and Compensation for Multiple-Exposure Correction
abstract
Images captured with improper exposures usually bring unsatisfactory visual effects. Previous works mainly focus on either underexposure or overexposure correction, resulting in poor generalization to various exposures. An alternative solution is to mix the multiple exposure data for training a single network. However, the procedures of correcting underexposure and overexposure to normal exposures are much different from each other, leading to large discrepancies for the network in correcting multiple-exposures, thus resulting in poor performance. The key point to address this issue lies in bridging different exposure representations. To achieve this goal, we design a multiple exposure correction framework based on an Exposure Normalization and Compensation (ENC) module. Specifically, the ENC module consists of an exposure normalization part for mapping different exposure features to the exposure-invariant feature space, and a compensation part for integrating the initial features unprocessed by the exposure normalization part to ensure the completeness of information. Besides, to further alleviate the imbalanced performance caused by variations in the optimization process, we introduce a parameter regularization fine-tuning strategy to improve the performance of the worst-performed exposure without degrading other exposures. Our model empowered by ENC outperforms the existing methods by more than 2dB and is robust to multiple image enhancement tasks, demonstrating its effectiveness and generalization capability for real-world applications. Code: https://github.com/KevinJ-Huang/ExposureNorm-Compensation.
Jie Huang 0017, Xueyang Fu, Man Zhou 0003, Yang Wang 0015, Feng Zhao 0004, Zhiwei Xiong
CVPR6
2022 Unleashing Potential of Unsupervised Pre-Training with Intra-Identity Regularization for Person Re-Identification
abstract
Existing person reidentification (ReID) methods typically load the pretrained ImageNet weights for initialization directly. However, as a fine-grained classification task, ReID is more challenging and there exists a large domain gap between ImageNet classification. Inspired by the great success of self-supervised representation learning with contrastive objectives, in this paper, we design an Unsupervised Pretraining framework for reidentification (UP-ReID) based on the contrastive learning (CL) pipeline. During the pre-training, we attempt to address two critical issues for learning fine-grained ReID features: (1) the augmentations in the CL pipeline usually distort the discriminative clues in person images, and (2) the fine-grained local features of person images are not fully-explored. Therefore, we introduce an intra-identity (12-) regularization in the UP-ReID, which is instantiated as two constraints coming from the global image and local patch aspects, respectively. A global consistency constraint is enforced between augmented and original person images to increase robustness to augmentation, while an intrinsic contrastive constraint among local patches of each image is employed to fully explore the local discriminative clues. Extensive experiments on multiple popular reid datasets, PersonX, Market1501, CUHK03, and MSMT17, demonstrate that our UP-ReID pretrained model can significantly benefit the downstream ReID fine-tuning and achieve state-of-the-art performance.
Zizheng Yang, Xin Jin 0014, Kecheng Zheng, Feng Zhao 0004
CVPR4
2022 Mutual Information-driven Pan-sharpening
abstract
Pan-sharpening aims to integrate the complementary information of texture-rich PAN images and multi-spectral (MS) images to produce the texture-rich MS images. Despite the remarkable progress, existing state-of-the-art Pansharpening methods don't explicitly enforce the complementary information learning between two modalities of PAN and MS images. This leads to information redundancy not being handled well, which further limits the performance of these methods. To address the above issue, we propose a novel mutual information-driven Pan-sharpening framework in this paper. To be specific, we first project the PAN and MS image into modality-aware feature space independently, and then impose the mutual information minimization over them to explicitly encourage the complementary information learning. Such operation is capable of reducing the information redundancy and improving the model performance. Extensive experimental results over multiple satellite datasets demonstrate that the proposed algorithm outperforms other state-of-the-art methods qualitatively and quantitatively with great generalization ability to real-world scenes.
Man Zhou 0003, Jie Huang 0017, Zihe Yang, Xueyang Fu, Feng Zhao 0004
CVPR6
2022 Bijective Mapping Network for Shadow Removal
abstract
Shadow removal, which aims to restore the background in the shadow regions, is challenging due to its highly ill-posed nature. Most existing deep learning-based methods individually remove the shadow by only considering the content of the matched paired images, barely taking into account the auxiliary supervision of shadow generation in the shadow removal procedure. In this work, we argue that shadow removal and generation are interrelated and could provide useful informative supervision for each other. Specifically, we propose a new Bijective Mapping Network (BMNet), which couples the learning procedures of shadow removal and shadow generation in a unified parameter-shared framework. With consistent two way constraints and synchronous optimization of the two procedures, BMNet could effectively recover the underlying background contents during the forward shadow removal procedure. In addition, through statistical analysis of real world datasets, we observe and verify that shadow appearances under different color spectrums are inconsistent. This motivates us to design a Shadow-Invariant Color Guidance Module (SICGM), which can explicitly utilize the learned shadow-invariant color information to guide network color restoration, thereby further reducing color-bias effects. Experiments on the representative ISTD, ISTD+ and SRD benchmarks show that our proposed network outperforms the state-of-the-art method [11] in de-shadowing performance, while only using its 0.25% network parameters and 6.25% floating point operations (FLOPs).
Yurui Zhu, Jie Huang 0017, Xueyang Fu, Feng Zhao 0004, Qibin Sun, Zhengjun Zha
CVPR4
2022 Deformable Feature Aggregation for Dynamic Multi-modal 3D Object Detection
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004
ECCV (8)6
2022 Deep Fourier-Based Exposure Correction Network with Spatial-Frequency Interaction
Jie Huang 0017, Feng Zhao 0004, Man Zhou 0003, Zhiwei Xiong
ECCV (19)3
2022 Frequency and Spatial Dual Guidance for Image Dehazing
Hu Yu 0001, Naishan Zheng, Man Zhou 0003, Jie Huang 0017, Zeyu Xiao 0002, Feng Zhao 0004
ECCV (19)6
2022 Spatial-Frequency Domain Information Integration for Pan-Sharpening
Man Zhou 0003, Jie Huang 0017, Hu Yu 0001, Xueyang Fu, Aiping Liu, Xian Wei, Feng Zhao 0004
ECCV (18)8
2022 RelCLIP: Adapting Language-Image Pretraining for Visual Relationship Detection via Relational Contrastive Learning
abstract
Conventional visual relationship detection models only use the numeric ids of relation labels for training, but ignore the semantic correlation between the labels, which leads to severe training biases and harms the generalization ability of representations.In this paper, we introduce compact language information of relation labels for regularizing the representation learning of visual relations.Specifically, we propose a simple yet effective visual Relationship prediction framework that transfers natural language knowledge learned from Contrastive Language-Image Pre-training (CLIP) models to enhance the relationship prediction, termed as RelCLIP.Benefiting from the powerful visual-semantic alignment ability of CLIP at image level, we introduce a novel Relational Contrastive Learning (RCL) approach that explores relation-level visual-semantic alignment via learning to match cross-modal relational embeddings.By collaboratively learning the semantic coherence and discrepancy from relation triplets, the model can generate more discriminative and robust representations.Experimental results on the Visual Genome dataset show that RelCLIP achieves significant improvements over strong baselines under full (providing accurate labels) and distant supervision (providing noise labels), demonstrating its powerful generalization ability in learning relationship representations.
Yi Zhu 0004, Zhaoqing Zhu, Bingqian Lin, Xiaodan Liang, Feng Zhao 0004, Jianzhuang Liu
EMNLP5
2022 CMANET: Curvature-Aware Soft Mask Guided Attention Fusion Network for 2D+3D Facial Expression Recognition
abstract
As 2D texture and 3D structural information can describe facial features complementarily, 2D+3D facial expression recognition (FER) has received widespread attention. Though recent methods for 2D+3D FER have reached excellent performance, they still face two challenges: the way for attending to critical face areas and the strategy for fusing multi-modal information. To address these issues, we propose a curvature-aware soft mask guided attention fusion network (CMANet), which mainly consists of two components: curvature-aware attention module and multi-modal attention fusion module. The former utilizes the curvature-aware soft mask guiding the homo-modal attention mechanism to focus on potentially important areas with soft weights, while the latter applies pixel-level fusion on multi-modal features to retain the significant information from different modalities and also allows multi-modal features to interact in a larger field of view. Extensive experimental results show that our CMANet achieves outstanding accuracies (90.24% on BU-3DFE and 89.36% on Bosphorus) and outperforms the state-of-the-art methods.
Zhaoqing Zhu, Mingzhe Sui, Hanting Li, Feng Zhao 0004
ICME4
2022 Dast-Net: Depth-Aware Spatio-Temporal Network for Video Deblurring
abstract
Video deblurring is a challenging task due to inevitable blurs caused by depth variation, object motion, and camera shake. Although several video deblurring methods resort to depth maps, they rarely produce visually appealing results since the information in the depth maps is used insufficiently. To address this issue, we propose a Depth-Aware Modulated (DAM) block for efficiently utilizing the depth map characteristics, in which the intensity and variation of depth are exploited according to the depth map value and edges. Based on the DAM block, we develop the Depth-Aware Spatio-Temporal Network (DAST-Net) tailored for video deblurring. Particularly, the Depth-Aware Temporal Alignment module uses the depth cues to guide the alignment of adjacent frames. The Depth-Modulated Spatial Fusion module then warps the aligned frames to maintain spatial invariance with the aligned features. The warped depth features are more effective in video deblurring, since they allow for the aggregation of multiple frames. Extensive quantitative and qualitative evaluations demonstrate that the proposed DAST-Net outperforms other state-of-the-art methods.
Qi Zhu 0010, Zeyu Xiao 0002, Jie Huang 0017, Feng Zhao 0004
ICME4
2022 AutoAlign: Pixel-Instance Feature Aggregation for Multi-Modal 3D Object Detection
abstract
Object detection through either RGB images or the LiDAR point clouds has been extensively explored in autonomous driving. However, it remains challenging to make these two data sources complementary and beneficial to each other. In this paper, we propose AutoAlign, an automatic feature fusion strategy for 3D object detection. Instead of establishing deterministic correspondence with camera projection matrix, we model the mapping relationship between the image and point clouds with a learnable alignment map. This map enables our model to automate the alignment of non-homogenous features in a dynamic and data-driven manner. Specifically, a cross-attention feature alignment module is devised to adaptively aggregate pixel-level image features for each voxel. To enhance the semantic consistency during feature alignment, we also design a self-supervised cross-modal feature interaction module, through which the model can learn feature aggregation with instance-level feature guidance. Extensive experimental results show that our approach can lead to 2.3 mAP and 7.0 mAP improvements on the KITTI and nuScenes datasets respectively. Notably, our best model reaches 70.9 NDS on the nuScenes testing leaderboard, achieving competitive performance among various state-of-the-arts.
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004, Bolei Zhou, Hang Zhao 0021
IJCAI6
2022 MMNet: Muscle Motion-Guided Network for Micro-Expression Recognition
abstract
Facial micro-expressions (MEs) are involuntary facial motions revealing people’s real feelings and play an important role in the early intervention of mental illness, the national security, and many human-computer interaction systems. However, existing micro-expression datasets are limited and usually pose some challenges for training good classifiers. To model the subtle facial muscle motions, we propose a robust micro-expression recognition (MER) framework, namely muscle motion-guided network (MMNet). Specifically, a continuous attention (CA) block is introduced to focus on modeling local subtle muscle motion patterns with little identity information, which is different from most previous methods that directly extract features from complete video frames with much identity information. Besides, we design a position calibration (PC) module based on the vision transformer. By adding the position embeddings of the face generated by the PC module at the end of the two branches, the PC module can help to add position information to facial muscle motion-pattern features for the MER. Extensive experiments on three public micro-expression datasets demonstrate that our approach outperforms state-of-the-art methods by a large margin. Code is available at https://github.com/muse1998/MMNet.
Hanting Li, Mingzhe Sui, Zhaoqing Zhu, Feng Zhao 0004
IJCAI4
2022 Graph-DETR3D: Rethinking Overlapping Regions for Multi-View 3D Object Detection
abstract
3D object detection from multiple image views is a fundamental and challenging task for visual scene understanding. However, accurately detecting objects through perspective views in the 3D space is extremely difficult due to the lack of depth information. Recently, DETR3D introduces a novel 3D-2D query paradigm in aggregating multi-view images for 3D object detection and achieves state-of-the-art performance. In this paper, with intensive pilot experiments, we quantify the objects located at different regions and find that the "truncated instances'' (i.e., at the border regions of each image) are the main bottleneck hindering the performance of DETR3D. Although it merges multiple features from two adjacent views in the overlapping regions, DETR3D still suffers from insufficient feature aggregation, thus missing the chance to fully boost the detection performance. In an effort to tackle the problem, we propose Graph-DETR3D to automatically aggregate multi-view imagery information through graph structure learning. It constructs a dynamic 3D graph between each object query and 2D feature maps to enhance the object representations, especially at the border regions. Besides, Graph-DETR3D benefits from a novel depth-invariant multi-scale training strategy, which maintains the visual depth consistency by simultaneously scaling the image size and the object depth. Extensive experiments on the nuScenes dataset demonstrate the effectiveness and efficiency of our Graph-DETR3D. Notably, our best model achieves 49.5 NDS on the nuScenes test leaderboard, achieving new state-of-the-art in comparison with various published image-view 3D object detectors.
Zhenyu Li 0007, Shiquan Zhang, Liangji Fang, Qinhong Jiang, Feng Zhao 0004
ACM Multimedia6
2022 Exposure-Consistency Representation Learning for Exposure Correction
abstract
Images captured under improper exposures including underexposure and overexposure often suffer from unsatisfactory visual effects. Since their correction procedures are quite different, it is challenging for a single network to correct various exposures. The key to addressing this issue is consistently learning underexposure and overexposure corrections. To achieve this goal, we propose an Exposure-Consistency Processing (ECP) module to consistently learn the representation of both underexposure and overexposure in the feature space. Specifically, the ECP module employs the bilateral activation mechanism that derives both underexposure and overexposure property features for exposure-consistency representation modeling, which is followed by two shared-weight branches to process these features. Based on the ECP module, we build the whole network by utilizing it as the basic unit. Additionally, to further assist the exposure-consistency learning, we develop an Exposure-Consistency Constraining (ECC) strategy that augments the various local region exposures and then constrains the feature representation change between the exposure augmented image and the original one. Our proposed network is lightweight and outperforms existing methods remarkably, while the ECP module can also be extended to other baselines, demonstrating its superiority and scalability. code: https://github.com/KevinJ-Huang/ECLNet.
Jie Huang 0017, Man Zhou 0003, Mingde Yao, Feng Zhao 0004, Zhiwei Xiong
ACM Multimedia5
2022 Customizing GAN Using Few-shot Sketches
abstract
Generative adversarial networks (GANs) have demonstrated remarkable success in image synthesis applications, but their performance deteriorates under limited data regimes. The fundamental challenge is that it is extremely difficult to synthesize photo-realistic and highly diversified images while capturing meaningful attributes of the targets under minimum supervision. Previous methods either fine-tune or rewrite the model weights to adapt to few-shot datasets. However, this either overfits or requires access to large-scale data on which they are trained. To tackle the problem, we propose a framework that repurposes the existing pre-trained generative models using only a few samples (e.g., <30) of sketches. Unlike previous works, we transfer the sample diversity and quality without accessing the source data using inter-domain distance consistency. By employing cross-domain adversarial learning, we encourage the model output to closely resemble the input sketches in both shape and pose. Extensive experiments show that our method significantly outperforms the existing approaches in terms of sample quality and diversity. The qualitative and quantitative results on various standard datasets also demonstrate its efficacy. On the most popularly used dataset, Gabled church, we achieve a Fréchet inception distance (FID) score of 15.63.
Syed Muhammad Israr, Feng Zhao 0004
ACM Multimedia2
2022 SIR-Former: Stereo Image Restoration Using Transformer
abstract
Stereo image pairs record the scene from two different views and introduce cross-view information for image restoration. However, there are two challenges in utilizing the cross-view information for stereo image restoration: cross-view alignment and information fusion. Most existing methods adopt convolutional neural networks to align the views and fuse the information locally, which has difficulty in capturing the global correspondence across stereo images for view alignment and makes it hard to integrate the long-term information across views. In this paper, we propose to address the stereo image restoration with transformer by leveraging its powerful capability of modeling long-range context dependencies. Specifically, we construct a stereo image restoration transformer (SIR-Former) to effectively exploit the cross-view correlations. First, to explore the global correspondence for view alignment effectively, we devise a stereo alignment transformer (SAT) module across stereo images, enabling robust alignment under the epipolar constraint. Then, we design a stereo fusion transformer (SFT) module for aggregating the cross-view information in a small horizontal neighborhood, aiming to enhance important features for succeeding restoration. Extensive experiments show that SIR-Former can remarkably boost quantitative and qualitative quality on various image restoration tasks (e.g., super-resolution, deblurring, deraining, and low-light enhancement), which demonstrate the effectiveness of the proposed framework.
Zizheng Yang, Mingde Yao, Jie Huang 0017, Man Zhou 0003, Feng Zhao 0004
ACM Multimedia5
2022 Source-Free Domain Adaptation for Real-World Image Dehazing
abstract
Deep learning-based source dehazing methods trained on synthetic datasets have achieved remarkable performance but suffer from dramatic performance degradation on real hazy images due to domain shift. Although certain Domain Adaptation (DA) dehazing methods have been presented, they inevitably require access to the source dataset to reduce the gap between the source synthetic and target real domains. To address these issues, we present a novel Source-Free Unsupervised Domain Adaptation (SFUDA) image dehazing paradigm, in which only a well-trained source model and an unlabeled target real hazy dataset are available. Specifically, we devise the Domain Representation Normalization (DRN) module to make the representation of real hazy domain features match that of the synthetic domain to bridge the gaps. With our plug-and-play DRN module, unlabeled real hazy images can adapt existing well-trained source networks. Besides, the unsupervised losses are applied to guide the learning of the DRN module, which consists of frequency losses and physical prior losses. Frequency losses provide structure and style constraints, while the prior loss explores the inherent statistic property of haze-free images. Equipped with our DRN module and unsupervised loss, existing source dehazing models are able to dehaze unlabeled real hazy images. Extensive experiments on multiple baselines demonstrate the validity and superiority of our method visually and quantitatively.
Hu Yu 0001, Jie Huang 0017, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004
ACM Multimedia6
2022 Structure- and Texture-Aware Learning for Low-Light Image Enhancement
abstract
Structure and texture information is critically important for low-light image enhancement, in terms of stable global adjustment and fine details recovery. However, most existing methods tend to learn the structure and texture of low-light images in a coupled manner, without well considering the heterogeneity between them, which challenges the capability of the model to learn both adequately. In this paper, we tackle this problem in a divide and conquer strategy, based on the observation that the structure and texture representations are highly separated in the frequency spectrum. Specifically, we propose a Structure and Texture Aware Network (STAN) for low-light image enhancement, which consists of a structure sub-network and a texture sub-network. The former exploits the low-pass characteristic of the transformer to capture low-frequency-related structural representation. While the latter builds upon central difference convolution to capture high-frequency-related texture representation. We establish the Multi-Spectrum Interaction (MSI) module between two sub-networks to bidirectionally provide complementary information. In addition, to further elevate the capability of the model, we introduce a dual distillation scheme that assists the learning process of two sub-networks via counterparts' normal-light structure and texture representations. Comprehensive experiments show that the proposed STAN outperforms the state-of-the-art methods qualitatively and quantitatively.
Jie Huang 0017, Mingde Yao, Man Zhou 0003, Feng Zhao 0004
ACM Multimedia5
2022 Enhancement by Your Aesthetic: An Intelligible Unsupervised Personalized Enhancer for Low-Light Images
abstract
Low-light image enhancement is an inherently subjective process whose targets vary with the user's aesthetic. Motivated by this, several personalized enhancement methods have been investigated. However, the enhancement process based on user preferences in these techniques is invisible, i.e., a "black box". In this work, we propose an intelligible unsupervised personalized enhancer (iUP-Enhancer) for low-light images, which establishes the correlations between the low-light and the unpaired reference images with regard to three user-friendly attributions (brightness, chromaticity, and noise). The proposed iUP-Enhancer is trained with the guidance of these correlations and the corresponding unsupervised loss functions. Rather than a "black box" process, our iUP-Enhancer presents an intelligible enhancement process with the above attributions. Extensive experiments demonstrate that the proposed algorithm produces competitive qualitative and quantitative results while maintaining excellent flexibility and scalability. This can be validated by personalization with single/multiple references, cross-attribution references, or merely adjusting parameters.
Naishan Zheng, Jie Huang 0017, Qi Zhu 0010, Man Zhou 0003, Feng Zhao 0004, Zhengjun Zha
ACM Multimedia5
2022 Adaptively Learning Low-high Frequency Information Integration for Pan-sharpening
abstract
Pan-sharpening aims to generate high-spatial resolution multi-spectral (MS) image by fusing high-spatial resolution panchromatic (PAN) image and its corresponding low-spatial resolution MS image. Despite the remarkable progress, most existing pan-sharpening methods only work in the spatial domain and rarely explore the potential solutions in the frequency domain. In this paper, we propose a novel pan-sharpening framework by adaptively learning low-high frequency information integration in the spatial and frequency dual domains. It consists of three key designs: mask prediction sub-network, low-frequency learning sub-network and high-frequency learning sub-network. Specifically, the first is responsible for measuring the modality-aware frequency information difference of PAN and MS images and further predicting the low-high frequency boundary in the form of a two-dimensional mask. In view of the mask, the second adaptively picks out the corresponding low-frequency components of different modalities and then restores the expected low-frequency one by spatial and frequency dual domains information integration while the third combines the above refined low-frequency and the original high-frequency for the latent high-frequency reconstruction. In this way, the low-high frequency information is adaptively learned, thus leading to the pleasing results. Extensive experiments validate the effectiveness of the proposed network and demonstrate the favorable performance against other state-of-the-art methods. The source code will be released at https://github.com/manman1995/pansharpening.
Man Zhou 0003, Jie Huang 0017, Chongyi Li, Hu Yu 0001, Naishan Zheng, Feng Zhao 0004
ACM Multimedia7
2022 Normalization-based Feature Selection and Restitution for Pan-sharpening
abstract
Pan-sharpening is essentially a panchromatic (PAN) image-guided low-spatial resolution MS image super-resolution problem. The commonly challenging issue of pan-sharpening is how to correctly select consistent features and propagate them, and properly handle inconsistent ones between PAN and MS modalities. To solve this issue, we propose a Normalization-based Feature Selection and Restitution mechanism, which is capable of filtering out the inconsistent features and promoting to learn the consistent ones. Specifically, we first modulate the PAN feature as the MS style in feature space by AdaIN operation \citeAdaIN. However, such operation inevitably removes the favorable features. We thus propose to distill the effective information from the removed part and restitute it back to the modulated part. To better distillation, we enforce a contrastive learning constraint to close the distance between the restituted feature and the ground truth, and push the removed part away from the ground truth. In this way, the consistent features of PAN images are correctly selected and the inconsistent ones are filtered out, thus relieving the over-transferred artifacts in the process of PAN-guided MS super-resolution. Extensive experiments validate the effectiveness of the proposed network and demonstrate its favorable performance against other state-of-the-art methods. The source code will be released at https://github.com/manman1995/pansharpening.
Man Zhou 0003, Jie Huang 0017, Aiping Liu, Chongyi Li, Feng Zhao 0004
ACM Multimedia7
2022 Roadblocks for Temporarily Disabling Shortcuts and Learning New Knowledge
abstract
Deep learning models have been found with a tendency of relying on shortcuts, i.e., decision rules that perform well on standard benchmarks but fail when transferred to more challenging testing conditions. Such reliance may hinder deep learning models from learning other task-related features and seriously affect their performance and robustness. Although recent studies have shown some characteristics of shortcuts, there are few investigations on how to help the deep learning models to solve shortcut problems. This paper proposes a framework to address this issue by setting up roadblocks on shortcuts. Specifically, roadblocks are placed when the model is urged to learn to complete a gently modified task to ensure that the learned knowledge, including shortcuts, is insufficient the complete the task. Therefore, the model trained on the modified task will no longer over-rely on shortcuts. Extensive experiments demonstrate that the proposed framework significantly improves the training of networks on both synthetic and real-world datasets in terms of both classification accuracy and feature diversity. Moreover, the visualization results show that the mechanism behind the proposed our method is consistent with our expectations. In summary, our approach can effectively disable the shortcuts and thus learn more robust features.
Hongjing Niu, Hanting Li, Feng Zhao 0004, Bin Li 0025
NeurIPS3
2022 Panchromatic and Multispectral Image Fusion via Alternating Reverse Filtering Network
abstract
Panchromatic (PAN) and multi-spectral (MS) image fusion, named Pan-sharpening, refers to super-resolve the low-resolution (LR) multi-spectral (MS) images in the spatial domain to generate the expected high-resolution (HR) MS images, conditioning on the corresponding high-resolution PAN images. In this paper, we present a simple yet effective alternating reverse filtering network for pan-sharpening. Inspired by the classical reverse filtering that reverses images to the status before filtering, we formulate pan-sharpening as an alternately iterative reverse filtering process, which fuses LR MS and HR MS in an interpretable manner. Different from existing model-driven methods that require well-designed priors and degradation assumptions, the reverse filtering process avoids the dependency on pre-defined exact priors. To guarantee the stability and convergence of the iterative process via contraction mapping on a metric space, we develop the learnable multi-scale Gaussian kernel module, instead of using specific filters. We demonstrate the theoretical feasibility of such formulations. Extensive experiments on diverse scenes to thoroughly verify the performance of our method, significantly outperforming the state of the arts.
Man Zhou 0003, Jie Huang 0017, Feng Zhao 0004, Chengjun Xie, Chongyi Li, Danfeng Hong
NeurIPS4
2022 Deep Fourier Up-Sampling
abstract
Existing convolutional neural networks widely adopt spatial down-/up-sampling for multi-scale modeling. However, spatial up-sampling operators (e.g., interpolation, transposed convolution, and un-pooling) heavily depend on local pixel attention, incapably exploring the global dependency. In contrast, the Fourier domain is in accordance with the nature of global modeling according to the spectral convolution theorem. Unlike the spatial domain that easily performs up-sampling with the property of local similarity, up-sampling in the Fourier domain is more challenging as it does not follow such a local property. In this study, we propose a theoretically feasible Deep Fourier Up-Sampling (FourierUp) to solve these issues. We revisit the relationships between spatial and Fourier domains and reveal the transform rules on the features of different resolutions in the Fourier domain, which provide key insights for FourierUp's designs. FourierUp as a generic operator consists of three key components: 2D discrete Fourier transform, Fourier dimension increase rules, and 2D inverse Fourier transform, which can be directly integrated with existing networks. Extensive experiments across multiple computer vision tasks, including object detection, image segmentation, image de-raining, image dehazing, and guided image super-resolution, demonstrate the consistent performance gains obtained by introducing our FourierUp. Code will be publicly available.
Man Zhou 0003, Hu Yu 0001, Jie Huang 0017, Feng Zhao 0004, Jinwei Gu, Chen Change Loy, Deyu Meng, Chongyi Li
NeurIPS4
2022 Effective Pan-Sharpening With Transformer and Invertible Neural Network
abstract
In remote sensing imaging systems, pan-sharpening is an important technique to obtain high-resolution multispectral images from a high-resolution panchromatic image and its corresponding low-resolution multispectral image. Due to the powerful learning capability of convolution neural networks (CNNs), CNN-based methods have dominated this field. However, due to the limitation of the convolution operator, long-range spatial features are often not accurately obtained, thus limiting the overall performance. To this end, we propose a novel and effective method by exploiting a customized transformer architecture and information-lossless invertible neural module for long-range dependencies modeling and effective feature fusion in this article. Specifically, the customized transformer formulates the panchromatic (PAN) and multispectral (MS) features as queries and keys to encourage joint feature learning across two modalities, while the designed invertible neural module enables effective feature fusion to generate the expected pan-sharpened results. To the best of our knowledge, this is the first attempt to introduce a transformer and a invertible neural network into the pan-sharpening field. Extensive experiments over different kinds of satellite datasets demonstrate that our method outperforms state-of-the-art algorithms both visually and quantitatively with fewer parameters and flops. Furthermore, the ablation experiments also prove the effectiveness of the proposed customized long-range transformer and effective invertible neural feature fusion module for pan-sharpening.
Man Zhou 0003, Xueyang Fu, Jie Huang 0017, Feng Zhao 0004, Aiping Liu, Rujing Wang
IEEE Trans. Geosci. Remote. Sens.4
2022 Effective Pan-Sharpening by Multiscale Invertible Neural Network and Heterogeneous Task Distilling
abstract
As recognized, the ground truth multi-spectral (MS) images possess the complementary information (e.g., high-frequency component) of low-resolution (LR) MS images, which can be considered as privileged information to alleviate the spectral distortion and insufficient spatial texture enhancement. Since existing supervised pan-sharpening methods only utilize the ground truth MS image to supervise the network training, its potential value has not been fully explored. To accomplish this, we propose a heterogeneous knowledge-distilling pan-sharpening framework that distills pan-sharpening by imitating the ground truth reconstruction task in both the feature space and network output. In our work, the teacher network performs as a variational auto-encoder to extract effective features of the ground truth MS. The student network, acting as pan-sharpening, is trained by the assistance of the teacher network with the process-oriented feature imitation learning. Moreover, we design a customized information-lossless multi-scale invertible neural module to effectively fuse LR-MS and panchromatic (PAN) images, producing expected pan-sharpened results. To reduce the artifacts generated by the knowledge distillation process, a knowledge-driven refinement sub-network is further devised according to the pan-sharpening imaging model. Extensive experimental results on different satellite datasets validate that the proposed network outperforms the state-of-the-art methods both visually and quantitatively. The source code will be released at https://github.com/manman1995/pansharpening.
Man Zhou 0003, Jie Huang 0017, Xueyang Fu, Feng Zhao 0004, Danfeng Hong
IEEE Trans. Geosci. Remote. Sens.4
2021 Unsupervised Person Re-Identification Via Global-Level And Patch-Level Discriminative Feature Learning
abstract
Due to the lack of labeled data, it is usually difficult for an unsupervised person re-identification (re-ID) model to learn discriminative features. To address this issue, we propose a global-level and patch-level unsupervised feature learning framework that utilizes both global and local information to obtain more discriminative features. For global-level learning, we design a global similarity-based loss (GSL) to leverage the similarities between whole images. Along with a memory-based non-parametric classifier, the GSL pulls credible samples closer to help train a discriminative model. For patch-level learning, we use a patch generation module to produce different patches. Applying the patch-based discriminative feature learning loss and image-level feature learning loss, the patch branch in the network can learn better representative patch features. Combining the global-level learning with patch-level learning, we obtain a more distinguishable re-ID model. Experimental results obtained on Market-1501 and DukeMTMC-reID datasets validate that our method has great superiority and effectiveness in unsupervised person re-ID.
Zongzhe Sun, Feng Zhao 0004, Feng Wu 0001
ICIP2
2021 FFNet-M: Feature Fusion Network with Masks for Multimodal Facial Expression Recognition
abstract
Compared with 2D facial expression recognition (FER) and 3D FER, 2D+3D FER can handle the effects of illumination changes and pose variations. The combination of 2D texture and 3D attribute information can further improve the performance. However, most existing approaches still face two challenges: the selection of proper networks for extracting multimodal features, and the significance of local features in salient regions for expression classification. To address these challenges, we propose an efficient feature fusion network with masks (FFNet-M) for 2D+3D FER. Each 3D scan is rep-resented by three types of attribute maps (i.e., depth map, normal map, and texture image), which are then fed into FFNet-M with different networks to extract both 2D and 3D features. Moreover, we design two masks to make FFNet-M focus on 2D local features while paying attention to 3D local features in salient regions. Experimental results show that our FFNet-M outperforms state-of-the-art methods on BU-3DFE dataset and also achieves a high accuracy on Bosphorus dataset.
Mingzhe Sui, Zhaoqing Zhu, Feng Zhao 0004, Feng Wu 0001
ICME3
2021 Disentangle Your Dense Object Detector
abstract
Deep learning-based dense object detectors have achieved great success in the past few years and have been applied to numerous multimedia applications such as video understanding. However, the current training pipeline for dense detectors is compromised to lots of conjunctions that may not hold. In this paper, we investigate three such important conjunctions: 1) only samples assigned as positive in classification head are used to train the regression head; 2) classification and regression share the same input feature and computational fields defined by the parallel head architecture; and 3) samples distributed in different feature pyramid layers are treated equally when computing the loss. We first carry out a series of pilot experiments to show disentangling such conjunctions can lead to persistent performance improvement. Then, based on these findings, we propose Disentangled Dense Object Detector (DDOD), in which simple and effective disentanglement mechanisms are designed and integrated into the current state-of-the-art dense object detectors. Extensive experiments on MS COCO benchmark show that our approach can lead to 2.0~mAP, 2.4~mAP and 2.2~mAP absolute improvements on RetinaNet, FCOS, and ATSS baselines with negligible extra overhead. Notably, our best model reaches 55.0 mAP on the COCOtest-dev set and 93.5 AP on the hard subset of WIDER FACE, achieving new state-of-the-art performance on these two competitive benchmarks. Code is available at https://github.com/zehuichen123/DDOD.
Chenhongyi Yang, Qiaofei Li, Feng Zhao 0004, Zhengjun Zha, Feng Wu 0001
ACM Multimedia4
2020 Deep Learning-Based Classification of Liver Cancer Histopathology Images Using Only Global Labels
abstract
Liver cancer is a leading cause of cancer deaths worldwide due to its high morbidity and mortality. Histopathological image analysis (HIA) is a crucial step in the early diagnosis of liver cancer and is routinely performed manually. However, this process is time-consuming, error-prone, and easily affected by the expertise of pathologists. Recently, computer-aided methods have been widely applied to medical image analysis; however, the current medical image analysis studies have not yet focused on the histopathological morphology of liver cancer due to its complex features and the insufficiency of training images with detailed annotations. This paper proposes a deep learning method for liver cancer histopathological image classification using only global labels. To compensate for the lack of detailed cancer region annotations in those images, patch features are extracted and fully utilized. Transfer learning is used to obtain the patch-level features and then combined with multiple-instance learning to acquire the image-level features for classification. The method proposed here solves the processing of large-scale images and training sample insufficiency in liver cancer histopathological images for image classification. The proposed method can distinguish and classify liver histopathological images as abnormal or normal with high accuracy, thus providing support for the early diagnosis of liver cancer.
Chunli Sun, Dong Liu 0002, Zhiwei Xiong, Feng Zhao 0004, Weiping Ding 0002
IEEE J. Biomed. Health Informatics5
2019 PredMP: a web server for de novo prediction and visualization of membrane proteins
abstract
MOTIVATION: PredMP is the first web service, to our knowledge, that aims at de novo prediction of the membrane protein (MP) 3D structure followed by the embedding of the MP into the lipid bilayer for visualization. Our approach is based on a high-throughput Deep Transfer Learning (DTL) method that first predicts MP contacts by learning from non-MPs and then predicts the 3D model of the MP using the predicted contacts as distance restraints. This algorithm is derived from our previous Deep Learning (DL) method originally developed for soluble protein contact prediction, which has been officially ranked No. 1 in CASP12. The DTL framework in our approach overcomes the challenge that there are only a limited number of solved MP structures for training the deep learning model. There are three modules in the PredMP server: (i) The DTL framework followed by the contact-assisted folding protocol has already been implemented in RaptorX-Contact, which serves as the key module for 3D model generation; (ii) The 1D annotation module, implemented in RaptorX-Property, is used to predict the secondary structure and disordered regions; and (iii) the visualization module to display the predicted MPs embedded in the lipid bilayer guided by the predicted transmembrane topology. RESULTS: Tested on 510 non-redundant MPs, our server predicts correct folds for ∼290 MPs, which significantly outperforms existing methods. Tested on a blind and live benchmark CAMEO from September 2016 to January 2018, PredMP can successfully model all 10 MPs belonging to the hard category. AVAILABILITY AND IMPLEMENTATION: PredMP is freely accessed on the web at http://www.predmp.com. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online.
Sheng Wang 0001, Shiyang Fei, Zongan Wang, Yu Li 0006, Jinbo Xu, Feng Zhao 0004, Xin Gao 0001
Bioinform.6
2015 Structure Learning Constrained by Node-Specific Degree Distribution
Jianzhu Ma, Feng Zhao 0004, Jinbo Xu
UAI2
2013 Protein threading using context-specific alignment potential
abstract
MOTIVATION: Template-based modeling, including homology modeling and protein threading, is the most reliable method for protein 3D structure prediction. However, alignment errors and template selection are still the main bottleneck for current template-base modeling methods, especially when proteins under consideration are distantly related. RESULTS: We present a novel context-specific alignment potential for protein threading, including alignment and template selection. Our alignment potential measures the log-odds ratio of one alignment being generated from two related proteins to being generated from two unrelated proteins, by integrating both local and global context-specific information. The local alignment potential quantifies how well one sequence residue can be aligned to one template residue based on context-specific information of the residues. The global alignment potential quantifies how well two sequence residues can be placed into two template positions at a given distance, again based on context-specific information. By accounting for correlation among a variety of protein features and making use of context-specific information, our alignment potential is much more sensitive than the widely used context-independent or profile-based scoring function. Experimental results confirm that our method generates significantly better alignments and threading results than the best profile-based methods on several large benchmarks. Our method works particularly well for distantly related proteins or proteins with sparse sequence profiles because of the effective integration of context-specific, structure and global information. AVAILABILITY: http://raptorx.uchicago.edu/download/.
Jianzhu Ma, Sheng Wang 0001, Feng Zhao 0004, Jinbo Xu
Bioinform.3
2010 Binary SIPPER plankton image classification using random subspace
Feng Zhao 0004, Feng Lin 0002, Seah Hock Soon
Neurocomputing1
2009 Feature-based registration of confocal fluorescence endomicroscopy images
abstract
In this paper, we propose a feature-based registration algorithm for the confocal fluorescence images captured by an endomicroscopy system. We first extract a number of feature points from the endomicroscopy images by applying the binarization-thinning process and using the Rutovitz crossing number. These feature points are then post-processed to eliminate the spurious ones. After that we use the purified feature points to complete the image registration between every two consecutive slice images. The aligned image stack will be finally utilized to reconstruct and visualize the 3D structure of the living cell and tissue in real time, which provides the opportunity for the clinicians to diagnose various diseases including the early-stage cancers.
Feng Zhao 0004, Feng Lin 0002, Kemao Qian, Seah Hock Soon, Sun-Yuan Kung
ICIP1
2009 Bagging based plankton image classification
abstract
Plankton image classification plays an important role in ocean biological research. In this paper, we present an approach based on the bagging technique to classify the marine plankton images captured by the shadowed image particle profiling and evaluation recorder. The difficulty of such classification is multifold because the data set is much noisier, and the plankton images are deformable, projection-variant, and often in partial occlusion. In addition, the images in our experiments are binary, thus are lack of pixel-depth information. By random sampling with replacement on the original training set, a number of independent bootstrap replicates are generated. Using these replicates as new training sets, we construct multiple classifiers that are complementary of one another. While such individual classifiers are less effective than a single classifier trained on the whole training set, the fusion of them using majority voting produces an improved tenfold cross-validation accuracy by more than 93%.
Feng Zhao 0004, Feng Lin 0002, Seah Hock Soon
ICIP1
2007 A Two-Stage Fusion Scheme using Multiple Fingerprint Impressions
abstract
In this paper, we propose a two-stage fusion scheme that takes full advantage of the complementary information among multiple fingerprint impressions. While comparing the query fingerprint with a template impression, all the other impressions are also transformed using the 2D warping model to register with the query fingerprint so that the additive matched minutiae pairs can be detected to improve the matching result with a subset combination scheme. Then a matching score level fusion or decision level fusion is performed to integrate the improved matching results corresponding to different impressions. Experiments conducted on FVC2002 show that the proposed method produces a much better performance for fingerprint matching.
Lifeng Sha, Feng Zhao 0004, Xiaoou Tang
ICIP (2)2
2007 Preprocessing and postprocessing for skeleton-based fingerprint minutiae extraction
Feng Zhao 0004, Xiaoou Tang
Pattern Recognit.1
2005 Fingerprint matching using minutiae and interpolation-based square tessellation fingercode
abstract
To improve the overall accuracy, a hybrid fingerprint-matching scheme using both minutiae and square-tessellation-fingercode has been proposed in the literature. However, for identification applications, the matching process is time-consuming since the fingercode of the query fingerprint is repeatedly extracted when it is compared with different template fingerprints in a large database. In addition, the matching accuracy is influenced by nonlinear distortions in fingerprint images. In this paper, we propose a new approach to solve the problem. We extract the fignercode of the query fingerprint only once. When compared with the template fingerprints, the corresponding fingercodes are generated by interpolation and resampling on the extracted fingercode according to the optimally estimated mapping functions with respect to the minutiae matching results. Experimental results on NIST-4 and FVC2002 demonstrate that our algorithm outperforms the original approach in terms of both accuracy and running time.
Lifeng Sha, Feng Zhao 0004, Xiaoou Tang
ICIP (2)2
2005 Binary plankton image classification using random subspace
abstract
In this paper, we implement a random subspace based algorithm to classify the plankton images detected in real time by the shadowed image particle profiling and evaluation recorder. The difficulty of such classification is compounded because the data sets are not only much noisier but the plankton are deformable, projection-variant, and often in partial occlusion. In addition, the images in our experiments are binary thus are lack of texture information. Using random sampling, we construct a set of stable classifiers to take full advantage of nearly all the discriminative information in the feature space of plankton images. The combination of multiple stable classifiers is better than a single classifier. We achieve over 93% classification accuracy on a collection of more than 3000 images, making it comparable with what a trained biologist can achieve by using conventional manual techniques.
Feng Zhao 0004, Xiaoou Tang, Feng Lin 0002, Scott Samson, Andrew Remsen
ICIP (1)1
2003 Improved fingercode for filterbank-based fingerprint matching
abstract
FingerCode has been shown to be an effective representation to capture both the local and global information in a fingerprint. However, the performance of fingercode is influenced by the reference point detection process, and the AAD features cannot fully extract the discriminating information in fingerprints. In this paper, we first propose a new rotation-invariant reference point location method, and then combine the direction features with the AAD features to form an oriented fingercode. Experiments conducted on a large fingerprint database (NIST-4) show that the proposed method produces a much improved matching performance.
Lifeng Sha, Feng Zhao 0004, Xiaoou Tang
ICIP (2)2