VLDB 2026 Research / reviewers in the wild / expert
Yan Zhang 0109
dblp:04/3348-109
· DBLP profile ↗
36ranked-venue papers
2as first author
34since 2021 · last 2026
0000-0003-1642-0758ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 22 · 2 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 1 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FedPSAWA: Federated personalization with state aware weighting aggregation for cross subject seizure prediction
Peipei Gu, Jibin Shou, Yuping Zhao, Meiyan Xu, Jiayang Guo, Yan Zhang 0109, Jianbin Jiao, Jingzhu Li |
Neurocomputing | 7 |
| 2026 | GraphMSR: A graph foundation model-based approach for MRI image super-resolution with multimodal semantic integration
Zhiquan Qin, Yan Zhang 0109, Yunhang Shen, Ke Li 0015 |
Pattern Recognit. | 3 |
| 2025 | BUFF: Bayesian Uncertainty Guided Diffusion Probabilistic Model for Single Image Super-ResolutionabstractSuper-resolution (SR) techniques are critical for enhancing image quality, particularly in scenarios where high-resolution imagery is essential yet limited by hardware constraints. Existing diffusion models for SR have relied predominantly on Gaussian models for noise generation, which often fall short when dealing with the complex and variable texture inherent in natural scenes. To address these deficiencies, we introduce the Bayesian Uncertainty Guided Diffusion Probabilistic Model (BUFF). BUFF distinguishes itself by incorporating a Bayesian network to generate high-resolution uncertainty masks. These masks guide the diffusion process, allowing for the adjustment of noise intensity in a manner that is both context-aware and adaptive. This novel approach not only enhances the fidelity of super-resolved images to their original high-resolution counterparts but also significantly mitigates artifacts and blurring in areas characterized by complex textures and fine details. The model demonstrates exceptional robustness against complex noise patterns and showcases superior adaptability in handling textures and edges within images. Empirical evidence, supported by visual results, illustrates the model's robustness, especially in challenging scenarios, and its effectiveness in addressing common SR issues such as blurring. Experimental evaluations conducted on the DIV2K dataset reveal that BUFF achieves a notable improvement, with a +0.61 increase compared to baseline in SSIM on BSD100, surpassing traditional diffusion approaches by an average additional +0.20dB PSNR gain. These findings underscore the potential of Bayesian methods in enhancing diffusion processes for SR, paving the way for future advancements in the field. Shengchuan Zhang, Runze Hu, Yunhang Shen, Yan Zhang 0109 |
AAAI | 5 |
| 2025 | Feature Denoising Diffusion Model for Blind Image Quality AssessmentabstractBlind Image Quality Assessment (BIQA) aims to evaluate image quality in line with human perception, without reference benchmarks. Currently, deep learning BIQA methods typically depend on using features from high-level tasks for transfer learning. However, the inherent differences between BIQA and these high-level tasks inevitably introduce noise into the quality-aware features. In this paper, we take an initial step toward exploring the diffusion model for feature denoising in BIQA, namely Perceptual Feature Diffusion for IQA (PFD-IQA), which aims to remove noise from quality-aware features. Specifically, 1) we propose a Perceptual Prior Discovery and Aggregation module to establish two auxiliary tasks to discover potential low-level features in images that are used to aggregate perceptual textual prompt conditions for the diffusion model. 2) we propose a Perceptual Conditional Feature Refinement strategy, which matches noisy features to predefined denoising trajectories and then performs exact feature denoising based on textual prompt conditions. By incorporating a lightweight denoiser and requiring only a few feature denoising steps (e.g., just five iterations), our PFD-IQA framework achieves superior performance across eight standard BIQA datasets, validating its effectiveness. Yan Zhang 0109, Yunhang Shen, Ke Li 0015, Runze Hu, Xiawu Zheng, Sicheng Zhao |
AAAI | 2 |
| 2025 | CFBench: A Comprehensive Constraints-Following Benchmark for LLMsabstractTao Zhang, ChengLIn Zhu, Yanjun Shen, Wenjing Luo, Yan Zhang, Hao Liang, Tao Zhang, Fan Yang, Mingan Lin, Yujing Qiao, Weipeng Chen, Bin Cui, Wentao Zhang, Zenan Zhou. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Tao Zhang 0194, Chenglin Zhu, Wenjing Luo, Yan Zhang 0109, Hao Liang 0017, Fan Yang 0132, Yujing Qiao, Weipeng Chen, Bin Cui 0001, Wentao Zhang 0001, Zenan Zhou |
ACL (1) | 5 |
| 2025 | Distilling Spatially-Heterogeneous Distortion Perception for Blind Image Quality AssessmentabstractIn the Blind Image Quality Assessment (BIQA) field, accurately assessing the quality of authentically distorted images presents a substantial challenge due to the diverse distortion types in natural settings. Existing state-of-the-art IQA methods mix a sequence of distortions into entire images to establish global distortion priors, but are inadequate for authentic images with spatially varied distortions. To address this, we introduce a novel IQA framework that employs knowledge distillation tailored to perceive spatially heterogeneous distortions, enhancing quality-distortion awareness. Specifically, we introduce a novel Block-wise Degradation Modelling approach that applies distinct distortions to different spatial blocks of an image, thereby expanding local distortion priors. Following this, we present a Block-wise Aggregation and Filtering module that enables fine-grained attention to the quality information within different distortion areas of the image. Furthermore, to effectively capture the complex relationships between distortions across different regions while preserving overall quality perception, we introduce Contrastive Knowledge Distillation to enhance the model’s ability to discriminate between different types of distortions and Affinity Knowledge Distillation to model the correlation among distortions in different regions. Extensive experiments on standard BIQA datasets demonstrate the effectiveness and competitiveness of the proposed method. Wenjie Nie, Yan Zhang 0109, Runze Hu, Ke Li 0015, Xiawu Zheng, Liujuan Cao |
CVPR | 3 |
| 2025 | UCOD-DPL: Unsupervised Camouflaged Object Detection via Dynamic Pseudo-label LearningabstractUnsupervised Camoflaged Object Detection (UCOD) has gained attention since it doesn't need to rely on extensive pixel-level labels. Existing UCOD methods typically generate pseudo-labels using fixed strategies and train 1 × 1 convolutional layers as a simple decoder, leading to low performance compared to fully-supervised methods. We emphasize two drawbacks in these approaches: 1). The model is prone to fitting incorrect knowledge due to the pseudo-label containing substantial noise. 2). The simple decoder fails to capture and learn the semantic features of camouflaged objects, especially for small-sized objects, due to the low-resolution pseudo-labels and severe confusion between foreground and background pixels. To this end, we propose a UCOD method with a teacher-student framework via Dynamic Pseudo-label Learning called UCOD-DPL, which contains an Adaptive Pseudo-label Module (APM), a Dual-Branch Adversarial (DBA) decoder, and a Look-Twice mechanism. The APM module adaptively combines pseudo-labels generated by fixed strategies and the teacher model to prevent the model from overfitting incorrect knowledge while preserving the ability for self-correction; the DBA decoder takes adversarial learning of different segmentation objectives, guides the model to overcome the foreground-background confusion of camouflaged objects, and the Look-Twice mechanism mimics the human tendency to zoom in on camouflaged objects and performs secondary refinement on small-sized objects. Extensive experiments show that our method demonstrates outstanding performance, even surpassing some existing fully supervised methods. The code is available now1. Weiqi Yan 0005, Lvhai Chen, Huaijia Kou, Shengchuan Zhang, Yan Zhang 0109, Liujuan Cao |
CVPR | 5 |
| 2025 | Few-Shot Image Quality Assessment via Adaptation of Vision-Language Models
Yan Zhang 0109, Yunhang Shen, Ke Li 0015, Xiawu Zheng, Liujuan Cao, Rongrong Ji |
ICCV | 3 |
| 2025 | ESCNet: Edge-Semantic Collaborative Network for Camouflaged Object Detection
Xin Chen 0032, Yan Zhang 0109, Xianming Lin, Liujuan Cao |
ICCV | 3 |
| 2025 | SysBench: Can LLMs Follow System Message?abstractLarge Language Models (LLMs) have become instrumental across various applications, with the customization of these models to specific scenarios becoming increasingly critical. System message, a fundamental component of LLMs, is consist of carefully crafted instructions that guide the behavior of model to meet intended goals. Despite the recognized potential of system messages to optimize AI-driven solutions, there is a notable absence of a comprehensive benchmark for evaluating how well LLMs follow system messages. To fill this gap, we introduce SysBench, a benchmark that systematically analyzes system message following ability in terms of three limitations of existing LLMs: constraint violation, instruction misjudgement and multi-turn instability. Specifically, we manually construct evaluation dataset based on six prevalent types of constraints, including 500 tailor-designed system messages and multi-turn user conversations covering various interaction relationships. Additionally, we develop a comprehensive evaluation protocol to measure model performance. Finally, we conduct extensive evaluation across various existing LLMs, measuring their ability to follow specified constraints given in system messages. The results highlight both the strengths and weaknesses of existing models, offering key insights and directions for future research. Yanzhao Qin, Tao Zhang 0194, Wenjing Luo, Haoze Sun, Yan Zhang 0109, Yujing Qiao, Weipeng Chen, Zenan Zhou, Wentao Zhang 0001, Bin Cui 0001 |
ICLR | 7 |
| 2025 | SCOUT: Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data SelectionabstractThe difficulty of pixel-level annotation has significantly hindered the development of the Camouflaged Object Detection (COD) field. To save on annotation costs, previous works leverage the semi-supervised COD framework that relies on a small number of labeled data and a large volume of unlabeled data. We argue that there is still significant room for improvement in the effective utilization of unlabeled data. To this end, we introduce a Semi-supervised Camouflaged Object Detection by Utilizing Text and Adaptive Data Selection (SCOUT). It includes an Adaptive Data Augment and Selection (ADAS) module and a Text Fusion Module (TFM). The ADSA module selects valuable data for annotation through an adversarial augment and sampling strategy. The TFM module further leverages the selected valuable data by combining camouflage-related knowledge and text-visual interaction. To adapt to this work, we build a new dataset, namely RefTextCOD. Extensive experiments show that the proposed method surpasses previous semi-supervised methods in the COD field and achieves state-of-the-art performance. Our code will be released at https://github.com/Heartfirey/UCOD-DPL. Weiqi Yan 0005, Lvhai Chen, Shengchuan Zhang, Yan Zhang 0109, Liujuan Cao |
IJCAI | 4 |
| 2025 | Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMsabstractMulti-modal Large Language Models (MLLMs) excel at single-image tasks but struggle with multi-image understanding due to cross-modal misalignment, leading to hallucinations (context omission, conflation, and misinterpretation). Existing methods using Direct Preference Optimization (DPO) constrain optimization to a solitary image reference within the input sequence, neglecting holistic context modeling. To address this, we propose Context-to-Cue Direct Preference Optimization (CcDPO), a multi-level preference optimization framework that enhances per-image perception in multi-image settings by zooming into visual clues—from sequential context to local details. Our approach features two sequentially dependent components: (i) Context-Level Optimization: By introducing low-cost sequence preference pairs, we optimize the model to distinguish between complete and disrupted multi-image contexts, thereby correcting cognitive biases in MLLMs’ multi-image understanding. (ii) Needle-Level Optimization: By integrating region-specific visual prompts with multimodal preference supervision, we direct the model’s attention to critical visual details, effectively suppressing perceptual biases toward fine-grained visual information. To support scalable optimization, we also construct MultiScope-42k, an automatically generated multi-image dataset with hierarchical preference pairs. Experiments show that CcDPO significantly reduces hallucinations and yields consistent performance gains across general single- and multi-image tasks. Codes are available at https://github.com/LXDxmu/CcDPO. Mengdan Zhang, Peixian Chen, Xiawu Zheng, Yan Zhang 0109, Jingyuan Zheng, Yunhang Shen, Ke Li 0015, Chaoyou Fu, Xing Sun 0001, Rongrong Ji |
NeurIPS | 5 |
| 2025 | LTD-Bench: Evaluating Large Language Models by Letting Them DrawabstractCurrent evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research—relying on opaque numerical metrics that conceal fundamental limitations in spatial reasoning while providing no intuitive understanding of model capabilities. This deficiency creates a dangerous disconnect between reported performance and practical abilities, particularly for applications requiring physical world understanding. We introduce LTD-Bench, a breakthrough benchmark that transforms LLM evaluation from abstract scores to directly observable visual outputs by requiring models to generate drawings through dot matrices or executable code. This approach makes spatial reasoning limitations immediately apparent even to non-experts, bridging the fundamental gap between statistical performance and intuitive assessment. LTD-Bench implements a comprehensive methodology with complementary generation tasks (testing spatial imagination) and recognition tasks (assessing spatial perception) across three progressively challenging difficulty levels, methodically evaluating both directions of the critical language-spatial mapping. Our extensive experiments with state-of-the-art models expose an alarming capability gap: even LLMs achieving impressive results on traditional benchmarks demonstrate profound deficiencies in establishing bidirectional mappings between language and spatial concepts—a fundamental limitation that undermines their potential as genuine world models. Furthermore, LTD-Bench's visual outputs enable powerful diagnostic analysis, offering a potential approach to investigate model similarity. Our dataset and codes are available at https://github.com/walktaster/LTD-Bench. Liuhao Lin, Ke Li 0015, Yulei Qin, Yan Zhang 0109, Xing Sun 0001, Rongrong Ji |
NeurIPS | 6 |
| 2025 | DPCA: Dynamic multi-prototype cross-attention for change detection unsupervised domain adaptation of remote sensing images
Rongbo Fan, Jialin Xie, Junmin Liu, Yan Zhang 0109, Hong Hou, Jianhua Yang 0005 |
Knowl. Based Syst. | 4 |
| 2025 | HGTL: A hypergraph transfer learning framework for survival prediction of ccRCC
Xiangmin Han, Wuchao Li, Yan Zhang 0109, Pinhao Li, Jianguo Zhu 0003, Tijiang Zhang, Rongpin Wang, Yue Gao 0002 |
Medical Image Anal. | 3 |
| 2025 | Attention-driven acoustic properties learning for underwater target ranging
Xiaohui Chu, Hantao Zhou, Yan Zhang 0109, Yachao Zhang 0001, Runze Hu, Haoran Duan 0001, Yawen Huang, Yefeng Zheng 0001, Rongrong Ji |
Pattern Recognit. | 3 |
| 2025 | NAPG: Neighborhood-Assisted Multiprototype Group Model for Cross-Domain Semantic Segmentation of Remote Sensing ImagesabstractUnsupervised domain adaptation (UDA) is crucial for semantic segmentation of remote sensing images (RS-SS), particularly when data distributions differ between source and target domains. Existing prototype-based UDA methods struggle with complex land cover class distributions and spatial information capture. To address these limitations, the Neighborhood-Assisted Multi-Prototype Group (NAPG) model is proposed. This model enhances cross-domain adaptability and spatial context richness by dynamically determining the number of prototype features and integrating neighborhood similarity gradients. Specifically, NAPG employs the Cross-Domain Representation of Multi-Prototype Group (CDR-MPG) module to generate multi-prototype group, capturing land cover complexity more effectively. Additionally, the Gradient Neighborhood Consistency Estimation (GNCE) module improves spatial representation by reducing intra-class variance and alleviating feature inconsistency. Experiments demonstrate that the proposed NAPG model outperforms state-of-the-art UDA methods across multiple datasets, achieving an mean intersection over union (mIoU) improvement of 3%. The source code is publicly available at https://github.com/Fanrongbo/NAPG-UDA-RS-SS. Rongbo Fan, Jialin Xie, Junmin Liu, Jun Zhang 0018, Yan Zhang 0109, Hong Hou, Jianhua Yang 0005 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | MMICT: Boosting Multi-Modal Fine-Tuning with In-Context ExamplesabstractAlthough In-Context Learning (ICL) brings remarkable performance gains to Large Language Models (LLMs), the improvements remain lower than fine-tuning on downstream tasks. This paper introduces Multi-Modal In-Context Tuning (MMICT), a novel multi-modal fine-tuning paradigm that boosts multi-modal fine-tuning by fully leveraging the promising ICL capability of Multi-Modal LLMs (MM-LLMs). We propose the Multi-Modal Hub (M-Hub), a unified module that captures various multi-modal features according to different inputs and objectives. Based on M-Hub, MMICT enables MM-LLMs to learn from in-context visual-guided textual features and subsequently generate outputs conditioned on the textual-guided visual features. Moreover, leveraging the flexibility of M-Hub, we design a variety of in-context demonstrations. Extensive experiments on a diverse range of downstream multi-modal tasks demonstrate that MMICT significantly outperforms traditional fine-tuning strategy and the vanilla ICT method that directly takes the concatenation of all information from different modalities as input. Our implementation is available at: https://github.com/KDEGroup/MMICT . Enwei Zhang, Ke Li 0015, Xing Sun 0001, Yan Zhang 0109, Hui Li 0057, Rongrong Ji |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | Semi-Supervised Blind Image Quality Assessment through Knowledge Distillation and Incremental LearningabstractBlind Image Quality Assessment (BIQA) aims to simulate human assessment of image quality. It has a great demand for labeled data, which is often insufficient in practice. Some researchers employ unsupervised methods to address this issue, which is challenging to emulate the human subjective system. To this end, we introduce a unified framework that combines semi-supervised and incremental learning to address the mentioned issue. Specifically, when training data is limited, semi-supervised learning is necessary to infer extensive unlabeled data. To facilitate semi-supervised learning, we use knowledge distillation to assign pseudo-labels to unlabeled data, preserving analytical capability. To gradually improve the quality of pseudo labels, we introduce incremental learning. However, incremental learning can lead to catastrophic forgetting. We employ Experience Replay by selecting representative samples during multiple rounds of semi-supervised learning, to alleviate forgetting and ensure model stability. Experimental results show that the proposed approach achieves state-of-the-art performance across various benchmark datasets. After being trained on the LIVE dataset, our method can be directly transferred to the CSIQ dataset. Compared with other methods, it significantly outperforms unsupervised methods on the CSIQ dataset with a marginal performance drop (-0.002) on the LIVE dataset. In conclusion, our proposed method demonstrates its potential to tackle the challenges in real-world production processes. Wensheng Pan, Timin Gao, Yan Zhang 0109, Xiawu Zheng, Yunhang Shen, Ke Li 0015, Runze Hu, Yutao Liu 0002, Pingyang Dai |
AAAI | 3 |
| 2024 | Adaptive Feature Selection for No-Reference Image Quality Assessment by Mitigating Semantic Noise SensitivityabstractThe current state-of-the-art No-Reference Image Quality Assessment (NR-IQA) methods typically rely on feature extraction from upstream semantic backbone networks, assuming that all extracted features are relevant. However, we make a key observation that not all features are beneficial, and some may even be harmful, necessitating careful selection. Empirically, we find that many image pairs with small feature spatial distances can have vastly different quality scores, indicating that the extracted features may contain quality-irrelevant noise. To address this issue, we propose a Quality-Aware Feature Matching IQA Metric (QFM-IQM) that employs an adversarial perspective to remove harmful semantic noise features from the upstream task. Specifically, QFM-IQM enhances the semantic noise distinguish capabilities by matching image pairs with similar quality scores but varying semantic features as adversarial semantic noise and adaptively adjusting the upstream task’s features by reducing sensitivity to adversarial noise perturbation. Furthermore, we utilize a distillation framework to expand the dataset and improve the model’s generalization ability. Extensive experiments conducted on eight standard IQA datasets have demonstrated the effectiveness of our proposed QFM-IQM. Timin Gao, Runze Hu, Yan Zhang 0109, Shengchuan Zhang, Xiawu Zheng, Jingyuan Zheng, Yunhang Shen, Ke Li 0015, Yutao Liu 0002, Pingyang Dai, Rongrong Ji |
ICML | 4 |
| 2024 | Integrating Global Context Contrast and Local Sensitivity for Blind Image Quality AssessmentabstractBlind Image Quality Assessment (BIQA) mirrors subjective made by human observers. Generally, humans favor comparing relative qualities over predicting absolute qualities directly. However, current BIQA models focus on mining the "local" context, i.e., the relationship between information among individual images and the absolute quality of the image, ignoring the "global" context of the relative quality contrast among different images in the training data. In this paper, we present the Perceptual Context and Sensitivity BIQA (CSIQA), a novel contrastive learning paradigm that seamlessly integrates "global” and "local” perspectives into the BIQA. Specifically, the CSIQA comprises two primary components: 1) A Quality Context Contrastive Learning module, which is equipped with different contrastive learning strategies to effectively capture potential quality correlations in the global context of the dataset. 2) A Quality-aware Mask Attention Module, which employs the random mask to ensure the consistency with visual local sensitivity, thereby improving the model’s perception of local distortions. Extensive experiments on eight standard BIQA datasets demonstrate the superior performance to the state-of-the-art BIQA methods. Runze Hu, Jingyuan Zheng, Yan Zhang 0109, Shengchuan Zhang, Xiawu Zheng, Ke Li 0015, Yunhang Shen, Yutao Liu 0002, Pingyang Dai, Rongrong Ji |
ICML | 4 |
| 2024 | Cantor: Inspiring Multimodal Chain-of-Thought of MLLMabstractWith the advent of large language models(LLMs) enhanced by the chain-of-thought(CoT) methodology, the visual reasoning problem is usually decomposed into manageable sub-tasks and tackled sequentially with various external tools. However, such a paradigm faces the challenge of the potential "determining hallucinations" in decision generation due to insufficient visual information and the limitation of low-level perception tools that fail to provide abstract summaries necessary for comprehensive reasoning. We argue that converging visual context acquisition and logical reasoning is pivotal for tackling visual reasoning tasks. This paper delves into the realm of multimodal CoT to solve intricate visual reasoning tasks with multimodal large language models(MLLMs) and their cognitive capability. To this end, we propose an innovative multimodal CoT framework, termed Cantor, characterized by a perception-decision architecture. Cantor first acts as a decision generator and integrates visual inputs to analyze the image and problem, ensuring a closer alignment with the actual context. Furthermore, Cantor leverages the advanced cognitive functions of MLLMs to perform as multifaceted experts for deriving higher-level information, enhancing the CoT generation process. Our extensive experiments demonstrate the efficacy of the proposed framework, showing significant improvements in multimodal CoT performance across two complex visual reasoning datasets, without necessitating fine-tuning or ground-truth rationales. Project Page: https://ggg0919.github.io/cantor/. Timin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu, Yunhang Shen, Yan Zhang 0109, Shengchuan Zhang, Xiawu Zheng, Xing Sun 0001, Liujuan Cao, Rongrong Ji |
ACM Multimedia | 6 |
| 2024 | Adaptive Selection based Referring Image SegmentationabstractReferring image segmentation (RIS) aims to segment a particular region based on a specific expression. Existing one-stage methods have explored various fusion strategies, yet they encounter two significant issues. Primarily, most methods rely on manually selected visual features from the visual encoder layers. Moreover, the direct fusion of word-level features into coarse aligned features disrupts the established vision-language alignment. In this paper, we introduce an innovative framework for RIS that seeks to overcome these challenges with adaptive alignment of vision and language features, termed the Adaptive Selection with Dual Alignment (ASDA). ASDA innovates in two aspects. Firstly, we design an Adaptive Feature Selection and Fusion (AFSF) module to dynamically select visual features focusing on different regions related to various descriptions. AFSF is equipped with scale-wise feature aggregator to provide hierarchically coarse features that preserve crucial low-level details. Secondly, a Word Guided Dual-Branch Aligner (WGDA) is leveraged to integrate coarse features with linguistic cues by word-guided attention, which effectively addresses the common issue of vision-language misalignment. Extensive experimental results demonstrate that our ASDA framework surpasses state-of-the-art methods on RefCOCO, RefCOCO+ and G-Ref benchmark. Pengfei Yue, Jianghang Lin, Shengchuan Zhang, Jie Hu 0018, Hongwei Niu, Haixin Ding, Yan Zhang 0109, Guannan Jiang, Liujuan Cao, Rongrong Ji |
ACM Multimedia | 8 |
| 2024 | RLE: A Unified Perspective of Data Augmentation for Cross-Spectral Re-IdentificationabstractThis paper makes a step towards modeling the modality discrepancy in the cross-spectral re-identification task. Based on the Lambertain model, we observe that the non-linear modality discrepancy mainly comes from diverse linear transformations acting on the surface of different materials. From this view, we unify all data augmentation strategies for cross-spectral re-identification as mimicking such local linear transformations and categorize them into moderate transformation and radical transformation. By extending the observation, we propose a Random Linear Enhancement (RLE) strategy which includes Moderate Random Linear Enhancement (MRLE) and Radical Random Linear Enhancement (RRLE) to push the boundaries of both types of transformation. Moderate Random Linear Enhancement is designed to provide diverse image transformations that satisfy the original linear correlations under constrained conditions, whereas Radical Random Linear Enhancement seeks to generate local linear transformations directly without relying on external information. The experimental results not only demonstrate the superiority and effectiveness of RLE but also confirm its great potential as a general-purpose data augmentation for cross-spectral re-identification. Keke Han, Pingyang Dai, Yan Zhang 0109, Yongjian Wu 0001, Rongrong Ji |
NeurIPS | 5 |
| 2024 | CPE COIN++: Towards Optimized Implicit Neural Representation Compression Via Chebyshev Positional Encoding
Haocheng Chu, Shaohui Dai, Wenqi Ding, Tianshuo Xu, Pingyang Dai, Shengchuan Zhang, Yan Zhang 0109, Xiang Chang, Chih-Min Lin, Fei Chao 0001, Changjiang Shang, Qiang Shen 0001 |
PRCV (9) | 8 |
| 2023 | Data-Efficient Image Quality Assessment with Attention-Panel DecoderabstractBlind Image Quality Assessment (BIQA) is a fundamental task in computer vision, which however remains unresolved due to the complex distortion conditions and diversified image contents. To confront this challenge, we in this paper propose a novel BIQA pipeline based on the Transformer architecture, which achieves an efficient quality-aware feature representation with much fewer data. More specifically, we consider the traditional fine-tuning in BIQA as an interpretation of the pre-trained model. In this way, we further introduce a Transformer decoder to refine the perceptual information of the CLS token from different perspectives. This enables our model to establish the quality-aware feature manifold efficiently while attaining a strong generalization capability. Meanwhile, inspired by the subjective evaluation behaviors of human, we introduce a novel attention panel mechanism, which improves the model performance and reduces the prediction uncertainty simultaneously. The proposed BIQA method maintains a light-weight design with only one layer of the decoder, yet extensive experiments on eight standard BIQA datasets (both synthetic and authentic) demonstrate its superior performance to the state-of-the-art BIQA methods, i.e., achieving the SRCC values of 0.875 (vs. 0.859 in LIVEC) and 0.980 (vs. 0.969 in LIVE). Checkpoints, logs and code will be available at https://github.com/narthchin/DEIQT. Guanyi Qin, Runze Hu, Yutao Liu 0002, Xiawu Zheng, Xiu Li 0001, Yan Zhang 0109 |
AAAI | 7 |
| 2023 | A Novel Neighbor Aggregation Function for Medical Point Cloud Analysis
Fan Wu 0006, Yumeng Qian, Haozhun Zheng, Yan Zhang 0109, Xiawu Zheng |
CGI (4) | 4 |
| 2023 | Event-Diffusion: Event-Based Image Reconstruction and Restoration with Diffusion ModelsabstractEvent cameras offer the advantages of low latency, high temporal resolution and HDR compared to conventional cameras. Due to the asynchronous and sparse nature of events, many existing algorithms cannot be directly applied, necessitating the reconstruction of intensity frames. However, existing reconstruction methods often result in artifacts and edge blurring due to noise and event accumulation. In this paper, we argue that the key to event-based image reconstruction is to enhance the edge information of objects and restore the artifacts in the reconstructed images. To explain, edge information is one of the most important features in the event stream, providing information on the shape and contour of objects. Considering the extraordinary capabilities of Denoising Diffusion Probabilistic Models (DDPMs) in image generation, reconstruction, and restoration, we propose a new framework which incorporate it into the reconstruction pipeline to obtain high-quality results which effectively remove artifacts and blur in reconstructed images. Specifically, we first extract edge information from the event stream using the proposed event-based denoising method. It employs the contrast maximization framework to remove noise from the event stream and extract clear object edge information. And then, the edge information is further adopted to our diffusion model, which is used to enhance the edges of objects in the reconstructed images, thus improving the restoration effect. Experimental results show that our method achieves significant improvements in the mean squared error (MSE), the structural similarity (SSIM), and the perceptual similarity (LPIPS) metrics, with average improvements of 40%, 15%, and 25%, respectively, compared to previous state-of-the-art models, and has good generalization performance. Quanmin Liang, Xiawu Zheng, Kai Huang 0001, Yan Zhang 0109, Jie Chen 0001, Yonghong Tian 0001 |
ACM Multimedia | 4 |
| 2023 | Learning Occlusion Disentanglement with Fine-grained Localization for Occluded Person Re-identificationabstractPerson re-identification (Re-ID) has been extensively investigated in recent years. However, many existing paradigms rely on holistic person regions for matching, disregarding the challenges posed by occlusions in real-world scenarios. Recent methods have explored occlusion augmentation or external semantic cues. Nevertheless, these approaches tend to be coarse-grained, discarding valuable semantic information in local regions when determining them as occlusions. In this paper, we propose a Fine-grained Occlusion Disentanglement Network (FODN) that can extract more information from limited person regions. Specifically, we propose a fine-grained occlusion augmentation scheme to generate diverse occlusion data and employ bilinear interpolation and downsampling strategies to obtain fine-grained occlusion labels. We then design an occlusion feature disentanglement Module that decouples norm and angle from features and supervises the occlusion-aware task using the aforementioned occlusion labeling and person re-identification tasks, respectively, resulting in more robust features. Additionally, we propose a dynamic local weight controller to balance the relative importance of various human body parts, thereby improving the model's ability to mine more effective local features from limited human body regions after occlusion removal. Comprehensive experiments on various person Re-ID benchmarks demonstrate the superiority of FODN over state-of-the-art methods. Yan Zhang 0109, Pingyang Dai, Yongjian Wu 0001, Rongrong Ji |
ACM Multimedia | 4 |
| 2023 | Classifier Decoupled Training for Black-Box Unsupervised Domain Adaptation
Xiangchuang Chen, Yunhang Shen, Yan Zhang 0109, Ke Li 0015, Shaohui Lin |
PRCV (3) | 4 |
| 2023 | Cross-Dataset Distillation with Multi-tokens for Image Quality Assessment
Timin Gao, Weixuan Jin, Bokai Lai, Runze Hu, Yan Zhang 0109, Pingyang Dai |
PRCV (6) | 6 |
| 2023 | Data-Free Low-Bit Quantization via Dynamic Multi-teacher Knowledge Distillation
Shaohui Lin, Yan Zhang 0109, Ke Li 0015, Baochang Zhang 0001 |
PRCV (8) | 3 |
| 2023 | Quality-Aware CLIP for Blind Image Quality Assessment
Wensheng Pan, Zhifu Yang, DingMing Liu, Chenxin Fang, Yan Zhang 0109, Pingyang Dai |
PRCV (6) | 5 |
| 2023 | Prompt Based Lifelong Person Re-identification
Chengde Yang, Yan Zhang 0109, Pingyang Dai |
PRCV (12) | 2 |
| 2016 | Search-Based Depth Estimation via Coupled Dictionary Learning with Large-Margin Structure Inference
Yan Zhang 0109, Rongrong Ji, Xiaopeng Fan 0001, Yan Wang 0059, Feng Guo 0005, Yue Gao 0002, Debin Zhao |
ECCV (5) | 1 |
| 2016 | 3D object retrieval with multi-feature collaboration and bipartite graph matching
Yan Zhang 0109, Feng Jiang 0001, Seungmin Rho, Shaohui Liu, Debin Zhao, Rongrong Ji |
Neurocomputing | 1 |