Yongbing Zhang 0002

dblp:95/5329-2 · DBLP profile ↗
← Back
135ranked-venue papers
25as first author
58since 2021 · last 2026
0000-0003-3320-2904ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 110 · 24 first-author · 37 since 2021Artificial intelligence and machine learning · 22 · 18 since 2021Applied, interdisciplinary, general and emerging computing · 21 · 21 since 2021Databases, data management, data science and information retrieval · 7 · 3 first-author · 1 since 2021Computer networks · 4 · 3 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Spectral Property-Driven Data Augmentation for Hyperspectral Single-Source Domain Generalization
abstract
While hyperspectral images (HSI) benefit from numerous spectral channels that provide rich information for classification, the increased dimensionality and sensor variability make them more sensitive to distributional discrepancies across domains, which in turn can affect classification performance. To tackle this issue, hyperspectral single-source domain generalization (SDG) typically employs data augmentation to simulate potential domain shifts and enhance model robustness under the condition of single-source domain training data availability. However, blind augmentation may produce samples misaligned with real-world scenarios, while excessive emphasis on realism can suppress diversity, highlighting a tradeoff between realism and diversity that limits generalization to target domains. To address this challenge, we propose a spectral property-driven data augmentation (SPDDA) that explicitly accounts for the inherent properties of HSI, namely the device-dependent variation in the number of spectral channels and the mixing of adjacent channels. Specifically, SPDDA employs a spectral diversity module that resamples data from the source domain along the spectral dimension to generate samples with varying spectral channels, and constructs a channel-wise adaptive spectral mixer by modeling inter-channel similarity, thereby avoiding fixed augmentation patterns. To further enhance the realism of the augmented samples, we propose a spatial-spectral co-optimization mechanism, which jointly optimizes a spatial fidelity constraint and a spectral continuity self-constraint. Moreover, the weight of the spectral self-constraint is adaptively adjusted based on the spatial counterpart, thus preventing over-smoothing in the spectral dimension and preserving spatial structure. Extensive experiments conducted on three remote sensing benchmarks demonstrate that SPDDA outperforms state-of-the-art methods.
Taiqin Chen, Yifeng Wang 0001, Xiaochen Feng, Hao Sha 0001, Yongbing Zhang 0002
AAAI7
2026 PathFLIP: Fine-grained Language-Image Pretraining for Versatile Computational Pathology
abstract
While Vision-Language Models (VLMs) have achieved notable progress in computational pathology (CPath), the gigapixel scale and spatial heterogeneity of Whole Slide Images (WSIs) continue to pose challenges for multimodal understanding. Existing alignment methods struggle to capture fine-grained correspondences between textual descriptions and visual cues across thousands of patches from a slide, compromising their performance on downstream tasks. In this paper, we propose PathFLIP (Pathology Fine-grained Language-Image Pretraining), a novel framework for holistic WSI interpretation. PathFLIP decomposes slide-level captions into region-level sub-captions and generates text-conditioned region embeddings to facilitate precise visual-language grounding. By harnessing Large Language Models (LLMs), PathFLIP can seamlessly follow diverse clinical instructions and adapt to varied diagnostic contexts. Furthermore, it exhibits versatile capabilities across multiple paradigms, efficiently handling slide-level classification and retrieval, fine-grained lesion localization, and instruction following. Extensive experiments demonstrate that PathFLIP outperforms existing large-scale pathological VLMs on four representative benchmarks while requiring significantly less training data, paving the way for fine-grained, instruction-aware WSI interpretation in research and clinical practice.
Fengchun Liu, Songhan Jiang, Linghan Cai, Ziyue Wang 0005, Yongbing Zhang 0002
AAAI5
2026 Reconstructing Temporal Heterogeneity: A Multidomain Collaborative Analysis Framework for Robust Time-Series Forecasting
Hengrui Li, Wenxue Cui, Yifeng Wang 0001, Chunshan Dong, Wenju Li, Jiangpeng Shi, Yongbing Zhang 0002, Shaohui Liu
IEEE Internet Things J.9
2026 Multi-Beholder: Biomarker Prediction for Low-Grade Glioma With Multiple Instance Learning and One-Class Classification
abstract
Biomarker detection is an indispensable part of the diagnosis and treatment of low-grade glioma (LGG). However, current LGG biomarker detection methods rely on expensive and complex molecular genetic testing, for which professionals are required to analyze the results, and intra-rater variability is often reported. To overcome these challenges, we propose an interpretable deep learning pipeline, named Multi-Biomarker Histomorphology Discoverer (Multi-Beholder), to predict the status of five biomarkers in LGG using only hematoxylin and eosin-stained whole slide images. Specifically, Multi-Beholder incorporates one-class classification into the multiple instance learning framework to achieve accurate instance-level pseudo-labeling, thereby complementing slide-level labels and improving prediction performance. Multi-Beholder demonstrates high performance on two LGG cohorts with diverse races and scanning protocols, with area under the receiver operating characteristic curve up to 0.973 on the internal-validated TCGA-LGG dataset and 0.820 on the external-validated Xiangya cohort. Moreover, the interpretability of Multi-Beholder allows for discovering quantitative and qualitative correlations between biomarker status and histomorphology characteristics. Our pipeline not only provides a novel approach for biomarker prediction, enhancing the applicability of molecular treatments for LGG patients but also facilitates the discovery of new mechanisms in molecular functionality and LGG progression. Code can be accessed athttps://github.com/Vison307/Multi-Beholder.
Zijie Fang, Yifeng Wang 0001, Yang Chen 0036, Changjing Cai, Yiyang Lin, Zhi Wang 0001, Shan Zeng, Yongbing Zhang 0002
IEEE Trans. Comput. Biol. Bioinform.12
2026 SeCoMIL: Semantic Anchor-Based Context-Aware Multiple Instance Learning for Whole Slide Image Classification
abstract
Context-aware Multiple Instance Learning (MIL) is gaining popularity in Whole Slide Image (WSI) classification. Existing methods typically convert instances in a WSI into one-dimensional sequences and learn the long-range contextual dependencies among instances. However, due to the extremely large size of WSIs and the morphological similarities within tissue structures, the enormous number of redundant instances significantly increases computational overhead in the context learning paradigm. Additionally, the rearrangement of instances into one dimension loses the inherent spatial information involved in image patches, further compromising the classification performance of pathological images. Consequently, efficiently modeling contextual dependencies in WSIs remains a crucial challenge. In this paper, we propose a novel Semantic Anchor-based Context-aware Multiple Instance Learning (SeCoMIL) framework. This framework partitions the WSI into a series of regions and encodes the coordinates of instances within these regions to preserve their spatial relationships. Subsequently, SeCoMIL identifies the most representative instances from each region as semantic anchors. By capturing both the local context around these anchors and the global context across different anchors, the framework efficiently summarizes the critical pathological information of the WSI, enabling precise classification. Extensive experiments on four public datasets (CAMELYON16, CAMELYON17, TCGA-NSCLC, and TCGA-RCC) demonstrate the robustness of our method, with superior performance compared to state-of-the-art methods.
Shenjin Huang, GuoJun Liu, Linghan Cai, Hailun Cheng, Zichun Huang, Yongbing Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.6
2026 Mining Temporal Priors for Template-Generated Video Compression
abstract
The popularity of template-generated videos has recently experienced a significant increase on social media platforms. In general, videos from the same template share similar temporal characteristics, which are unfortunately ignored in the current compression schemes. In view of this, we aim to examine how such temporal priors from templates can be effectively utilized during the compression process for template-generated videos. First, a comprehensive statistical analysis is conducted, revealing that the coding decisions, including the merge, non-affine, and motion information, across template-generated videos are strongly correlated. Subsequently, leveraging such correlations as prior knowledge, a simple yet effective prior-driven compression scheme for template-generated videos is proposed. In particular, a mode decision pruning algorithm is devised to dynamically skip unnecessarily advanced motion vector prediction (AMVP) or affine AMVP decisions. Moreover, an improved AMVP motion estimation algorithm is applied to further accelerate reference frame selection and the motion estimation process. Experimental results on the versatile video coding (VVC) platform VTM-23.0 demonstrate that the proposed scheme achieves moderate time reductions of 14.31% and 14.99% under the Low-Delay P (LDP) and Low-Delay B (LDB) configurations, respectively, while maintaining negligible increases in Bjøntegaard Delta Rate (BD-Rate) of 0.15% and 0.18%, respectively.
Feng Xing, Yingwen Zhang, Meng Wang 0017, Hengyu Man, Yongbing Zhang 0002, Shiqi Wang 0001, Xiaopeng Fan 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2026 Uncertainty-Aware Survival Analysis With Dirichlet Distribution for Multi-Scale Pathology and Genomics
abstract
Over the last few decades, the integration of AI-driven computational techniques into digital pathology has revolutionized survival prediction tasks. However, most existing methods in survival analysis discretize the entire survival period into predefined intervals, overlooking the inherent uncertainty in event occurrence and the heterogeneity of patient survival times. The censored data further exacerbate these challenges, amplifying uncertainty and variability. To address these limitations, we introduce the Dirichlet distribution to model discretized outputs as continuous probability distributions, providing a more accurate representation of uncertainty awareness. Building upon this foundation, we propose a universal multi-modal survival analysis loss function that leverages uncertainty-driven fusion. Our Uncertainty-Aware Multi-Modal Survival Analysis (UMSA) framework further explores the interactions between multi-scale pathological images and genomic data, providing promising insights into multi-modal survival analysis. Experimental evaluations on five publicly available datasets demonstrate that UMSA achieves state-of-the-art performance, validating its effectiveness and scalability in survival prediction tasks.
Songhan Jiang, Linghan Cai, Zhengyu Gan, Yifeng Wang 0001, Guo Tang, Yongbing Zhang 0002
IEEE Trans. Medical Imaging6
2025 Efficient Self-Supervised Video Hashing with Selective State Spaces
abstract
Self-supervised video hashing (SSVH) is a practical task in video indexing and retrieval. Although Transformers are predominant in SSVH for their impressive temporal modeling capabilities, they often suffer from computational and memory inefficiencies. Drawing inspiration from Mamba, an advanced state-space model, we explore its potential in SSVH to achieve a better balance between efficacy and efficiency. We introduce S5VH, a Mamba-based video hashing model with an improved self-supervised learning paradigm. Specifically, we design bidirectional Mamba layers for both the encoder and decoder, which are effective and efficient in capturing temporal relationships thanks to the data-dependent selective scanning mechanism with linear complexity. In our learning strategy, we transform global semantics in the feature space into semantically consistent and discriminative hash centers, followed by a center alignment loss as a global learning signal. Our self-local-global (SLG) paradigm significantly improves learning efficiency, leading to faster and better convergence. Extensive experiments demonstrate S5VH's improvements over state-of-the-art methods, superior transferability, and scalable advantages in inference efficiency.
Jinpeng Wang 0002, Niu Lian, Jun Li 0131, Bin Chen 0011, Yongbing Zhang 0002, Shutao Xia
AAAI7
2025 OT-StainNet: Optimal Transport Driven Semantic Matching for Weakly Paired H&E-to-IHC Stain Transfer
abstract
Immunohistochemistry (IHC) examination is essential for characterizing tumor subtypes, providing prognostic information, and developing personalized treatment plans. However, IHC staining preparation is more complex and expensive compared to Hematoxylin and Eosin (H&E) staining, limiting its widespread clinical application. Transforming H&E images into IHC images presents a promising solution. In this paper, we propose OT-StainNet, a novel virtual IHC staining method. OT-StainNet employs a pre-trained diffusion model with richer prior knowledge as the generator and fine-tunes it with LoRA adapters through adversarial training. Given that adjacent images of the same tissue stained with H&E and IHC are not precisely aligned at the pixel level, existing methods struggle to fully utilize the supervisory information from weakly paired IHC images. To address this issue, we propose an optimal transport-driven semantic matching (OTSM) mechanism, establishing accurate semantic correspondences between H&E-IHC image pairs. By leveraging the real IHC features obtained through the OTSM mechanism, we design a semantic consistency constraint (SCC) to ensure that the correlations among virtual IHC features remain consistent with those among real IHC features, thereby preserving valuable correlation information during stain transfer. We validate OT-StainNet using four types of IHC staining across two datasets. Extensive experiments demonstrate the effectiveness of our method compared to state-of-the-art approaches.
Xianchao Guan, Yifeng Wang 0001, Ye Zhang 0043, Zheng Zhang 0006, Yongbing Zhang 0002
AAAI5
2025 Category Prompt Mamba Network for Nuclei Segmentation and Classification
abstract
Nuclei segmentation and classification provide an essential basis for tumor immune microenvironment analysis. The previous nuclei segmentation and classification models require splitting large images into smaller patches for training, leading to two significant issues. First, nuclei at the borders of adjacent patches often misalign during inference. Second, this patch-based approach significantly increases the model's training and inference time. Recently, Mamba has garnered attention for its ability to model large-scale images with linear time complexity and low memory consumption. It offers a promising solution for training nuclei segmentation and classification models on full-sized images. However, the Mamba orientation-based scanning method lacks account for category-specific features, resulting in suboptimal performance in scenarios with imbalanced class distributions. To address these challenges, this paper introduces a novel scanning strategy based on category probability sorting, which independently ranks and scans features for each category according to confidence from high to low. This approach enhances the feature representation of uncertain samples and mitigates the issues caused by imbalanced distributions. Extensive experiments conducted on four public datasets demonstrate that our method outperforms state-of-the-art approaches, delivering superior performance in nuclei segmentation and classification tasks.
Ye Zhang 0043, Zijie Fang, Yifeng Wang 0001, Lingbo Zhang, Xianchao Guan, Yongbing Zhang 0002
AAAI6
2025 CA-GAN: Context-Aware Generative Adversarial Networks for Pathological Image Super-Resolution
abstract
High-quality pathology images are essential for accurate clinical diagnosis and treatment. However, acquiring high-resolution (HR) pathology images is often hindered by equipment limitations, limited expert availability, and complex slide preparation procedures. Image super-resolution (SR), which reconstructs HR images from low-resolution (LR) inputs, offers a practical solution. However, most existing SR methods are designed for natural images and often struggle to capture the distinct structural characteristics of pathology data. In this paper, we propose Context-Aware Generative Adversarial Network(CAGAN), a novel SR framework tailored for pathological images. It introduces a context path that effectively leverages the rich spatial context in whole slide images (WSIs) while maintaining computational efficiency. In addition, considering the significant differences in staining patterns and reconstruction difficulty between the nucleus and cytoplasm, we propose a Nucleus-Enhanced Hematoxylin Channel (NEHC) loss. This loss imposes targeted constraints on nuclei to better preserve morphological consistency. Experiments on two pathological datasets demonstrate that CA-GAN achieves state-of-the-art performance in both quantitative metrics and perceptual quality. Code will be available soon.
Zhiyuan Fan, Xianchao Guan, Yifeng Wang 0001, Yongbing Zhang 0002
BIBM4
2025 Fast and Accurate Gigapixel Pathological Image Classification with Hierarchical Distillation Multi-Instance Learning
abstract
Although multi-instance learning (MIL) has succeeded in pathological image classification, it faces the challenge of high inference costs due to processing numerous patches from gigapixel whole slide images (WSIs). To address this, we propose HDMIL, a hierarchical distillation multi-instance learning framework that achieves fast and accurate classification by eliminating irrelevant patches. HD-MIL consists of two key components: the dynamic multi-instance network (DMIN) and the lightweight instance pre-screening network (LIPN). DMIN operates on high-resolution WSIs, while LIPN operates on the corresponding low-resolution counterparts. During training, DMIN are trained for WSI classification while generating attention-score-based masks that indicate irrelevant patches. These masks then guide the training of LIPN to predict the relevance of each low-resolution patch. During testing, LIPN first determines the useful regions within low-resolution WSIs, which indirectly enables us to eliminate irrelevant regions in high-resolution WSIs, thereby reducing inference time without causing performance degradation. In addition, we further design the first Chebyshev-polynomials-based Kolmogorov-Arnold classifier in computational pathology, which enhances the performance of HDMIL through learnable activation layers. Extensive experiments on three public datasets demonstrate that HDMIL outperforms previous state-of-the-art methods, e.g., achieving improvements of 3.13% in AUC while reducing inference time by 28.6% on the Camelyon16 dataset. The project is available at https://github.com/JiuyangDong/HDMIL.
Jiuyang Dong, Junjun Jiang, Kui Jiang, Jiahan Li, Yongbing Zhang 0002
CVPR5
2025 An Efficient Hidden Markov Model-Based Sample Adaptive Offset Mode Decision Algorithm for Versatile Video Coding
abstract
This paper proposes a highly efficient sample adaptive offset (SAO) mode decision algorithm. By leveraging both the directional correlations between the SAO and intra-prediction decisions, and the SAO decisions' spatial correlations, the SAO mode candidates are effectively pruned during the rate-distortion optimization process, accelerating the SAO encoding process with negligible BD-rate loss.
Feng Xing, Yingwen Zhang, Meng Wang 0017, Hengyu Man, Yongbing Zhang 0002, Shiqi Wang 0001, Xiaopeng Fan 0001
DCC5
2025 Correlated Multiple IHC Virtual Staining for Breast Histopathological Images
abstract
Immunohistochemistry (IHC) examination is essential for determining breast cancer subtypes and provides critical prognostic factors to guide treatment decisions. However, the complex and expensive preparation of IHC staining limits its widespread use in clinical practice. Recent advancements in generative models have introduced virtual staining as a promising alternative, yet obtaining pixel-level paired data in clinical settings remains a significant challenge. In this paper, we propose Multi-IHC Net, which utilizes unpaired data to simultaneously generate Ki67, ER, PR, and HER2 images from H&E-stained breast tissue. Specifically, a general encoder extracts generalized features from H&E images, while interactive decoders reconstruct the four types of IHC images. Additionally, a feature alignment module also models the correlations among the different IHC stains. To enhance accuracy, we introduce a pathology consistency mechanism between H&E and adjacent IHC images. Extensive experiments demonstrate the superiority of our method compared to state-of-the-art approaches.
Xianchao Guan, Zheng Zhang 0006, Yifeng Wang 0001, Ye Zhang 0043, Danling Jiang, Yongbing Zhang 0002
ICASSP6
2025 Multi-scale Context Intertwining for Panoramic Renal Pathology Segmentation
abstract
Panoramic segmentation of renal pathological tissues plays a crucial role in diagnosing renal carcinoma and other kidney-related diseases. The multi-scale nature of kidney tissues, which requires different magnification levels for accurate analysis, presents a significant challenge for segmentation models. In this work, we propose a Multi-scale Context Intertwining Network (MCINet) to address this issue. Our approach utilizes an auxiliary interaction network to enhance feature interaction between different scales and generate pseudo-labels for unannotated structures. By incorporating exponential moving average strategies, we ensure seamless feature integration across scales. Extensive experiments demonstrate that MCINet outperforms state-of-the-art models in key metrics such as Dice and Hausdorff Distance, proving its efficacy in renal tissue segmentation tasks.
Ye Zhang 0043, Xianchao Guan, Hengrui Li, Xiangming Yan, Ziyue Wang 0005, Yongbing Zhang 0002
ICASSP6
2025 AMKD: Adaptive Multi-modality Knowledge Distillation for Pathological Survival Analysis
Yangfan Xu, Linghan Cai, Yifeng Wang 0001, Hailun Cheng, Fengchun Liu, Runming Wang, Yongbing Zhang 0002
ICIC (27)7
2025 ROMA: Regularization for Out-of-distribution Detection with Masked Autoencoders
abstract
Existing out-of-distribution (OOD) detection methods without outlier exposure learn effective in-distribution (ID) representations distinguishable for OOD samples, showing promising performance in many OOD detection tasks. However, we observe a performance degradation in some challenging OOD detection scenarios, where pre-trained networks tend to perform worse during the fine-tuning process, exhibiting the over-fitting of ID representations. Motivated by this, we propose the critical task of hidden OOD detection, wherein ID representations offer limited, or even counterproductive, assistance in OOD detection. To address this issue, we introduce a novel Regularization framework for OOD detection with Masked Autoencoders (ROMA), which notably outperforms previous OOD detection methods in hidden OOD detection. Moreover, the robustness of ROMA is further evidenced by its state-of-the-art performance on benchmarks for other challenging OOD detection tasks.
Xiaochen Feng, Hao Sha 0001, Yongbing Zhang 0002
ICME4
2025 Relation-Aware Graph Attention Network for Nuclei Classification
abstract
Nuclei classification plays a pivotal role in pathological research. Recent advances in graph neural networks (GNNs) have shown great promise in modeling cell-cell interactions. However, many existing methods overlook tissue context, which is crucial for accurate nuclei identification, as nuclei exhibit distinct patterns within specific tissue structures. To address this limitation, we propose a novel Relation-Aware Graph AT-tention network (RAGAT) that effectively leverages nucleus-related features for precise classification. RAGAT constructs a cell graph based on spatial proximity and visual feature similarity, while also introducing a tissue-aware graph by sampling regions around each nucleus to capture the tissue microenvironment and depict local cellular contexts. Furthermore, RAGAT employs a hybrid graph attention module to integrate cell-cell and tissue-cell interactions, enabling a comprehensive understanding of the nuclear context. Experimental results on three benchmark datasets demonstrate that our method significantly outperforms state-of-the-art approaches, offering valuable insight into the analysis of nuclear microenvironments. Our code is available at https://github.com/lingboboo/RAGAT.
Lingbo Zhang, Ye Zhang 0043, Linghan Cai, Xianchao Guan, Kai Zhang 0012, Yongbing Zhang 0002
ICME6
2025 The Four Color Theorem for Cell Instance Segmentation
abstract
Cell instance segmentation is critical to analyzing biomedical images, yet accurately distinguishing tightly touching cells remains a persistent challenge. Existing instance segmentation frameworks, including detection-based, contour-based, and distance mapping-based approaches, have made significant progress, but balancing model performance with computational efficiency remains an open problem. In this paper, we propose a novel cell instance segmentation method inspired by the four-color theorem. By conceptualizing cells as countries and tissues as oceans, we introduce a four-color encoding scheme that ensures adjacent instances receive distinct labels. This reformulation transforms instance segmentation into a constrained semantic segmentation problem with only four predicted classes, substantially simplifying the instance differentiation process. To solve the training instability caused by the non-uniqueness of four-color encoding, we design an asymptotic training strategy and encoding transformation method. Extensive experiments on various modes demonstrate our approach achieves state-of-the-art performance. The code is available at https://github.com/zhangye-zoe/FCIS.
Ye Zhang 0043, Yifeng Wang 0001, Ziyue Wang 0005, Yongbing Zhang 0002, Jianxu Chen 0001
ICML6
2025 Prototype-Guided Cross-Modal Knowledge Enhancement for Adaptive Survival Prediction
Fengchun Liu, Linghan Cai, Zhikang Wang, Zhiyuan Fan, Jin-gang Yu, Hao Chen 0011, Yongbing Zhang 0002
MICCAI (6)7
2025 Counting by Points: Density-Guided Weakly-Supervised Nuclei Segmentation in Histopathological Images
Lingbo Zhang, Bingqian Sun, Linghan Cai, Yifeng Wang 0001, Ye Zhang 0043, Songhan Jiang, Kai Zhang 0012, Yongbing Zhang 0002
ACM Multimedia8
2025 Self-supervised Scalable Deep Compressed Sensing
Bin Chen 0006, Xuanyu Zhang 0003, Yongbing Zhang 0002, Jian Zhang 0018
Int. J. Comput. Vis.4
2025 AttriMIL: Revisiting attention-based multiple instance learning for whole-slide pathological image classification from a perspective of instance attributes
abstract
Multiple instance learning (MIL) is a powerful approach for whole-slide pathological image (WSI) analysis, particularly suited for processing gigapixel-resolution images with slide-level labels. Recent attention-based MIL architectures have significantly advanced weakly supervised WSI classification, facilitating both clinical diagnosis and localization of disease-positive regions. However, these methods often face challenges in differentiating between instances, leading to tissue misidentification and a potential degradation in classification performance. To address these limitations, we propose AttriMIL, an attribute-aware multiple instance learning framework. By dissecting the computational flow of attention-based MIL models, we introduce a multi-branch attribute scoring mechanism that quantifies the pathological attributes of individual instances. Leveraging these quantified attributes, we further establish region-wise and slide-wise attribute constraints to dynamically model instance correlations both within and across slides during training. These constraints encourage the network to capture intrinsic spatial patterns and semantic similarities between image patches, thereby enhancing its ability to distinguish subtle tissue variations and sensitivity to challenging instances. To fully exploit the two constraints, we further develop a pathology adaptive learning technique to optimize pre-trained feature extractors, enabling the model to efficiently gather task-specific features. Extensive experiments on five public datasets demonstrate that AttriMIL consistently outperforms state-of-the-art methods across various dimensions, including bag classification accuracy, generalization ability, and disease-positive region localization. The implementation code is available at https://github.com/MedCAI/AttriMIL.
Linghan Cai, Shenjin Huang, Ye Zhang 0043, Jinpeng Lu, Yongbing Zhang 0002
Medical Image Anal.5
2025 AMH-Net: Adaptive Multi-Band Hybrid-Aware Network for Emotion Recognition in Speech
abstract
Speech emotion recognition (SER) technology analyzes speech signals to automatically identify the speaker's emotional state. However, existing methods overlook feature extraction based on human acoustic characteristics. In this paper, we propose AMH-Net, an Adaptive Multi-band Hybridaware Network designed for SER. The model leverages formant (F1, F2, F3) from speech science, which describe the human vocal tract, to partition speech signals into multiple frequency bands. A variable-depth residual network structure is employed for more precise extraction of emotional characteristics. In addition, a hybrid attention mechanism is integrated to combine information, resulting in a more comprehensive emotional representation. Experimental evaluations of six diverse datasets show that AMH-Net outperforms state-of-the-art methods, achieving improvements of 2.11% and 2.64% in average UAR and WAR, respectively, on each corpus. The code is publicly available at https://github.com/hengruili1997/AMH-net
Hengrui Li, Yongbing Zhang 0002, Shaohui Liu
IEEE Signal Process. Lett.2
2025 Focus Your Attention: Multiple Instance Learning With Attention Modification for Whole Slide Pathological Image Classification
abstract
Computer-aided pathology diagnosis based on whole slide images, which is often formulated as a weakly supervised multiple instance learning (MIL) paradigm. Current approaches generally employ attention mechanisms to aggregate instance-level features. However, the weakly supervised signal and the imbalanced instance distribution often lead to inaccurate attention localization, compromising the performance and generalization capability of the MIL framework. To address these problems, this paper presents a novel MIL framework called FAMIL that focuses on inaccurate attention and refines them. FAMIL adopts a dual-branch structure and incorporates two innovative online data augmentation strategies: attention-based Mixup (ABMix) and attention-based Masking (ABMask). ABMix emphasizes the significance of positive instances, generalizing Mixup in the MIL scenarios, while ABMask flexibly identifies challenging positive instances to optimize the feature representation. Moreover, these two methods are plug-and-play and can be easily embedded into attention-based MIL methods. Extensive experiments on three public benchmarks demonstrate the superiority of our FAMIL, outperforming current state-of-the-art methods. The test AUC for the binary tumor classification can be up to 92.61% over CAMELYON16. And the AUC over the cancer subtype classification can be up to 93.81% and 98.41% on TCGA-NSCLC and TCGA-RCC datasets, respectively.
Hailun Cheng, Shenjin Huang, Linghan Cai, Yangfan Xu, Runming Wang, Yongbing Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.6
2025 DAWN: Domain-Adaptive Weakly Supervised Nuclei Segmentation via Cross-Task Interactions
abstract
Weakly supervised segmentation methods have garnered considerable attention due to their potential to alleviate the need for labor-intensive pixel-level annotations during model training. Traditional weakly supervised nuclei segmentation approaches typically involve a two-stage process: pseudo-label generation followed by network training. The performance of these methods is highly dependent on the quality of the generated pseudo-labels, which can limit their effectiveness. In this paper, we propose a novel domain-adaptive weakly supervised nuclei segmentation framework that addresses the challenge of pseudo-label generation through cross-task interaction strategies. Specifically, our approach leverages weakly annotated data to train an auxiliary detection task, which facilitates domain adaptation of the segmentation network. To improve the efficiency of domain adaptation, we introduce a consistent feature constraint module that integrates prior knowledge from the source domain. Additionally, we develop methods for pseudo-label optimization and interactive training to enhance domain transfer capabilities. We validate the effectiveness of our proposed method through extensive comparative and ablation experiments conducted on six datasets. The results demonstrate that our approach outperforms existing weakly supervised methods and achieves performance comparable to or exceeding that of fully supervised methods. Our code is available athttps://github.com/zhangye-zoe/DAWN.
Ye Zhang 0043, Yifeng Wang 0001, Zijie Fang, Hao Bian, Linghan Cai, Ziyue Wang 0005, Yongbing Zhang 0002
IEEE Trans. Circuits Syst. Video Technol.7
2025 Context-Aware Contrastive Learning for Virtual IHC Staining With Inconsistent Image Pairs
abstract
In recent years, virtual immunohistochemical (IHC) staining, which converts hematoxylin and eosin (H&E) images into IHC images, has emerged as a promising technology in digital histopathology. Most existing methods rely on paired H&E and IHC patches extracted from adjacent tissue sections for supervised training. However, tissue misalignment and tissue loss between adjacent sections lead to inconsistent training pairs, limiting the models' ability to produce accurate staining results. To address this issue, we propose ConCLR, a two-stage virtual IHC staining framework based on context-aware contrastive learning, designed to handle inconsistently paired patches. Our method is built on the assumption that for a given mini-patch in the H&E patch, there may exist a corresponding mini-patch in the reference IHC patch exhibiting a similar Pos/Neg pathological pattern. If such a mini-patch exists, it is typically located spatially close to the H&E mini-patch due to the local consistency of tissue structure. In the first stage, we leverage this assumption to design a similarity-guided mini-patch sampling (SGMS) module. For each mini-patch anchor in the staining results, SGMS searches within the real IHC patch to find the most similar mini-patch to serve as the positive sample for contrastive learning, enabling effective supervision despite mild tissue misalignment. In the second stage, we design a context-aware adaptive refinement module, which addresses significant inconsistencies between training pairs caused by potential tissue loss, by expanding the search range of positive samples to include neighboring patches. Extensive experiments on two network backbones across four virtual IHC staining tasks demonstrate the effectiveness of our ConCLR. Evaluations include qualitative and quantitative assessments of staining results, as well as downstream diagnostic performance. In addition to experiments on existing public datasets, we collected a PanCK-NSCLC dataset by acquiring H&E and pan-cytokeratin staining images from the same lung tissue sections via destaining and restaining. This dataset offers significantly improved tissue alignment compared to those derived from adjacent sections, with the aim of facilitating further progress in virtual IHC staining.
Jiahan Li, Jiuyang Dong, Yongbing Zhang 0002, Haiyu Zhou, Xiaopeng Fan 0001
IEEE Trans. Image Process.4
2025 SEINE: Structure Encoding and Interaction Network for Nuclei Instance Segmentation
abstract
Nuclei instance segmentation in histopathological images is crucial for biological analysis and cancer diagnosis. However, it faces two significant challenges: (1) poorly stained nuclei can lead to under-segmentation, as the background may be mistakenly identified as the foreground; and (2) deep textures within nuclei often result in fragmented instance predictions, as these textures can be misinterpreted as contours. To address these problems, this paper proposes a Structure Encoding and Interaction NEtwork, termed SEINE, which develops the nuclei structure modeling scheme and takes advantage of the similarity between nuclei structure to improve the integrality of instance segmentation. Specifically, SEINE introduces a contour-based structure encoding mechanism that integrates the correlation between nuclear structure and semantics, enabling a more accurate structural representation. Building on this encoding, we propose a structure-guided attention module, which uses clear nuclei as prototypes to guide the structural learning of unclear nuclei, thereby addressing the under-segmentation problem. Additionally, a position enhancement strategy applies a centroid distance constraint to reduce contour prediction errors, effectively mitigating fragmented instance segmentation. Extensive experiments demonstrate the effectiveness of SEINE, achieving state-of-the-art performance across four benchmark datasets.
Ye Zhang 0043, Linghan Cai, Ziyue Wang 0005, Yongbing Zhang 0002
IEEE J. Biomed. Health Informatics4
2025 Disentangled Pseudo-Bag Augmentation for Whole Slide Image Multiple Instance Learning
abstract
As the predominant approach for pathological whole slide image (WSI) classification, multiple instance learning (MIL) methods struggle with limited labeled WSIs. Although MIL has achieved notable progress with pseudo-bag-oriented augmentation methods, their effectiveness is often constrained by noisy pseudo-labels and low-quality pseudo-bags. To overcome these problems, we revisit the use of pseudo-bags for WSI data augmentation and propose a new pseudo-bag generation paradigm, dubbed DPBAug. Its distinctive features can be summarized as: i) We develop an intra-slide pseudo-bag generation module, which separates the heterogeneous instances within each slide through phenotype partitioning. Moreover, to ensure accurate label inheritance when generating pseudo-bags, we propose an instance sampling algorithm with replacement. ii) An inter-slide pseudo-bag fusion module is designed to integrate heterogeneous information across multiple WSIs, producing high-quality training samples that better leverage the potential of neural networks. iii) A pseudo-bag memory update module prioritizes valuable synthetic pseudo-bags. This further enhances the network's classification performance. Extensive experiments demonstrate that DPBAug surpasses existing augmentation methods, enhancing the classification performance and reliability of multiple MIL baselines across various public datasets. DPBAug also improves the generalization and data efficiency of existing MIL methods, facilitating their adoption in clinical practice and rare cancer research The project is available at: https://github.com/JiuyangDong/DPBAug.
Jiuyang Dong, Junjun Jiang, Kui Jiang, Jiahan Li, Linghan Cai, Yongbing Zhang 0002
IEEE Trans. Medical Imaging6
2025 HisynSeg: Weakly-Supervised Histopathological Image Segmentation via Image-Mixing Synthesis and Consistency Regularization
abstract
Tissue semantic segmentation is one of the key tasks in computational pathology. To avoid the expensive and laborious acquisition of pixel-level annotations, a wide range of studies attempt to adopt the class activation map (CAM), a weakly-supervised learning scheme, to achieve pixel-level tissue segmentation. However, CAM-based methods are prone to suffer from under-activation and over-activation issues, leading to poor segmentation performance. To address this problem, we propose a novel weakly-supervised semantic segmentation framework for histopathological images based on image-mixing synthesis and consistency regularization, dubbed HisynSeg. Specifically, synthesized histopathological images with pixel-level masks are generated for fully-supervised model training, where two synthesis strategies are proposed based on Mosaic transformation and Bézier mask generation. Besides, an image filtering module is developed to guarantee the authenticity of the synthesized images. In order to further avoid the model overfitting to the occasional synthesis artifacts, we additionally propose a novel self-supervised consistency regularization, which enables the real images without segmentation masks to supervise the training of the segmentation model. By integrating the proposed techniques, the HisynSeg framework successfully transforms the weakly-supervised semantic segmentation problem into a fully-supervised one, greatly improving the segmentation accuracy. Experimental results on three datasets prove that the proposed method achieves a state-of-the-art performance. Code is available at https://github.com/Vison307/HisynSeg.
Zijie Fang, Yifeng Wang 0001, Peizhang Xie, Zhi Wang 0001, Yongbing Zhang 0002
IEEE Trans. Medical Imaging5
2025 Supervised Information Mining From Weakly Paired Images for Breast IHC Virtual Staining
abstract
Immunohistochemistry (IHC) examination is essential to determine the tumour subtypes, provide key prognostic factors, and develop personalized treatment plans for breast cancer. However, compared to Hematoxylin and Eosin (H&E) staining, the preparation process of IHC staining is more complex and expensive, which limits its application in clinical practice. Therefore, H&E to IHC stain transfer may be an ideal solution to obtain IHC staining. To ensure high transferring quality, it would be much more desirable to exploit the supervised information between adjacent layer images of the same tissue, which are stained by H&E and IHC stainings, respectively. Nevertheless, adjacent layer tissue images are not accurately paired at the pixel level, which poses significant challenges to network training. To address this problem, we propose a generative adversarial network for breast IHC virtual staining, which contains an optimal transport-based supervised information mining (OT-SIM) mechanism and a pathological correlation-based supervised information mining (PC-SIM) mechanism. The OT-SIM guides the network in mining matching consistency between H&E images and the adjacent layer's real IHC images, providing as much instance-level supervision as possible. The PC-SIM further explores the consistency between the correlation among virtual IHC images and the correlation among real IHC images, providing batch-level supervision. Extensive experiments show the superiority of our method on two breast tissue benchmark datasets compared to the state-of-the-art methods both quantitatively and qualitatively. The code is available at https://github.com/xianchaoguan/SIM-GAN.
Xianchao Guan, Zheng Zhang 0006, Yifeng Wang 0001, Yueheng Li, Yongbing Zhang 0002
IEEE Trans. Medical Imaging5
2025 A Multi-Perspective Self-Supervised Generative Adversarial Network for FS to FFPE Stain Transfer
abstract
In clinical practice, frozen section (FS) images can be utilized to obtain the immediate pathological results of the patients in operation due to their fast production speed. However, compared with the formalin-fixed and paraffin-embedded (FFPE) images, the FS images greatly suffer from poor quality. Thus, it is of great significance to transfer the FS image to the FFPE one, which enables pathologists to observe high-quality images in operation. However, obtaining the paired FS and FFPE images is quite hard, so it is difficult to obtain accurate results using supervised methods. Apart from this, the FS to FFPE stain transfer faces many challenges. Firstly, the number and position of nuclei scattered throughout the image are hard to maintain during the transfer process. Secondly, transferring the blurry FS images to the clear FFPE ones is quite challenging. Thirdly, compared with the center regions of each patch, the edge regions are harder to transfer. To overcome these problems, a multi-perspective self-supervised GAN, incorporating three auxiliary tasks, is proposed to improve the performance of FS to FFPE stain transfer. Concretely, a nucleus consistency constraint is designed to enable the high-fidelity of nuclei, an FFPE guided image deblurring is proposed for improving the clarity, and a multi-field-of-view consistency constraint is designed to better generate the edge regions. Objective indicators and pathologists' evaluation for experiments on the five datasets across different countries have demonstrated the effectiveness of our method. In addition, the validation in the downstream task of microsatellite instability prediction has also proved the performance improvement by transferring the FS images to FFPE ones. Our code link is https://github.com/linyiyang98/Self-Supervised-FS2FFPE.git.
Yiyang Lin, Yifeng Wang 0001, Zijie Fang, Xianchao Guan, Danling Jiang, Yongbing Zhang 0002
IEEE Trans. Medical Imaging7
2024 MamMIL: Multiple Instance Learning for Whole Slide Images with State Space Models
abstract
Recently, pathological diagnosis has achieved superior performance by combining deep learning models with the multiple instance learning (MIL) framework using whole slide images (WSIs). However, the giga-pixeled nature of WSIs poses a great challenge for efficient MIL. Existing studies either do not consider global dependencies among instances, or use approximations such as linear attentions to model the pair-to-pair instance interactions, which inevitably brings performance bottlenecks. To tackle this challenge, we propose a framework named MamMIL for WSI analysis by cooperating the selective structured state space model (i.e., Mamba) with MIL, enabling the modeling of global instance dependencies while maintaining linear complexity. Specifically, considering the irregularity of the tissue regions in WSIs, we represent each WSI as an undirected graph. To address the problem that Mamba can only process 1D sequences, we further propose a topology-aware scanning mechanism to serialize the WSI graphs while preserving the topological relationships among the instances. Finally, in order to further perceive the topological structures among the instances and incorporate short-range feature interactions, we propose an instance aggregation block based on graph neural networks. Experiments show that MamMIL can achieve advanced performance than the state-of-the-art frameworks. The code can be accessed at https://github.com/Vison307/MamMIL.
Zijie Fang, Yifeng Wang 0001, Ye Zhang 0043, Zhi Wang 0001, Jian Zhang 0018, Xiangyang Ji, Yongbing Zhang 0002
BIBM7
2024 Virtual Immunohistochemistry Staining for Histological Images Assisted by Weakly-supervised Learning
abstract
Recently, virtual staining technology has greatly promoted the advancement of histopathology. Despite the practical successes achieved, the outstanding performance of most virtual staining methods relies on hard-to-obtain paired images in training. In this paper, we propose a method for virtual immunohistochemistry (IHC) staining, named confusion-GAN, which does not require paired images and can achieve comparable performance to supervised algorithms. Specifically, we propose a multi-branch discriminator, which judges if the features of generated images can be embedded into the feature pool of target domain images, to improve the visual quality of generated images. Meanwhile, we also propose a novel patch-level pathology information extractor, which is assisted by multiple instance learning, to ensure pathological consistency during virtual staining. Extensive experiments were conducted on three types of IHC images, including a high-resolution hepatocel-lular carcinoma immunohistochemical dataset proposed by us. The results demonstrated that our proposed confusion-GAN can generate highly realistic images that are capable of deceiving even experienced pathologists. Furthermore, compared to using H&E images directly, the downstream diagnosis achieved higher accuracy when using images generated by confusion-GAN. Our dataset and codes will be available at https://github.com/jiahanli2022/confusion-GAN.
Jiahan Li, Jiuyang Dong, Shenjin Huang, Junjun Jiang, Xiaopeng Fan 0001, Yongbing Zhang 0002
CVPR7
2024 Multimodal Cross-Task Interaction for Survival Analysis in Whole Slide Pathological Images
Songhan Jiang, Zhengyu Gan, Linghan Cai, Yifeng Wang 0001, Yongbing Zhang 0002
MICCAI (4)5
2024 Exploiting Supervision Information in Weakly Paired Images for IHC Virtual Staining
Yueheng Li, Xianchao Guan, Yifeng Wang 0001, Yongbing Zhang 0002
MICCAI (4)4
2024 H2ASeg: Hierarchical Adaptive Interaction and Weighting Network for Tumor Segmentation in PET/CT Images
Jinpeng Lu, Jingyun Chen, Linghan Cai, Songhan Jiang, Yongbing Zhang 0002
MICCAI (8)5
2024 Dynamic Pseudo Label Optimization in Point-Supervised Nuclei Segmentation
Ziyue Wang 0005, Ye Zhang 0043, Yifeng Wang 0001, Linghan Cai, Yongbing Zhang 0002
MICCAI (8)5
2024 GS-Hider: Hiding Messages into 3D Gaussian Splatting
abstract
3D Gaussian Splatting (3DGS) has already become the emerging research focus in the fields of 3D scene reconstruction and novel view synthesis. Given that training a 3DGS requires a significant amount of time and computational cost, it is crucial to protect the copyright, integrity, and privacy of such 3D assets. Steganography, as a crucial technique for encrypted transmission and copyright protection, has been extensively studied. However, it still lacks profound exploration targeted at 3DGS. Unlike its predecessor NeRF, 3DGS possesses two distinct features: 1) explicit 3D representation; and 2) real-time rendering speeds. These characteristics result in the 3DGS point cloud files being public and transparent, with each Gaussian point having a clear physical significance. Therefore, ensuring the security and fidelity of the original 3D scene while embedding information into the 3DGS point cloud files is an extremely challenging task. To solve the above-mentioned issue, we first propose a steganography framework for 3DGS, dubbed GS-Hider, which can embed 3D scenes and images into original GS point clouds in an invisible manner and accurately extract the hidden messages. Specifically, we design a coupled secured feature attribute to replace the original 3DGS's spherical harmonics coefficients and then use a scene decoder and a message decoder to disentangle the original RGB scene and the hidden message. Extensive experiments demonstrated that the proposed GS-Hider can effectively conceal multimodal messages without compromising rendering quality and possesses exceptional security, robustness, capacity, and flexibility. Our project is available at: https://xuanyuzhang21.github.io/project/gshider.
Xuanyu Zhang 0003, Jiarui Meng, Runyi Li, Zhipei Xu, Yongbing Zhang 0002, Jian Zhang 0018
NeurIPS5
2024 Domain generalization across tumor types, laboratories, and species - Insights from the 2022 edition of the Mitosis Domain Generalization Challenge
Marc Aubreville, Nikolas Stathonikos, Taryn A. Donovan, Robert Klopfleisch, Jonas Ammeling, Jonathan Ganz, Frauke Wilm, Mitko Veta, Samir Jabari, Markus Eckstein, Jonas Annuscheit, Christian Krumnow, Engin Bozaba, Sercan Cayir, Hongyan Gu, Xiang 'Anthony' Chen, Mostafa Jahanifar, Adam J. Shephard, Satoshi Kondo, Satoshi Kasai, Sujatha Kotte, Vangala Saipradeep, Maxime W. Lafarge, Viktor H. Koelzer, Ziyue Wang 0005, Yongbing Zhang 0002, Sen Yang 0006, Katharina Breininger, Christof Bertram
Medical Image Anal.26
2024 Know your orientation: A viewpoint-aware framework for polyp segmentation
abstract
Automatic polyp segmentation in endoscopic images is critical for the early diagnosis of colorectal cancer. Despite the availability of powerful segmentation models, two challenges still impede the accuracy of polyp segmentation algorithms. Firstly, during a colonoscopy, physicians frequently adjust the orientation of the colonoscope tip to capture underlying lesions, resulting in viewpoint changes in the colonoscopy images. These variations increase the diversity of polyp visual appearance, posing a challenge for learning robust polyp features. Secondly, polyps often exhibit properties similar to the surrounding tissues, leading to indistinct polyp boundaries. To address these problems, we propose a viewpoint-aware framework named VANet for precise polyp segmentation. In VANet, polyps are emphasized as a discriminative feature and thus can be localized by class activation maps in a viewpoint classification process. With these polyp locations, we design a viewpoint-aware Transformer (VAFormer) to alleviate the erosion of attention by the surrounding tissues, thereby inducing better polyp representations. Additionally, to enhance the polyp boundary perception of the network, we develop a boundary-aware Transformer (BAFormer) to encourage self-attention towards uncertain regions. As a consequence, the combination of the two modules is capable of calibrating predictions and significantly improving polyp segmentation performance. Extensive experiments on seven public datasets across six metrics demonstrate the state-of-the-art results of our method, and VANet can handle colonoscopy images in real-world scenarios effectively. The source code is available at https://github.com/1024803482/Viewpoint-Aware-Network.
Linghan Cai, Lijiang Chen, Yifeng Wang 0001, Yongbing Zhang 0002
Medical Image Anal.5
2024 D3C2-Net: Dual-Domain Deep Convolutional Coding Network for Compressive Sensing
abstract
By mapping iterative optimization algorithms into neural networks (NNs), deep unfolding networks (DUNs) exhibit well-defined and interpretable structures and achieve remarkable success in the field of compressive sensing (CS). However, most existing DUNs solely rely on the image-domain unfolding, which restricts the information transmission capacity and reconstruction flexibility, leading to their loss of image details and unsatisfactory performance. To overcome these limitations, this paper develops a dual-domain optimization framework that combines the priors of (1) image- and (2) convolutional-coding-domains and offers generality to CS and other inverse imaging tasks. By converting this optimization framework into deep NN structures, we present a Dual-Domain Deep Convolutional Coding Network (D3C2-Net), which enjoys the ability to efficiently transmit high-capacity self-adaptive convolutional features across all its unfolded stages. Our theoretical analyses and experiments on simulated and real captured data, covering 2D and 3D natural, medical, and scientific signals, demonstrate the effectiveness, practicality, superior performance, and generalization ability of our method over other competing approaches and its significant potential in achieving a balance among accuracy, complexity, and interpretability. Code is available at https://github.com/lwq20020127/D3C2-Net.
Bin Chen 0006, Shijie Zhao 0001, Bowen Du 0002, Yongbing Zhang 0002, Jian Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.6
2024 Progressive Content-Aware Coded Hyperspectral Snapshot Compressive Imaging
abstract
Hyperspectral imaging plays a pivotal role across diverse applications, like remote sensing, medicine, and cytology. The utilization of 2D sensors to acquire 3D hyperspectral images (HSIs) via a coded aperture snapshot spectral imaging (CASSI) system has proven successful, owing to its hardware-friendly implementation and fast sampling speed. Nevertheless, for less spectrally sparse scenes, the use of a single snapshot and unreasonable coded aperture design limits the efficacy of CASSI systems and renders HSI reconstruction more ill-posed, leading to compromised spatial and spectral fidelity. This paper proposes a novel Progressive Content-Aware CASSI (PCA-CASSI) framework, which progressively captures HSIs using multiple optimized content-aware coded apertures and fuses all snapshot measurements for reconstruction. By unlocking the full potential of CASSI systems and elevating their performance ceilings, this framework offers researchers new avenues for improving imaging quality. Furthermore, we develop the RndHRNet, a Range-Null space Decomposition (RND)-inspired deep unfolding network with multiple iterative phases for HSI recovery. Each unfolded recovery phase efficiently exploits the physical information within the coded apertures via explicit RND and adaptively explores the spatial-spectral correlation by dual transformer blocks. Through comprehensive experiments, our approach demonstrates superior performance compared to existing state-of-the-art methods in both the multiple- and single-shot compressive HSI imaging tasks with substantial improvements. Code is available athttps://github.com/xuanyuzhang21/PCA-CASSI.
Xuanyu Zhang 0003, Bin Chen 0006, Wenzhen Zou, Yongbing Zhang 0002, Ruiqin Xiong, Jian Zhang 0018
IEEE Trans. Circuits Syst. Video Technol.5
2024 Unsupervised Multi-Domain Progressive Stain Transfer Guided by Style Encoding Dictionary
abstract
In histopathology, the tissue slides are usually stained by common H&E stain or special stains (MAS, PAS, and PASM, etc.) to clearly show specific tissue structures. The rapid development of deep learning provides a good solution to generate virtual staining images to significantly reduce the time and labor costs associated with histochemical staining. However, most existing methods need to train a special model for every two stains, which consumes a lot of computing resources with the increasing of staining types. To address this problem, we propose an unsupervised multi-domain stain transfer method, GramGAN, which realizes the progressive transfer through cascaded Style-Guided blocks. For each Style-Guided block, we design a style encoding dictionary to characterize and store all the staining style information. In addition, we propose a Rényi entropy-based regularization term to improve the discrimination ability of different styles. The experimental results show that our method can realize accurate transferring among multiple staining styles with better performance. Furthermore, we build and publish a special stained image dataset suitable for glomeruli segmentation (including H&E staining), where the accuracy of glomeruli detection and segmentation can be significantly improved after transferring H&E-stained images to PAS-stained and PASM-stained ones by our method. The code is publicly available at: https://github.com/xianchaoguan/GramGAN.
Xianchao Guan, Yifeng Wang 0001, Yiyang Lin, Yongbing Zhang 0002
IEEE Trans. Image Process.5
2023 Weakly-Supervised Semantic Segmentation for Histopathology Images Based on Dataset Synthesis and Feature Consistency Constraint
abstract
Tissue segmentation is a critical task in computational pathology due to its desirable ability to indicate the prognosis of cancer patients. Currently, numerous studies attempt to use image-level labels to achieve pixel-level segmentation to reduce the need for fine annotations. However, most of these methods are based on class activation map, which suffers from inaccurate segmentation boundaries. To address this problem, we propose a novel weakly-supervised tissue segmentation framework named PistoSeg, which is implemented under a fully-supervised manner by transferring tissue category labels to pixel-level masks. Firstly, a dataset synthesis method is proposed based on Mosaic transformation to generate synthesized images with pixel-level masks. Next, considering the difference between synthesized and real images, this paper devises an attention-based feature consistency, which directs the training process of a proposed pseudo-mask refining module. Finally, the refined pseudo-masks are used to train a precise segmentation model for testing. Experiments based on WSSS4LUAD and BCSS-WSSS validate that PistoSeg outperforms the state-of-the-art methods. The code is released at https://github.com/Vison307/PistoSeg.
Zijie Fang, Yang Chen 0036, Yifeng Wang 0001, Zhi Wang 0001, Xiangyang Ji, Yongbing Zhang 0002
AAAI6
2023 HVTSurv: Hierarchical Vision Transformer for Patient-Level Survival Prediction from Whole Slide Image
abstract
Survival prediction based on whole slide images (WSIs) is a challenging task for patient-level multiple instance learning (MIL). Due to the vast amount of data for a patient (one or multiple gigapixels WSIs) and the irregularly shaped property of WSI, it is difficult to fully explore spatial, contextual, and hierarchical interaction in the patient-level bag. Many studies adopt random sampling pre-processing strategy and WSI-level aggregation models, which inevitably lose critical prognostic information in the patient-level bag. In this work, we propose a hierarchical vision Transformer framework named HVTSurv, which can encode the local-level relative spatial information, strengthen WSI-level context-aware communication, and establish patient-level hierarchical interaction. Firstly, we design a feature pre-processing strategy, including feature rearrangement and random window masking. Then, we devise three layers to progressively obtain patient-level representation, including a local-level interaction layer adopting Manhattan distance, a WSI-level interaction layer employing spatial shuffle, and a patient-level interaction layer using attention pooling. Moreover, the design of hierarchical network helps the model become more computationally efficient. Finally, we validate HVTSurv with 3,104 patients and 3,752 WSIs across 6 cancer types from The Cancer Genome Atlas (TCGA). The average C-Index is 2.50-11.30% higher than all the prior weakly supervised methods over 6 TCGA datasets. Ablation study and attention visualization further verify the superiority of the proposed HVTSurv. Implementation is available at: https://github.com/szc19990412/HVTSurv.
Zhuchen Shao, Yang Chen 0036, Hao Bian, Jian Zhang 0018, Yongbing Zhang 0002
AAAI6
2023 LNPL-MIL: Learning from Noisy Pseudo Labels for Promoting Multiple Instance Learning in Whole Slide Image
abstract
Gigapixel Whole Slide Images (WSIs) aided patient diagnosis and prognosis analysis are promising directions in computational pathology. However, limited by expensive and time-consuming annotation costs, WSIs usually only have weak annotations, including 1) WSI-level Annotations (WA) and 2) Limited Patch-level Annotations (LPA). Currently, Multiple Instance Learning (MIL) often exploits WA, while LPA usually assign pseudo-labels for unlabeled data. Intuitively, pseudo-labels can serve as a practical guide for MIL, but the unreliable prediction caused by LPA inevitably introduce noise. Furthermore, WA-supervised MIL training inevitably suffers from the semantical unalignment between instances and bag-level labels. To address these problems, we design a framework called Learning from Noisy Pseudo Labels for promoting Multiple Instance Learning (LNPL-MIL), which considers both types of weak annotation. Specifically, for the LPA-trained weak classifier, we design a Super-Patch-based LNPL (SP-LNPL) method to reduce false positives in the noisy pseudo-labels and then select more accurate Top-K key instances. In MIL, we propose a Transformer aware of instance Order and Distribution (TOD-MIL) that strengthens instances correlation and weakens semantical unalignment in the bag. We validate our LNPL-MIL on Tumor Diagnosis and Survival Prediction, achieving state-of-the-art performance with at least 2.7%/2.9% AUC and 2.6%/2.3% C-Index improvement with the patches labeled for two scale. Ablation study and visualization analysis further verify the effectiveness.
Zhuchen Shao, Yifeng Wang 0001, Yang Chen 0036, Hao Bian, Shaohui Liu, Haoqian Wang, Yongbing Zhang 0002
ICCV7
2023 Accurate Image Restoration with Attention Retractable Transformer
Yulun Zhang 0001, Jinjin Gu, Yongbing Zhang 0002, Linghe Kong, Xin Yuan 0002
ICLR4
2023 dMIL-Transformer: Multiple Instance Learning Via Integrating Morphological and Spatial Information for Lymph Node Metastasis Classification
abstract
Automated classification of lymph node metastasis (LNM) plays an important role in the diagnosis and prognosis. However, it is very challenging to achieve satisfactory performance in LNM classification, because both the morphology and spatial distribution of tumor regions should be taken into account. To address this problem, this article proposes a two-stage dMIL-Transformer framework, which integrates both the morphological and spatial information of the tumor regions based on the theory of multiple instance learning (MIL). In the first stage, a double Max-Min MIL (dMIL) strategy is devised to select the suspected top-K positive instances from each input histopathology image, which contains tens of thousands of patches (primarily negative). The dMIL strategy enables a better decision boundary for selecting the critical instances compared with other methods. In the second stage, a Transformer-based MIL aggregator is designed to integrate all the morphological and spatial information of the selected instances from the first stage. The self-attention mechanism is further employed to characterize the correlation between different instances and learn the bag-level representation for predicting the LNM category. The proposed dMIL-Transformer can effectively deal with the thorny classification in LNM with great visualization and interpretability. We conduct various experiments over three LNM datasets, and achieve 1.79%-7.50% performance improvement compared with other state-of-the-art methods.
Yang Chen 0036, Zhuchen Shao, Hao Bian, Zijie Fang, Yifeng Wang 0001, Yuanhao Cai, Haoqian Wang, GuoJun Liu, Yongbing Zhang 0002
IEEE J. Biomed. Health Informatics10
2022 Unpaired Multi-Domain Stain Transfer for Kidney Histopathological Images
abstract
As an essential step in the pathological diagnosis, histochemical staining can show specific tissue structure information and, consequently, assist pathologists in making accurate diagnoses. Clinical kidney histopathological analyses usually employ more than one type of staining: H&E, MAS, PAS, PASM, etc. However, due to the interference of colors among multiple stains, it is not easy to perform multiple staining simultaneously on one biological tissue. To address this problem, we propose a network based on unpaired training data to virtually generate multiple types of staining from one staining. Our method can preserve the content of input images while transferring them to multiple target styles accurately. To efficiently control the direction of stain transfer, we propose a style guided normalization (SGN). Furthermore, a multiple style encoding (MSE) is devised to represent the relationship among different staining styles dynamically. An improved one-hot label is also proposed to enhance the generalization ability and extendibility of our method. Vast experiments have demonstrated that our model can achieve superior performance on a tiny dataset. The results exhibit not only good performance but also great visualization and interpretability. Especially, our method also achieves satisfactory results over cross-tissue, cross-staining as well as cross-task. We believe that our method will significantly influence clinical stain transfer and reduce the workload greatly for pathologists. Our code and Supplementary materials are available at https://github.com/linyiyang98/UMDST.
Yiyang Lin, Bowei Zeng, Yifeng Wang 0001, Yang Chen 0036, Zijie Fang, Jian Zhang 0018, Xiangyang Ji, Haoqian Wang, Yongbing Zhang 0002
AAAI9
2022 HerosNet: Hyperspectral Explicable Reconstruction and Optimal Sampling Deep Network for Snapshot Compressive Imaging
abstract
Hyperspectral imaging is an essential imaging modality for a wide range of applications, especially in remote sensing, agriculture, and medicine. Inspired by existing hyperspectral cameras that are either slow, expensive, or bulky, reconstructing hyperspectral images (HSIs) from a low-budget snapshot measurement has drawn wide attention. By mapping a truncated numerical optimization algorithm into a network with a fixed number of phases, recent deep unfolding networks (DUNs) for spectral snapshot compressive sensing (SCI) have achieved remarkable success. However, DUNs are far from reaching the scope of industrial applications limited by the lack of cross-phase feature interaction and adaptive parameter adjustment. In this paper, we propose a novel Hyperspectral Explicable Reconstruction and Optimal Sampling deep Network for SCI, dubbed HerosNet, which includes several phases under the ISTA-unfolding framework. Each phase can flexibly simulate the sensing matrix and contextually adjust the step size in the gradient descent step, and hierarchically fuse and interact the hidden states of previous phases to effectively recover current HSI frames in the proximal mapping step. Simultaneously, a hardware-friendly optimal binary mask is learned end-to-end to further improve the reconstruction performance. Finally, our HerosNet is validated to outperform the state-of-the-art methods on both simulation and real datasets by large margins. The source code is available at https://github.com/jianzhangcs/HerosNet.
Xuanyu Zhang 0003, Yongbing Zhang 0002, Ruiqin Xiong, Qilin Sun 0001, Jian Zhang 0018
CVPR2
2022 Multiple Instance Learning with Mixed Supervision in Gleason Grading
Hao Bian, Zhuchen Shao, Yang Chen 0036, Yifeng Wang 0001, Haoqian Wang, Jian Zhang 0018, Yongbing Zhang 0002
MICCAI (8)7
2022 Semi-supervised PR Virtual Staining for Breast Histopathological Images
Bowei Zeng, Yiyang Lin, Yifeng Wang 0001, Yang Chen 0036, Jiuyang Dong, Yongbing Zhang 0002
MICCAI (2)7
2022 Cross Aggregation Transformer for Image Restoration
abstract
Recently, Transformer architecture has been introduced into image restoration to replace convolution neural network (CNN) with surprising results. Considering the high computational complexity of Transformer with global attention, some methods use the local square window to limit the scope of self-attention. However, these methods lack direct interaction among different windows, which limits the establishment of long-range dependencies. To address the above issue, we propose a new image restoration model, Cross Aggregation Transformer (CAT). The core of our CAT is the Rectangle-Window Self-Attention (Rwin-SA), which utilizes horizontal and vertical rectangle window attention in different heads parallelly to expand the attention area and aggregate the features cross different windows. We also introduce the Axial-Shift operation for different window interactions. Furthermore, we propose the Locality Complementary Module to complement the self-attention mechanism, which incorporates the inductive bias of CNN (e.g., translation invariance and locality) into Transformer, enabling global-local coupling. Extensive experiments demonstrate that our CAT outperforms recent state-of-the-art methods on several image restoration applications. The code and models are available at https://github.com/zhengchen1999/CAT.
Zheng Chen 0014, Yulun Zhang 0001, Jinjin Gu, Yongbing Zhang 0002, Linghe Kong, Xin Yuan 0002
NeurIPS4
2021 TransMIL: Transformer based Correlated Multiple Instance Learning for Whole Slide Image Classification
abstract
Multiple instance learning (MIL) is a powerful tool to solve the weakly supervised classification in whole slide image (WSI) based pathology diagnosis. However, the current MIL methods are usually based on independent and identical distribution hypothesis, thus neglect the correlation among different instances. To address this problem, we proposed a new framework, called correlated MIL, and provided a proof for convergence. Based on this framework, we devised a Transformer based MIL (TransMIL), which explored both morphological and spatial information. The proposed TransMIL can effectively deal with unbalanced/balanced and binary/multiple classification with great visualization and interpretability. We conducted various experiments for three different computational pathology problems and achieved better performance and faster convergence compared with state-of-the-art methods. The test AUC for the binary tumor classification can be up to 93.09% over CAMELYON16 dataset. And the AUC over the cancer subtypes classification can be up to 96.03% and 98.82% over TCGA-NSCLC dataset and TCGA-RCC dataset, respectively. Implementation is available at: https://github.com/szc19990412/TransMIL.
Zhuchen Shao, Hao Bian, Yang Chen 0036, Yifeng Wang 0001, Jian Zhang 0018, Xiangyang Ji, Yongbing Zhang 0002
NeurIPS7
2021 A Distortion Propagation Oriented CU-tree Algorithm for x265
abstract
Rate-distortion optimization (RDO) is widely used in video coding to improve coding efficiency. Conventionally, RDO is applied to each block independently to avoid high computational complexity. However, various prediction techniques introduce spatio-temporal dependency between blocks, therefore the independent RDO is not optimal. Specifically, because of the motion compensation, the distortion of reference blocks will affect the quality of subsequent prediction blocks. And considering this temporal dependency in RDO can improve the global rate-distortion (R-D) performance. x265 leveraged on a lookahead module to analyze the temporal dependency between blocks, and weighted the quality of each block based on its reference strength. However, the original algorithm in x265 ignored the impacts of quantization, and this shortcoming degraded the R-D performance of x265. In this paper, we propose a new linear distortion propagation model to estimate the temporal dependency, which introduces the impacts of quantization. And from a perspective of global RDO, a corresponding adaptive quantization formula is presented. The proposed algorithm was conducted in x265 version 3.2. Experiments revealed that, the proposed algorithm achieved average 15.43% PSNR-based and 23.81% SSIM-based BD-rate reductions, which outperformed the original algorithm in x265 by 4.14% and 9.68%, respectively.
Xinye Jiang, Yongbing Zhang 0002, Xiangyang Ji
VCIP3
2021 Precise No-Reference Image Quality Evaluation Based on Distortion Identification
abstract
The difficulty of no-reference image quality assessment (NR IQA) often lies in the lack of knowledge about the distortion in the image, which makes quality assessment blind and thus inefficient. To tackle such issue, in this article, we propose a novel scheme for precise NR IQA, which includes two successive steps, i.e., distortion identification and targeted quality evaluation. In the first step, we employ the well-known Inception-ResNet-v2 neural network to train a classifier that classifies the possible distortion in the image into the four most common distortion types, i.e., Gaussian white noise (WN), Gaussian blur (GB), jpeg compression (JPEG), and jpeg2000 compression (JP2K). Specifically, the deep neural network is trained on the large-scale Waterloo Exploration database, which ensures the robustness and high performance of distortion classification. In the second step, after determining the distortion type of the image, we then design a specific approach to quantify the image distortion level, which can estimate the image quality specially and more precisely. Extensive experiments performed on LIVE, TID2013, CSIQ, and Waterloo Exploration databases demonstrate that (1) the accuracy of our distortion classification is higher than that of the state-of-the-art distortion classification methods, and (2) the proposed NR IQA method outperforms the state-of-the-art NR IQA methods in quantifying the image quality.
Chenggang Yan 0001, Tong Teng, Yutao Liu 0002, Yongbing Zhang 0002, Haoqian Wang, Xiangyang Ji
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Depth Image Denoising Using Nuclear Norm and Learning Graph Model
abstract
Depth image denoising is increasingly becoming the hot research topic nowadays, because it reflects the three-dimensional scene and can be applied in various fields of computer vision. But the depth images obtained from depth camera usually contain stains such as noise, which greatly impairs the performance of depth-related applications. In this article, considering that group-based image restoration methods are more effective in gathering the similarity among patches, a group-based nuclear norm and learning graph (GNNLG) model was proposed. For each patch, we find and group the most similar patches within a searching window. The intrinsic low-rank property of the grouped patches is exploited in our model. In addition, we studied the manifold learning method and devised an effective optimized learning strategy to obtain the graph Laplacian matrix, which reflects the topological structure of image, to further impose the smoothing priors to the denoised depth image. To achieve fast speed and high convergence, the alternating direction method of multipliers is proposed to solve our GNNLG. The experimental results show that the proposed method is superior to other current state-of-the-art denoising methods in both subjective and objective criterion.
Chenggang Yan 0001, Zhisheng Li, Yongbing Zhang 0002, Yutao Liu 0002, Xiangyang Ji, Yongdong Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2020 Fast confocal microscopy imaging based on deep learning
abstract
Confocal microscopy is the de-facto standard technique in bio-imaging for acquiring 3D images in the presence of tissue scattering. However, the point-scanning mechanism inherent in confocal microscopy implies that the capture speed is much too slow for imaging dynamic objects at sufficient spatial resolution and signal to noise ratio(SNR). In this paper, we propose an algorithm for super-resolution confocal microscopy that allows us to capture high-resolution, high SNR confocal images at an order of magnitude faster acquisition speed. The proposed Back-Projection Generative Adversarial Network (BPGAN) consists of a feature extraction step followed by a back-projection feedback module (BPFM) and an associated reconstruction network, these together allow for super-resolution of low-resolution confocal scans. We validate our method using real confocal captures of multiple biological specimens and the results demonstrate that our proposed BPGAN is able to achieve similar quality to high-resolution confocal scans while the imaging speed can be up to 64 times faster.
Xiu Li 0001, Jiuyang Dong, Yongbing Zhang 0002, Ashok Veeraraghavan, Xiangyang Ji
ICCP5
2020 Simple accurate model-based phase diversity phase retrieval algorithm for wavefront sensing in high-resolution optical imaging systems
abstract
In optical imaging systems, the aberration is an important factor that impedes realising diffraction‐limited imaging. Accurate wavefront sensing and control play important role in modern high‐resolution optical imaging systems nowadays. In this study, a simple model‐based phase retrieval algorithm is proposed for accurate efficient wavefront sensing with high dynamic range. In the authors’ algorithm, a wavefront is represented by the Zernike polynomials, and the Zernike coefficients are solved by the least‐squares‐based non‐linear optimisation method, i.e. the Lederberg–Marquardt algorithm, with multiple phase‐diversity images. The numerical results show that the proposed algorithm is capable of retrieving wavefront with a large dynamic range up to seven wavelength and robust to noise. In comparison, the proposed algorithm is more efficient than the existing model‐based technique and more accurate than existing Fourier ‐ transformation‐based iterative techniques.
Shun Qin, Yongbing Zhang 0002, Haoqian Wang, Wai Kin Chan
IET Image Process.2
2020 Unsupervised Blind Image Quality Evaluation via Statistical Measurements of Structure, Naturalness, and Perception
abstract
Most existing blind image quality assessment (BIQA) methods belong to supervised methods, which always need a large number of image samples and expensive subjective scores for training a quality prediction model. In this paper, we focus our attention on the unsupervised BIQA methods and put forward a novel unsupervised approach. The main idea of our method is to quantify the image quality degradation through measuring the structure, naturalness, and the perception quality variations of the distorted image from the pristine natural images. In specific, the structure variation is captured by the deviations of the image phase congruency and gradients distributions. The naturalness variation is characterized through the distributions variations of the locally mean subtracted and contrast normalized (MSCN) coefficients and the products of pairs of the adjacent MSCN coefficients. Compared with existing unsupervised methods, we initiatively introduce the perception quality measurement into the construction of unsupervised BIQA method, which is conducted by characterizing the prediction discrepancy between the image and its brain prediction based on the free-energy principle in the newly revealed brain theory. After feature extraction, we learn a pristine multivariate Gaussian (MVG) model with the extracted features from a set of pristine natural images. The quality of a new image is finally defined as the distance between its MVG model and the learned pristine MVG model. The extensive experiments conducted on LIVE, TID2013, CSIQ, Toyama, CID2013, and the Waterloo Exploration databases demonstrate that the proposed method achieves comparative prediction performance with the state-of-the-art BIQA methods.
Yutao Liu 0002, Ke Gu 0001, Yongbing Zhang 0002, Xiu Li 0001, Guangtao Zhai, Debin Zhao, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2020 Weighted Convolutional Motion-Compensated Frame Rate Up-Conversion Using Deep Residual Network
abstract
Frame rate up-conversion (FRUC) usually suffers from unreliable motion vectors due to the absence of the current frame to be interpolated. In addition, since the majority of video sequences are usually compressed by various coding standards to reduce the data volume, the quality of the generated frames in the FRUC will be further impaired. To address this problem, we proposed two FRUC algorithms based on deep residual network. We first present a deep residual network for the FRUC (DRNFRUC), which consists of feature extraction, feature recursive analysis, and image restoration parts with a skip connection between the input and the output of the network. The proposed DRNFRUC takes the result of an arbitrary existing FRUC method as the input and is able to significantly reduce the edge blurring and blocking artifacts when the motion of the block is violent. In addition, we proposed a deep residual network with weighted convolutional motion compensation (DRNWCMC) for the FRUC, where the convolution operations can be embedded into the motion compensation interpolation (MCI) in any existing MCI-based FRUC method. In DRNWCMC, we first devise two convolutional neural networks corresponding to the forward and backward motion compensated frames, respectively. And then, the adaptive interpolation coefficients for motion compensation are designed as two$1\times1$convolutional kernels. Finally, the interpolation result of WCMC is fed into another convolutional neural network to further improve the performance. All the parameters involved in the DRNWCMC are trained simultaneously under the same cost function. The experimental results show that the two proposed algorithms can remarkably improve both the objective and subjective quality of the interpolated frames.
Yongbing Zhang 0002, Lixin Chen, Chenggang Yan 0001, Peiwu Qin, Xiangyang Ji, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.1
2020 Color-Guided Depth Image Recovery With Adaptive Data Fidelity and Transferred Graph Laplacian Regularization
abstract
Depth images play an important role and are prevalently used in many computer vision and computational imaging tasks. However, due to the limitation of active sensing technology, the captured depth images in practice usually suffer from low resolution and noise, which prevents its further applications. To remedy this problem, in this paper, we first propose an adaptive data fidelity formulation to optimally generate each depth pixel from a mixture probability distribution, characterizing the similarity both in the depth map and the corresponding high-resolution guided color image. The proposed method is able to fit the distribution of the input depth signal as an optimization problem by maximizing the mixture probability. Furthermore, to promote the piecewise property that depth images exhibit, we propose a transferred graph Laplacian model as a regularization term, which is general and able to handle various depth recovery tasks such as super-resolution and denoising well. Specifically, each pixel within the recovered depth image is represented as a vertex in a graph with weights in connected edges representing the similarity between vertices. By minimizing the squared variations of the image signal, the task of depth image recovery can be converted to the problem of graph-based image filtering. Since the proposed graph Laplacian regularization model is able to fully exploit a priori information about the depth image, a much more accurate and robust estimation of the underlying depth can be obtained. Extensive experiment evaluations verify that the proposed method obtains recovered depth with higher quality in terms of both objective and subjective criteria, compared with most of the state-of-the-art methods.
Yongbing Zhang 0002, Yihui Feng, Xianming Liu 0005, Deming Zhai, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.1
2020 Multiple Cycle-in-Cycle Generative Adversarial Networks for Unsupervised Image Super-Resolution
abstract
With the help of convolutional neural networks (CNN), the single image super-resolution problem has been widely studied. Most of these CNN based methods focus on learning a model to map a low-resolution (LR) image to a highresolution (HR) image, where the LR image is downsampled from the HR image with a known model. However, in a more general case when the process of the down-sampling is unknown and the LR input is degraded by noises and blurring, it is difficult to acquire the LR and HR image pairs for traditional supervised learning. Inspired by the recent unsupervised imagestyle translation applications using unpaired data, we propose a multiple Cycle-in-Cycle network structure to deal with the more general case using multiple generative adversarial networks (GAN) as the basis components. The first network cycle aims at mapping the noisy and blurry LR input to a noise-free LR space, then a new cycle with a well-trained ×2 network model is orderly introduced to super-resolve the intermediate output of the former cycle. The number of total cycles depends on the different up-sampling factors (×2, ×4, ×8). Finally, all modules are trained in an end-to-end manner to get the desired HR output. Quantitative indexes and qualitative results show that our proposed method achieves comparable performance with the state-of-the-art supervised models.
Yongbing Zhang 0002, Chao Dong 0005, Xinfeng Zhang 0001, Yuan Yuan 0007
IEEE Trans. Image Process.1
2020 STAT: Spatial-Temporal Attention Mechanism for Video Captioning
abstract
Video captioning refers to automatic generate natural language sentences, which summarize the video contents. Inspired by the visual attention mechanism of human beings, temporal attention mechanism has been widely used in video description to selectively focus on important frames. However, most existing methods based on temporal attention mechanism suffer from the problems of recognition error and detail missing, because temporal attention mechanism cannot further catch significant regions in frames. In order to address above problems, we propose the use of a novel spatial-temporal attention mechanism (STAT) within an encoder-decoder neural network for video captioning. The proposed STAT successfully takes into account both the spatial and temporal structures in a video, so it makes the decoder to automatically select the significant regions in the most relevant temporal segments for word prediction. We evaluate our STAT on two well-known benchmarks: MSVD and MSR-VTT-10K. Experimental results show that our proposed STAT achieves the state-of-the-art performance with several popular evaluation metrics: BLEU-4, METEOR, and CIDEr.
Chenggang Yan 0001, Yunbin Tu, Xingzheng Wang, Yongbing Zhang 0002, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.4
2020 Corrections to "STAT: Spatial-Temporal Attention Mechanism for Video Captioning"
abstract
Presents corrections to affiliations in the above named paper.
Chenggang Yan 0001, Yunbin Tu, Xingzheng Wang, Yongbing Zhang 0002, Xinhong Hao, Yongdong Zhang 0001, Qionghai Dai
IEEE Trans. Multim.4
2020 Blind Image Quality Assessment by Natural Scene Statistics and Perceptual Characteristics
abstract
Opinion-unaware blind image quality assessment (OU BIQA) refers to establishing a blind quality prediction model without using the expensive subjective quality scores, which is a highly promising direction in the BIQA research. In this article, we focus on OU BIQA and propose a novel OU BIQA method. Specifically, in our proposed method, we deeply investigate the natural scene statistics (NSS) and the perceptual characteristics of the human brain for visual perception. Accordingly, a set of quality-aware NSS and perceptual characteristics-related features are designed to characterize the image quality effectively. For inferring the image quality, we learn a pristine multivariate Gaussian (MVG) model on a collection of pristine images, which serves as the reference information for quality evaluation. At last, the quality of a new given image is defined by measuring the divergence between its MVG model and the learned pristine MVG model. Thorough experiments performed on seven popular image databases demonstrate that the proposed OU BIQA method delivers superior performance to the state-of-the-art OU BIQA methods. The Matlab source code of the proposed method will be made publicly available at https://github.com/YT2015?tab=;repositories.
Yutao Liu 0002, Ke Gu 0001, Xiu Li 0001, Yongbing Zhang 0002
ACM Trans. Multim. Comput. Commun. Appl.4
2019 Second-Order Attention Network for Single Image Super-Resolution
abstract
Recently, deep convolutional neural networks (CNNs) have been widely explored in single image super-resolution (SISR) and obtained remarkable performance. However, most of the existing CNN-based SISR methods mainly focus on wider or deeper architecture design, neglecting to explore the feature correlations of intermediate layers, hence hindering the representational power of CNNs. To address this issue, in this paper, we propose a second-order attention network (SAN) for more powerful feature expression and feature correlation learning. Specifically, a novel train- able second-order channel attention (SOCA) module is developed to adaptively rescale the channel-wise features by using second-order feature statistics for more discriminative representations. Furthermore, we present a non-locally enhanced residual group (NLRG) structure, which not only incorporates non-local operations to capture long-distance spatial contextual information, but also contains repeated local-source residual attention groups (LSRAG) to learn increasingly abstract feature representations. Experimental results demonstrate the superiority of our SAN network over state-of-the-art SISR methods in terms of both quantitative metrics and visual quality.
Tao Dai 0001, Jianrui Cai, Yongbing Zhang 0002, Shutao Xia, Lei Zhang 0006
CVPR3
2019 Recovering Extremely Degraded Faces by Joint Super-Resolution and Facial Composite
abstract
In the past a few years, we witnessed rapid advancement in face super-resolution from very low resolution(VLR) images. However, most of the previous studies focus on solving such problem without explicitly considering the impact of severe real-life image degradation (e.g. blur and noise). We can show that robustly recover details from VLR images is a task beyond the ability of current state-of-the-art method. In this paper, we borrow ideas from "facial composite" and propose an alternative approach to tackle this problem. We endow the degraded VLR images with additional cues by integrating existing face components from multiple reference images into a novel learning pipeline with both low level and high level semantic loss function as well as a specialized adversarial based training scheme. We show that our method is able to effectively and robustly restore relevant facial details from 16x16 images with extreme degradation. We also tested our approach against real-life images and our method performs favorably against previous methods.
Xiu Li 0001, Guichun Duan, Zhouxia Wang, Jimmy S. J. Ren, Yongbing Zhang 0002, Jiawei Zhang 0002, Kaixiang Song
ICTAI5
2019 CG-Cast: Scalable Wireless Image SoftCast Using Compressive Gradient
abstract
G-Cast is a wireless visual communication scheme that conveys visual information via image gradient. It is inspired by the characteristics of human vision systems and can provide improved perceptual quality. G-Cast is power efficient but bandwidth demanding, because gradient data have double the size of the original image. This paper presents a scheme named CG-Cast for scalable image transmission in bandwidth-limited wireless scenarios. It employs a compressive-gradient-based image representation to describe perceptually sensitive image details and reduce the bandwidth requirement at the same time, combining gradient-based visual representation with compressive sensing techniques. The compressive gradient data are transmitted in a pseudo-analog way so that it achieves elegant quality transition in a wide channel signal-to-noise ratio (CSNR) range. CG-Cast also sends a small set of low-frequency data in digital a way to provide the global and local luminance of the image. We developed an effective optimization algorithm for the decoder to reconstruct the original image from the received noisy compressive gradient and the low-frequency part of the image. Experimental results demonstrate that the proposed scheme improves the quality of received images remarkably under different CSNR and channel bandwidth conditions.
Hangfan Liu, Ruiqin Xiong, Xiaopeng Fan 0001, Debin Zhao, Yongbing Zhang 0002, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2019 Collaborative Representation Cascade for Single-Image Super-Resolution
abstract
Most recent learning-based single-image superresolution methods first interpolate the low-resolution (LR) input, from which overlapped LR features are then extracted to reconstruct their high-resolution (HR) counterparts and the final HR image. However, most of them neglect to take advantage of the intermediate recovered HR image to enhance image quality further. We conduct principal component analysis (PCA) to reduce LR feature dimension. Then we find that the number of principal components after conducting PCA in the LR feature space from the reconstructed images is larger than that from the interpolated images by using bicubic interpolation. Based on this observation, we present an unsophisticated yet effective framework named collaborative representation cascade (CRC) that learns multilayer mapping models between LR and HR feature pairs. In particular, we extract the features from the intermediate recovered image to upscale and enhance LR input progressively. In the learning phase, for each cascade layer, we use the intermediate recovered results and their original HR counterparts to learn single-layer mapping model. Then, we use this single-layer mapping model to super-resolve the original LR inputs. And the intermediate HR outputs are regarded as training inputs for the next cascade layer, until we obtain multilayer mapping models. In the reconstruction phase, we extract multiple sets of LR features from the LR image and intermediate recovered. Then, in each cascade layer, mapping model is utilized to pursue HR image. Our experiments on several commonly used image SR testing datasets show that our proposed CRC method achieves state-of-the-art image SR results.
Yongbing Zhang 0002, Yulun Zhang 0001, Jian Zhang 0018, Dong Xu 0001, Yun Fu 0001, Yisen Wang 0001, Xiangyang Ji, Qionghai Dai
IEEE Trans. Syst. Man Cybern. Syst.1
2018 When Deep Fool Meets Deep Prior: Adversarial Attack on Super-Resolution Network
abstract
This paper investigates the vulnerability of the deep prior used in deep learning based image restoration. In particular, the image super-resolution, which relies on the strong prior information to regularize the solution space and plays important roles in the image pre-processing for future viewing and analysis, is shown to be vulnerable to the well-designed adversarial examples. We formulate the adversarial example generation process as an optimization problem, and given super-resolution model three different types of attack are designed based on the subsequent tasks: (i) style transfer attack; (ii) classification attack; (iii) caption attack. Another interesting property of our design is that the attack is hidden behind the super-resolution process, such that the utilization of low resolution images is not significantly influenced. We show that the vulnerability to adversarial examples could bring risks to the pre-processing modules such as super-resolution deep neural network, which is also of paramount significance for the security of the whole system. Our results also shed light on the potential security issues of the pre-processing modules, and raise concerns regarding the corresponding countermeasures for adversarial examples.
Minghao Yin, Yongbing Zhang 0002, Xiu Li 0001, Shiqi Wang 0001
ACM Multimedia2
2018 Fast, Robust, and Accurate Image Denoising via Very Deeply Cascaded Residual Networks
abstract
Patch based image modelings have shown great potential in image denoising. They mainly exploit the nonlocal self-similarity (NSS) of either input degraded images or clean natural ones when training models, while failing to learn the mappings between them. More seriously, these algorithms have very high time complexity and poor robustness when handling images with different noise variances and resolutions. To address these problems, in this paper, we propose very deeply cascaded residual networks (VDCRN) to build the precise relationships between the noisy images and their corresponding noise-free ones. It adopts a new residual unit with an identity skip connection (shortcut) to make training easy and improve generalization. The introduction of shortcut is helpful to avoid the problem of gradient vanishing and preserve more image details. By cascading three such residual units, we build the VDCRN to deploy deeper and larger convolutional networks. Based on such a residual network, our VDCRN achieves very fast speed and good robustness. Experimental results demonstrate that our model outperforms a lot of state-of-the-art denoising algorithms quantitively and qualitively.
Yongbing Zhang 0002, Xingzheng Wang, Haoqian Wang, Qionghai Dai
MMSP2
2018 Referenceless quality metric of multiply-distorted images based on structural degradation
Tao Dai 0001, Ke Gu 0001, Li Niu 0002, Yongbing Zhang 0002, Weizhi Lu, Shutao Xia
Neurocomputing4
2018 Accurate saliency detection based on depth feature of 3D images
Haoqian Wang, Xingzheng Wang, Yongbing Zhang 0002
Multim. Tools Appl.4
2018 Residual Highway Convolutional Neural Networks for in-loop Filtering in HEVC
abstract
High efficiency video coding (HEVC) standard achieves half bit-rate reduction while keeping the same quality compared with AVC. However, it still cannot satisfy the demand of higher quality in real applications, especially at low bit rates. To further improve the quality of reconstructed frame while reducing the bitrates, a residual highway convolutional neural network (RHCNN) is proposed in this paper for in-loop filtering in HEVC. The RHCNN is composed of several residual highway units and convolutional layers. In the highway units, there are some paths that could allow unimpeded information across several layers. Moreover, there also exists one identity skip connection (shortcut) from the beginning to the end, which is followed by one small convolutional layer. Without conflicting with deblocking filter (DF) and sample adaptive offset (SAO) filter in HEVC, RHCNN is employed as a high-dimension filter following DF and SAO to enhance the quality of reconstructed frames. To facilitate the real application, we apply the proposed method to I frame, P frame, and B frame, respectively. For obtaining better performance, the entire quantization parameter (QP) range is divided into several QP bands, where a dedicated RHCNN is trained for each QP band. Furthermore, we adopt a progressive training scheme for the RHCNN where the QP band with lower value is used for early training and their weights are used as initial weights for QP band of higher values in a progressive manner. Experimental results demonstrate that the proposed method is able to not only raise the PSNR of reconstructed frame but also prominently reduce the bit-rate compared with HEVC reference software.
Yongbing Zhang 0002, Xiangyang Ji, Yun Zhang 0002, Ruiqin Xiong, Qionghai Dai
IEEE Trans. Image Process.1
2018 Adaptive Residual Networks for High-Quality Image Restoration
abstract
Image restoration methods based on convolutional neural networks have shown great success in the literature. However, since most of networks are not deep enough, there is still some room for the performance improvement. On the other hand, though some models are deep and introduce shortcuts for easy training, they ignore the importance of location and scaling of different inputs within the shortcuts. As a result, existing networks can only handle one specific image restoration application. To address such problems, we propose a novel adaptive residual network (ARN) for high-quality image restoration in this paper. Our ARN is a deep residual network, which is composed of convolutional layers, parametric rectified linear unit layers, and some adaptive shortcuts. We assign different scaling parameters to different inputs of the shortcuts, where the scaling is considered as part parameters of the ARN and trained adaptively according to different applications. Due to the special construction of ARN, it can solve many image restoration problems and have superior performance. We demonstrate its capabilities with three representative applications, including Gaussian image denoising, single image super resolution, and JPEG image deblocking. Experimental results prove that our model greatly outperforms numerous state-of-the-art restoration methods in terms of both peak signal-to-noise ratio and structure similarity index metrics, e.g., it achieves 0.2-0.3 dB gain in average compared with the second best method at a wide range of situations.
Yongbing Zhang 0002, Chenggang Yan 0001, Xiangyang Ji, Qionghai Dai
IEEE Trans. Image Process.1
2017 Foveated nonlocal dual denoising
abstract
Recently developed dual domain image denoising (DDID) algorithm and its variants, such as dual domain filter (DDF), achieve remarkable results by combining bilateral filter with frequency-based method. However, this kind of algorithms require large patches to guarantee the denoising performance and most of them produce ringing artifacts due to the Gibbs phenomenon induced by high-contrast details. To address these issues, we propose a Foveated Nonlocal Dual Denoising (FNDD) algorithm by unifying foveated nonlocal means and frequency-based methods. In this way, the ability to preserve the high-contrast details is noticeably improved by exploiting foveated self-similarity (patch similarity) instead of pixel similarity, thus leading to void of artifacts. Moreover, we propose an entropy-based back projection step for compensating the detail loss to further improve the performance. Experimental results validate that FNDD significantly outperforms DDID in terms of both quantitative metrics and subjective visual quality under much smaller patches, and even achieves comparable results against state-of-the-art competitors.
Tao Dai 0001, Ke Gu 0001, Qingtao Tang, Kwok-Wai Hung, Yongbing Zhang 0002, Weizhi Lu, Shutao Xia
ICIP5
2017 Blind quality assessment of multiply-distorted images based on structural degradation
abstract
It is known that images available usually undergo some stages of processing (e.g., acquisition, compression, transmission and display), and each stage may introduce certain type of distortion. Hence, images distorted by multiple types of distortions are common in real applications. Research in human visual perception has evidenced that the human visual system (HVS) is sensitive to image structural information. This fact inspires us to design a new blind/no-reference (NR) image quality assessment (IQA) method to evaluate the visual quality of multiply-distorted images based on structural degradation. Specifically, quality-aware features are extracted from both the first- and high-order image structures by local binary pattern (LBP) operators. Experimental results on two well-known multiply-distorted image databases demonstrate the outstanding performance of the proposed method.
Tao Dai 0001, Ke Gu 0001, Zhiya Xu, Qingtao Tang, Haoyi Liang, Yongbing Zhang 0002, Shutao Xia
ICIP6
2017 Single depth image super-resolution and denoising based on sparse graphs via structure tensor
abstract
The existing single depth image super-resolution (SR) methods suppose that the image to be interpolated is noise free. However, the supposition is invalid in practice because noise will be inevitably introduced in the depth image acquisition process. In this paper, we address the problem of image denoising and SR jointly based on designing sparse graphs that are useful for describing the geometric structures of data domains. In our method, we first cluster similar patches in a noisy depth image and compute an average patch. Different from the majority of the graph Fourier transform (GFT) that assumed an underlying 4-connected graph structure with vertical and horizontal edges only, we select more general sparse graph structures and edges weights based on the difference of the blocks' structure tensors. For the average patch, a graph template with edges orthogonal to the principal gradient is designed. Finally, the graph based transform (GBT) dictionary is learned from the derived correlation graph for signal representation. As shown in our experimental results, the proposed method obtains a lot of improvement in performance.
Yihui Feng, Xianming Liu 0005, Yongbing Zhang 0002, Qionghai Dai
ICIP3
2017 An accurate saliency prediction method based on generative adversarial networks
abstract
In this paper, we propose a saliency prediction algorithm utilizing generative adversarial networks. The proposed system contains two parts: saliency network and adversarial networks. The saliency network is the basis for saliency prediction, which calculates an Euclidean cost function on the grayscale values between the predicted saliency map and the ground truth. In order to improve the accuracy of the algorithm, adversarial networks are subsequently utilized to extract the features of input data by coordinating the learning rates of the two sub-networks contained in the networks. Experimental results validate the high accuracy of the proposed approach compared with the state-of-the-art models on three public datasets, SALICON, MIT1003 and Cerf.
Haoqian Wang, Xingzheng Wang, Yongbing Zhang 0002
ICIP4
2017 Nonlocal Gradient Sparsity Regularization for Image Restoration
abstract
Total variation (TV) regularization is widely used in image restoration to exploit the local smoothness of image content. Essentially, the TV model assumes a zero-mean Laplacian distribution for the gradient at all pixels. However, real-world images are nonstationary in general, and the zero-mean assumption of pixel gradient might be invalid, especially for regions with strong edges or rich textures. This paper introduces a nonlocal (NL) extension of TV regularization, which models the sparsity of the image gradient with pixelwise content-adaptive distributions, reflecting the nonstationary nature of image statistics. Taking advantage of the NL similarity of natural images, the proposed approach estimates the image gradient statistics at a particular pixel from a group of nonlocally searched patches, which are similar to the patch located at the current pixel. The gradient data in these NL similar patches are regarded as the samples of the gradient distribution to be learned. In this way, more accurate estimation of gradient is achieved. Experimental results demonstrate that the proposed method outperforms the conventional TV and several other anchors remarkably and produces better objective and subjective image qualities.
Hangfan Liu, Ruiqin Xiong, Xinfeng Zhang 0001, Yongbing Zhang 0002, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.4
2017 Depth Estimation by Parameter Transfer With a Lightweight Model for Single Still Images
abstract
In this paper, we propose a novel method for automatic depth estimation from color images using parameter transfer. By modeling the correlation between color images and their depth maps with a set of parameters, we get a database of parameter sets. Given an input image, we extract the high-level features to find the best matched image sets from the database. Then the set of parameters corresponding to the best match are used to estimate the depth of the input image. Compared with the past learning-based methods, our trained model consists only of trained features and parameter sets, which occupy little space. We evaluate our depth estimation method on several benchmark RGB-D (RGB + depth) data sets. The experimental results are comparable to the state-of-the-art results, while the model size is very small and very suitable for mobile devices, demonstrating the promising performance of our proposed method.
Hongwei Qin, Xiu Li 0001, Yangang Wang 0001, Yongbing Zhang 0002, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.4
2017 Multi-Task Rank Learning for Image Quality Assessment
abstract
In practice, images are distorted by more than one distortion. For image quality assessment (IQA), existing machine learning (ML)-based methods generally establish a unified model for all the distortion types, or each model is trained independently for each distortion type, which is therefore distortion aware. In distortion-aware methods, the common features among different distortions are not exploited. In addition, there are fewer training samples for each model training task, which may result in overfitting. To address these problems, we propose a multi-task learning framework to train multiple IQA models together, where each model is for each distortion type; however, all the training samples are associated with each model training task. Thus, the common features among different distortion types and the said underlying relatedness among all the learning tasks are exploited, which would benefit the generalization ability of trained models and prevent overfitting possibly. In addition, pairwise image quality ranking instead of image quality rating is optimized in our learning task, which is fundamentally departed from traditional ML-based IQA methods toward better performance. The experimental results confirm that the proposed multi-task rank-learning-based IQA metric is prominent against all state-of-the-art nonreference IQA approaches.
Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yihua Yan
IEEE Trans. Circuits Syst. Video Technol.4
2017 Light-Field Depth Estimation via Epipolar Plane Image Analysis and Locally Linear Embedding
abstract
In this paper, we propose a novel method for 4D light-field (LF) depth estimation exploiting the special linear structure of an epipolar plane image (EPI) and locally linear embedding (LLE). Without high computational complexity, depth maps are locally estimated by locating the optimal slope of each line segmentation on the EPIs, which are projected by the corresponding scene points. For each pixel to be processed, we build and then minimize the matching cost that aggregates the intensity pixel value, gradient pixel value, spatial consistency, as well as reliability measure to select the optimal slope from a predefined set of directions. Next, a subangle estimation method is proposed to further refine the obtained optimal slope of each pixel. Furthermore, based on a local reliability measure, all the pixels are classified into reliable and unreliable pixels. For the unreliable pixels, LLE is employed to propagate the missing pixels by the reliable pixels based on the assumption of manifold preserving property maintained by natural images. We demonstrate the effectiveness of our approach on a number of synthetic LF examples and real-world LF data sets, and show that our experimental results can achieve higher performance than the typical and recent state-of-the-art LF stereo matching methods.
Yongbing Zhang 0002, Huijin Lv, Yebin Liu, Haoqian Wang, Xingzheng Wang, Qian Huang 0008, Xinguang Xiang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.1
2017 Reducing Image Compression Artifacts by Structural Sparse Representation and Quantization Constraint Prior
abstract
The block discrete cosine transform (BDCT) has been widely used in current image and video coding standards, owing to its good energy compaction and decorrelation properties. However, because of independent quantization of DCT coefficients in each block, BDCT usually gives rise to visually annoying blocking compression artifacts, especially at low bit rates. In this paper, to reduce blocking artifacts and obtain high-quality images, image deblocking is cast as an optimization problem within maximum a posteriori framework, and a novel algorithm for image deblocking by using structural sparse representation (SSR) prior and quantization constraint (QC) prior is proposed. The SSR prior is utilized to simultaneously enforce the intrinsic local sparsity and the nonlocal self-similarity of natural images, while QC is explicitly incorporated to ensure a more reliable and robust estimation. A new split Bregman iteration-based method with an adaptively adjusted regularization parameter is developed to solve the proposed optimization problem, which makes the entire algorithm more practical. Experiments demonstrate that the proposed image-deblocking algorithm combining SSR and QC outperforms the current state-of-the-art methods in both peak signal-to-noise ratio and visual perception.
Chen Zhao 0002, Jian Zhang 0018, Siwei Ma 0001, Xiaopeng Fan 0001, Yongbing Zhang 0002, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.5
2016 Deep Convolutional Neural Network for Decompressed Video Enhancement
abstract
Block-wise intra/inter prediction, transformation and quantization used in block-based hybrid video coding will inevitably result in blocking artifacts, especially at the low bit rate. To address this problem, this paper employs a deep convolutional neural network (CNN) to approximate the reverse function of video compression, motived by the great success of deep learning in computer vision fields recently. The proposed method establishes an end-to-end mapping, represented as the CNN, which takes the decompressed frame as input and outputs the enhanced one. Employing numerous sequences compressed by H.264 and HEVC reference software, the proposed CNN learns the connections between the lossy frame and the original one in an implicit way under different quantization parameters (QP). Figure 1 shows the architecture of our CNN and the pipeline of the network training. We build our network with convolution layers and ReLU layer and the weights and biases of all the convolution layers in our model are updated by minimizing the loss using stochastic gradient descent with the standard backpropagation. We implement the CNN as a post-loop deblocking filter and explore varying CNN parameters for different QPs. Various experimental results demonstrate that the proposed method is able to significantly improve the quality of enhanced frames in terms of both objective and subjective criterions.
Rongqun Lin, Yongbing Zhang 0002, Haoqian Wang, Xingzheng Wang, Qionghai Dai
DCC2
2016 Fourier ptychographic reconstruction using weighted replacement in the fourier domain
abstract
Fourier ptychographic microscopy (FPM) is an attractive method to extend the resolution beyond the conventional limit defined by a microscope optics, sharing properties with ptychographic, synthetic aperture imaging and phase retrieval. The algorithm uses a sequence of low-resolution (LR) images acquired under angularly varying illumination to reconstruct a high-resolution (HR) image. However, traditional FPM may trap to the sub-optimal solution, since brute-force replacement in the Fourier domain is applied. To address this problem, we propose here a weighted replacement for Fourier ptychographic microscopy (WFPM). We employ the weighted average of spectrums corresponding to different illumination angle to replace the overlapped regions in the Fourier domain. A series of experimental results demonstrate that the reconstructed image using the proposed WFPM shows a better quality and a faster convergence compared with the results obtained by FPM.
Pengming Song, Weixin Jiang, Yongbing Zhang 0002, Qionghai Dai
ICIP3
2016 Parameterized reconstruction based Fourier Ptychography
abstract
Fourier ptychography (FP) is recently proposed as a computational imaging technique, which aims at enhancing the space-bandwidth product (SBP) of the optical imaging system. Specifically, the FP recovery routine iteratively stitches together a number of low-resolution images, which are captured under angularly varying illumination, to produce a wide-field, high-resolution image. However, the reconstruction procedure of the FP recovery routine is based on a low efficient phase retrieval algorithm and may severely degrade quality of the reconstructions. To address this problem, in this paper, we develop and test a Parameterized Reconstruction based Fourier Ptychography (PR-FP), which parameterizes the reconstruction procedure by introducing an updating-coefficient. Meanwhile, a convergence-related metric, which measures how good the reconstruction matches the input dataset, is proposed to help determine a proper updating-coefficient. Extensive experimental results demonstrate that the proposed PR-FP algorithm achieves superior reconstructions both on simulated dataset and real captured dataset.
Weixin Jiang, Yongbing Zhang 0002, Qionghai Dai
ICME2
2016 Depth Feature Based Accurate Saliency Detection for 3D Images
abstract
In this paper, we present an accurate saliency detection algorithm based on depth feature for 3D images. We first calculate depth cue based on the sharp regions' positions within the depth ranges. Then, the coarse saliency map is computed based on the background and location prior. Finally, we employ the contrast information in the coarse saliency map to obtain the final result. Experimental evaluation by comparison with existed methods verifies the effectiveness of our proposed algorithm in terms of precision, recall and F-Measure.
Haoqian Wang, Xingzheng Wang, Yongbing Zhang 0002
PDCAT4
2016 Single image super-resolution via projective dictionary learning with anchored neighborhood regression
abstract
We propose a novel single image super-resolution (SR) algorithm based on the projective dictionary pair learning with anchored neighborhood regression. Different from previous dictionary learning methods that aim to learn only a synthesis or an analysis dictionary, our method would learn both types of dictionaries jointly for regression to achieve image SR. We first cluster the training features into K clusters in order to learn synthesis and analysis dictionaries. Moreover, we learn the regressions with the training samples at training phase and use them on reconstruction stage. As shown in our experimental results, the proposed method obtains high-quality SR results quantitatively and visually against state-of-the-art methods.
Yihui Feng, Yongbing Zhang 0002, Yulun Zhang 0001, Qionghai Dai
VCIP2
2016 Decompressed video enhancement via accurate regression prior
abstract
There is an increasing need for high-quality multimedia applications based on block-based hybrid video coding. Inevitably, the frame will degrade during the process of block-wise intra/inter prediction, transformation, and quantization, especially when the bit rate is low. In this paper, we propose an efficient decompressed video enhancement algorithm based on the adjusted anchored neighborhood regression (A+) method. In our work, first, we learn offline linear regressors, i.e. projection matrices from the decompressed to original video frames in the training phase. For grouping anchored neighborhoods more accurately, we adopt MI-KSVD rather than KSVD to learn the dictionary. Moreover, we exploit the mutual coherence between dictionary atoms and training samples to find the nearest neighbors. Second, in the enhancement phase, we boost the quality of input decompressed videos offline by learned regression priors. To verify the robustness of our enhancement method, extensive experiments are conducted. As shown in our experimental results, the proposed enhancement method yields superior performance both objectively and subjectively.
Yulun Zhang 0001, Yongbing Zhang 0002, Xingzheng Wang, Haoqian Wang, Qionghai Dai
VCIP3
2016 A Polynomial Approximation Motion Estimation Model for Motion-Compensated Frame Interpolation
abstract
Motion-compensated frame interpolation (MCFI) usually finds the most matched blocks by minimizing pixel intensity discrepancies between neighboring frames along the motion trajectory. However, the quality of interpolated frames is susceptible to inaccurate motion vectors possibly for regions with complex texture patterns, irregularly shaped objects, repeated patterns, motion blurring or aliasing, and so on. It is believed that pixel intensity across adjacent frames varies gradually and smoothly, and therefore can be modeled mathematically by a continuous and differentiable function. Thus, the pixel intensity within one frame can be expressed as either a forward polynomial approximation (FWPA) or a backward polynomial approximation (BWPA) by the Taylor expansion in this paper. The discrepancy between the FWPA and BWPA is employed to find the best motion vector. In addition, a motion-aligned partial derivative is proposed to calculate the Taylor expansion along the motion trajectory. The proposed method is applicable to any existing MCFI schemes and achieve superior performance by consuming relatively more buffer memories and computational resources. Extensive experimentation with comparison with previous techniques validates our method in terms of both objective and subjective criteria.
Yongbing Zhang 0002, Long Xu 0001, Xiangyang Ji, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.1
2016 CONCOLOR: Constrained Non-Convex Low-Rank Model for Image Deblocking
abstract
Due to independent and coarse quantization of transform coefficients in each block, block-based transform coding usually introduces visually annoying blocking artifacts at low bitrates, which greatly prevents further bit reduction. To alleviate the conflict between bit reduction and quality preservation, deblocking as a post-processing strategy is an attractive and promising solution without changing existing codec. In this paper, in order to reduce blocking artifacts and obtain high-quality image, image deblocking is formulated as an optimization problem within maximum a posteriori framework, and a novel algorithm for image deblocking using constrained non-convex low-rank model is proposed. The ℓ(p) (0 < p < 1) penalty function is extended on singular values of a matrix to characterize low-rank prior model rather than the nuclear norm, while the quantization constraint is explicitly transformed into the feasible solution space to constrain the non-convex low-rank optimization. Moreover, a new quantization noise model is developed, and an alternatively minimizing strategy with adaptive parameter adjustment is developed to solve the proposed optimization problem. This parameter-free advantage enables the whole algorithm more attractive and practical. Experiments demonstrate that the proposed image deblocking algorithm outperforms the current state-of-the-art methods in both the objective quality and the perceptual quality.
Jian Zhang 0018, Ruiqin Xiong, Chen Zhao 0002, Yongbing Zhang 0002, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Image Process.4
2016 High-Efficiency 3D Depth Coding Based on Perceptual Quality of Synthesized Video
abstract
In 3D video systems, imperfect depth images often induce annoying temporal noise, e.g., flickering, to the synthesized video. However, the quality of synthesized view is usually measured with peak signal-to-noise ratio or mean squared error, which mainly focuses on pixelwise frame-by-frame distortion regardless of the obvious temporal artifacts. In this paper, a novel full reference synthesized video quality metric (SVQM) is proposed to measure the perceptual quality of the synthesized video in 3D video systems. Based on the proposed SVQM, an improved rate-distortion optimization (RDO) algorithm is developed with the target of minimizing the perceptual distortion of synthesized view at given bit rate. Then, the improved RDO algorithm is incorporated into the 3D High Efficiency Video Coding (3D-HEVC) software to improve the 3D depth video coding efficiency. Experimental results show that the proposed SVQM metric has better consistency with human perception on evaluating the synthesized view compared with the state-of-the-art image/video quality assessment algorithms. Meanwhile, this SVQM metric maintains low complexity and easy integration to the current video codec. In addition, the proposed SVQM-based depth coding scheme can achieve approximately 15.27% and 17.63% overall bit rate reduction or 0.42- and 0.46-dB gain in terms of SVQM quality score on average as compared with the latest 3D-HEVC reference model and the state-of-the-art depth coding algorithm, respectively.
Yun Zhang 0002, Xiaoxiang Yang, Xiangkai Liu, Yongbing Zhang 0002, Gangyi Jiang, Sam Kwong
IEEE Trans. Image Process.4
2016 Free-Energy Principle Inspired Video Quality Metric and Its Use in Video Coding
abstract
In this paper, we extend the free-energy principle to video quality assessment (VQA) by incorporating with the recent psychophysical study on human visual speed perception (HVSP). A novel video quality metric, namely the free-energy principle inspired video quality metric (FePVQ), is therefore developed and applied to perceptual video coding optimization. The free-energy principle suggests that the human visual system (HVS) can actively predict “orderly” information and avoid “disorderly” information for image perception. Basically, “orderly” is associated with the skeletons and edges of objects, and “disorderly” mostly concerns textures in images. Based on this principle, an image is separated into orderly and disorderly regions, and processed differently in image quality assessment. For videos, visual attention, or fixation, is associated with the objects with significant motion according to HVSP, resulting in a motion strength factor in the FePVQ so that the free-energy principle is extended into spatio-temporal domain for VQA. In addition, we investigate the application of the FePVQ in perceptual rate distortion optimization (RDO). For this purpose, the FePVQ is realized with low computational cost by using the relative total variation model and the block-wise motion vectors of video coding to simulate the free-energy principle and the HVSP, respectively. The experimental results indicate that the proposed FePVQ is highly consistent with the HVS perception. The linear correlation coefficient and Spearman's rank-order correlation coefficient are up to 0.8324 and 0.8281 on the LIVE video database. Better perceptual quality of encoded video sequences is achieved by FePVQ-motivated RDO in video coding.
Long Xu 0001, Weisi Lin, Lin Ma 0002, Yongbing Zhang 0002, Yuming Fang 0001, King Ngi Ngan, Songnan Li, Yihua Yan
IEEE Trans. Multim.4
2016 CCR: Clustering and Collaborative Representation for Fast Single Image Super-Resolution
abstract
Clustering and collaborative representation (CCR) have recently been used in fast single image super-resolution (SR). In this paper, we propose an effective and fast single image super-resolution (SR) algorithm by combining clustering and collaborative representation. In particular, we first cluster the feature space of low-resolution (LR) images into multiple LR feature subspaces and group the corresponding high-resolution (HR) feature subspaces. The local geometry property learned from the clustering process is used to collect numerous neighbor LR and HR feature subsets from the whole feature spaces for each cluster center. Multiple projection matrices are then computed via collaborative representation to map LR feature subspaces to HR subspaces. For an arbitrary input LR feature, the desired HR output can be estimated according to the projection matrix, whose corresponding LR cluster center is nearest to the input. Moreover, by learning statistical priors from the clustering process, our clustering-based SR algorithm would further decrease the computational time in the reconstruction phase. Extensive experimental results on commonly used datasets indicate that our proposed SR algorithm obtains compelling SR images quantitatively and qualitatively against many state-of-the-art methods.
Yongbing Zhang 0002, Yulun Zhang 0001, Jian Zhang 0018, Qionghai Dai
IEEE Trans. Multim.1
2015 Image colorization using hybrid domain transform
abstract
Image colorization is the process of spreading the user specified colors to all the expected regions in image. It is a great challenge to distinguish between texture edge and object boundary. To improve the performance of colorization, we incorporate depth and texture to accurately extract boundary information in an implicit way. Inspired by the low complexity and high efficiency properties of recently proposed domain transform, we proposed a hybrid domain transform, taking corresponding depth image of the processed texture image into account, to perform colorization and recoloring. Various experimental results demonstrate that hybrid domain transform is able to achieve better edge-aware property while maintaining the property of low complexity.
Hongbo Ao, Yongbing Zhang 0002, Qionghai Dai
ICASSP2
2015 Multi-task rank learning for image quality assessment
abstract
In practice, multiple types of distortions are associated with an image quality degradation process. The existing machine learning (ML) based image quality assessment (IQA) approaches generally established a unified model for all distortion types, or each model is trained independently for each distortion type by using single-task learning, which lead to the poor generalization ability of the models as applied to practical image processing. There are often the underlying cross relatedness amongst these single-task learnings in IQA, which is ignored by the previous approaches. To solve this problem, we propose a multi-task learning framework to train IQA models simultaneously across individual tasks each of which concerns one distortion type. These relatedness can be therefore exploited to improve the generalization ability of IQA models from single-task learning. In addition, pairwise image quality rank instead of image quality rating is optimized in learning task. By mapping image quality rank to image quality rating, a novel no-reference (NR) IQA approach can be derived. The experimental results confirm that the proposed Multi-task Rank Learning based IQA (MRLIQ) approach is prominent among all state-of-the-art NR-IQA approaches.
Long Xu 0001, Jia Li 0003, Weisi Lin, Yongbing Zhang 0002, Lin Ma 0002, Yuming Fang 0001, Yun Zhang 0002, Yihua Yan
ICASSP4
2015 Image super-resolution based on dictionary learning and anchored neighborhood regression with mutual incoherence
abstract
In this paper, we employ unified mutual coherence between the dictionary atoms and atoms/samples when learning the dictionary and sampling anchored neighborhoods respectively for image super-resolution (SR) application algorithm. On one hand, an incoherence promoting term in dictionary learning for SR is introduced to encourage dictionary atoms, associated to different anchored regressors, to be as independent as possible, while still allowing for different regressors to share same samples. On the other hand, a unified form with mutual coherence between dictionary atoms and training samples is proposed when we group neighborhoods of samples centered on each atom and find the nearest neighbors for input samples in image super-resolution. Extensive experimental results on commonly used datasets demonstrate that our method outperforms state-of-the-art methods by obtaining compelling results with improved quality, such as sharper edges, finer textures and higher structural similarity.
Yulun Zhang 0001, Kaiyu Gu, Yongbing Zhang 0002, Jian Zhang 0018, Qionghai Dai
ICIP3
2015 Image deblocking using group-based sparse representation and quantization constraint prior
abstract
To alleviate the conflict between bit reduction and quality preservation, deblocking as a post-processing strategy is an attractive and promising solution without changing existing codec. In this paper, in order to reduce blocking artifacts and obtain high-quality image, image deblocking is formulated as an optimization problem via maximum a posteriori framework, and a novel algorithm for image deblocking using group-based sparse representation (GSR) and quantization constraint (QC) prior is proposed. GSR prior is utilized to simultaneously enforce the intrinsic local sparsity and the nonlocal self-similarity of natural images, while QC prior is explicitly incorporated to ensure a more reliable and robust estimation. A new split Bregman iteration based method with adaptively adjusted regularization parameter is developed to solve the proposed optimization problem for image deblocking. The parameter-adaptive advantage enables the whole algorithm more attractive and practical. Experiments manifest that the proposed image deblocking algorithm improves current state-of-the-art results by a large margin in both PSNR and visual perception.
Jian Zhang 0018, Siwei Ma 0001, Yongbing Zhang 0002, Wen Gao 0001
ICIP3
2015 A novel light field super-resolution framework based on hybrid imaging system
abstract
We propose a novel light field super-resolution framework based on hybrid imaging system, which combines two different imaging mechanisms: conventional imaging and current art-of-the-state imaging - light field imaging. We take advantage of conventional imaging in spatial resolution to make up light field and reconstruct a higher quality light field. In our method, we classify the points of the 3D scene: First, for highlight and occlusion, dictionary learning based interpolation is utilized, Second, for other areas, an improved patch matching algorithm is applied. As shown in experimental results, compared with four methods, which include the art-of-the-state algorithms, our approach is effective.
Judong Wu, Haoqian Wang, Xingzheng Wang, Yongbing Zhang 0002
VCIP4
2015 Accurate image specular highlight removal based on light field imaging
abstract
Specular reflection removal is indispensable to many computer vision tasks. However, most existing methods fail or degrade in complex real scenarios for their individual drawbacks. Benefiting from the light field imaging technology, this paper proposes a novel and accurate approach to remove specularity and improve image quality. We first capture images with specularity by the light field camera (Lytro ILLUM). After accurately estimating the image depth, a simple and concise threshold strategy is adopted to cluster the specular pixels into "unsaturated" and "saturated" category. Finally, a color variance analysis of multiple views and a local color refinement are individually conducted on these two categories to recover diffuse color information. Experimental evaluation by comparison with existed methods verifies the effectiveness of our proposed algorithm.
Chenxue Xu, Xingzheng Wang, Haoqian Wang, Yongbing Zhang 0002
VCIP4
2015 Adaptive local nonparametric regression for fast single image super-resolution
abstract
We propose a fast single image super-resolution algorithm based on adaptive local nonparametric regression. Making use of dictionary learning and regression, we learn multiple projection matrices mapping low-resolution features to their corresponding high-resolution ones directly. Different from previous linear regression that needs some constant parameters, our method would not use extra parameters for regression. We use the mutual coherence between dictionary atom and low-resolution feature as a label to reconstruct more sophisticated high-resolution feature. As we use the same form of mutual coherence as labels in both training and testing phases, our method would lead to an adaptive local linear regression model. Moreover, we investigate the statistical property of the dictionary atoms from the training features. Utilizing the learned statistical priors, our method would not only obtain more useful dictionary atoms, but also further decrease the computational time. As shown in our experimental results, the proposed method yields high-quality super-resolution images quantitatively and visually against state-of-the-art methods.
Yulun Zhang 0001, Yongbing Zhang 0002, Jian Zhang 0018, Haoqian Wang, Xingzheng Wang, Qionghai Dai
VCIP2
2014 DEPT: Depth Estimation by Parameter Transfer for Single Still Images
Xiu Li 0001, Hongwei Qin, Yangang Wang 0001, Yongbing Zhang 0002, Qionghai Dai
ACCV (2)4
2014 Automatic foreground extraction in video
abstract
This paper presents an automatic and efficient system for extracting dynamic objects of interest from videos. We take advantage of a saliency map and an optimization-based segmentation algorithm to extract the foreground objects automatically in some key frames. Then, the segmentation results in those key frames are propagated to other frames via an error map-based propagation scheme. Finally, a Bayesian matting-based refinement approach is employed to to handle the topology changes. Experiments show that our system is able to generate high quality results at a low computation cost.
Haoqian Wang, Kai Li 0016, Yongbing Zhang 0002, Lei Zhang 0006
ICASSP4
2014 Depth map super-resolution via iterative joint-trilateral-upsampling
abstract
In this paper, we propose a new approach to solve the depth map super-resolution (SR) and denoising problems simultaneously. Inspired by joint-bilateral-upsampling (JBU), we devised the joint-trilateral-upsampling (JTU), which takes edge of the initial depth map, texture of the corresponding high-resolution color image and the values of the surrounding depth pixels, into consideration during the process of SR. To preserve the sharp edge of the up-sampled depth map and remove the noise, we introduce an iterative implementation, where current up-sampled depth map is fed into the next iteration, to refine the filter coefficients of JTU. The iterative JTU presents a high performance at many aspects such as sharping edge, denoising and none texture copying, etc. To demonstrate the superiority of the proposed method, we carry out various experiments and show an across-the-board quality improvement by both of subjective and objective evaluations compared with previous state-of-art methods.
Lei Zhang 0006, Yongbing Zhang 0002, Huiming Xuan, Qionghai Dai
VCIP3
2014 Synthesis-guided depth super resolution
abstract
Depth map, as important auxiliary information in 3D procession, is used to synthesize virtual view rather than exhibition. Inspired by this, a synthesis-guided depth super resolution (SGDSR) algorithm is proposed. Employing the synthesis error between virtual view and corresponding original one as the criteria, the best super-resolved result is selected among numerous candidate super resolution (SR) results. To fully exploit varying property within different regions of an image, a patch-based SGDSR is further devised in this paper. Experimental results demonstrate the effectiveness of our method subjectively and objectively on both single view and two views platform based on depth-image-based rendering (DIBR).
Huijin Lv, Yongbing Zhang 0002, Kai Li 0016, Xingzheng Wang, Huiming Xuan, Qionghai Dai
VCIP2
2014 Real-time air quality estimation based on color image processing
abstract
This paper address the problem of efficient, realtime estimation of the particulate mass concentration, exactly PM2.5 (particles with aerodynamic diameters less than 2.5 μm) from a superb view image. And the proposed method is to achieve high degree of accuracy at the cost of only modest user's effort by analyzing the relationship between the PM2.5 and the degradation of the observed image. With the fitting algorithm with experimental data, the PM2.5 could be real-time estimated by a general camera with little artificial participation, and the correlation coefficient produced by our data set and the standard observation will be as high as 0.8219, as the MSE (Mean Squared Error) value 51.2324 μg/m3.
Haoqian Wang, Xin Yuan 0002, Xingzheng Wang, Yongbing Zhang 0002, Qionghai Dai
VCIP4
2014 Texture aided depth frame interpolation
Yongbing Zhang 0002, Jian Zhang 0018, Qionghai Dai
Signal Process. Image Commun.1
2013 A novel depth propagation algorithm with color guided motion estimation
abstract
Depth propagation is an effective and efficient way to produce depth maps for a video sequence. Motion estimation in most existing depth propagation schemes is performed only based on the estimated depth maps without consideration for color information. This paper presents a novel key frame depth propagation algorithm combining bilateral filtering and motion estimation. A color guided motion estimation process is proposed by taking both color and depth information into account when estimating the motion vectors. In addition, a bidirectional propagation strategy is adopted to reduce the accumulation of depth errors. Experimental results show that the proposed algorithm outperforms most of the existing techniques in obtaining high quality depth maps indicating a better effect of the synthesized stereoscopic video.
Haoqian Wang, Yushi Tian, Yongbing Zhang 0002
VCIP3
2013 Effective stereo matching using reliable points based graph cut
abstract
In this paper, we propose an effective stereo matching algorithm using reliable points and region-based graph cut. Firstly, the initial disparity maps are calculated via local windowbased method. Secondly, the unreliable points are detected according to the DSI(Disparity Space Image) and the estimated disparity values of each unreliable point are obtained by considering its surrounding points. Then, the scheme of reliable points is introduced in region-based graph cut framework to optimize the initial result. Finally, remaining errors in the disparity results are effectively handled in a multi-step refinement process. Experiment results show that the proposed algorithm achieves a significant reduction in computation cost and guarantee high matching quality.
Haoqian Wang, Yongbing Zhang 0002, Lei Zhang 0006
VCIP3
2013 Up-sampling oriented frame rate reduction
Yongbing Zhang 0002, Haoqian Wang, Debin Zhao
Signal Process. Image Commun.1
2013 Stereo Interleaving Video Coding With Content Adaptive Image Subsampling
abstract
Stereo interleaving video coding, in which both left and right view frames are subsampled into half size and multiplexed into one single frame before being encoded by a traditional 2-D video encoder, is an efficient encoding scenario for stereoscopic video. Many existing stereo interleaving video coding methods subsample each frame by utilizing fixed subsampling filter coefficients. Such methods are easy to implement; however, the varying property of the frame signal is ignored. By jointly considering the influences of subsampling and compression, a rate and distortion analysis about stereo interleaving video coding is proposed. The final distortion in stereo interleaving video coding is the summation of errors caused by subsampling (causing distortion between subsampling-interpolated image and the original full resolution one) and by quantization during compression. Based on the provided rate distortion analysis, a content adaptive image subsampling (CAIS) is also proposed. In CAIS, the half-size frames are generated by the optimal subsampling filters, which are calculated based on frame contents and the targeted interpolation coefficients. Experimental results demonstrate that the proposed CAIS is able to greatly improve compression efficiency of stereo interleaving video coding.
Yongbing Zhang 0002, Xiangyang Ji, Haoqian Wang, Qionghai Dai
IEEE Trans. Circuits Syst. Video Technol.1
2012 A Single Frame Super-Resolution Method Based on Matrix Completion
abstract
Efficiently exploring the linear relationship among neighboring pixels is a pervasive way to reconstruct high-resolution image from low-resolution one. However, it is a challenge to determine the order of linear model. According to the theory of matrix completion, we propose a single frame super-resolution algorithm by minimizing the sum of all the augmented matrices' rank, which can reflect the order of the region aware linear model. Various experiments demonstrate the images reconstructed by the proposed method have superior PSNR and visual quality, benefitting from its desirable ability of depressing the ringing noise and other artifacts.
Changjun Fu, Xiangyang Ji, Yongbing Zhang 0002, Qionghai Dai
DCC3
2012 Packet Video Error Concealment Based on Compressed Sensing and Regularized Least Squares
abstract
Error concealment (EC) is an important post processing technique to deal with the packet loss during the transmission of compressed video stream. This paper aims to address the problem of recovering the missing block in the decoded video stream from the perspective of compressed sensing. The missing block is assumed to be sparsely represented by a dictionary of prototype signal atoms. The atoms are generated by the motion-compensated blocks with a range of motion displacements from the temporally previously reconstructed frame. To avoid inefficient exploration for the prior of sparsity due to the potential coherency among atoms, the regularized least square is incorporated into the compressed sensing reconstruction for the recovery of the missing block. Experimental results demonstrate the superiority of the proposed EC method in terms of objective (PSNR) and subjective quality compared to the existing methods.
Changjun Fu, Xiangyang Ji, Yongbing Zhang 0002, Qionghai Dai
DCC3
2012 Content Adaptive Subsampling for Stereo Interleaving Video Coding
abstract
Stereo interleaving video coding receives considerable attention due to its desirable property of being compatible with 2D video coding standards. The errors caused by sub sampling (causing distortion between subsampling interpolated image and the original full resolution one) and by quantization during compression lead to the final distortion in stereo interleaving video coding. In this paper, the rate and distortion analysis in stereo interleaving video coding is provided. It proves that appropriate sub sampling in stereo interleaving video coding is able to obtain good compression performance. Subsequently, a content adaptive sub sampling (CAS) is proposed. In CAS, the half resolution frames are generated by decimation, where the down sampling filter coefficients are calculated based on frame contents and the targeted interpolation coefficients. Experiment results demonstrate that the CAS is able to achieve high compression efficiency of stereo interleaving encoding scheme for stereoscopic videos.
Yongbing Zhang 0002, Xiangyang Ji, Haoqian Wang, Lei Zhang 0006, Qionghai Dai
DCC1
2012 Geometric mapping assisted multi-view depth video coding
abstract
Multi-view plus depth (MVD), as a video representation supporting view synthesis based on depth video, has attracted more and more attention for the free view video (FVV) application. It is a challenge to efficiently compress the multi-view depth data in MVD format. In this paper, we explore the geometric relationships in 3D space and propose a geometric mapping assisted (GMA) multi-view depth video coding algorithm. The proposed GMA utilizes the mapped depth image as a reference candidate during prediction. Furthermore, the inpainting method is employed to fill in the holes in mapped depth images. Experimental results demonstrate the gains of up to 2.45 dB for the depth coding, as well as better quality of synthesized views.
Qiong Liu 0001, Yongbing Zhang 0002, Xiangyang Ji, Qionghai Dai
ICASSP2
2012 Robust joint reconstruction in compressed multi-view imaging
abstract
The newly emerging sampling methodology of compressed sensing opens a door to obtain compressed data directly. How to efficiently reconstruct the original signal from the compressed data becomes a new challenge. Many reconstruction works have been proposed on mono-view images by exploring the sparsity of the original image. However, it is a challenge to efficiently explore the correlations among different views in compressed multi-view imaging systems. With the aid of inter-view disparity information at receiver end, a joint reconstruction approach is presented for independently captured view-point images via compressed imaging. In the proposed approach, a robust reconstruction is obtained by formulating the occurrences of outliers, usually caused by illumination change, mismatch and discontinuity in disparity estimation, as a sparse model, which can be efficiently solved by a proximal sub-gradient algorithm bas ed on l1-norm minimization. Experimental results show that the joint reconstruction of compressed multi-view images can achieve significantly better recovery quality than the independently reconstructed ones.
Qionghai Dai, Changjun Fu, Xiangyang Ji, Yongbing Zhang 0002
PCS4
2012 Side information generation with auto regressive model for low-delay distributed video coding
Yongbing Zhang 0002, Debin Zhao, Hongbin Liu 0004, Yongpeng Li, Siwei Ma 0001, Wen Gao 0001
J. Vis. Commun. Image Represent.1
2012 Packet Video Error Concealment With Auto Regressive Model
abstract
In this paper, auto regressive (AR) model is applied to error concealment for block-based packet video coding. In the proposed error concealment scheme, the motion vector for each corrupted block is first derived by any kind of recovery algorithms. Then each pixel within the corrupted block is replenished as the weighted summation of pixels within a square centered at the pixel indicated by the derived motion vector in a regression manner. Two block-dependent AR coefficient derivation algorithms under spatial and temporal continuity constraints are proposed respectively. The first one derives the AR coefficients via minimizing the summation of the weighted square errors within all the available neighboring blocks under the spatial continuity constraint. The confidence weight of each pixel sample within the available neighboring blocks is inversely proportional to the distance between the sample and the corrupted block. The second one derives the AR coefficients by minimizing the summation of the weighted square errors within an extended block in the previous frame along the motion trajectory under the temporal continuity constraint. The confidence weight of each extended sample is inversely proportional to the distance toward the corresponding motion aligned block whereas the confidence weight of each sample within the motion aligned block is set to be one. The regression results generated by the two algorithms are then merged to form the ultimate restorations. Various experimental results demonstrate that the proposed error concealment strategy is able to improve both the objective and subjective quality of the replenished blocks compared to other methods.
Yongbing Zhang 0002, Xinguang Xiang, Debin Zhao, Siwei Ma 0001, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2011 Up-sampling Dependent Frame Rate Reduction for Low Bit-Rate Video Coding
abstract
Summary form only given. In low bit rate video coding, the frame rate of input sequence can be reduced to the half or even smaller portion by skipping or deleting frames before compression, and then the temporal resolution is restored via up-sampling at the decoder side. Numerous algorithms have been developed to address the problem of temporal resolution improvement. Actually, the quality of up-sampled frames depends on not only the performance of up-sampling method but also the information maintained in the down-sampled video sequence. To improve the quality of up-sampled frames and smooth the quality between the up-sampled and decompressed frames, this paper proposes an up-sampling dependent frame rate reduction, which is shown in Fig. 1. The proposed low bit rate video coding scheme is composed of up-sampling dependent frame rate reduction, compression, decompression and up-sampling components. The proposed frame rate reduction method is hinged to the temporal up-sampling. It is noted that there is a feedback between frame rate reduction and up-sampling in the proposed up-sampling dependent frame rate reduction, of which the goal is to obtain a down-sampled sequence maintaining more information about the frames to be up-sampled at the decoder side.
Yongbing Zhang 0002, Haoqian Wang, Debin Zhao
DCC1
2011 Stereoscopic video coding in AVS
abstract
This paper is an overview for AVS stereoscopic video coding technology, including two channels based inter-view prediction coding and stereo packing mode coding. The first one utilizes inter-view prediction to efficiently exploit the redundancy between the two channels of stereoscopic video. The superior coding performance of the inter-view prediction scheme benefits from an enhanced block prediction algorithm, which includes an improvement of direct mode for B-picture and motion vector prediction for P-picture. In addition, stereo packing mode, including side by side and top bottom, is adopted in AVS stereoscopic video coding to support the stereoscopic video service deployments based on the frame-compatible approach. Furthermore, an enhanced stereo packing mode is also developed to allow the prediction between signals coming from different channels within one packed frame. The simulation results demonstrate that the adopted techniques in AVS stereoscopic video are able to improve the compression efficiency of stereoscopic videos compared to simulcast one.
Xiangyang Ji, Yongbing Zhang 0002, Lu Yu 0003, Gwo Giun Lee
VCIP2
2011 Interpolation-Dependent Image Downsampling
abstract
Traditional methods for image downsampling commit to remove the aliasing artifacts. However, the influences on the quality of the image interpolated from the downsampled one are usually neglected. To tackle this problem, in this paper, we propose an interpolation-dependent image downsampling (IDID), where interpolation is hinged to downsampling. Given an interpolation method, the goal of IDID is to obtain a downsampled image that minimizes the sum of square errors between the input image and the one interpolated from the corresponding downsampled image. Utilizing a least squares algorithm, the solution of IDID is derived as the inverse operator of upsampling. We also devise a content-dependent IDID for the interpolation methods with varying interpolation coefficients. Numerous experimental results demonstrate the viability and efficiency of the proposed IDID.
Yongbing Zhang 0002, Debin Zhao, Jian Zhang 0018, Ruiqin Xiong, Wen Gao 0001
IEEE Trans. Image Process.1
2010 Auto Regressive Model and Weighted Least Squares Based Packet Video Error Concealment
abstract
In this paper, auto regressive (AR) model is applied to error concealment for block-based packet video encoding. Each pixel within the corrupted block is restored as the weighted summation of corresponding pixels within the previous frame in a linear regression manner. Two novel algorithms using weighted least squares method are proposed to derive the AR coefficients. First, we present a coefficient derivation algorithm under the spatial continuity constraint, in which the summation of the weighted square errors within the available neighboring blocks is minimized. The confident weight of each sample is inversely proportional to the distance between the sample and the corrupted block. Second, we provide a coefficient derivation algorithm under the temporal continuity constraint, where the summation of the weighted square errors around the target pixel within the previous frame is minimized. The confident weight of each sample is proportional to the similarity of geometric proximity as well as the intensity gray level. The regression results generated by the two algorithms are then merged to form the ultimate restorations. Various experimental results demonstrate that the proposed error concealment strategy is able to increase the peak signal-to-noise ratio (PSNR) compared to other methods.
Yongbing Zhang 0002, Xinguang Xiang, Siwei Ma 0001, Debin Zhao, Wen Gao 0001
DCC1
2010 Context-adaptive pixel based prediction for intra frame encoding
abstract
Intra prediction is one effective method to remove the spatial redundancies in intra frame coding. Better intra prediction will result in the residual with less energy, which will decrease the number of bits needed to reconstruct the signal at decoder. To improve the accuracy of intra prediction, a context-adaptive pixel based prediction (CAPBP) algorithm is proposed in this paper. For each pixel within the target block to be encoded, the prediction is calculated as the linear weighted summation of the reconstructed pixels within the left column and the above row. Based on the assumption that pixels having the same coordinates within one block own the same prediction weights, we calculate the corresponding weights for each pixel within the target block by the least square method. The same processing is also performed at decoder; hence the prediction weights do not need to be sent to the decoder. Experimental results verify that the proposed algorithm is able to improve the efficiency of intra frame coding up to 0.5dB.
Yongbing Zhang 0002, Li Zhang 0006, Siwei Ma 0001, Debin Zhao, Wen Gao 0001
ICASSP1
2010 Low bit-rate image coding via interpolation oriented adaptive down-sampling
abstract
An interpolation oriented adaptive down-sampling algorithm is proposed for low bit-rate image coding in this paper. Given an image, the proposed algorithm is able to obtain a low resolution image, from which a high quality image with the same resolution as the input image can be interpolated. Different from the traditional down-sampling algorithms, which are independent from the interpolation process, the proposed down-sampling algorithm hinges the down-sampling to the interpolation process. Consequently, the proposed down-sampling algorithm is able to maintain the original information of the input image to the largest extent. The down-sampled image is then fed into JPEG. A total variation (TV) based post processing is then applied to the decompressed low resolution image. Ultimately, the processed image is interpolated to maintain the original resolution of the input image. Experimental results verify that utilizing the downsampled image by the proposed algorithm, an interpolated image with much higher quality can be achieved. Besides, the proposed algorithm is able to achieve superior performance than JPEG for low bit rate image coding.
Yongbing Zhang 0002, Jian Zhang 0018, Ruiqin Xiong, Debin Zhao, Siwei Ma 0001
VCIP1
2010 Corrections to "A Spatio-Temporal Auto Regressive Model for Frame Rate Up-Conversion" [Sep 09 1289-1301]
abstract
In the above titled paper (ibid., vol. 19, no. 9, pp. 1289-1301, Sep. 09), the column under 3-DRS in Table II is incorrect due to the author's table editing error. The correct table is presented here.
Yongbing Zhang 0002, Debin Zhao, Xiangyang Ji, Ronggang Wang, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2010 A Motion-Aligned Auto-Regressive Model for Frame Rate Up Conversion
abstract
In this paper, a motion-aligned auto-regressive (MAAR) model is proposed for frame rate up conversion, where each pixel is interpolated as the average of the results generated by one forward MAAR (Fw-MAAR) model and one backward MAAR (Bw-MAAR) model. In the Fw-MAAR model, each pixel in the to-be-interpolated frame is generated as a linear weighted summation of the pixels within a motion-aligned square neighborhood in the previous frame. To derive more accurate interpolation weights, the aligned actual pixels in the following frame are also estimated as a linear weighted summation of the newly interpolated pixels in the to-be-interpolated frame by the same weights. Consequently, the backward-aligned actual pixels in the following frame can be estimated as a weighted summation of the corresponding pixels within an enlarged square neighborhood in the previous frame. The Bw-MAAR is performed likewise except that it is operated in the reverse direction. A damping Newton algorithm is then proposed to compute the adaptive interpolation weights for the Fw-MAAR and Bw-MAAR models. Extensive experiments demonstrate that the proposed MAAR model is able to achieve superior performance than the traditional frame interpolation methods such as MCI, OBMC, and AOBMC, and it is even better than STAR model for the most test sequences with moderate or large motions.
Yongbing Zhang 0002, Debin Zhao, Siwei Ma 0001, Ronggang Wang, Wen Gao 0001
IEEE Trans. Image Process.1
2009 Joint learning for side information and correlation model based on linear regression model in distributed video coding
abstract
The coding efficiency of distributed video coding system is significantly determined by the side information quality and correlation model. Motivated by theoretical analysis of the maximum likelihood treatment for linear regression model, we propose a novel joint online learning model for side information generation and correlation model estimation in this paper. In our proposed scheme, each pixel in the side information is approximated as the linear weighted combination of samples within a local spatio-temporal neighboring space. Weights are trained in a self-feedback fashion, during which the correlation model parameters can also be achieved. The efficiency of the proposed joint learning model is confirmed experimentally.
Xianming Liu 0005, Debin Zhao, Yongbing Zhang 0002, Siwei Ma 0001, Qingming Huang, Wen Gao 0001
ICIP3
2009 Local adaptive learning and fusion for side information interpolation in distributed video coding
abstract
Motivated by theoretical analysis of the curve fitting problem based on equivalent kernel, in this paper we propose a local adaptive learning and fusion model for side information interpolation in distributed video coding. In the proposed model, each pixel in the interpolated frame is approximated as the linear combination of samples within a local spatio-temporal window using kernel parameters as weight. The size of training window can be adaptive to the motion characteristic of video, from samples in which the kernel parameters can be locally learned. In order to further improve the quality of interpolated frames, we introduce a belief-projection based fusion strategy with adaptive weights for multiple interpolated results which are with the same time index. Experimental results demonstrate that the proposed learning and fusion model is effective in performance for side information interpolation in distributed video coding.
Xianming Liu 0005, Yongbing Zhang 0002, Yongpeng Li, Hongbin Liu 0004, Siwei Ma 0001, Debin Zhao
PCS2
2009 A high efficient error concealment scheme based on auto-regressive model for video coding
abstract
In this paper, a high efficient temporal error concealment scheme based on auto-regressive (AR) model is proposed for video coding. The proposed AR based error concealment scheme includes a forward AR model for P slice, and a bi-direction AR model for B slice. First, we utilize the block matching algorithm (BMA) to select the best motions for lost blocks from the motions of available neighboring blocks. Then, the proposed AR model coefficients are computed according to the spatial neighboring pixels and their temporal-correlated pixels indicated by the selected best motions. Finally, applying the AR model, each pixel of the lost block is interpolated as a weighted summation of pixels in the reference frame along the selected best motions. Simulation results show that the performance of the proposed scheme is superior to conventional temporal error concealment methods.
Xinguang Xiang, Yongbing Zhang 0002, Debin Zhao, Siwei Ma 0001, Wen Gao 0001
PCS2
2009 An auto-regressive model for checkboard splitting based Wyner-Ziv coding
abstract
An auto-regressive (AR) based side information (SI) generation is proposed in this paper for block based chessboard pattern Wyner-Ziv (WZ) coding, where each WZ frame is split into two sets at encoder and then encoded separately. At the decoder, one set of the WZ frame will be firstly reconstructed, and then proposed AR model is used to generate the SI of the other set, where each pixel is generated as a linear weighted summation of pixels within two square windows in the previous and following reconstructed WZ/key frames along the motion trajectory. To obtain high quality SI for the second set, reconstructed pixels in the four neighboring blocks of the first set are employed to derive accurate AR coefficients. Several experimental results demonstrate that the proposed AR model is able to improve the quality of the SI for the second set of the WZ frame, which leads to the improvement of the rate-distortion performance of the WZ coding.
Yongbing Zhang 0002, Debin Zhao, Siwei Ma 0001, Ronggang Wang, Wen Gao 0001
PCS1
2009 A Spatio-Temporal Auto Regressive Model for Frame Rate Upconversion
abstract
This paper proposes a spatio-temporal auto regressive (STAR) model for frame rate upconversion. In the STAR model, each pixel in the interpolated frame is approximated as the weighted combination of a sample space including the pixels within its two temporal neighborhoods from the previous and following original frames as well as the available interpolated pixels within its spatial neighborhood in the current to-be-interpolated frame. To derive accurate STAR weights, an iterative self-feedback weight training algorithm is proposed. In each iteration, first the pixels of each training window in the interpolated frames are approximated by the sample space from the previous and following original frames and the to-be-interpolated frame. And then the actual pixels of each training window in the original frame are approximated by the sample space from the previous and following interpolated frames and the current original frame with the same weights. The weights of each training window are calculated by jointly minimizing the distortion between the interpolated frames in the current and previous iterations as well as the distortion between the original frame and its interpolated one. Extensive simulation results demonstrate that the proposed STAR model is able to yield the interpolated frames with high performance in terms of both subjective and objective qualities.
Yongbing Zhang 0002, Debin Zhao, Xiangyang Ji, Ronggang Wang, Wen Gao 0001
IEEE Trans. Circuits Syst. Video Technol.1
2007 A Spatio-Temporal Autoregressive Frame Rate Up Conversion Scheme
abstract
A spatio-temporal autoregressive model is proposed in this paper to address the problem of frame rate up conversion. Every pixel in a skipped frame is generated as a linear combination of pixel values from forward and backward reference frames. At the beginning of the presented scheme, the coarse model parameters are computed according to the given initial pixel values for skipped frames. Then the coarse parameters are refined by an iteration process, during which we also interpolate the original low rate frames by the two closest generated skipped frames to derive more accurate parameters. Experimental results verify that the proposed algorithm significantly improves both the subjective and objective quality of the interpolated frames.
Yongbing Zhang 0002, Debin Zhao, Xiangyang Ji, Ronggang Wang, Xilin Chen 0001
ICIP (1)1