Tieyong Zeng

dblp:63/2745 · DBLP profile ↗
← Back
150ranked-venue papers
4as first author
108since 2021 · last 2026
0000-0002-0688-202XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 83 · 4 first-author · 48 since 2021Artificial intelligence and machine learning · 60 · 53 since 2021Applied, interdisciplinary, general and emerging computing · 19 · 18 since 2021Databases, data management, data science and information retrieval · 5 · 5 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RPE-PAD: Relative Pose Estimation for Pose-agnostic Anomaly Detection
abstract
Pose-agnostic Anomaly Detection (PAD) aims to detect anomalies when the poses of query images are unknown and differ from those in the training set. Therefore, accurately estimating the camera poses for the query images in the test set is critical for this task. Existing query-specific framework methods require re-optimizing a new set of parameters for each query image, limiting their generalization and increasing computational burden. To overcome these limitations, we propose a novel method, Relative Pose Estimation for Pose-agnostic Anomaly Detection (RPE-PAD), which enhances both generalization and efficiency with a query-independent framework. Specifically, we propose a Random View Synthesis Scheme (RVSS) that generates new poses by adding Gaussian perturbations to the original poses, then renders the corresponding views to augment the dataset. To estimate the relative camera pose between two input images, we introduce an Iterative Relative Pose Refinement Network (IRPRN), which incorporates a hierarchical coarse-to-fine refinement strategy. Furthermore, we employ a Multi-Pair Training Strategy (MPTS) to train the proposed IRPRN, leveraging multiple image pairs to expand the relative pose transformation space during training. Extensive experiments demonstrate that our method achieves robust anomaly detection performance while significantly improving inference efficiency.
Mengzan Qi, Rongkang Ma, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
AAAI6
2026 An efficient algorithm for vertex enumeration of arrangement
Zelin Dong, Fenglei Fan, Huan Xiong, Tieyong Zeng
Discret. Appl. Math.4
2026 Recurrent Mamba for efficient and high-fidelity MRI reconstruction
Zhaoming Hou, Boying Wu, Qiyu Jin, Dong Liang 0001, Tieyong Zeng
Expert Syst. Appl.6
2026 Hyper-Compression: Model Compression via Hyperfunction
abstract
The rapid growth of large models' size has far outpaced that of computing resources. To bridge this gap, encouraged by the parsimonious relationship between genotype and phenotype in the brain's growth and development, we propose the so-called Hyper-Compression that turns the model compression into the issue of parameter representation via a hyperfunction. Specifically, it is known that the trajectory of some low-dimensional dynamic systems can fill the high-dimensional space eventually. Thus, Hyper-Compression, using these dynamic systems as the hyperfunctions, represents the parameters of the target network by their corresponding composition number or trajectory length. This suggests a novel mechanism for model compression, substantially different from the existing pruning, quantization, distillation, and decomposition. Along this direction, we methodologically identify a suitable dynamic system with the irrational winding as the hyperfunction and theoretically derive its associated error bound. Next, guided by our theoretical insights, we propose several engineering twists to make the Hyper-Compression pragmatic and effective. Lastly, systematic and comprehensive experiments on NLP models such as LLaMA and Qwen series and vision models confirm that Hyper-Compression enjoys the following PNAS merits: 1) Preferable compression ratio; 2) No post-hoc retraining; 3) Affordable inference time; and 4) Short compression time. It compresses LLaMA2-7B in an hour and achieves close-to-int4-quantization performance, without retraining and with a performance drop of less than 1%.
Fenglei Fan, Juntong Fan, Dayang Wang, Jingbo Zhang 0002, Zelin Dong, Ge Wang 0001, Tieyong Zeng
IEEE Trans. Pattern Anal. Mach. Intell.8
2026 Quaternion adaptive approximation normalization graph guided implicit low rank for robust matrix completion
Yu Guo 0008, Tieyong Zeng, Qiyu Jin, Michael Kwok-Po Ng
Pattern Recognit.4
2026 Localized simple multiple kernel k-means with matrix regularization
Boying Wu, Qiyu Jin, Tieyong Zeng
Pattern Recognit.5
2026 Effective Solutions to Robust Orthogonal Nonnegative Matrix Factorization via Oblique Manifold Transformation
abstract
Abstract. In this paper, we address the problem of robust orthogonal nonnegative matrix factorization (RONMF), a crucial challenge in data analysis. We first propose a RONMF model that explicitly handles both dense and sparse noise, making it suitable for a wide range of real-world applications. To circumvent the computational complexity associated with the Stiefel manifold and effectively solve the proposed model, we introduce an exact penalty method that transforms the optimization problem from the Stiefel manifold to the Oblique manifold. To achieve this, we develop the EP-RONMF algorithm, which seeks a point satisfying the weak second-order optimality conditions through an alternating proximal method and iterative updates of the penalty parameter. This approach successfully addresses the nonconvex nature of the ONMF problem and ensures convergence to a stable solution. To validate the efficacy of our method, we conducted extensive experiments on diverse datasets, including image, text, and hyperspectral data. The results clearly demonstrate the superiority of our approach compared to existing techniques.
Fan Jia 0007, Yuxiang Hui, Tieyong Zeng
SIAM J. Imaging Sci.3
2026 A structure adaptivity variation-based segmentation model for image with retinex and noise
Tieyong Zeng, Zhi-Feng Pang
Signal Process.2
2026 GIGAS: Adversarial Attacks on Visual Question Answering With Multi-Modal Generative Models
abstract
VQA models, which answer questions about images by combining both visual and textual information, have been proven susceptible to adversarial attacks. These attacks introduce subtle perturbations to the input data to manipulate the model’s predictions. This paper focuses on adversarial attacks targeting VQA models that follow the “pre-training & fine-tuning” paradigm, an area that remains under-explored. We have identified two key issues in the current field. On one hand, existing multi-modal attacks have low ASR due to inter-modal semantic inconsistency from insufficient cross-modal interaction. On the other hand, the dilemma between attack effectiveness and stealthiness limits the practical applicability of adversarial texts. To address these issues, we propose GIGAS, an innovative attack that uses multi-modal generative models to explore multi-modal interaction through three key modules tailored to solve above-mentioned problems. MIGA aligns adversarial visual features with semantics of misleading images generated by multi-modal generative models to mitigate cross-modal inconsistencies. GSA employs MLLMs to generate natural adversarial texts with greater variation and evaluate similarity to filter based on clean images, balancing effectiveness and stealthiness. Iteration Allocation dynamically adjusts attack iterations based on image-text similarity, maximizing the utility of the limited iterations. Experiments conducted on various VL models and VQA datasets demonstrate superior attack performance, with an average ASR of 89.09% on VQAv2.0. Furthermore, our GIGAS exhibits outstanding transferability, around 60% ASR, across diverse models and specific domains. Our code will be available at: https://github.com/Yvonna-cloud/GIGAS.
Yunxuan Li, Jing Yu 0007, Tieyong Zeng, Liyan Ma
IEEE Trans. Circuits Syst. Video Technol.3
2026 MSAFed: Generalized Multi-Stage and Adaptive Federated Learning for Test-Time Medical Segmentation
abstract
Federated learning (FL) enables collaborative model training across multiple medical centers without sharing data, offering significant promise for privacy-preserving AI in healthcare. However, FL models often lack generalization across all participating clients (inside FL) and perform poorly when deployed to unseen clients (outside FL), particularly in heterogeneous domains. Current test-time adaptation methods for outside FL fail to address biases in personalized models toward source distributions, limiting their clinical applications. To tackle these challenges, we propose MSAFed, a generalized multi-stage adaptive FL framework that enhances both inside generalization and outside test-time adaptation. During pretraining, intra-client and inter-client contrastive learning with prototype-aware aggregation produces a generalized global model. An adaptive learning rate strategy further improves inside FL generalization. For unseen clients, source knowledge, including adaptive learning rates and prototypes, is leveraged to dynamically adapt the network architecture during test time. Experiments on three real-world multi-center medical datasets demonstrate the effectiveness of MSAFed, achieving superior performance on both inside and outside FL tasks.
Jiajie Jin, Xuanmin Chen, Liyan Ma, Shihui Ying, Guang Yang 0006, Tieyong Zeng
IEEE J. Biomed. Health Informatics6
2026 Msa-Splatting: Multi-Scale Adaptive Gaussian Splatting for High-Fidelity View Synthesis
abstract
3D Gaussian Splatting (3DGS) has revolutionized novel view synthesis with real-time rendering and high-quality reconstruction. However, its performance degrades under varying scales, causing aliasing and dilation artifacts. To address these challenges, we introduce Msa-Splatting, an enhanced adaptive Gaussian splatting method designed for high-fidelity multi-scale novel view synthesis. Our approach reimagines the Gaussian adaptive control process with a novel multi-fold splitting and cloning mechanism, which optimizes Gaussian properties and reduces spurious artifacts by halving the densification frequency. We further mitigate aliasing in scaled-down renderings with a weighted alpha blending sampling technique, ensuring accurate representation of high-frequency details. To eliminate scale-mismatch artifacts, we introduce 2D variable dilation Gaussians, which dynamically adjust dilation based on the rendering scale. Msa-Splatting significantly enhances the fidelity and adaptability of Gaussian-based rendering, achieving state-of-the-art performance in multi-scale novel view synthesis. Extensive quantitative and qualitative evaluations demonstrate the efficacy of our approach, showcasing its ability to produce perceptually superior renderings with reduced artifacts across varying scales, thus establishing a robust framework for future advancements in scalable rendering technologies.
Yaoyong Zhao, Boying Wu, Qiyu Jin, Tieyong Zeng
IEEE Trans. Multim.5
2026 Gradient-Refined Federated Learning on Head-Tail Imbalanced Data
abstract
Federated learning has emerged as a transformative paradigm for distributed data collaboration, facilitating knowledge aggregation across multiple local clients through a global server while rigorously preserving data privacy. However, its performance is significantly hindered by the global head-tail imbalance, where tail classes with scarce data are often dominated by head classes. This challenge, known as federated long-tailed learning, arises from the intrinsic conflict between class knowledge acquisition and privacy preservation. Existing methodologies falter in resolving this conflict, as the abstraction of data knowledge in federated communication complicates the extraction of class-level knowledge, resulting in imbalanced global models and diminished performance. To simultaneously address this imbalance and uphold privacy, we introduce FedGRE, a gradient-refined federated learning approach that constructs global gradients and facilitates refined global gradient descent. FedGRE enhances gradients through two pivotal mechanisms: accumulation diffusion and accumulation refinement. The former amalgamates accumulated gradients with stochastic gradient perturbations to alleviate class imbalance, while the latter utilizes the accumulation as an anchor to calibrate global gradient updates, ensuring consistency and mitigating oscillations. Additionally, we implement a consistency integration technique to incorporate the refined accumulation into the global model, guaranteeing privacy-preserving and class-balanced global optimization. Extensive experiments on six datasets demonstrate that FedGRE significantly outperforms 14 state-of-the-art (SOTA) methods in federated long-tailed classification while maintaining robust privacy protection.
Heye Zhang, Chenchu Xu, Lin Gu 0003, Jingfeng Zhang, Tieyong Zeng, Zhifan Gao
IEEE Trans. Neural Networks Learn. Syst.6
2025 OS Agents: A Survey on MLLM-based Agents for Computer, Phone and Browser Use
abstract
Xueyu Hu, Tao Xiong, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao, Yuhuai Li, Shengze Xu, Shenzhi Wang, Xinchen Xu, Shuofei Qiao, Zhaokai Wang, Kun Kuang, Tieyong Zeng, Liang Wang, Jiwei Li, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang, Keting Yin, Zhou Zhao, Hongxia Yang, Fan Wu, Shengyu Zhang, Fei Wu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Xueyu Hu, Biao Yi, Zishu Wei, Ruixuan Xiao, Yurun Chen 0004, Jiasheng Ye, Meiling Tao, Xiangxin Zhou, Ziyu Zhao 0001, Yuhuai Li, Shengze Xu, Shenzhi Wang, Shuofei Qiao, Zhaokai Wang, Kun Kuang 0001, Tieyong Zeng, Liang Wang 0001, Jiwei Li 0001, Yuchen Eleanor Jiang, Wangchunshu Zhou, Guoyin Wang 0002, Keting Yin, Zhou Zhao 0001, Hongxia Yang, Fan Wu 0006, Shengyu Zhang 0001, Fei Wu 0001
ACL (1)18
2025 Blind Noisy Image Deblurring Using Residual Guidance Strategy
Heyan Liu, Jun Liu 0012, Xi-Le Zhao, Tingting Wu 0001, Tieyong Zeng
ICCV6
2025 Learning Cocoercive Conservative Denoisers via Helmholtz Decomposition for Poisson Imaging Inverse Problems
abstract
Plug-and-play (PnP) methods with deep denoisers have shown impressive results in imaging problems. They typically require strong convexity or smoothness of the fidelity term and a (residual) non-expansive denoiser for convergence. These assumptions, however, are violated in Poisson inverse problems, and non-expansiveness can hinder denoising performance. To address these challenges, we propose a cocoercive conservative (CoCo) denoiser, which may be (residual) expansive, leading to improved denoising performance. By leveraging the generalized Helmholtz decomposition, we introduce a novel training strategy that combines Hamiltonian regularization to promote conservativeness and spectral regularization to ensure cocoerciveness. We prove that CoCo denoiser is a proximal operator of a weakly convex function, enabling a restoration model with an implicit weakly convex prior. The global convergence of PnP methods to a stationary point of this restoration model is established. Extensive experimental results demonstrate that our approach outperforms closely related methods in both visual quality and quantitative metrics.
Deliang Wei, Peng Chen 0045, Jiale Yao, Fang Li 0004, Tieyong Zeng
NeurIPS6
2025 Progressive Tasks Guided Multi-Source Network for Customer Lifetime Value Prediction in Online Advertising
abstract
Customer lifetime value (LTV) is crucial to companies who are intending to adopt personalized promoting strategies to optimize the profits. However, LTV prediction in the scenario of online App advertising usually suffers from label sparsity issue, towards which existing methods designed complex model structures but ignored the information contained in intermediate user behaviors. Moreover, previous works mainly focus on fitting the overall LTV distribution, overlooking the fact that LTV in online App advertising consists of sources with diverse data distributions and thus resulting in sub-optimal solutions. In this paper, we propose a novel Progressive Tasks guided Multi-Source Network (PTMSN) to tackle the aforementioned problems. Specifically, a Cascaded Sub-task Module (CSM) is introduced to alleviate data sparsity by modeling reliance between explicit interactions and implicit monetization. In addition, as the overall LTV is assembled from multiple sources, we propose a divide-and-conquer scheme named Multi-source Integrating Module (MIM) to disentangle the original single target into several source distributions and model in a fine-grained manner. Extensive offline experiments on real-world industrial datasets compared to state-of-the-art baseline models validate the effectiveness of our approach. PTMSN has been successfully deployed in industrial online advertising system, serving various business scenarios and acquiring 2.97% absolute ROI gains.
Xingyu Lou, Chiye Ou, Feng Liu 0047, Tieyong Zeng, Chengwei He, Lilong Wei, Jun Wang 0020
WSDM6
2025 Randomly Projected Convex Clustering Model: Motivation, Realization, and Cluster Recovery Guarantees
abstract
In this paper, we propose a randomly projected convex clustering model for clustering a collection of $n$ high dimensional data points in $\mathbb{R}^d$ with $K$ hidden clusters. Compared to the convex clustering model for clustering original data with dimension $d$, we prove that, under some mild conditions, the perfect recovery of the cluster membership assignments of the convex clustering model, if exists, can be preserved by the randomly projected convex clustering model with embedding dimension $m = O(\epsilon^{-2}\log(n))$, where $\epsilon > 0$ is some given parameter. We further prove that the embedding dimension can be improved to be $O(\epsilon^{-2}\log(K))$, which is independent of the number of data points. We also establish the recovery guarantees of our proposed model with uniform weights for clustering a mixture of spherical Gaussians. Extensive numerical results demonstrate the robustness and superior performance of the randomly projected convex clustering model. The numerical results will also demonstrate that the randomly projected convex clustering model can outperform other popular clustering models on the dimension-reduced data, including the randomly projected K-means model.
Yancheng Yuan, Jiaming Ma, Tieyong Zeng, Defeng Sun
J. Mach. Learn. Res.4
2025 Quaternion deep matrix factorization and non-local Laplacian regularization for matrix completion
Yu Guo 0008, Qiyu Jin, Tieyong Zeng, Michael Kwok-Po Ng
Knowl. Based Syst.4
2025 MPGB: Learning discriminative embeddings with multi-prototype and gradient balancing strategy for multi-modal 3D open world object detection
Liyan Ma, Zhi Li 0080, Tieyong Zeng
Knowl. Based Syst.4
2025 Don't fear peculiar activation functions: EUAF and beyond
Qianchao Wang, Dong Zeng, Zhaoheng Xie, Hengtao Guo, Tieyong Zeng, Fenglei Fan
Neural Networks6
2025 One Neuron Saved is One Neuron Earned: On Parametric Efficiency of Quadratic Networks
abstract
Inspired by neuronal diversity in the biological neural system, a plethora of studies proposed to design novel types of artificial neurons and introduce neuronal diversity into artificial neural networks. Recently proposed quadratic neuron, which replaces the inner-product operation in conventional neurons with a quadratic one, have achieved great success in many essential tasks. Despite the promising results of quadratic neurons, there is still an unresolved issue: Is the superior performance of quadratic networks simply due to the increased parameters or due to the intrinsic expressive capability? Without clarifying this issue, the performance of quadratic networks is always suspicious. Additionally, resolving this issue is reduced to finding killer applications of quadratic networks. In this paper, with theoretical and empirical studies, we show that quadratic networks enjoy parametric efficiency, thereby confirming that the superior performance of quadratic networks is due to the intrinsic expressive capability. This intrinsic expressive ability comes from that quadratic neurons can easily represent nonlinear interaction, while it is hard for conventional neurons. Theoretically, we derive the approximation efficiency of quadratic networks over conventional ones in terms of real space and manifolds. Moreover, from the perspective of the Barron space, we demonstrate that there exists a functional space whose functions can be approximated by quadratic networks in a dimension-free error, but the approximation error of conventional networks is dependent on dimensions. Empirically, experimental results on synthetic data, classic benchmarks, and real-world applications show that quadratic models broadly enjoy parametric efficiency, and the gain of efficiency depends on the task.
Fenglei Fan, Hangcheng Dong, Zhongming Wu, Lecheng Ruan, Tieyong Zeng, Yiming Cui 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Quaternion Nuclear Norm Minus Frobenius Norm Minimization for color image reconstruction
Yu Guo 0008, Tieyong Zeng, Qiyu Jin, Michael Kwok-Po Ng
Pattern Recognit.3
2025 A Contrast-Saturation Adaptive Model for Low-Light Image Enhancement
abstract
Abstract. Previous Retinex-based methods simultaneously estimate illumination and reflectance, complicating model design and potentially impacting image enhancement due to flawed prior assumptions. In this paper, we explore the physical basis of low-light image enhancement, focusing on contrast, saturation, and brightness, and propose an adaptive contrast-saturation (ConSat) model. Our ConSat innovation simplifies model design by focusing solely on a brightness enhancement function, allowing for precise image brightness control while keeping colors natural and realistic. We design a new saturation metric that assesses pixel dispersion from the mean of RGB channels, accurately depicting exposure variations across the image. On this basis, we propose a smart contrast stretching function that adapts contrast adjustments to enhance images under varying light. It boosts contrast in dark, low-saturation areas to clarify details and textures, while curbing it in bright, high-saturation areas to prevent overexposure and color distortion. Finally, we employ a pretrained convolutional neural network (CNN)-based denoiser to achieve satisfactory visual appeal. Numerical experiments show that our ConSat is superior to other state-of-the-art methods in brightness improvement, noise removal, and artifact elimination.
Qianting Ma, Yang Wang 0020, Tieyong Zeng
SIAM J. Imaging Sci.3
2025 An Adaptive Nonuniform Map Re-weighting Method for Image Dehazing
abstract
Abstract. Nonhomogeneous image dehazing has long been a highly challenging problem due to the uneven distribution of haze in real-world images. This challenge is further exacerbated given complex weather conditions, making it even more difficult to effectively remove haze and restore image clarity. To address this issue, we propose a novel variational model based on the Atmospheric Scattering Model, incorporating an adaptive nonuniform map reweighting technique. Our approach dynamically adjusts to different haze levels across an image, ensuring effective region-specific dehazing. Meanwhile, we introduce a localized spatial-aware regularization, which enables us to improve consistency in regions with heavier haze while providing more flexibility in areas with rich details. To solve the proposed model, we introduce a Block Coordinate Proximal Gradient (BCPG) method and rigorously analyze the existence and convergence properties of this algorithm, ensuring that it reliably produces optimal solutions. Extensive experiments have been conducted to validate the effectiveness of our method. The results demonstrate that our approach significantly improves image clarity and detail, even given challenging nonuniform haze conditions. Our findings indicate that the proposed method is robust and effective, making it a valuable tool for addressing nonhomogeneous image dehazing given complex weather conditions.
Hao Zhang 0154, Te Qi, Hok Shing Wong, Tieyong Zeng
SIAM J. Imaging Sci.4
2025 Scene recovery with detail-preserving
Tingting Wu 0001, Jun Liu 0012, Tieyong Zeng
Signal Process. Image Commun.4
2025 Adaptive Superpixel-Guided Non-Homogeneous Image Dehazing
abstract
Image dehazing is regarded as a fundamental image processing task with a major impact on higher-level imaging tasks. Many existing haze removal methods are designed for homogeneous haze, but in real-world cases, the haze is normally non-homogeneous. Superpixels, which segment an image into a set of closely spaced regions, can be employed in real-world scenarios to deal with non-homogeneous haze. In our paper, an adaptive non-homogeneous image dehazing approach that utilizes the superpixel-guided algorithm is designed to segment different hazy regions. Considering that both ambient light and transmission map estimation have a significant impact on the results, our research focuses on the development of a variational dehazing model that takes into account non-uniform ambient light and non-uniform transmission maps to address varying levels of haze. A series of numerical results illustrate the superiority and efficacy of our method.
Hao Zhang 0154, Ping Luo 0002, Te Qi, Tieyong Zeng
IEEE Signal Process. Lett.5
2025 Boosting Geometric Invariants for Discriminative Forensics of Large-Scale Generated Visual Content
abstract
Generative artificial intelligence has shown great success in visual content synthesis such that humans struggle to distinguish between real and synthesized images. Forensic research seeks to reveal artifacts in such generated images, ensuring information security or improving generation capability. In this regard, the robustness and interpretability are important for the trustworthy purpose of forensic tasks. However, typical forensic models and their underlying data representations rely on empirical learning algorithms, which cannot effectively handle the high robustness and interpretability requirements beyond experience. As an effective solution, we extend the classical geometric invariants to the forensic research of large-scale generated images. Invariants are handcrafted representations with robust and interpretable geometric principles. However, their discriminability is far from the large scale of today's forensic tasks. We boost the discriminability by extending the classical invariants to the hierarchical architecture of convolutional neural networks. The resulting overcompleteness allows for an automatic selection of task-discriminative features, while retaining the previous advantages of robustness and interpretability. From generative adversarial networks to diffusion models, the forensic with our boosted invariants demonstrates state-of-the-art discriminability against large-scale content diversity. It also exhibits high efficiency on training examples, intrinsic invariance to geometric variations, and better interpretability of the forensic process.
Chao Wang 0028, Yushu Zhang 0001, Xiangyu Chen 0006, Yi Zhang 0018, Tieyong Zeng, Fenglei Fan
IEEE Trans. Image Process.7
2025 High-Frequency Modulated Transformer for Multi-Contrast MRI Super-Resolution
abstract
Accelerating the MRI acquisition process is always a key issue in modern medical practice, and great efforts have been devoted to fast MR imaging. Among them, multi-contrast MR imaging is a promising and effective solution that utilizes and combines information from different contrasts. However, existing methods may ignore the importance of the high-frequency priors among different contrasts. Moreover, they may lack an efficient method to fully utilize the information from the reference contrast. In this paper, we propose a lightweight and accurate High-frequency Modulated Transformer (HFMT) for multi-contrast MRI super-resolution. The key ideas of HFMT are high-frequency prior enhancement and its fusion with global features. Specifically, we employ an enhancement module to enhance and amplify the high-frequency priors in the reference and target modalities. In addition, we utilize the Rectangle Window Transformer Block (RWTB) to capture global information in the target contrast. Meanwhile, we propose a novel cross-attention mechanism to fuse the high-frequency enhanced features with the global features sequentially, which assists the network in recovering clear texture details from the low-resolution inputs. Extensive experiments show that our proposed method can reconstruct high-quality images with fewer parameters and faster inference time.
Juncheng Li 0003, Hanhui Yang, Qiaosi Yi, Minhua Lu, Jun Shi 0004, Tieyong Zeng
IEEE Trans. Medical Imaging6
2025 Auxiliary Representation Guided Network for Visible-Infrared Person Re-Identification
abstract
Visible-Infrared Person Re-identification aims to retrieve images of specific identities across modalities. To relieve the large cross-modality discrepancy, researchers introduce the auxiliary modality within the image space to assist modality-invariant representation learning. However, the challenge persists in constraining the inherent quality of generated auxiliary images, further leading to a bottleneck in retrieval performance. In this paper, we propose a novel Auxiliary Representation Guided Network (ARGN) to explore the potential of auxiliary representations, which are directly generated within the modality-shared embedding space. In contrast to the original visible and infrared representations, which contain information solely from their respective modalities, these auxiliary representations integrate cross-modality information by fusing both modalities. In our framework, we utilize these auxiliary representations as modality guidance to reduce the cross-modality discrepancy. First, we propose a High-quality Auxiliary Representation Learning (HARL) framework to generate identity-consistent auxiliary representations. The primary objective of our HARL is to ensure that auxiliary representations capture diverse modality information from both modalities while concurrently preserving identity-related discrimination. Second, guided by these auxiliary representations, we design an Auxiliary Representation Guided Constraint (ARGC) to optimize the modality-shared embedding space. By incorporating this constraint, the modality-shared embedding space is optimized to achieve enhanced intra-identity compactness and inter-identity separability, further improving the retrieval performance. In addition, to improve the robustness of our framework against the modality variation, we introduce a Part-based Adaptive Gaussian Module (PAGM) to adaptively extract discriminative information across modalities. Finally, extensive experiments are conducted to demonstrate the superiority of our method over state-of-the-art approaches on three VI-ReID datasets.
Mengzan Qi, Sixian Chan 0001, Chen Hang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
IEEE Trans. Multim.5
2025 SAB Net: A Semantic Attention Boosting Framework for Semantic Segmentation
abstract
Semantic segmentation has achieved great progress by effectively fusing features of contextual information. In this article, we propose an end-to-end semantic attention boosting (SAB) framework to adaptively fuse the contextual information iteratively across layers with semantic regularization. Specifically, we first propose a pixelwise semantic attention (SAP) block, with a semantic metric representing the pixelwise category relationship, to aggregate the nonlocal contextual information. In addition, we improve the computation complexity of SAP block from to for images with size . Second, we present a categorywise semantic attention (SAC) block to adaptively balance the nonlocal contextual dependencies and the local consistency with a categorywise weight, overcoming the contextual information confusion caused by the feature imbalance within intra-category. Furthermore, we propose the SAB module to refine the segmentation with SAC and SAP blocks. By applying the SAB module iteratively across layers, our model shrinks the semantic gap and enhances the structure reasoning by fully utilizing the coarse segmentation information. Extensive quantitative evaluations demonstrate that our method significantly improves the segmentation results and achieves superior performance on the PASCAL VOC 2012, Cityscapes, PASCAL Context, and ADE20K datasets.
Xiaofeng Ding 0003, Chaomin Shen 0001, Tieyong Zeng, Yaxin Peng
IEEE Trans. Neural Networks Learn. Syst.3
2025 Fast and Reliable Score-Based Generative Model for Parallel MRI
abstract
The score-based generative model (SGM) can generate high-quality samples, which have been successfully adopted for magnetic resonance imaging (MRI) reconstruction. However, the recent SGMs may take thousands of steps to generate a high-quality image. Besides, SGMs neglect to exploit the redundancy in space. To overcome the above two drawbacks, in this article, we propose a fast and reliable SGM (FRSGM). First, we propose deep ensemble denoisers (DEDs) consisting of SGM and the deep denoiser, which are used to solve the proximal problem of the implicit regularization term. Second, we propose a spatially adaptive self-consistency (SASC) term as the regularization term of the -space data. We use the alternating direction method of multipliers (ADMM) algorithm to solve the minimization model of compressed sensing (CS)-MRI incorporating the image prior term and the SASC term, which is significantly faster than the related works based on SGM. Meanwhile, we can prove that the iterating sequence of the proposed algorithm has a unique fixed point. In addition, the DED and the SASC term can significantly improve the generalization ability of the algorithm. The features mentioned above make our algorithm reliable, including the fixed-point convergence guarantee, the exploitation of the space, and the powerful generalization ability.
Ruizhi Hou, Fang Li 0004, Tieyong Zeng
IEEE Trans. Neural Networks Learn. Syst.3
2024 Triple Feature Disentanglement for One-Stage Adaptive Object Detection
abstract
In recent advancements concerning Domain Adaptive Object Detection (DAOD), unsupervised domain adaptation techniques have proven instrumental. These methods enable enhanced detection capabilities within unlabeled target domains by mitigating distribution differences between source and target domains. A subset of DAOD methods employs disentangled learning to segregate Domain-Specific Representations (DSR) and Domain-Invariant Representations (DIR), with ultimate predictions relying on the latter. Current practices in disentanglement, however, often lead to DIR containing residual domain-specific information. To address this, we introduce the Multi-level Disentanglement Module (MDM) that progressively disentangles DIR, enhancing comprehensive disentanglement. Additionally, our proposed Cyclic Disentanglement Module (CDM) facilitates DSR separation. To refine the process further, we employ the Categorical Features Disentanglement Module (CFDM) to isolate DIR and DSR, coupled with category alignment across scales for improved source-target domain alignment. Given its practical suitability, our model is constructed upon the foundational framework of the Single Shot MultiBox Detector (SSD), which is a one-stage object detection approach. Experimental validation highlights the effectiveness of our method, demonstrating its state-of-the-art performance across three benchmark datasets.
Haoan Wang, Shilong Jia, Tieyong Zeng, Guixu Zhang, Zhi Li 0080
AAAI3
2024 ACTIVE: A Deep Network for Sperm and Impurity Detection in Microscopic Videos
abstract
The accurate detection of sperms and impurities is a very challenging task, facing problems such as the small size of targets, indefinite target morphologies, low contrast and resolution of the video, and similarity of sperms and impurities. So far, the detection of sperms and impurities still largely relies on the traditional image processing and detection techniques which only yield limited performance and often require manual intervention in the detection process, thus unfavorably escalating the time cost and injecting the subjective bias into the analysis. Encouraged by the success of deep learning methods in numerous object detection tasks, here we report a deep learning network: ACTIVE based on Double Branch Feature Extraction Network (DBFEN) and Cross-conjugate Feature Pyramid Network (CCFPN). DBFEN extracts visual features from tiny objects with a double branch structure, and CCFPN fuses the features extracted by DBFEN to enhance the description of the position and high-level semantic information. Our work is the pioneer of introducing deep learning approaches to the detection of sperms and impurities. Experiments show that the highest AP50of the sperm and impurity detection is 91.13% and 59.64%, which lead its competitors by a substantial margin and establish the state-of-the-art results in this problem. Code is available for readers’ free evaluation at https://github.com/anheqiao-neu/ACTIVE.
Ao Chen 0001, Fenglei Fan, Md Mamunur Rahaman, Tao Jiang 0014, Tieyong Zeng, Marcin Grzegorzek, Chen Li 0022
BIBM7
2024 Navigating Beyond Dropout: An Intriguing Solution Towards Generalizable Image Super Resolution
abstract
Deep learning has led to a dramatic leap on Single Image Super-Resolution (SISR) performances in recent years. While most existing work assumes a simple and fixed degradation model (e.g., bicubic downsampling), the research of Blind SR seeks to improve model generalization ability with unknown degradation. Recently, Kong et al. [37] pioneer the investigation of a more suitable training strategy for Blind SR using Dropout [63]. Although such method indeed brings substantial generalization improvements via mitigating overfitting, we argue that Dropout simultaneously introduces undesirable side-effect that compromises model's capacity to faithfully reconstruct fine details. We show both the theoretical and experimental analyses in our paper, and furthermore, we present another easy yet effective training strategy that enhances the generalization ability of the model by simply modulating its first and second-order features statistics. Experimental results have shown that our method could serve as a model-agnostic regularization and outperforms Dropout on seven benchmark datasets including both synthetic and real-world scenarios.
Hongjun Wang 0007, Jiyuan Chen, Yinqiang Zheng, Tieyong Zeng
CVPR4
2024 EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
abstract
The vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However, these methods often use a few metrics, e.g., FVD [56] or IS [45], to evaluate the performance. We argue that it is hard to judge the large conditional generative models from the simple metrics since these models are often trained on very large datasets with multi-aspect abilities. Thus, we propose a novel framework and pipeline for exhaustively evaluating the performance of the generated videos. Our approach involves generating a diverse and comprehensive list of 700 prompts for text-to-video generation, which is based on an analysis of real-world user data and generated with the assistance of a large language model. Then, we evaluate the state-of-the-art video generative models on our carefully designed benchmark, in terms of visual qualities, content qualities, motion qualities, and text-video alignment with 17 well-selected objective metrics. To obtain the finalleaderboard of the models, we further fit a series of coefficients to align the objective metrics to the users' opinions. Based on the proposed human alignment method, our final score shows a higher correlation than simply averaging the metrics, showing the effectiveness of the proposed evaluation method.
Yaofang Liu, Xiaodong Cun, Xuebo Liu 0002, Xintao Wang 0002, Yong Zhang 0034, Haoxin Chen, Yang Liu 0005, Tieyong Zeng, Raymond Chan 0001, Ying Shan
CVPR8
2024 Purified Distillation: Bridging Domain Shift and Category Gap in Incremental Object Detection
abstract
Incremental Object Detection (IOD) simulates the dynamic data flow in real-world applications, which require detectors to learn new classes or adapt to new domains while retaining knowledge from previous tasks. Most existing IOD methods focus only on class incremental learning, assuming all data comes from the same domain. However, this is hardly achievable in practical applications, as images collected under different conditions often exhibit completely different characteristics, such as lighting, weather, style, etc. Class IOD methods suffer from performance degradation in these scenarios with domain shifts. To bridge domain shifts and category gaps in IOD, we propose Purified Distillation (PD), where we use a set of trainable queries to transfer the teacher's attention on old tasks to the student and adopt the gradient reversal layer to guide the student to learn the teacher's feature space structure from a micro perspective, which has not been extensively studied in previous works. Meanwhile, PD combines classification confidence with localization confidence to purify the most meaningful output nodes, so that the student model inherits a more comprehensive teacher knowledge. Extensive experiments across various IOD settings on six widely used datasets show that PD significantly outperforms state-of-the-art methods. Even after five steps of incremental learning, our method can preserve 60.6% mAP on the first task, while compared methods can only maintain up to 55.9%.
Shilong Jia, Tingting Wu 0001, Yingying Fang, Tieyong Zeng, Guixu Zhang, Zhi Li 0080
ACM Multimedia4
2024 Dynamic Multimodal Information Bottleneck for Multimodality Classification
abstract
Effectively leveraging multimodal data such as various images, laboratory tests and clinical information is becoming increasingly attractive in a variety of AI-based medical diagnosis and prognosis tasks. Most existing multi-modal techniques only focus on enhancing their performance by leveraging the differences or shared features from various modalities and fusing feature across different modalities. These approaches are generally not optimal for clinical settings, which pose the additional challenges of limited training data, as well as being rife with redundant data or noisy modality channels, leading to subpar performance. To address this gap, we study the robustness of existing methods to data redundancy and noise and propose a generalized dynamic multimodal information bottleneck framework for attaining a robust fused feature representation. Specifically, our information bottleneck module serves to filter out the task-irrelevant information and noises in the fused feature, and we further introduce a sufficiency loss to prevent dropping of task-relevant information, thus explicitly preserving the sufficiency of prediction information in the distilled feature. We validate our model on an in-house and a public COVID19 dataset for mortality prediction as well as two public biomedical datasets for diagnostic tasks. Extensive experiments show that our method surpasses the state-of-the-art and is significantly more robust, being the only method to remain performance when large-scale noisy channels exist. Our code is publicly available at https://github.com/ayanglab/DMIB.
Yingying Fang, Shuang Wu 0002, Sheng Zhang 0024, Chaoyan Huang, Tieyong Zeng, Xiaodan Xing, Simon Walsh, Guang Yang 0006
WACV5
2024 Enhanced dual contrast representation learning with cell separation and merging for breast cancer diagnosis
Yang Liu 0119, Yiqi Zhu, Zhehao Gu, Jinshan Pan, Juncheng Li 0003, Ming Fan 0003, Lihua Li 0002, Tieyong Zeng
Comput. Vis. Image Underst.8
2024 Few sampling meshes-based 3D tooth segmentation via region-aware graph convolutional network
Bodong Cheng, Najun Niu, Jun Wang 0024, Tieyong Zeng, Guixu Zhang, Jun Shi 0004, Juncheng Li 0003
Expert Syst. Appl.5
2024 Efficient private SCO for heavy-tailed data via averaged clipping
abstract
Abstract We consider stochastic convex optimization for heavy-tailed data with the guarantee of being differentially private (DP). Most prior works on differentially private stochastic convex optimization for heavy-tailed data are either restricted to gradient descent (GD) or performed multi-times clipping on stochastic gradient descent (SGD), which is inefficient for large-scale problems. In this paper, we consider a one-time clipping strategy and provide principled analyses of its bias and private mean estimation. We establish new convergence results and improved complexity bounds for the proposed algorithm called AClipped-dpSGD for constrained and unconstrained convex problems. We also extend our convergent analysis to the strongly convex case and non-smooth case (which works for generalized smooth objectives with H $$\ddot{\text {o}}$$ o ¨ lder-continuous gradients). All the above results are guaranteed with a high probability for heavy-tailed data. Numerical experiments are conducted to justify the theoretical improvement.
Chenhan Jin, Kaiwen Zhou 0001, Bo Han 0003, James Cheng, Tieyong Zeng
Mach. Learn.5
2024 EWT: Efficient Wavelet-Transformer for single image denoising
Juncheng Li 0003, Bodong Cheng, Guangwei Gao, Jun Shi 0004, Tieyong Zeng
Neural Networks6
2024 HFGN: High-Frequency residual Feature Guided Network for fast MRI reconstruction
Faming Fang, Le Hu, Qiaosi Yi, Tieyong Zeng, Guixu Zhang
Pattern Recognit.5
2024 Kernel correlation-dissimilarity for Multiple Kernel k-Means clustering
Rina Su, Yu Guo 0008, Caiying Wu, Qiyu Jin, Tieyong Zeng
Pattern Recognit.5
2024 Scene recovery: Combining visual enhancement and resolution improvement
Hao Zhang 0154, Te Qi, Tieyong Zeng
Pattern Recognit.3
2024 A Variational Model for Nonuniform Low-Light Image Enhancement
abstract
Abstract. Low-light image enhancement plays an important role in computer vision applications, which is a fundamental low-level task and can affect high-level computer vision tasks. To solve this ill-posed problem, a lot of methods have been proposed to enhance low-light images. However, their performance degrades significantly under nonuniform lighting conditions. Due to the rapid variation of illuminance in different regions in natural images, it is challenging to enhance low-light parts and retain normal-light parts simultaneously in the same image. Commonly, either the low-light parts are underenhanced or the normal-light parts are overenhanced, accompanied by color distortion and artifacts. To overcome this problem, we propose a simple and effective Retinex-based model with reflectance map reweighting for images under nonuniform lighting conditions. An alternating proximal gradient (APG) algorithm is proposed to solve the proposed model, in which the illumination map, the reflectance map, and the weighting map are updated iteratively. To make our model applicable to a wide range of light conditions, we design an initialization scheme for the weighting map. A theoretical analysis of the existence of the solution to our model and the convergence of the APG algorithm are also established. A series of experiments on real-world low-light images are conducted, which demonstrate the effectiveness of our method.
Fan Jia 0007, Shen Mao, Xue-Cheng Tai, Tieyong Zeng
SIAM J. Imaging Sci.4
2024 Extrapolated Plug-and-Play Three-Operator Splitting Methods for Nonconvex Optimization with Applications to Image Restoration
abstract
Abstract. This paper investigates the convergence properties and applications of the three-operator splitting method, also known as the Davis–Yin splitting (DYS) method, integrated with extrapolation and plug-and-play (PnP) denoiser within a nonconvex framework. We first propose an extrapolated DYS method to effectively solve a class of structural nonconvex optimization problems that involve minimizing the sum of three possibly nonconvex functions. Our approach provides an algorithmic framework that encompasses both extrapolated forward–backward splitting and extrapolated Douglas–Rachford splitting methods. To establish the convergence of the proposed method, we rigorously analyze its behavior based on the Kurdyka–Łojasiewicz property, subject to some tight parameter conditions. Moreover, we introduce two extrapolated PnP-DYS methods with convergence guarantee, where the traditional regularization step is replaced by a gradient step–based denoiser. This denoiser is designed using a differentiable neural network and can be reformulated as the proximal operator of a specific nonconvex functional. We conduct extensive experiments on image deblurring and image superresolution problems, where our numerical results showcase the advantage of the extrapolation strategy and the superior performance of the learning-based model that incorporates the PnP denoiser in terms of achieving high-quality recovery images.
Zhongming Wu, Chaoyan Huang, Tieyong Zeng
SIAM J. Imaging Sci.3
2024 Image Segmentation Using Bayesian Inference for Convex Variant Mumford-Shah Variational Model
abstract
Abstract. The Mumford–Shah model is a classical segmentation model, but its objective function is nonconvex. The smoothing and thresholding (SaT) approach is a convex variant of the Mumford–Shah model, which seeks a smoothed approximation solution to the Mumford–Shah model. The SaT approach separates the segmentation into two stages: first, a convex energy function is minimized to obtain a smoothed image; then, a thresholding technique is applied to segment the smoothed image. The energy function consists of three weighted terms and the weights are called the regularization parameters. Selecting appropriate regularization parameters is crucial to achieving effective segmentation results. Traditionally, the regularization parameters are chosen by trial-and-error, which is a very time-consuming procedure and is not practical in real applications. In this paper, we apply a Bayesian inference approach to infer the regularization parameters and estimate the smoothed image. We analyze the convex variant Mumford–Shah variational model from a statistical perspective and then construct a hierarchical Bayesian model. A mean field variational family is used to approximate the posterior distribution. The variational density of the smoothed image is assumed to have a Gaussian density, and the hyperparameters are assumed to have Gamma variational densities. All the parameters in the Gaussian density and Gamma densities are iteratively updated. Experimental results show that the proposed approach is capable of generating high-quality segmentation results. Although the proposed approach contains an inference step to estimate the regularization parameters, it requires less CPU running time to obtain the smoothed image than previous methods.
You-Wei Wen, Raymond Chan 0001, Tieyong Zeng
SIAM J. Imaging Sci.4
2024 Deep Inertia $L_{p}$ Half-Quadratic Splitting Unrolling Network for Sparse View CT Reconstruction
abstract
Sparse view computed tomography (CT) reconstruction poses a challenging ill-posed inverse problem, necessitating effective regularization techniques. In this letter, we employ$L_{p}$-norm ($0< p< 1$) regularization to induce sparsity and introduce inertial steps, leading to the development of the inertial$L_{p}$-norm half-quadratic splitting algorithm. We rigorously prove the convergence of this algorithm. Furthermore, we leverage deep learning to initialize the conjugate gradient method, resulting in a deep unrolling network with theoretical guarantees. Our extensive numerical experiments demonstrate that our proposed algorithm surpasses existing methods, particularly excelling in fewer scanned views and complex noise conditions.
Yu Guo 0008, Caiying Wu, Qiyu Jin, Tieyong Zeng
IEEE Signal Process. Lett.5
2024 DQDG: Data-Free Quantization With Dual Generators for Keyword Spotting
abstract
Data-free quantization effectively compresses deep learning models with privacy guarantees. However, previous data-free quantization methods applied to keyword spotting models have the following issues: (1) The synthesized samples are excessively similar, leading to severe homogenization problems; (2) The low-quality samples during the initial training hinder model fine-tuning. To address these issues, this paper proposes a novel framework called Data-Free Quantization with Dual Generator (DQDG). Our framework introduces Dual Generators with Center Distance Constraint (DGCDC) to enhance the intra-class heterogeneity of synthesized samples, and utilizes a selector to select high-quality samples to assist in model fine-tuning. Additionally, allowing the quantized model to infer complete data from masked data, we adopt Time Masking Quantization Distillation (TMQD) to improve the understanding of the data distribution. Experimental results demonstrate that DQDG outperforms existing data-free quantization methods by a large margin.
Xinbiao Xu, Liyan Ma, Fan Jia 0007, Tieyong Zeng
IEEE Signal Process. Lett.4
2024 Deep Multi-Dictionary Learning for Survival Prediction With Multi-Zoom Histopathological Whole Slide Images
abstract
Survival prediction based on histopathological whole slide images (WSIs) is of great significance for risk-benefit assessment and clinical decision. However, complex microenvironments and heterogeneous tissue structures in WSIs bring challenges to learning informative prognosis-related representations. Additionally, previous studies mainly focus on modeling using mono-scale WSIs, which commonly ignore useful subtle differences existed in multi-zoom WSIs. To this end, we propose a deep multi-dictionary learning framework for cancer survival prediction with multi-zoom histopathological WSIs. The framework can recognize and learn discriminative clusters (i.e., microenvironments) based on multi-scale deep representations for survival analysis. Specifically, we learn multi-scale features based on multi-zoom tiles from WSIs via stacked deep autoencoders network followed by grouping different microenvironments by cluster algorithm. Based on multi-scale deep features of clusters, a multi-dictionary learning method with a post-pruning strategy is devised to learn discriminative representations from selected prognosis-related clusters in a task-driven manner. Finally, a survival model (i.e., EN-Cox) is constructed to estimate the risk index of an individual patient. The proposed model is evaluated on three datasets derived from The Cancer Genome Atlas (TCGA), and the experimental results demonstrate that it outperforms several state-of-the-art survival analysis approaches.
Chao Tu, Denghui Du, Tieyong Zeng, Yu Zhang 0064
IEEE ACM Trans. Comput. Biol. Bioinform.3
2024 WeaFU: Weather-Informed Image Blind Restoration via Multi-Weather Distribution Diffusion
abstract
The extraction of distribution from images with diverse weather conditions is crucial for enhancing the robustness of visual algorithms. When addressing image degradation caused by different weather, accurately perceiving the data distribution of weather-informed degradation becomes a fundamental challenge. However, given the highly stochastic nature, modelling weather distribution poses a formidable task. In this paper, we propose a novel multi-Weather distribution difFUsion blind restoration model, named WeaFU. Firstly, the model employs representation learning to map image distribution into a latent space. Subsequently, WeaFU utilizes a diffusion-based approach, with the assistance of Diffusion Distribution Generator (DDG), to perceive and extract corresponding weather distribution. This strategy ingeniously injects data distribution into the recovery process, significantly enhancing the robustness of the model in diverse weather scenarios. Finally, a Conditional Distribution-Aware Transformer (CDAT) is constructed to align the distribution information with pixels, thereby obtaining clear images. Extensive experiments on real and synthetic datasets demonstrate that WeaFU achieves superior performance.
Bodong Cheng, Juncheng Li 0003, Jun Shi 0004, Yingying Fang, Guixu Zhang, Tieyong Zeng, Zhi Li 0080
IEEE Trans. Circuits Syst. Video Technol.7
2024 The Illusion of Visual Security: Reconstructing Perceptually Encrypted Images
abstract
Perceptual image encryption degrades image quality by selectively encrypting some key information of the plain images. The encrypted images are partially perceptible according to the security or quality requirements. Although several types of attacks have tried to infer privacy information from the encrypted images, they can only either extract statistical information or enhance image sketch. In this paper, we take one step further and fully recover the plain images from perceptually encrypted counterparts by designing a non-local attack network (NL-ANet). NL-ANet is composed of densely cascaded multiscale non-local modules (MSNL) and a hierarchical attention fusion module (HAFM). In particular, to better reconstruct encryption distortion, we introduce MSNL to capture powerful hierarchical features from different scales, and propose HAFM to adaptively aggregate and enhance informative hierarchical features for reconstruction. We also propose a new instantiation of the multi-head non-local block with channel attention (MHCA) to explore the long-range dependencies of global contextual information. Extensive experiments show that NL-ANet is encryption-agnostic and superior on different perceptual encryption schemes under different encryption strengths. NL-ANet also achieves better performance than state-of-the-art image restoration methods.
Ying Yang 0019, Tao Xiang 0001, Shangwei Guo, Tieyong Zeng
IEEE Trans. Circuits Syst. Video Technol.5
2024 FEFA: Frequency Enhanced Multi-Modal MRI Reconstruction With Deep Feature Alignment
abstract
Integrating complementary information from multiple magnetic resonance imaging (MRI) modalities is often necessary to make accurate and reliable diagnostic decisions. However, the different acquisition speeds of these modalities mean that obtaining information can be time consuming and require significant effort. Reference-based MRI reconstruction aims to accelerate slower, under-sampled imaging modalities, such as T2-modality, by utilizing redundant information from faster, fully sampled modalities, such as T1-modality. Unfortunately, spatial misalignment between different modalities often negatively impacts the final results. To address this issue, we propose FEFA, which consists of cascading FEFA blocks. The FEFA block first aligns and fuses the two modalities at the feature level. The combined features are then filtered in the frequency domain to enhance the important features while simultaneously suppressing the less essential ones, thereby ensuring accurate reconstruction. Furthermore, we emphasize the advantages of combining the reconstruction results from multiple cascaded blocks, which also contributes to stabilizing the training process. Compared to existing registration-then-reconstruction and cross-attention-based approaches, our method is end-to-end trainable without requiring additional supervision, extensive parameters, or heavy computation. Experiments on the public fastMRI, IXI and in-house datasets demonstrate that our approach is effective across various under-sampling patterns and ratios.
Xuanmin Chen, Liyan Ma, Shihui Ying, Dinggang Shen, Tieyong Zeng
IEEE J. Biomed. Health Informatics5
2024 Double Transformer Super-Resolution for Breast Cancer ADC Images
abstract
Diffusion-weighted imaging (DWI) has been extensively explored in guiding the clinic management of patients with breast cancer. However, due to the limited resolution, accurately characterizing tumors using DWI and the corresponding apparent diffusion coefficient (ADC) is still a challenging problem. In this paper, we aim to address the issue of super-resolution (SR) of ADC images and evaluate the clinical utility of SR-ADC images through radiomics analysis. To this end, we propose a novel double transformer-based network (DTformer) to enhance the resolution of ADC images. More specifically, we propose a symmetric U-shaped encoder-decoder network with two different types of transformer blocks, named as UTNet, to extract deep features for super-resolution. The basic backbone of UTNet is composed of a locally-enhanced Swin transformer block (LeSwin-T) and a convolutional transformer block (Conv-T), which are responsible for capturing long-range dependencies and local spatial information, respectively. Additionally, we introduce a residual upsampling network (RUpNet) to expand image resolution by leveraging initial residual information from the original low-resolution (LR) images. Extensive experiments show that DTformer achieves superior SR performance. Moreover, radiomics analysis reveals that improving the resolution of ADC images is beneficial for tumor characteristic prediction, such as histological grade and human epidermal growth factor receptor 2 (HER2) status.
Ying Yang 0019, Tao Xiang 0001, Lihua Li 0002, Lok Ming Lui, Tieyong Zeng
IEEE J. Biomed. Health Informatics6
2024 Multi-Prototypes Convex Merging Based K-Means Clustering Algorithm
abstract
K-Means algorithm is a popular clustering method. However, it has two limitations: 1) it gets stuck easily in spurious local minima, and 2) the number of clusters$k$has to be given a priori. To solve these two issues, a multi-prototypes convex merging based K-Means clustering algorithm (MCKM) is presented. First, based on the structure of the spurious local minima of the K-Means problem, a multi-prototypes sampling (MPS) is designed to select the appropriate number of multi-prototypes for data with arbitrary shapes. Then, a merging technique, called convex merging (CM), merges the multi-prototypes to get a better local minima without$k$being given a priori. Specifically, CM can obtain the optimal merging and estimate the correct$k$. By integrating these two techniques with K-Means algorithm, the proposed MCKM is an efficient and explainable clustering algorithm for escaping the undesirable local minima of K-Means problem without given$k$first. Two theoretical proofs are given to guarantee that the cost of MCKM (MPS+CM) can achieve a constant factor approximation to the optimal cost of the K-Means problem. Experimental results performed on synthetic and real-world data sets have verified the effectiveness of the proposed algorithm.
Shuisheng Zhou, Tieyong Zeng, Raymond Chan 0001
IEEE Trans. Knowl. Data Eng.3
2024 Flow Guidance Deformable Compensation Network for Video Frame Interpolation
abstract
Flow-based and deformable convolution (DConv)-based methods are two mainstream approaches for solving the video frame interpolation (VFI) problem, which have made remarkable progress with the development of deep convolutional networks over the past years. However, flow-based VFI methods often suffer from the inaccuracy of flow map estimation, especially in dealing with complex and irregular real-world motions. DConv-based VFI methods have advantages in handling complex motions, while the increased degree of freedom makes the training of the DConv model difficult. To address these problems, in this article, we propose a flow guidance deformable compensation network (FGDCN) for the VFI task. FGDCN decomposes the frame sampling process into two steps: a flow step and a deformation step. Specifically, the flow step utilizes a coarse-to-fine flow estimation network to directly estimate the intermediate flows and synthesizes an anchor frame simultaneously. To ensure the accuracy of the estimated flow, a distillation loss and a task-oriented loss are jointly employed in this step. Under the guidance of the flow priors learned in step one, the deformation step designs a new pyramid deformable compensation network to compensate for the missing details of the flow step. In addition, a pyramid loss is proposed to supervise the model in both the image and frequency domains. Experimental results show that the proposed algorithm achieves excellent performance on various datasets with fewer parameters.
Pengcheng Lei, Faming Fang, Tieyong Zeng, Guixu Zhang
IEEE Trans. Multim.3
2024 Retinex Image Enhancement Based on Sequential Decomposition With a Plug-and-Play Framework
abstract
The Retinex model is one of the most representative and effective methods for low-light image enhancement. However, the Retinex model does not explicitly tackle the noise problem and shows unsatisfactory enhancing results. In recent years, due to the excellent performance, deep learning models have been widely used in low-light image enhancement. However, these methods have two limitations. First, the desirable performance can only be achieved by deep learning when a large number of labeled data are available. However, it is not easy to curate massive low-/normal-light paired data. Second, deep learning is notoriously a black-box model. It is difficult to explain their inner working mechanism and understand their behaviors. In this article, using a sequential Retinex decomposition strategy, we design a plug-and-play framework based on the Retinex theory for simultaneous image enhancement and noise removal. Meanwhile, we develop a convolutional neural network-based (CNN-based) denoiser into our proposed plug-and-play framework to generate a reflectance component. The final image is enhanced by integrating the illumination and reflectance with gamma correction. The proposed plug-and-play framework can facilitate both post hoc and ad hoc interpretability. Extensive experiments on different datasets demonstrate that our framework outcompetes the state-of-the-art methods in both image enhancement and denoising.
Tingting Wu 0001, Wenna Wu, Ying Yang 0019, Fenglei Fan, Tieyong Zeng
IEEE Trans. Neural Networks Learn. Syst.5
2023 Uncertainty-Aware Unsupervised Image Deblurring with Deep Residual Prior
abstract
Non-blind deblurring methods achieve decent performance under the accurate blur kernel assumption. Since the kernel uncertainty (i.e. kernel error) is inevitable in practice, semi-blind deblurring is suggested to handle it by introducing the prior of the kernel (or induced) error. However, how to design a suitable prior for the kernel (or induced) error remains challenging. Hand-crafted prior, incorporating domain knowledge, generally performs well but may lead to poor performance when kernel (or induced) error is complex. Data-driven prior, which excessively depends on the diversity and abundance of training data, is vulnerable to out-of-distribution blurs and images. To address this challenge, we suggest a dataset-free deep residual prior for the kernel induced error (termed as residual) expressed by a customized untrained deep neural network, which allows us to flexibly adapt to different blurs and images in real scenarios. By organically integrating the respective strengths of deep priors and hand-crafted priors, we propose an unsupervised semi-blind deblurring model which recovers the clear image from the blurry image and inaccurate blur kernel. To tackle the formulated model, an efficient alternating minimization algorithm is developed. Extensive experiments demonstrate the favorable performance of the proposed method as compared to model-driven and data-driven methods in terms of image quality and the robustness to different types of kernel error.
Xiaole Tang, Xi-Le Zhao, Jun Liu 0012, Jianli Wang, Yuchun Miao, Tieyong Zeng
CVPR6
2023 PFT-SSR: Parallax Fusion Transformer for Stereo Image Super-Resolution
abstract
Stereo image super-resolution aims to boost the performance of image super-resolution by exploiting the supplementary information provided by binocular systems. Although previous methods have achieved promising results, they did not fully utilize the information of cross-view and intra-view. To further unleash the potential of binocular images, in this letter, we propose a novel Transformer-based parallax fusion module called Parallax Fusion Transformer (PFT). PFT employs a Cross-view Fusion Transformer (CVFT) to utilize cross-view information and an Intra-view Refinement Transformer (IVRT) for intra-view feature refinement. Meanwhile, we adopted the Swin Transformer as the backbone for feature extraction and SR reconstruction to form a pure Transformer architecture called PFT-SSR. Extensive experiments and ablation studies show that PFT-SSR achieves competitive results and outperforms most SOTA methods. Source code is available at https://github.com/MIVRC/PFT-PyTorch.
Hansheng Guo, Juncheng Li 0013, Guangwei Gao, Zhi Li 0080, Tieyong Zeng
ICASSP5
2023 Decomposition-Based Variational Network for Multi-Contrast MRI Super-Resolution and Reconstruction
abstract
Multi-contrast MRI super-resolution (SR) and reconstruction methods aim to explore complementary information from the reference image to help the reconstruction of the target image. Existing deep learning-based methods usually manually design fusion rules to aggregate the multi-contrast images, fail to model their correlations accurately and lack certain interpretations. Against these issues, we propose a multi-contrast variational network (MC-VarNet) to explicitly model the relationship of multi-contrast images. Our model is constructed based on an intuitive motivation that multi-contrast images have consistent (edges and structures) and inconsistent (contrast) information. We thus build a model to reconstruct the target image and decompose the reference image as a common component and a unique component. In the feature interaction phase, only the common component is transferred to the target image. We solve the variational model and unfold the iterative solutions into a deep network. Hence, the proposed method combines the good interpretability of model-based methods with the powerful representation ability of deep learning-based methods. Experimental results on the multi-contrast MRI reconstruction and SR demonstrate the effectiveness of the proposed model. Especially, since we explicitly model the multi-contrast images, our model is more robust to the reference images with noises and large inconsistent structures. The code is available at https://github.com/lpcccccv/MC-VarNet.
Pengcheng Lei, Faming Fang, Guixu Zhang, Tieyong Zeng
ICCV4
2023 Recognizable Information Bottleneck
abstract
Information Bottlenecks (IBs) learn representations that generalize to unseen data by information compression. However, existing IBs are practically unable to guarantee generalization in real-world scenarios due to the vacuous generalization bound. The recent PAC-Bayes IB uses information complexity instead of information compression to establish a connection with the mutual information generalization bound. However, it requires the computation of expensive second-order curvature, which hinders its practical application. In this paper, we establish the connection between the recognizability of representations and the recent functional conditional mutual information (f-CMI) generalization bound, which is significantly easier to estimate. On this basis we propose a Recognizable Information Bottleneck (RIB) which regularizes the recognizability of representations through a recognizability critic optimized by density ratio matching under the Bregman divergence. Extensive experiments on several commonly used datasets demonstrate the effectiveness of the proposed method in regularizing the model and estimating the generalization gap.
Yilin Lyu, Xin Liu 0086, Yaxin Peng, Tieyong Zeng, Liping Jing
IJCAI6
2023 Conditional Physics-Informed Graph Neural Network for Fractional Flow Reserve Assessment
Baihong Xie, Xiujian Liu, Heye Zhang, Chenchu Xu, Tieyong Zeng, Yixuan Yuan, Guang Yang 0006, Zhifan Gao
MICCAI (7)5
2023 Not All Tasks Are Equal: A Parameter-Efficient Task Reweighting Method for Few-Shot Learning
Xin Liu 0086, Yilin Lyu, Liping Jing, Tieyong Zeng, Jian Yu 0001
ECML/PKDD (2)4
2023 vMF Loss: Exploring a Scattered Intra-class Hypersphere for Few-Shot Learning
Xin Liu 0086, Shijing Wang, Kairui Zhou, Yilin Lyu, Liping Jing, Tieyong Zeng, Jian Yu 0001
ECML/PKDD (2)7
2023 Snow Mask Guided Adaptive Residual Network for Image Snow Removal
Bodong Cheng, Juncheng Li 0003, Tieyong Zeng
Comput. Vis. Image Underst.4
2023 Proximal linearized alternating direction method of multipliers algorithm for nonconvex image restoration with impulse noise
abstract
Abstract Image restoration with impulse noise is an important task in image processing. Taking into account the statistical distribution of impulse noise, the ℓ 1 ‐norm data fidelity and total variation () model has been widely used in this area. However, the model usually performs worse when the noise level is high. To overcome this drawback, several nonconvex models have been proposed. In this paper, an efficient iterative algorithm is proposed to solve nonconvex models arising in impulse noise. Compared to existing algorithms, the proposed algorithm is a completely explicit algorithm in which every subproblem has a closed‐form solution. The key idea is to transform the original nonconvex models into an equivalent constrained minimization problem with two separable objective functions, where one is differentiable but nonconvex. As a consequence, the proximal linearized alternating direction method of multipliers is employed to solve it. Extensive numerical experiments are presented to demonstrate the efficiency and effectiveness of the proposed algorithm.
Yuchao Tang, Shirong Deng, Tieyong Zeng
IET Image Process.4
2023 Rank-One Prior: Real-Time Scene Recovery
abstract
Scene recovery is a fundamental imaging task with several practical applications, including video surveillance and autonomous vehicles, etc. In this article, we provide a new real-time scene recovery framework to restore degraded images under different weather/imaging conditions, such as underwater, sand dust and haze. A degraded image can actually be seen as a superimposition of a clear image with the same color imaging environment (underwater, sand or haze, etc.). Mathematically, we can introduce a rank-one matrix to characterize this phenomenon, i.e., rank-one prior (ROP). Using the prior, a direct method with the complexity$O(N)$is derived for real-time recovery. For general cases, we develop ROP$^+$to further improve the recovery performance. Comprehensive experiments of the scene recovery illustrate that our method outperforms competitively several state-of-the-art imaging methods in terms of efficiency and robustness.
Jun Liu 0012, Ryan Wen Liu, Tieyong Zeng
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Deep Generative Mixture Model for Robust Imbalance Classification
abstract
Discovering hidden pattern from imbalanced data is a critical issue in various real-world applications. Existing classification methods usually suffer from the limitation of data especially for minority classes, and result in unstable prediction and low performance. In this paper, a deep generative classifier is proposed to mitigate this issue via both model perturbation and data perturbation. Specially, the proposed generative classifier is derived from a deep latent variable model where two variables are involved. One variable is to capture the essential information of the original data, denoted as latent codes, which are represented by a probability distribution rather than a single fixed value. The learnt distribution aims to enforce the uncertainty of model and implement model perturbation, thus, lead to stable predictions. The other variable is a prior to latent codes so that the codes are restricted to lie on components in Gaussian Mixture Model. As a confounder affecting generative processes of data (feature/label), the latent variables are supposed to capture the discriminative latent distribution and implement data perturbation. Extensive experiments have been conducted on widely-used real imbalanced image datasets. Experimental results demonstrate the superiority of our proposed model by comparing with popular imbalanced classification baselines on imbalance classification task.
Liping Jing, Yilin Lyu, Mingzhe Guo, Jiaqi Wang 0006, Huafeng Liu 0001, Jian Yu 0001, Tieyong Zeng
IEEE Trans. Pattern Anal. Mach. Intell.8
2023 Single-particle reconstruction in cryo-EM based on three-dimensional weighted nuclear norm minimization
abstract
Single-particle reconstruction (SPR) in cryogenic electron microscopy (cryo-EM) aims at aligning and averaging two-dimensional micrographs to reconstruct a three-dimensional particle. How to reconstruct micrographs from heavy noise is a crucial point for achieving better micrograph quality, and thus many methods focus on noise removal. However, new problems such as over-smoothing often occur in their results due to failure in handling heavy noise well. This paper proposes a three-dimensional weighted nuclear norm minimization (3DWNNM) model for SPR in the cryo-EM task to address these issues. Specifically, we design a minimization solver based on the forward-backward splitting algorithm to tackle our model efficiently. Under certain conditions, this solution has an energy-decaying feature and performs exceptionally well in reconstruction. Numerical experiments fully demonstrate the effectiveness and the robustness of the proposed method.
Chaoyan Huang, Tingting Wu 0001, Juncheng Li 0013, Tieyong Zeng
Pattern Recognit.5
2023 A reflectance re-weighted Retinex model for non-uniform and low-light image enhancement
Fan Jia 0007, Hok Shing Wong, Tieyong Zeng
Pattern Recognit.4
2023 Spherical Image Inpainting with Frame Transformation and Data-Driven Prior Deep Networks
abstract
Abstract. Spherical image processing has been widely applied in many important fields, such as omnidirectional vision for autonomous cars, global climate modeling, and medical imaging. It is nontrivial to extend an algorithm developed for flat images to the spherical ones. In this work, we focus on the challenging task of spherical image inpainting with a deep learning-based regularizer. Instead of a naive application of existing models for planar images, we employ a fast directional spherical Haar framelet transform and develop a novel optimization framework based on a sparsity assumption of the framelet transform. Furthermore, by employing progressive encoder-decoder architecture, a new and better-performed deep CNN denoiser is carefully designed and works as an implicit regularizer. Finally, we use a plug-and-play method to handle the proposed optimization model, which can be implemented efficiently by training the CNN denoiser prior. Numerical experiments are conducted and show that the proposed algorithms can greatly recover damaged spherical images and achieve the best performance over purely using a deep learning denoiser and a plug-and-play model.
Jianfei Li, Chaoyan Huang, Raymond Chan 0001, Michael Kwok-Po Ng, Tieyong Zeng
SIAM J. Imaging Sci.6
2023 Single image noise level estimation by artificial noise
Fang Li 0004, Faming Fang, Zhi Li 0080, Tieyong Zeng
Signal Process.4
2023 Adaptive weighted curvature-based active contour for ultrasonic and 3T/5T MR image segmentation
Zhi-Feng Pang, Mengxiao Geng, Yanru Zhou, Tieyong Zeng, Liyun Zheng, Na Zhang 0001, Dong Liang 0001, Hairong Zheng, Yongming Dai, Zhenxing Huang, Zhanli Hu
Signal Process.5
2023 Efficient SAV Algorithms for Curvature Minimization Problems
abstract
The curvature regularization method is well-known for its good geometric interpretability and strong priors in the continuity of edges, which has been applied to various image processing tasks. However, due to the non-convex, non-smooth, and highly non-linear intrinsic limitations, most existing algorithms lack a convergence guarantee. This paper proposes an efficient yet accurate scalar auxiliary variable (SAV) scheme for solving both mean curvature and Gaussian curvature minimization problems. The SAV-based algorithms are shown unconditionally energy diminishing, fast convergent, and very easy to be implemented for different image applications. Numerical experiments on noise removal, image deblurring, and single image super-resolution are presented on both gray and color image datasets to demonstrate the robustness and efficiency of our method. Source codes are made publicly available athttps://github.com/Duanlab123/SAV-curvature.
Chenxin Wang, Zhenwei Zhang 0002, Zhichang Guo, Tieyong Zeng, Yuping Duan
IEEE Trans. Circuits Syst. Video Technol.4
2023 CTCNet: A CNN-Transformer Cooperation Network for Face Image Super-Resolution
abstract
Recently, deep convolution neural networks (CNNs) steered face super-resolution methods have achieved great progress in restoring degraded facial details by joint training with facial priors. However, these methods have some obvious limitations. On the one hand, multi-task joint learning requires additional marking on the dataset, and the introduced prior network will significantly increase the computational cost of the model. On the other hand, the limited receptive field of CNN will reduce the fidelity and naturalness of the reconstructed facial images, resulting in suboptimal reconstructed images. In this work, we propose an efficient CNN-Transformer Cooperation Network (CTCNet) for face super-resolution tasks, which uses the multi-scale connected encoder-decoder architecture as the backbone. Specifically, we first devise a novel Local-Global Feature Cooperation Module (LGCM), which is composed of a Facial Structure Attention Unit (FSAU) and a Transformer block, to promote the consistency of local facial detail and global facial structure restoration simultaneously. Then, we design an efficient Feature Refinement Module (FRM) to enhance the encoded features. Finally, to further improve the restoration of fine facial details, we present a Multi-scale Feature Fusion Unit (MFFU) to adaptively fuse the features from different stages in the encoder procedure. Extensive evaluations on various datasets have assessed that the proposed CTCNet can outperform other state-of-the-art methods significantly. Source code will be available at https://github.com/IVIPLab/CTCNet.
Guangwei Gao, Zixiang Xu, Juncheng Li 0003, Jian Yang 0003, Tieyong Zeng, Guo-Jun Qi
IEEE Trans. Image Process.5
2023 Cross-Parametric Generative Adversarial Network-Based Magnetic Resonance Image Feature Synthesis for Breast Lesion Classification
abstract
Dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) contains information on tumor morphology and physiology for breast cancer diagnosis and treatment. However, this technology requires contrast agent injection with more acquisition time than other parametric images, such as T2-weighted imaging (T2WI). Current image synthesis methods attempt to map the image data from one domain to another, whereas it is challenging or even infeasible to map the images with one sequence into images with multiple sequences. Here, we propose a new approach of cross-parametric generative adversarial network (GAN)-based feature synthesis (CPGANFS) to generate discriminative DCE-MRI features from T2WI with applications in breast cancer diagnosis. The proposed approach decodes the T2W images into latent cross-parameter features to reconstruct the DCE-MRI and T2WI features by balancing the information shared between the two. A Wasserstein GAN with a gradient penalty is employed to differentiate the T2WI-generated features from ground-truth features extracted from DCE-MRI. The synthesized DCE-MRI feature-based model achieved significantly (p = 0.036) higher prediction performance (AUC = 0.866) in breast cancer diagnosis than that based on T2WI (AUC = 0.815). Visualization of the model shows that our CPGANFS method enhances the predictive power by levitating attention to the lesion and the surrounding parenchyma areas, which is driven by the interparametric information learned from T2WI and DCE-MRI. Our proposed CPGANFS provides a framework for cross-parametric MR image feature generation from a single-sequence image guided by an information-rich, time-series image with kinetic information. Extensive experimental results demonstrate its effectiveness with high interpretability and improved performance in breast cancer diagnosis.
Ming Fan 0003, Guangyao Huang 0002, Junhong Lou, Xin Gao 0001, Tieyong Zeng, Lihua Li 0002
IEEE J. Biomed. Health Informatics5
2023 Learning Common and Task-Specific Radiomic Features via Graph Regularized NMF for the Joint Prediction of Multiple Clinical Indicators in Breast Cancer
abstract
Assessments of multiple clinical indicators based on radiomic analysis of magnetic resonance imaging (MRI) are beneficial to the diagnosis, prognosis and treatment of breast cancer patients. Many machine learning methods have been designed to jointly predict multiple indicators for more accurate assessments while using original clinical labels directly without considering the noisy and redundant information among them. To this end, we propose a multilabel learning method based on label space dimensionality reduction (LSDR), which learns common and task-specific features via graph regularized nonnegative matrix factorization (CTFGNMF) for the joint prediction of multiple indicators in breast cancer. A nonnegative matrix factorization (NMF) is adopted to map original clinical labels to a low-dimensional latent space. The latent labels are employed to exploit task correlations by using a least square loss function with [Formula: see text]-norm regularization to identify common features, which help to improve the generalization performance of correlated tasks. Furthermore, task-specific features were retained by a multitask regression formulation to increase the discrimination power for different tasks. Common and task-specific features are incorporated by dynamic graph Laplacian regularization into a unified model to learn complementary features. Then, a multilabel classification is built to predict multiple clinical indicators including human epidermal growth factor receptor 2 (HER2), Ki-67, and histological grade. Experimental results show that CTFGNMF achieves AUCs of 0.823, 0.691 and 0.776 in the three indicator predictions, outperforming other counterparts that consider only task-independent features or common features. It indicates CTFGNMF is a promising application for multiple classification tasks in breast cancer.
Jian Guan 0007, Ming Fan 0003, Tieyong Zeng, Lihua Li 0002
IEEE J. Biomed. Health Informatics3
2023 Frequency Learning via Multi-Scale Fourier Transformer for MRI Reconstruction
abstract
Since Magnetic Resonance Imaging (MRI) requires a long acquisition time, various methods were proposed to reduce the time, but they ignored the frequency information and non-local similarity, so that they failed to reconstruct images with a clear structure. In this article, we propose Frequency Learning via Multi-scale Fourier Transformer for MRI Reconstruction (FMTNet), which focuses on repairing the low-frequency and high-frequency information. Specifically, FMTNet is composed of a high-frequency learning branch (HFLB) and a low-frequency learning branch (LFLB). Meanwhile, we propose a Multi-scale Fourier Transformer (MFT) as the basic module to learn the non-local information. Unlike normal Transformers, MFT adopts Fourier convolution to replace self-attention to efficiently learn global information. Moreover, we further introduce a multi-scale learning and cross-scale linear fusion strategy in MFT to interact information between features of different scales and strengthen the representation of features. Compared with normal Transformers, the proposed MFT occupies fewer computing resources. Based on MFT, we design a Residual Multi-scale Fourier Transformer module as the main component of HFLB and LFLB. We conduct several experiments under different acceleration rates and different sampling patterns on different datasets, and the experiment results show that our method is superior to the previous state-of-the-art method.
Qiaosi Yi, Faming Fang, Guixu Zhang, Tieyong Zeng
IEEE J. Biomed. Health Informatics4
2023 Hierarchical Perception Adversarial Learning Framework for Compressed Sensing MRI
abstract
The long acquisition time has limited the accessibility of magnetic resonance imaging (MRI) because it leads to patient discomfort and motion artifacts. Although several MRI techniques have been proposed to reduce the acquisition time, compressed sensing in magnetic resonance imaging (CS-MRI) enables fast acquisition without compromising SNR and resolution. However, existing CS-MRI methods suffer from the challenge of aliasing artifacts. This challenge results in the noise-like textures and missing the fine details, thus leading to unsatisfactory reconstruction performance. To tackle this challenge, we propose a hierarchical perception adversarial learning framework (HP-ALF). HP-ALF can perceive the image information in the hierarchical mechanism: image-level perception and patch-level perception. The former can reduce the visual perception difference in the entire image, and thus achieve aliasing artifact removal. The latter can reduce this difference in the regions of the image, and thus recover fine details. Specifically, HP-ALF achieves the hierarchical mechanism by utilizing multilevel perspective discrimination. This discrimination can provide the information from two perspectives (overall and regional) for adversarial learning. It also utilizes a global and local coherent discriminator to provide structure information to the generator during training. In addition, HP-ALF contains a context-aware learning block to effectively exploit the slice information between individual images for better reconstruction performance. The experiments validated on three datasets demonstrate the effectiveness of HP-ALF and its superiority to the comparative methods.
Zhifan Gao, Yifeng Guo, Tieyong Zeng, Guang Yang 0006
IEEE Trans. Medical Imaging4
2023 SRRNet: A Semantic Representation Refinement Network for Image Segmentation
abstract
Semantic context has raised concerns in semantic segmentation. In most cases, it is applied to guide feature learning. Instead, this paper applies it to extract the semantic representation, which records the global feature information of each category with a memory tensor. Specifically, we propose a novel semantic representation (SR) module, which consists of semantic embedding (SE) and semantic attention (SA) blocks. The SE block adaptively embeds features into the semantic representation by calculating the memory similarity, and the SA block aggregates the embedded features with semantic attention. The main advantages of the SR module lie in three aspects: i) it enhances the representation ability of semantic context by employing global (cross-image) semantic information; ii) it improves the consistency of intraclass features by aggregating global features of the same categories; and iii) it can be extended to build a semantic representation refinement network (SRRNet) by iteratively applying the SR module across multiple scales, shrinking the semantic gap and enhancing the structural reasoning of the model. Extensive experiments demonstrate that our method significantly improves the segmentation results and achieves superior performance on the PASCAL VOC 2012, Cityscapes, and PASCAL Context datasets.
Xiaofeng Ding 0003, Tieyong Zeng, Jian Tang 0008, Zhengping Che, Yaxin Peng
IEEE Trans. Multim.2
2023 Retinex-Based Variational Framework for Low-Light Image Enhancement and Denoising
abstract
Low-light image enhancement is an important task in the domain of computer vision. Images taken under insufficient lighting conditions manifest low visibility and unknown noises which disrupt image contents and pose considerable challenges for low-light image enhancement. Most of Retinex-based methods usually attempt to design different priors on the gradient of both illumination and reflectance. However, noises can be involved in the Retinex-based models. To address the problem, we explore the problem of low-light image restoration through joint contrast enhancement and denoising. We propose a Retinex-based variational model for low-light image enhancement that effectively generates a noise-free image, yet proves to generalize well to diverse light-conditions. First, we present a simple constraint on the fidelity term between the fractional derivative of an observed image and the fractional derivative of the recomposed one which is the product of the reflectance and illumination. This strategy aims to model spatial consistency to preserve natural variation. Second, we introduce a weighted regularization term for the reflectance that can remove noise with a adaptive texture map. We evaluate our proposed approach using three challenging datasets: NPE, LOL and GladNet. Extensive experiments demonstrate that our proposed method outperforms other competing methods in terms of visual quality and quantitative comparisons.
Qianting Ma, Yang Wang 0020, Tieyong Zeng
IEEE Trans. Multim.3
2022 Pixel screening based intermediate correction for blind deblurring
abstract
Blind deblurring has attracted much interest with its wide applications in reality. The blind deblurring problem is usually solved by estimating the intermediate kernel and the intermediate image alternatively, which will finally converge to the blurring kernel of the observed image. Numerous works have been proposed to obtain intermediate images with fewer undesirable artifacts by designing delicate regularization on the latent solution. However, these methods still fail while dealing with images containing saturations and large blurs. To address this problem, we propose an intermediate image correction method which utilizes Bayes posterior estimation to screen through the intermediate image and exclude those unfavorable pixels to reduce their influence for kernel estimation. Extensive experiments have proved that the proposed method can effectively improve the accuracy of the final derived kernel against the state-of-the-art methods on benchmark datasets by both quantitative and qualitative comparisons.
Meina Zhang, Yingying Fang, Guoxi Ni, Tieyong Zeng
CVPR4
2022 A Graph-Based Dual Convolutional Network for Automatic Road Extraction from High Resolution Remote Sensing Images
abstract
Recently, deep-learning-based methods, especially deep convolutional neural networks (DCNNs), have effectively shown state-of-the-art performance in road extraction from high resolution remote sensing images (HRSI). However, due to the loss of location information and global context information, most existing DCNNs are inadequate for extracting tiny roads or roads which are severely occluded, leading to incomplete and discontinuous results. To address this problem, this paper proposes a graph-based dual convolutional network (GDCNet), which combines graph convolutional network (GCN) and convolutional neural network (CNN). In this model, GCN and CNN branches perform feature learning on large-scale irregular regions and small-scale regular regions, and generate complementary spatial-spectral features at superpixel and pixel levels, respectively. Then, a graph decoder is utilized to propagate features between graph nodes and image pixels, enabling the GCN and CNN to collaborate in a single network. Extensive experiments on two benchmark datasets demonstrate that the proposed GDCNet is competitive compared with other state-of-the-art methods both qualitatively and quantitatively, and is effective against the incomplete and discontinuous problems of the extracted roads.
Fumin Cui, Yichang Shi, Ruyi Feng, Lizhe Wang 0001, Tieyong Zeng
IGARSS5
2022 Graph Laplacian Regularized Spectral-Spatial-Sparse Unmixing for Hyperspectral Imagery
abstract
Sparse unmixing aims at finding the optimal subset of endmembers in a spectral library to approximate the observed data, and has received increasing attention as it can circumvent the estimation of the endmember. In this paper, a graph Laplacian regularized spectral-spatial-sparse unmixing algorithm is proposed, namely, gLapS3U, incorporating the graph Laplacian regularization to consider the similarity between pixels of the whole image, and enforcing the spectral-spatial-sparse constraints to enhance the local spatial information as well as the sparsity of the abundance solution jointly. Experimental results on simulated and real data show the superiority of the proposed algorithm compared with state-of-the-art existing methods.
Zhi Li 0080, Ruyi Feng, Yichang Shi, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Tieyong Zeng
IGARSS7
2022 Remote Sensing Image Super-Resolution via Dilated Convolution Network with Gradient Prior
abstract
Due to the limitations of the imaging sensor, the spatial resolution of satellite imagery is often insufficient, namely, low resolution (LR). Therefore, super-resolution (SR) is proposed, which strives to improve image resolution, perfectly to compensate for the shortcomings of satellite sensor imaging. In this study, we develop a unique dilated convolution network with gradient prior (DCNG) for remote sensing SR, aiming to extract powerful low-level features with gradient prior and efficitive network and then reconstruct the high-level feature details. The DCNG is built of two components: the Multi-Scale Feature Extraction Network and the Feature Reconstruction Network. In the Multi-Scale Feature Extraction Network, the Double-Path Dilated Residual Block (DPDRB) is designed with the dilation convolution operation to obtain the multi-scale features and increase the receptive field, the Global Self-attention Module (GSA) to catch the long-range dependency among picture patches, and a Gradient Propagation Network (GPN) is proposed to extract high-level gradient information. In the Feature Reconstruction Network, the Pixel Shuffle is introduced to reconstruct the feature by combining characteristics of different frequency bands. Experiments using Massachusetts_Roads and 3K VEHICLE_SR data sets indicate that our DCNG surpasses state-of-the-art algorithms in terms of quantitative and qualitative evaluations.
Ruyi Feng, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Tieyong Zeng
IGARSS6
2022 Lightweight Bimodal Network for Single-Image Super-Resolution via Symmetric CNN and Recursive Transformer
abstract
Single-image super-resolution (SISR) has achieved significant breakthroughs with the development of deep learning. However, these methods are difficult to be applied in real-world scenarios since they are inevitably accompanied by the problems of computational and memory costs caused by the complex operations. To solve this issue, we propose a Lightweight Bimodal Network (LBNet) for SISR. Specifically, an effective Symmetric CNN is designed for local feature extraction and coarse image reconstruction. Meanwhile, we propose a Recursive Transformer to fully learn the long-term dependence of images thus the global information can be fully used to further refine texture details. Studies show that the hybrid of CNN and Transformer can build a more efficient model. Extensive experiments have proved that our LBNet achieves more prominent performance than other state-of-the-art methods with a relatively low computational cost and memory consumption. The code is available at https://github.com/IVIPLab/LBNet.
Guangwei Gao, Zhengxue Wang, Juncheng Li 0003, Yi Yu 0001, Tieyong Zeng
IJCAI6
2022 Exploring Latent Sparse Graph for Large-Scale Semi-supervised Learning
Li Wang 0033, Raymond Chan 0001, Tieyong Zeng
ECML/PKDD (4)4
2022 Adjustable super-resolution network via deep supervised learning and progressive self-distillation
Juncheng Li 0003, Faming Fang, Tieyong Zeng, Guixu Zhang, Xizhao Wang
Neurocomputing3
2022 A survey on epistemic (model) uncertainty in supervised learning: Recent advances and applications
Xinlei Zhou, Han Liu 0002, Farhad Pourpanah, Tieyong Zeng, Xizhao Wang
Neurocomputing4
2022 Quaternion-based weighted nuclear norm minimization for color image restoration
Chaoyan Huang, Zhi Li 0080, Yubing Liu, Tingting Wu 0001, Tieyong Zeng
Pattern Recognit.5
2022 Phase retrieval from incomplete data via weighted nuclear norm minimization
Zhi Li 0080, Ming Yan 0006, Tieyong Zeng, Guixu Zhang
Pattern Recognit.3
2022 Efficient Boosted DC Algorithm for Nonconvex Image Restoration with Rician Noise
abstract
Image deblurring under Rician noise has attracted considerable attention in imaging science. Frequently appearing in medical imaging, Rician noise leads to an interesting nonconvex optimization problem, termed as the MAP-Rician model, which is based on the Maximum a Posteriori (MAP) estimation approach. As the MAP-Rician model is deeply rooted in Bayesian analysis, we want to understand its mathematical analysis carefully. Moreover, one needs to properly select a suitable algorithm for tackling this nonconvex problem to get the best performance. This paper investigates both issues. Indeed, we first present a theoretical result about the existence of a minimizer for the MAP-Rician model under mild conditions. Next, we aim to adopt an efficient boosted difference of convex functions algorithm (BDCA) to handle this challenging problem. Basically, BDCA combines the classical difference of convex functions algorithm (DCA) with a backtracking line search, which utilizes the point generated by DCA to define a search direction. In particular, we apply a smoothing scheme to handle the nonsmooth total variation (TV) regularization term in the discrete MAP-Rician model. Theoretically, using the Kurdyka--Lojasiewicz (KL) property, the convergence of the numerical algorithm can be guaranteed. We also prove that the sequence generated by the proposed algorithm converges to a stationary point with the objective function values decreasing monotonically. Numerical simulations are then reported to clearly illustrate that our BDCA approach outperforms some state-of-the-art methods for both medical and natural images in terms of image recovery capability and CPU-time cost.
Tingting Wu 0001, Xiaoyu Gu, Zhi Li 0080, Jianwei Niu 0005, Tieyong Zeng
SIAM J. Imaging Sci.6
2022 Learning multi-level structural information for small organ segmentation
Yueyun Liu, Yuping Duan, Tieyong Zeng
Signal Process.3
2022 Quaternion Screened Poisson Equation for Low-Light Image Enhancement
abstract
Image enhancement is a technique to enhance the illumination of dark images while keeping the reality and the naturalness of the enhanced images at the same time. For color images, most methods tackle different color channels in a separate way, which overlooks the connection between the color channels. Therefore, in this paper, we consider a quaternion-based model to reserve the color connectivity, which integrates the color information of a pixel by a quaternion number. Moreover, we propose a regularizer based on the gamma-correction function and incorporate it into a screened Poisson equation for the image enhancement task. The uniqueness and existence of the solution of the proposed model are analyzed. The numerical results also prove the superiority of our scheme for color image enhancement.
Chaoyan Huang, Yingying Fang, Tingting Wu 0001, Tieyong Zeng, Yonghua Zeng
IEEE Signal Process. Lett.4
2022 Deep Tensor CCA for Multi-View Learning
abstract
We present Deep Tensor Canonical Correlation Analysis (DTCCA), a method to learn complex nonlinear transformations of multiple views (more than two) of data such that the resulting representations are linearly correlated in high order. The high-order correlation of given multiple views is modeled by covariance tensor, which is different from most CCA formulations relying solely on the pairwise correlations. Parameters of transformations of each view are jointly learned by maximizing the high-order canonical correlation. To solve the resulting problem, we reformulate it as the best sum of rank-1 approximation, which can be efficiently solved by existing tensor decomposition method. DTCCA is a nonlinear extension of tensor CCA (TCCA) via deep networks. Comparing with kernel TCCA, DTCCA not only can deal with arbitrary dimensions of the input data, but also does not need to maintain the training data for computing representations of any given data point. Hence, DTCCA as a unified model can efficiently overcome the scalable issue of TCCA for either high-dimensional multi-view data or a large amount of views, and it also naturally extends TCCA for learning nonlinear representation. Extensive experiments on four multi-view data sets demonstrate the effectiveness of the proposed method.
Hok Shing Wong, Li Wang 0033, Raymond Chan 0001, Tieyong Zeng
IEEE Trans. Big Data4
2022 Local Spatial Constraint and Total Variation for Hyperspectral Anomaly Detection
abstract
Hyperspectral anomaly detection, which is aimed at locating anomaly, has received widespread attention. In this article, a new anomaly detector, named local spatial constraint and total variation (LSC-TV), is proposed for hyperspectral imagery. In anomaly detection methods based on low-rank representation, background pixels are usually considered to have a global low-dimensional structure. However, the complex background distribution in hyperspectral images (HSIs) means that this global low-dimensional structure rarely occurs. In LSC-TV, the effective local spatial information is extracted by superpixel segmentation, and the regularization based on the F-norm is used to force the background within the same superpixel to show uniform spectral features. Moreover, each pixel is given a penalty based on the degree of anomaly determined during model iteration, while the anomaly is not considered by the background constraint. In addition, the background pixels in the neighborhood often show a high correlation, whereas the anomaly does not possess this feature. Nonisotropic TV is introduced into the proposed LSC model using the correlation of first-order neighborhoods to make it easier for anomalies to be separated. The proposed LSC-TV method and current state-of-the-art methods are tested on a set of simulated data and four sets of real data. The experimental results demonstrate that the proposed method is superior to the comparative method in terms of both color map detection and quantitative evaluation.
Ruyi Feng, Hao Li 0058, Lizhe Wang 0001, Yanfei Zhong, Liangpei Zhang 0001, Tieyong Zeng
IEEE Trans. Geosci. Remote. Sens.6
2022 Dual Learning-Based Graph Neural Network for Remote Sensing Image Super-Resolution
abstract
High-resolution (HR) remote sensing imagery plays a critical role in remote sensing image interpretation, and single image super-resolution (SISR) reconstruction technology is becoming increasingly valuable and significant. The state-of-the-art deep-learning-based SISR methods have demonstrated remarkable advantages, while reconstructing complex texture details still remains a big challenge. Besides, as a typical ill-posed inverse problem, how to determine the optimal solution is another important topic. To address these problems, in this work, a dual learning-based graph neural network (DLGNN) is proposed, in which the GNN is utilized to consider the self-similarity patches in remote sensing imagery by aggregating cross-scale neighboring feature patches, and dual learning strategy is adopted to refine the reconstruction results by constraining the mapping process in terms of the loss function, transferring the typical ill-posed problem to a well-posed one. Abundant experiments on 3K VEHICLE_SR datasets and Massachusetts Roads demonstrate the validity and outstanding performance for remote sensing image super-resolution tasks compared with other state-of-the-art super-resolution construction methods. Code is available at https://github.com/CUG-RS/DLGNN.
Ruyi Feng, Lizhe Wang 0001, Wei Han 0006, Tieyong Zeng
IEEE Trans. Geosci. Remote. Sens.5
2022 An O-Shape Neural Network With Attention Modules to Detect Junctions in Biomedical Images Without Segmentation
abstract
Junction plays an important role in biomedical research such as retinal biometric identification, retinal image registration, eye-related disease diagnosis and neuron reconstruction. However, junction detection in original biomedical images is extremely challenging. For example, retinal images contain many tiny blood vessels with complicated structures and low contrast, which makes it challenging to detect junctions. In this paper, we propose an O-shape Network architecture with Attention modules (Attention O-Net), which includes Junction Detection Branch (JDB) and Local Enhancement Branch (LEB) to detect junctions in biomedical images without segmentation. In JDB, the heatmap indicating the probabilities of junctions is estimated and followed by choosing the positions with the local highest value as the junctions, whereas it is challenging to detect junctions when the images contain weak filament signals. Therefore, LEB is constructed to enhance the thin branch foreground and make the network pay more attention to the regions with low contrast, which is helpful to alleviate the imbalance of the foreground between thin and thick branches and to detect the junctions of the thin branch. Furthermore, attention modules are utilized to introduce the feature maps of LEB to JDB, which can establish a complementary relationship and further integrate local features and contextual information between these two branches. The proposed method achieves the highest average F1-scores of 0.82, 0.73 and 0.94 in two retinal datasets and one neuron dataset, respectively. The experimental results confirm that Attention O-Net outperforms other state-of-the-art detection methods, and is helpful for retinal biometric identification.
Min Liu 0008, Fuhao Yu, Tieyong Zeng, Yaonan Wang 0001
IEEE J. Biomed. Health Informatics4
2022 Quaternion-Based Dictionary Learning and Saturation-Value Total Variation Regularization for Color Image Restoration
abstract
Color image restoration is a critical task in imaging sciences. Most variational methods regard the color image as a Euclidean vector or the direct combination of three monochrome images and completely ignore the inherent color structures within channels. To better describe the relationship of color channels, we represent the color image as the so-called pure quaternion matrix. Note that the celebrated dictionary learning method has attracted considerable attention for image recovery in the past decade. Following this idea, we propose a novel quaternion-based color image recovery method. This model combines the advantages of dictionary learning and the total variation method for color image restoration. The new strategy used in the proposed model manages to handle the color image restoration problem in the quaternion space. Moreover, the new proposed model can be easily solved by the classical alternating direction method of multipliers (ADMM) algorithm. Numerical results demonstrate clearly that the performance of our proposed dictionary learning method is better than some state-of-the-art color image dictionary learning and total variation methods in terms of some criteria and visual quality.
Chaoyan Huang, Michael Kwok-Po Ng, Tingting Wu 0001, Tieyong Zeng
IEEE Trans. Multim.4
2021 Rank-One Prior: Toward Real-Time Scene Recovery
abstract
Scene recovery is a fundamental imaging task for several practical applications, e.g., video surveillance and autonomous vehicles, etc. To improve visual quality under different weather/imaging conditions, we propose a real-time light correction method to recover the degraded scenes in the cases of sandstorms, underwater, and haze. The heart of our work is that we propose an intensity projection strategy to estimate the transmission. This strategy is motivated by a straightforward rank-one transmission prior. The complexity of transmission estimation is O(N ) where N is the size of the single image. Then we can recover the scene in real-time. Comprehensive experiments on different types of weather/imaging conditions illustrate that our method outperforms competitively several state-of-the-art imaging methods in terms of efficiency and robustness.
Jun Liu 0012, Ryan Wen Liu, Tieyong Zeng
CVPR4
2021 Structure-Preserving Deraining with Residue Channel Prior Guidance
abstract
Single image deraining is important for many high-level computer vision tasks since the rain streaks can severely degrade the visibility of images, thereby affecting the recognition and analysis of the image. Recently, many CNN-based methods have been proposed for rain removal. Although these methods can remove part of the rain streaks, it is difficult for them to adapt to real-world scenarios and restore high-quality rain-free images with clear and accurate structures. To solve this problem, we propose a Structure-Preserving Deraining Network (SPDNet) with RCP guidance. SPDNet directly generates high-quality rain-free images with clear and accurate structures under the guidance of RCP but does not rely on any rain-generating assumptions. Specifically, we found that the RCP of images contains more accurate structural information than rainy images. Therefore, we introduced it to our deraining network to protect structure information of the rain-free image. Meanwhile, a Wavelet-based Multi-Level Module (WMLM) is proposed as the backbone for learning the background information of rainy images and an Interactive Fusion Module (IFM) is designed to make full use of RCP information. In addition, an iterative guidance strategy is proposed to gradually improve the accuracy of RCP, refining the result in a progressive path. Extensive experimental results on both synthetic and real-world datasets demonstrate that the proposed model achieves new state-of-the-art results. Code: https://github.com/Joyies/SPDNet
Qiaosi Yi, Juncheng Li 0003, Qinyan Dai, Faming Fang, Guixu Zhang, Tieyong Zeng
ICCV6
2021 Local distribution-based adaptive minority oversampling for imbalanced data classification
Tieyong Zeng, Liping Jing
Neurocomputing3
2021 Smooth Soft-Balance Discriminative Analysis for imbalanced data
Liping Jing, Yilin Lyu, Mingzhe Guo, Tieyong Zeng
Knowl. Based Syst.5
2021 Surface-Aware Blind Image Deblurring
abstract
Blind image deblurring is a conundrum because there are infinitely many pairs of latent image and blur kernel. To get a stable and reasonable deblurred image, proper prior knowledge of the latent image and the blur kernel is urgently required. Different from the recent works on the statistical observations of the difference between the blurred image and the clean one, our method is built on the surface-aware strategy arising from the intrinsic geometrical consideration. This approach facilitates the blur kernel estimation due to the preserved sharp edges in the intermediate latent image. Extensive experiments demonstrate that our method outperforms the state-of-the-art methods on deblurring the text and natural images. Moreover, our method can achieve attractive results in some challenging cases, such as low-illumination images with large saturated regions and impulse noise. A direct extension of our method to the non-uniform deblurring problem also validates the effectiveness of the surface-aware prior.
Jun Liu 0012, Ming Yan 0006, Tieyong Zeng
IEEE Trans. Pattern Anal. Mach. Intell.3
2021 Adaptive total variation based image segmentation with semi-proximal alternating minimization
Tingting Wu 0001, Xiaoyu Gu, Youguo Wang, Tieyong Zeng
Signal Process.4
2021 Pixel-Attention CNN With Color Correlation Loss for Color Image Denoising
abstract
Convolutional neural networks (CNNs) have been applied to many image processing tasks and achieve great successes. In order to extract common features, every pixel in an image shares the same filters. However, pixels in different regions of an image varies dramatically and shared filters may lose some important local information. Rather than shared filters, smart filters which can be adapted to image context should be designed to better remove noise which occurs randomly in noisy image. Meanwhile, current CNN architectures compute the loss of each color channel independently, regardless of the potential color information. In this letter, we proposed a pixel-attention convolutional neural network (PACNN) with color correlation loss for the color image denoising task. The pixel-attention mechanism could generate pixel-wise attention maps which help remove random noise. The color correlation loss exploits color correlation to further improve denoising performance on color noisy images. The experimental results on several standard datasets demonstrate the state-of-the-art (SOTA) performance and the superiority of the proposed method.
Fan Jia 0007, Liyan Ma, Yijin Yang, Tieyong Zeng
IEEE Signal Process. Lett.4
2021 Multilevel Edge Features Guided Network for Image Denoising
abstract
Image denoising is a challenging inverse problem due to complex scenes and information loss. Recently, various methods have been considered to solve this problem by building a well-designed convolutional neural network (CNN) or introducing some hand-designed image priors. Different from previous works, we investigate a new framework for image denoising, which integrates edge detection, edge guidance, and image denoising into an end-to-end CNN model. To achieve this goal, we propose a multilevel edge features guided network (MLEFGN). First, we build an edge reconstruction network (Edge-Net) to directly predict clear edges from the noisy image. Then, the Edge-Net is embedded as part of the model to provide edge priors, and a dual-path network is applied to extract the image and edge features, respectively. Finally, we introduce a multilevel edge features guidance mechanism for image denoising. To the best of our knowledge, the Edge-Net is the first CNN model specially designed to reconstruct image edges from the noisy image, which shows good accuracy and robustness on natural images. Extensive experiments clearly illustrate that our MLEFGN achieves favorable performance against other methods and plenty of ablation studies demonstrate the effectiveness of our proposed Edge-Net and MLEFGN. The code is available at https://github.com/MIVRC/MLEFGN-PyTorch.
Faming Fang, Juncheng Li 0003, Yiting Yuan, Tieyong Zeng, Guixu Zhang
IEEE Trans. Neural Networks Learn. Syst.4
2021 Probabilistic Semi-Supervised Learning via Sparse Graph Structure Learning
abstract
We present a probabilistic semi-supervised learning (SSL) framework based on sparse graph structure learning. Different from existing SSL methods with either a predefined weighted graph heuristically constructed from the input data or a learned graph based on the locally linear embedding assumption, the proposed SSL model is capable of learning a sparse weighted graph from the unlabeled high-dimensional data and a small amount of labeled data, as well as dealing with the noise of the input data. Our representation of the weighted graph is indirectly derived from a unified model of density estimation and pairwise distance preservation in terms of various distance measurements, where latent embeddings are assumed to be random variables following an unknown density function to be learned, and pairwise distances are then calculated as the expectations over the density for the model robustness to the data noise. Moreover, the labeled data based on the same distance representations are leveraged to guide the estimated density for better class separation and sparse graph structure learning. A simple inference approach for the embeddings of unlabeled data based on point estimation and kernel representation is presented. Extensive experiments on various data sets show promising results in the setting of SSL compared with many existing methods and significant improvements on small amounts of labeled data.
Li Wang 0033, Raymond Chan 0001, Tieyong Zeng
IEEE Trans. Neural Networks Learn. Syst.3
2020 Learning deep edge prior for image denoising
Yingying Fang, Tieyong Zeng
Comput. Vis. Image Underst.2
2020 Residual network with detail perception loss for single image super-resolution
Zhijie Wen, Jiawei Guan, Tieyong Zeng, Ying Li 0028
Comput. Vis. Image Underst.3
2020 A Three-Stage Variational Image Segmentation Framework Incorporating Intensity Inhomogeneity Information
abstract
In this paper, we propose a new three-stage segmentation framework based on a convex variant of the Mumford--Shah model and the intensity inhomogeneity information of an image. The first stage in our framework is to perform a dimension lifting method. An intensity inhomogeneity image is added as an additional channel, which results in a vector-valued image. In the second stage, a convex variant of the Mumford--Shah model is applied to each channel of the vector-valued image to obtain a smooth approximation. We use the semi--proximal alternating direction method of multipliers (sPADMM) to solve this model and prove that the sPADMM for solving this convex model has Q-linear convergence rate. In the last stage, we apply a thresholding method to the smoothed vector-valued image to get the final segmentation. Experiments demonstrate clearly that the proposed methods can provide more accurate segmentation results in comparison with five state-of-the-art methods including a deep learning approach.
Xiaoping Yang 0001, Tieyong Zeng
SIAM J. Imaging Sci.3
2020 Image denoising based on the adaptive weighted TVp regularization
Zhi-Feng Pang, Hui-Li Zhang, Shousheng Luo, Tieyong Zeng
Signal Process.4
2020 A weighted bounded Hessian variational model for image labeling and segmentation
Qiuxiang Zhong, Yuping Duan, Tieyong Zeng
Signal Process.4
2020 Deep Multi-Level Wavelet-CNN Denoiser Prior for Restoring Blurred Image With Cauchy Noise
abstract
Cauchy noise, as a typical non-Gaussian noise, appears frequently in many important fields, such as radar, medical, and biomedical imaging. In this letter, we focus on image recovery under Cauchy noise. Instead of the celebrated total variation or low-rank prior, we adopt a novel deep-learning-based image denoiser prior to effectively remove Cauchy noise with blur. To preserve more detailed texture and better balance between the receptive field size and the computational cost, we apply the multi-level wavelet convolutional neural network (MWCNN) to train this denoiser. We use the forward-backward splitting (FBS) method to handle the proposed model, which can be implemented efficiently without introducing auxiliary variables. Moreover, the multi-noise-levels strategy is employed to train a series of denoisers to restore the image corrupted by Cauchy noise and blur. Numerical experiments demonstrate clearly that our method has better performance than the existing image restoration methods for removing Cauchy noise in terms of the quantitative index and visual quality.
Tingting Wu 0001, Wei Li 0146, Shilong Jia, Yiqiu Dong, Tieyong Zeng
IEEE Signal Process. Lett.5
2020 Soft-Edge Assisted Network for Single Image Super-Resolution
abstract
The task of single image super-resolution (SISR) is a highly ill-posed inverse problem since reconstructing the highfrequency details from a low-resolution image is challenging. Most previous CNN-based super-resolution (SR) methods tend to directly learn the mapping from the low-resolution image to the high-resolution image through some complex convolutional neural networks. However, the method of blindly increasing the depth of the network is not the best choice because the performance improvement of such methods is marginal but the computational cost is huge. A more efficient method is to integrate the image prior knowledge into the model to assist the image reconstruction. Indeed, the soft-edge has been widely applied in many computer vision tasks as the role of an important image feature. In this paper, we propose a Soft-edge assisted Network (SeaNet) to reconstruct the high-quality SR image with the help of image soft-edge. The proposed SeaNet consists of three sub-nets: a rough image reconstruction network (RIRN), a soft-edge reconstruction network (Edge-Net), and an image refinement network (IRN). The complete reconstruction process consists of two stages. In Stage-I, the rough SR feature maps and the SR soft-edge are reconstructed by the RIRN and Edge-Net, respectively. In Stage-II, the outputs of the previous stages are fused and then feed to the IRN for high-quality SR image reconstruction. Extensive experiments show that our SeaNet converges rapidly and achieves excellent performance under the assistance of image soft-edge. The code is available at https://gitlab.com/junchenglee/seanet-pytorch.
Faming Fang, Juncheng Li 0003, Tieyong Zeng
IEEE Trans. Image Process.3
2020 Variational Single Image Dehazing for Enhanced Visualization
abstract
In this paper, we investigate the challenging task of removing haze from a single natural image. The analysis on the haze formation model shows that the atmospheric veil has much less relevance to chrominance than luminance, which motivates us to neglect the haze in the chrominance channel and concentrate on the luminance channel in the dehazing process. Besides, the experimental study illustrates that the YUV color space is most suitable for image dehazing. Accordingly, a variational model is proposed in the Y channel of the YUV color space by combining the reformulation of the haze model and the two effective priors. As we mainly focus on the Y channel, most of the chrominance information of the image is preserved after dehazing. The numerical procedure based on the alternating direction method of multipliers (ADMM) scheme is presented to obtain the optimal solution. Extensive experimental results on real-world hazy images and synthetic dataset demonstrate clearly that our method can unveil the details and recover vivid color information, which is competitive among many existing dehazing algorithms. Further experiments show that our model also can be applied for image enhancement.
Faming Fang, Tingting Wang 0007, Yang Wang 0020, Tieyong Zeng, Guixu Zhang
IEEE Trans. Multim.4
2020 A Superpixel-Based Variational Model for Image Colorization
abstract
Image colorization refers to a computer-assisted process that adds colors to grayscale images. It is a challenging task since there is usually no one-to-one correspondence between color and local texture. In this paper, we tackle this issue by exploiting weighted nonlocal self-similarity and local consistency constraints at the resolution of superpixels. Given a grayscale target image, we first select a color source image containing similar segments to target image and extract multi-level features of each superpixel in both images after superpixel segmentation. Then a set of color candidates for each target superpixel is selected by adopting a top-down feature matching scheme with confidence assignment. Finally, we propose a variational approach to determine the most appropriate color for each target superpixel from color candidates. Experiments demonstrate the effectiveness of the proposed method and show its superiority to other state-of-the-art methods. Furthermore, our method can be easily extended to color transfer between two color images.
Faming Fang, Tingting Wang 0007, Tieyong Zeng, Guixu Zhang
IEEE Trans. Vis. Comput. Graph.3
2019 Image Smoothing Via Gradient Sparsity and Surface Area Minimization
abstract
Image smoothing is a very important topic in image processing. Among these image smoothing methods, the L0gradient minimization method is one of the most popular ones. However, the L0gradient minimization method suffers from the staircasing effect and over-sharpening issue, which highly degrade the quality of the smoothed image. To overcome these issues, we use not only the L0gradient term for finding edges, but also a surface area based term for the purpose of smoothing the inside of each region. An alternating minimization algorithm is suggested to efficiently solve the proposed model, where each subproblem has a closed-form solution. Leveraging the introduced surface area term, the proposed method can effectively alleviate the staircasing effect and the over-sharpening issue. The superiority of our method over the state-of-the-art methods is demonstrated by a series of experiments.
Jun Liu 0012, Ming Yan 0006, Jinshan Zeng, Tieyong Zeng
ICIP4
2018 Weighted variational model for selective image segmentation with application to medical images
Michael Kwok-Po Ng, Tieyong Zeng
Pattern Recognit.3
2018 Variational Phase Retrieval with Globally Convergent Preconditioned Proximal Algorithm
abstract
We reformulate the original phase retrieval problem into two variational models (with and without regularization), both containing a globally Lipschitz differentiable term. These two models can be efficiently solved via the proposed Partially Preconditioned Proximal Alternating Linearized Minimization (P${}^3$ALM) for masked Fourier measurements. Thanks to the Lipschitz differentiable term, we prove the global convergence of P${}^3$ALM for solving the nonconvex phase retrieval problems. Extensive experiments are conducted to show the effectiveness of the proposed methods.
Huibin Chang, Stefano Marchesini, Yifei Lou, Tieyong Zeng
SIAM J. Imaging Sci.4
2016 A New Algorithm Framework for Image Inpainting in Transform Domain
abstract
In this paper, we focus on variational approaches for image inpainting in transform domain and propose two new algorithms, iterative coupled transform domain inpainting (ICTDI) and iterative decoupled transform domain inpainting. In the derivation of ICTDI, we use operator splitting and the quadratic penalty technique to get a new approximate problem of the basic model. By the alternating minimization method, the approximate problem can be decomposed as three relatively simple subproblems with closed-form solutions. However, ICTDI is not efficient when some adaptive regularization operator is used, such as the learned BM3D frame. To overcome this drawback, with some modifications, we decouple our framework into three relatively independent parts: denoising, linear combination in the transform domain, and linear combination in the image domain. Therefore, we can use any existing denoising method in the denoising step. We consider three choices for regularization operators in our approach: gradient operator, tight framelet transform, and learned BM3D frame. The numerical experiments and comparisons on various images demonstrate the effectiveness of the proposed methods. The convergence of the numerical algorithms is proved under some assumptions.
Fang Li 0004, Tieyong Zeng
SIAM J. Imaging Sci.2
2015 A Weighted Difference of Anisotropic and Isotropic Total Variation Model for Image Processing
abstract
We propose a weighted difference of anisotropic and isotropic total variation (TV) as a regularization for image processing tasks, based on the well-known TV model and natural image statistics. Due to the form of our model, it is natural to compute via a difference of convex algorithm (DCA). We draw its connection to the Bregman iteration for convex problems and prove that the iteration generated from our algorithm converges to a stationary point with the objective function values decreasing monotonically. A stopping strategy based on the stable oscillatory pattern of the iteration error from the ground truth is introduced. In numerical experiments on image denoising, image deblurring, and magnetic resonance imaging (MRI) reconstruction, our method improves on the classical TV model consistently and is on par with representative state-of-the-art methods.
Yifei Lou, Tieyong Zeng, Stanley J. Osher, Jack Xin
SIAM J. Imaging Sci.2
2015 Variational Approach for Restoring Blurred Images with Cauchy Noise
abstract
The restoration of images degraded by blurring and noise is one of the most important tasks in image processing. In this paper, based on the total variation (TV) we propose a new variational method for recovering images degraded by Cauchy noise and blurring. In order to obtain a strictly convex model, we add a quadratic penalty term, which guarantees the uniqueness of the solution. Due to the convexity of our model, the primal dual algorithm is employed to solve the minimization problem. Experimental results show the effectiveness of the proposed method for simultaneously deblurring and denoising images corrupted by Cauchy noise. Comparison with other existing and well-known methods is provided as well.
Federica Sciacchitano, Yiqiu Dong, Tieyong Zeng
SIAM J. Imaging Sci.3
2014 A Two-Stage Image Segmentation Method Using Euler's Elastica Regularized Mumford-Shah Model
abstract
As one of the most important image segmentation models, the Mumford-Shah functional was developed to pursue a piecewise smooth approximation of a given image based on the regularization on the total length of curves. In this paper, we modify the Mumford-Shah model using Euler's elastic a as the regularization. A two-stage segmentation method is applied the Euler's elastic a regularized Mumford-Shah model. The first stage is to find a smooth solution of the variant Mumford-Shah functional based on augmented Lagrangian method while a thresholding is performed in the second stage to obtain different phases for the segmentation. The K-means clustering method is used as the technique to find the thresholds for the segmentation. For intensity inhomogeneous images, we eliminate the effect of the bias field by bias-corrected fuzzy c-means method. Experimental results show that as the regularization, Euler's elastic a makes the Mumford-Shah model perform better for many kinds of images, including tubular and irregular shaped, CT Angiography (CTA) and MRI images in different noise level.
Yuping Duan, Weimin Huang 0002, Jiayin Zhou, Huibin Chang, Tieyong Zeng
ICPR5
2014 A Two-Stage Image Segmentation Method for Blurry Images with Poisson or Multiplicative Gamma Noise
abstract
In this paper, a two-stage method for segmenting blurry images in the presence of Poisson or multiplicative Gamma noise is proposed. The method is inspired by a previous work on two-stage segmentation and the usage of an I-divergence term to handle the noise. The first stage of our method is to find a smooth solution $u$ to a convex variant of the Mumford--Shah model where the $\ell_2$ data-fidelity term is replaced by an I-divergence term. A primal-dual algorithm is adopted to efficiently solve the minimization problem. We prove the convergence of the algorithm and the uniqueness of the solution $u$. Once $u$ is obtained, in the second stage, the segmentation is done by thresholding $u$ into different phases. The thresholds can be given by the users or can be obtained automatically by using any clustering method. In our method, we can obtain any $K$-phase segmentation ($K\geq 2$) by choosing $(K-1)$ thresholds after $u$ is found. Changing $K$ or the thresholds does not require $u$ to be recomputed. Experimental results show that our two-stage method performs better than many standard two-phase or multiphase segmentation methods for very general images, including antimass, tubular, magnetic resonance imaging, and low-light images.
Raymond Chan 0001, Hongfei Yang, Tieyong Zeng
SIAM J. Imaging Sci.3
2014 Single Image Dehazing and Denoising: A Fast Variational Approach
abstract
In this paper, we propose a new fast variational approach to dehaze and denoise simultaneously. The proposed method first estimates a transmission map using a windows adaptive method based on the celebrated dark channel prior. This transmission map can significantly reduce the edge artifact in the resulting image and enhance the estimation precision. The transmission map is then converted to a depth map, with which the new variational model can be built to seek the final haze- and noise-free image. The existence and uniqueness of a minimizer of the proposed variational model is further discussed. A numerical procedure based on the Chambolle--Pock algorithm is given, and the convergence of the algorithm is ensured. Extensive experimental results on real scenes demonstrate that our method can restore vivid and contrastive haze- and noise-free images effectively.
Faming Fang, Fang Li 0004, Tieyong Zeng
SIAM J. Imaging Sci.3
2014 A Universal Variational Framework for Sparsity-Based Image Inpainting
abstract
In this paper, we extend an existing universal variational framework for image inpainting with new numerical algorithms. Given certain regularization operator Φ and denoting u the latent image, the basic model is to minimize the l(p), (p=0,1) norm of Φu preserving the pixel values outside the inpainting region. Utilizing the operator splitting technique, the original problem can be approximated by a new problem with extra variable. With the alternating minimization method, the new problem can be decomposed as two subproblems with exact solutions. There are many choices for Φ in our approach such as gradient operator, wavelet transform, framelet transform, or other tight frames. Moreover, with slight modification, we can decouple our framework into two relatively independent parts: 1) denoising and 2) linear combination. Therefore, we can take any denoising method, including BM3D filter in the denoising step. The numerical experiments on various image inpainting tasks, such as scratch and text removal, randomly missing pixel filling, and block completion, clearly demonstrate the super performance of the proposed methods. Furthermore, the theoretical convergence of the proposed algorithms is proved.
Fang Li 0004, Tieyong Zeng
IEEE Trans. Image Process.2
2013 A Two-Stage Image Segmentation Method Using a Convex Variant of the Mumford-Shah Model and Thresholding
abstract
The Mumford--Shah model is one of the most important image segmentation models and has been studied extensively in the last twenty years. In this paper, we propose a two-stage segmentation method based on the Mumford--Shah model. The first stage of our method is to find a smooth solution $g$ to a convex variant of the Mumford--Shah model. Once $g$ is obtained, then in the second stage the segmentation is done by thresholding $g$ into different phases. The thresholds can be given by the users or can be obtained automatically using any clustering methods. Because of the convexity of the model, $g$ can be solved efficiently by techniques like the split-Bregman algorithm or the Chambolle--Pock method. We prove that our method is convergent and that the solution $g$ is always unique. In our method, there is no need to specify the number of segments $K$ ($K\geq2$) before finding $g$. We can obtain any $K$-phase segmentations by choosing $(K-1)$ thresholds after $g$ is found in the first stage, and in the second stage there is no need to recompute $g$ if the thresholds are changed to reveal different segmentation features in the image. Experimental results show that our two-stage method performs better than many standard two-phase or multiphase segmentation methods for very general images, including antimass, tubular, MRI, noisy, and blurry images.
Xiaohao Cai, Raymond Chan 0001, Tieyong Zeng
SIAM J. Imaging Sci.3
2013 A Convex Variational Model for Restoring Blurred Images with Multiplicative Noise
abstract
In this paper, a new variational model for restoring blurred images with multiplicative noise is proposed. Based on the statistical property of the noise, a quadratic penalty function technique is utilized in order to obtain a strictly convex model under a mild condition, which guarantees the uniqueness of the solution and the stabilization of the algorithm. For solving the new convex variational model, a primal-dual algorithm is proposed, and its convergence is studied. The paper ends with a report on numerical tests for the simultaneous deblurring and denoising of images subject to multiplicative noise. A comparison with other methods is provided as well.
Yiqiu Dong, Tieyong Zeng
SIAM J. Imaging Sci.2
2013 Sparse Representation Prior and Total Variation-Based Image Deblurring under Impulse Noise
abstract
In this paper, we study the image recovery problem where the observed image is simultaneously corrupted by blur and impulse noise. Our proposed patch-based model contains three terms: the sparse representation prior, the total variation regularization, and the data-fidelity term. We are interested in the two-phase approach. The first phase is to identify the possible impulse noise positions; the second phase is to recover the image via the patch-based model using noise position information. An alternating minimization method is then applied to solve the model. This approach works extremely well for image deblurring under salt-and-pepper noise. However, as the detection for random-valued noise is usually unreliable, extra work is then needed. Indeed, to get better recovery results for the latter case, we combine the two separate phases to simultaneously detect the random-valued noise positions and to recover the image. The numerical experiments clearly demonstrate the super performance of the proposed methods.
Liyan Ma, Jian Yu 0001, Tieyong Zeng
SIAM J. Imaging Sci.3
2013 General Framework to Histogram-Shifting-Based Reversible Data Hiding
abstract
Histogram shifting (HS) is a useful technique of reversible data hiding (RDH). With HS-based RDH, high capacity and low distortion can be achieved efficiently. In this paper, we revisit the HS technique and present a general framework to construct HS-based RDH. By the proposed framework, one can get a RDH algorithm by simply designing the so-called shifting and embedding functions. Moreover, by taking specific shifting and embedding functions, we show that several RDH algorithms reported in the literature are special cases of this general construction. In addition, two novel and efficient RDH algorithms are also introduced to further demonstrate the universality and applicability of our framework. It is expected that more efficient RDH algorithms can be devised according to the proposed framework by carefully designing the shifting and embedding functions.
Xiaolong Li 0001, Bin Li 0011, Bin Yang 0001, Tieyong Zeng
IEEE Trans. Image Process.4
2013 A Dictionary Learning Approach for Poisson Image Deblurring
abstract
The restoration of images corrupted by blur and Poisson noise is a key issue in medical and biological image processing. While most existing methods are based on variational models, generally derived from a maximum a posteriori (MAP) formulation, recently sparse representations of images have shown to be efficient approaches for image recovery. Following this idea, we propose in this paper a model containing three terms: a patch-based sparse representation prior over a learned dictionary, the pixel-based total variation regularization term and a data-fidelity term capturing the statistics of Poisson noise. The resulting optimization problem can be solved by an alternating minimization technique combined with variable splitting. Extensive experimental results suggest that in terms of visual quality, peak signal-to-noise ratio value and the method noise, the proposed algorithm outperforms state-of-the-art methods.
Liyan Ma, Lionel Moisan, Jian Yu 0001, Tieyong Zeng
IEEE Trans. Medical Imaging4
2013 Dictionary Learning-Based Subspace Structure Identification in Spectral Clustering
abstract
In this paper, we study dictionary learning (DL) approach to identify the representation of low-dimensional subspaces from high-dimensional and nonnegative data. Such representation can be used to provide an affinity matrix among different subspaces for data clustering. The main contribution of this paper is to consider both nonnegativity and sparsity constraints together in DL such that data can be represented effectively by nonnegative and sparse coding coefficients and nonnegative dictionary bases. In the algorithm, we employ the proximal point technique for the resulting DL and sparsity optimization problem. We make use of coding coefficients to perform spectral clustering (SC) for data partitioning. Extensive experiments on real-world high-dimensional and nonnegative data sets, including text, microarray, and image data demonstrate that the proposed method can discover their subspace structures. Experimental results also show that our algorithm is computationally efficient and effective for obtaining high SC performance and interpreting the clustering results compared with the other testing methods.
Liping Jing, Michael Kwok-Po Ng, Tieyong Zeng
IEEE Trans. Neural Networks Learn. Syst.3
2012 Lagrangian multipliers and split Bregman methods for minimization problems constrained on Sn-1
Fang Li 0004, Tieyong Zeng, Guixu Zhang
J. Vis. Commun. Image Represent.2
2012 Explicit Coherence Enhancing Filter With Spatial Adaptive Elliptical Kernel
abstract
The goal of this letter is to provide an elliptical filter to improve image coherence for the task of image smoothing and inpainting. The kernel of this filter is adaptively weighted and its shape is determined by local coherence estimation. The long axis of its ellipse is the same as the coherence direction and we put more weight there to enhance coherence. Compared with the related anisotropic partial differential equations (PDEs) or wavelet shrinkage methods, the proposed filter is extremely simple, instinctive and easy to code. Numerical examples and comparisons illustrate clearly the good performance of the proposed filter.
Fang Li 0004, Ling Pi, Tieyong Zeng
IEEE Signal Process. Lett.3
2012 Multiplicative Noise Removal via a Learned Dictionary
abstract
Multiplicative noise removal is a challenging image processing problem, and most existing methods are based on the maximum a posteriori formulation and the logarithmic transformation of multiplicative denoising problems into additive denoising problems. Sparse representations of images have shown to be efficient approaches for image recovery. Following this idea, in this paper, we propose to learn a dictionary from the logarithmic transformed image, and then to use it in a variational model built for noise removal. Extensive experimental results suggest that in terms of visual quality, peak signal-to-noise ratio, and mean absolute deviation error, the proposed algorithm outperforms state-of-the-art methods.
Yu-Mei Huang, Lionel Moisan, Michael Kwok-Po Ng, Tieyong Zeng
IEEE Trans. Image Process.4
2011 Restoration of images corrupted by mixed Gaussian-impulse noise via l1-l0 minimization
Tieyong Zeng, Jian Yu 0001, Michael Kwok-Po Ng
Pattern Recognit.2
2011 Matching pursuit shrinkage in Hilbert spaces
Tieyong Zeng, François Malgouyres
Signal Process.1
2011 Efficient Reversible Watermarking Based on Adaptive Prediction-Error Expansion and Pixel Selection
abstract
Prediction-error expansion (PEE) is an important technique of reversible watermarking which can embed large payloads into digital images with low distortion. In this paper, the PEE technique is further investigated and an efficient reversible watermarking scheme is proposed, by incorporating in PEE two new strategies, namely, adaptive embedding and pixel selection. Unlike conventional PEE which embeds data uniformly, we propose to adaptively embed 1 or 2 bits into expandable pixel according to the local complexity. This avoids expanding pixels with large prediction-errors, and thus, it reduces embedding impact by decreasing the maximum modification to pixel values. Meanwhile, adaptive PEE allows very large payload in a single embedding pass, and it improves the capacity limit of conventional PEE. We also propose to select pixels of smooth area for data embedding and leave rough pixels unchanged. In this way, compared with conventional PEE, a more sharply distributed prediction-error histogram is obtained and a better visual quality of watermarked image is observed. With these improvements, our method outperforms conventional PEE. Its superiority over other state-of-the-art methods is also demonstrated experimentally.
Xiaolong Li 0001, Bin Yang 0001, Tieyong Zeng
IEEE Trans. Image Process.3
2010 Reliable histogram features for detecting LSB matching
abstract
This paper proposes a novel steganalyzer for detecting one of the most popular steganography, LSB matching (also known as “±1 embedding”). The histogram of difference image (the differences of adjacent pixels), which is usually a generalized Gaussian distribution centered at 0, is exploited for deriving statistical features. We have proved theoretically that the peak-value of the histogram would decrease after LSB matching embedding, while the renormalized histogram (the ratio of the histogram to the peak-value) would increase. Then we take the peak-value and the renormalized histogram as features for classification. Extensive experimental results show that the proposed steganalytic method outperforms some previous ones.
Kaiwei Cai, Xiaolong Li 0001, Tieyong Zeng, Bin Yang 0001, Xiaoqing Lu
ICIP3
2010 Poisson noise removal via learned dictionary
abstract
In this paper, we address the restoration of images corrupted by Poisson noise. The proposed new model contains two terms: one is from the sparse representation of the transformed image via variance stabilizing transformation (VST); the other is a data-fidelity term caused by the statistical properties of Poisson noise. The main algorithm is efficient. We first learn a dictionary to sparsely represent the transformed image using a state-of-the-art dictionary learning method, and then solve the minimization of the variational form by Newton method. Comparative experiments are carried out to show the leading performance of our new model.
Tieyong Zeng
ICIP2
2010 A Multiphase Image Segmentation Method Based on Fuzzy Region Competition
abstract
The goal of this paper is to develop a multiphase image segmentation method based on fuzzy region competition. A new variational functional with constraints is proposed by introducing fuzzy membership functions which represent several different regions in an image. The existence of a minimizer of this functional is established. We propose three methods for handling the constraints of membership functions in the minimization. We also add auxiliary variables to approximate the membership functions in the functional such that Chambolle's fast dual projection method can be used. An alternate minimization method can be employed to find the solution, in which the region parameters and the membership functions have closed form solutions. Numerical examples using grayscale and color images are given to demonstrate the effectiveness of the proposed methods.
Fang Li 0004, Michael Kwok-Po Ng, Tieyong Zeng, Chunli Shen
SIAM J. Imaging Sci.3
2010 On the Total Variation Dictionary Model
abstract
The goal of this paper is to provide a theoretical study of a total variation (TV) dictionary model. Based on the properties of convex analysis and bounded variation functions, the existence of solutions of the TV dictionary model is proved. We then show that the dual form of the model can be given by the minimization of the sum of the l(1) -norm of the dual solution and the Bregman distance between the curvature of the primal solution and the subdifferential of TV norm of the dual solution. This theoretical result suggests that the dictionary must represent sparsely the curvatures of solution image in order to obtain a better denoising performance.
Tieyong Zeng, Michael Kwok-Po Ng
IEEE Trans. Image Process.1
2009 Improving embedding efficiency via matrix embedding: A case study
abstract
Matrix embedding is proved an effective way to improve embedding efficiency of steganography, and usually, higher dimensional matrix will provide better embedding efficiency. However, the sender might suffer huge computational complexity when employing high dimensional matrix. In this paper, a variation of matrix embedding is proposed. Instead of finding a coset leader as the modification of the cover image (which is the most time consuming step in matrix embedding), we turn to finding a vector in the coset which has relatively small Hamming weight. Such a vector can be found at reduced computational cost. Consequently, higher dimensional matrix can be used for practically implementable matrix embedding. The parity check matrix of random linear code is used in the experiments. It has shown that the proposed scheme can enhance the practicability and efficiency of the conventional matrix embedding.
Yunkai Gao 0002, Xiaolong Li 0001, Tieyong Zeng, Bin Yang 0001
ICIP3
2009 A Predual Proximal Point Algorithm Solving a Non Negative Basis Pursuit Denoising Model
François Malgouyres, Tieyong Zeng
Int. J. Comput. Vis.2
2009 A Generalization of LSB Matching
abstract
Recently, a significant improvement of the well-known least significant bit (LSB) matching steganography has been proposed, reducing the changes to the cover image for the same amount of embedded secret data. When the embedding rate is 1, this method decreases the expected number of modification per pixel (ENMPP) from 0.5 to 0.375. In this letter, we propose the so-called generalized LSB matching (G-LSB-M) scheme, which generalizes this method and LSB matching. The lower bound of ENMPP for G-LSB-M is investigated, and a construction of G-LSB-M is presented by using the sum and difference covering set of finite cyclic group. Compared with the previous works, we show that the suitable G-LSB-M can further reduce the ENMPP and lead to more secure steganographic schemes. Experimental results illustrate clearly the better resistance to steganalysis of G-LSB-M.
Xiaolong Li 0001, Bin Yang 0001, Daofang Cheng, Tieyong Zeng
IEEE Signal Process. Lett.4
2008 A further study on steganalysis of LSB matching by calibration
abstract
In this paper, based on a careful investigation on the calibration (downsample) technique, we improve two detectors for detecting LSB matching: calibrated HCF COM and calibrated adjacency HCF COM. Instead of using the COM (center of mass) of the HCF (histogram characteristic function), we consider the ratio of the histogram's DFT coefficients of the image to the corresponding coefficients of the down-sampled image. Moreover, we propose to down-sample only for non-oscillating pixels. With a same level of computational complexity, the new detectors thus obtained are better than the old ones, especially for uncompressed images.
Xiaolong Li 0001, Tieyong Zeng, Bin Yang 0001
ICIP2
2008 Incorporating known features into a total variation dictionary model for source separation
abstract
The goal of this paper is to investigate the impact of dictionary choosing for a total variation dictionary model. After theoretical analysis, we present the experiments in which the dictionary contains the curvatures of known forms (letters). The data-fidelity term of this model allows the appearance in the residue of all structures except forms being used to build the dictionary. Therefore, these forms will remain in the result image while the other structures will disappear. Our experiments are carried on the source separation problem and confirm this impression. The starting image contains letters (known) on a very structured background (an image). We show that it is possible, with this model, to obtain a reasonable separation of these structures. Finally, this work illustrates clearly that the dictionary must contain the curvature of elements which we seek to preserve.
Tieyong Zeng
ICIP1
2008 Improvement of the embedding efficiency of LSB matching by sum and difference covering set
abstract
As an important attribute directly influencing the steganographic scheme, the embedding efficiency is defined as the average number of random data bits per one embedding change. In this paper, we propose a novel approach to improve the embedding efficiency of LSB matching based on the sum and difference covering set (SDCS) of finite cyclic group. We show that the suitable choice of SDCS will lead to a new steganographic scheme which is more efficient than LSB matching. Then we illustrate that the new scheme keeps the statistical imperceptibility of LSB matching. The detailed constructions of SDCS are given and some related problems are also discussed.
Xiaolong Li 0001, Tieyong Zeng, Bin Yang 0001
ICME2
2006 Using Gabor Dictionaries in A TV - L∞Model, for Denoising
abstract
The goal of this paper is to report on experiments where we use Gabor dictionaries in a TV - linfinmodel for denoising. This allows many possible choices. Our conclusions are that the choice of the dictionary mostly impact the restoration of textures. Moreover, for most images, better results are obtained when the Gaussian term of the Gabor filters is close to isotropic
Tieyong Zeng, François Malgouyres
ICASSP (2)1