VLDB 2026 Research / reviewers in the wild / expert
Kaibing Zhang
dblp:91/10256
· DBLP profile ↗
99ranked-venue papers
17as first author
71since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 51 · 12 first-author · 33 since 2021Artificial intelligence and machine learning · 41 · 6 first-author · 32 since 2021Security and privacy · 3 · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Point-SRA: Self-Representation Alignment for 3D Representation LearningabstractMasked autoencoders (MAE) have become a dominant paradigm in 3D representation learning, setting new performance benchmarks across various downstream tasks. Existing methods with fixed mask ratios neglect multi-level representational correlations and intrinsic geometric structures, while relying on point-wise reconstruction assumptions that conflict with the diversity of point cloud. To address these issues, we propose a 3D representation learning method, termed Point-SRA, which aligns representations through self-distillation and probabilistic modeling. Specifically, we assign different masking ratios to the MAE to capture complementary geometric and semantic information, while the MeanFlow Transformer (MFT) leverages cross-modal conditional embeddings to enable diverse probabilistic reconstruction. Our analysis further reveals that representations at different time steps in MFT also exhibit complementarity. Therefore, a Dual Self-Representation Alignment mechanism is proposed at both the MAE and MFT levels. Finally, we design a Flow-Conditioned Fine-Tuning Architecture to fully exploit the point cloud distribution learned via MeanFlow. Point-SRA outperforms Point-MAE by 5.37% on ScanObjectNN. On intracranial aneurysm segmentation, it reaches 96.07% mean IoU for arteries and 86.87% for aneurysms. For 3D object detection, Point-SRA achieves 47.3% AP@50, surpassing MaskPoint by 5.12%. Lintong Wei, Haozhe Cheng, Jihua Zhu, Kaibing Zhang |
AAAI | 5 |
| 2026 | An interpretable spatiotemporal method with a composite attention mechanism for the prediction of air pollution with stable and dynamic spatial relationships
Yee Leung, Kaibing Zhang, Dinghua Xue, Shuyun Yang |
Expert Syst. Appl. | 3 |
| 2026 | ACENet: Contextual graph reasoning for semi-supervised crowd counting
Yalei Meng, Kaibing Zhang |
Expert Syst. Appl. | 6 |
| 2026 | Discriminative approximate low-rank projection with adaptive distance penalty for feature extraction
Shigang Liu, Di Wu 0058, Weihua Ou, Kaibing Zhang |
Inf. Process. Manag. | 6 |
| 2026 | Dynamic spatiotemporal air quality modeling with local and sparse spatial fusion and structured temporal modeling
Jiaxing Yu, Kaibing Zhang, Dinghua Xue, Pengfang Li, Minna Xiao, Shuyun Yang |
Inf. Sci. | 4 |
| 2026 | Scale-Customized Feature Learning Network for crowd counting
Kaibing Zhang, Yalei Meng |
J. Vis. Commun. Image Represent. | 2 |
| 2026 | DAF-VITON: Lightweight Diffusion Virtual Try-On via Efficient Dynamic Attention and Boundary-Aware Fusion
Dongchuang Zhao, Jinguang Chen, Kaibing Zhang |
Multim. Syst. | 4 |
| 2026 | Rethinking progressive low-light image enhancement: A frequency-aware tripartite multi-scale network
Kaibing Zhang, Zhouqiang Zhang |
Neural Networks | 2 |
| 2026 | TextSRFormer: Multi-head axial self-attention transformer for scene text image super-resolution
Aobin Cheng, Kaibing Zhang, Dinghua Xue |
Signal Process. | 3 |
| 2026 | FlowAdapt-GS: Dual Optical Flow Supervision With Adaptive Keyframe Sampling for Endoscopic 4D Reconstruction
Weihua Ou, Kaibing Zhang |
IEEE Signal Process. Lett. | 5 |
| 2025 | Text2Printing: Controllable Textile Digital Printing Pattern Generation with Attention Modulation
Minghui Ding, Kaibing Zhang |
PRCV (5) | 2 |
| 2025 | CAP-Shift: Domain Adaptive Nighttime Object Detection via Illumination Degradation and Confidence-Adaptive Pseudo Labeling
Dinghua Xue, Kaibing Zhang |
PRCV (17) | 5 |
| 2025 | DMR-YOLO: Dual-Modality Robust YOLO for Small Object Detection in Infrared-RGB UAV Imagery
Kaibing Zhang |
PRCV (16) | 2 |
| 2025 | LiteSpiralGCN: Lightweight 3D hand mesh reconstruction via spiral graph convolution
Yiteng Wang, Minqi Li, Kaibing Zhang, Xiangjian He |
Appl. Intell. | 3 |
| 2025 | SmartPoints: Enhanced local feature extraction and neighborhood diffusion network for 3D point cloud semantic segmentation
Xiaogai Chen, Kaibing Zhang |
Comput. Graph. | 5 |
| 2025 | VITON-DRR: Details retention virtual try-on via non-rigid registration
Minqi Li, Kaibing Zhang |
Comput. Graph. | 4 |
| 2025 | SASFNet: Soft-edge awareness and spatial-attention feedback deep network for blind image deblurring
Kaibing Zhang, Jiahui Hou |
Comput. Vis. Image Underst. | 2 |
| 2025 | ADT: Person re-identification based on efficient attention mechanism and single-channel dual-channel fusion with transformer features aggregation
Jiahui Xing, Kaibing Zhang, Xiaogai Chen |
Expert Syst. Appl. | 3 |
| 2025 | Unsupervised Vast-Receptive-Field attention for blind Super-Resolution
Pengfei Yin, Weihua Ou, Kaibing Zhang |
Expert Syst. Appl. | 4 |
| 2025 | FMFI: Transformer based four branches multi-granularity feature integration for person Re-ID
Jiahui Xing, Kaibing Zhang |
Expert Syst. Appl. | 4 |
| 2025 | Multiscale Transformer Hierarchically Embedded CNN Hybrid Network for Visible-Infrared Person ReidentificationabstractVisible-infrared person reidentification (VI-ReID) is considered a pivotal technology for intelligent security surveillance systems for the Internet of Things (IoT). For the VI-ReID task, one key challenge is extracting and fusing robust global and local pedestrian information to mitigate the intermodality discrepancy. Despite the significant success achieved by convolutional neural network (CNN)-based methods, the extraction of global pedestrian information is limited by their inherent properties, namely, local receptive fields and downsampling processes, making cross-modality information fusion difficult. While existing pure Transformer-based methods excel at capturing global pedestrian information, uniform-sized queries, keys, and values are employed by their core self-attention mechanism. This results in the acquisition of uniform-scale information only, thereby limiting the learning of multiscale information and preventing the full extraction of local pedestrian information. To address the aforementioned issues, a multiscale Transformer hierarchically embedded CNN hybrid network (MTECN) is proposed by us. MTECN enables the simultaneous extraction of pedestrian local and global information at different scales to mitigate the adverse impact on recognition caused by the discrepancy in features extracted across different modalities. Moreover, the effects of inherent factors, including camera viewpoint and illumination variations, are alleviated by incorporating a spatial consistency (SC) loss, which guides the network in exploring and discriminating the spatial structures of pedestrians across different modalities, consequently aligning the underlying spatial semantic information. Furthermore, in the low-light VI-ReID task, information insufficiency is encountered by the LLCM dataset due to low-light conditions. Consequently, a low-light enhancement (LLE) module is employed to restore the obscured detail information in low-light images, thereby further enhancing MTECN’s robust feature learning in complex backgrounds. To the best of our knowledge, this is the first work to use Transformer hierarchically embedded CNN networks for VI-ReID research, and the first to use LLE techniques for low-light VI-ReID task. Extensive experiments on the SYSU-MM01, RegDB, and LLCM datasets show that the proposed MTECN method excels over several state-of-the-art methods. Suixin Liang, Kaibing Zhang, Xiaogai Chen |
IEEE Internet Things J. | 3 |
| 2025 | BeyondPoints: Curve Fusion and Attention-Driven Local Feature Learning for 3-D Semantic SegmentationabstractEfficient semantic segmentation of large-scale point cloud scenes is regarded as a fundamental and essential task for perceiving and understanding 3-D environments. It is also recognized as a key technology for environmental perception and intelligent decision-making in Internet of Things (IoT) applications. However, the diversity of objects and occlusion issues within scenes often hinder the ability of existing networks to effectively represent varying object shapes, leading to point information ambiguity and loss caused by pose variations. To address these challenges, a novel curve fusion and attention-driven local feature learning network (BeyondPoints) is proposed for point cloud segmentation. The proposed network consists of three key modules: 1) a local feature enhancement (LFE) module; 2) a dual-axis attention (DAA) module; and 3) a hybrid curve fusion (HCF) module. Specifically, the LFE module explicitly models spatial relationships and decouples local aggregation, effectively integrating additional geometric information into local features to compensate for point information loss and enhance environmental perception. To further alleviate local spatial perception ambiguity, the DAA module is designed to extract critical information from different spatial positions, thereby enhancing the representation of significant regions in point clouds and achieving more precise semantic segmentation. Finally, the HCF module serializes point cloud data to reduce model computational complexity while enabling cross-domain feature fusion, effectively integrating contextual information and suppressing noise interference. Extensive experiments on benchmark datasets, such as S3DIS, ScanNetV2, and SemanticKITTI, validate the exceptional segmentation performance of the proposed BeyondPoints network. Particularly, its superior performance in large-scale point cloud scenes underscores its potential to enhance environmental information perception in IoT scenarios. Liguo Luo, Kaibing Zhang, Haozhe Cheng, Xiaogai Chen |
IEEE Internet Things J. | 3 |
| 2025 | Learning scalable Omni-scale distribution for crowd counting
Huake Wang, Xingsong Hou, Kaibing Zhang, Minqi Li, Wenke Sun, Xueming Qian |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | RMANet: Refined-mixed attention network for progressive low-light image enhancement
Kaibing Zhang, Feifei Pang, Xinbo Gao 0001 |
Signal Process. | 2 |
| 2025 | Extracting Noise and Darkness: Low-Light Image Enhancement via Dual Prior GuidanceabstractThe complex entanglement between darkness and noise hinders the advance of low-light image enhancement. Most existing methods adopted lightening-then-denoising or embedded a special denoising module into enhancement network without specific noise knowledge as supervision to restore low-light images. However, they either fail to remove the amplified noise or blur the detail information. Against above drawbacks, we propose a novel dual prior guidance method for low-light image enhancement that relights darkness and suppresses noise simultaneously. Concretely, the main novelties of our proposed method are three-fold. Firstly, our formulation originates from a statistic observation that darkness can be disentangled into luminance channel, yet noise still exists each channel when low-light images are transformed from RGB space to YCbCr space. It inspires us to design an ingenious method, extracting noise and darkness, termed END, to enhance low-light images. Secondly, we propose a prior extraction network with prior composition module to extract luminance and noise priors from different channels. Thirdly, an image enhancement network deployed with prior guidance module is proposed to progressively lighten the darkness and remove noise. Extensive experiments on multiple benchmarks demonstrate that our proposed method achieves remarkable performance compared to other state-of-the-art low-light image enhancement methods. The source code and trained model can be found inhttps://github.com/WHK-Huake/END. Huake Wang, Xiaoyang Yan, Xingsong Hou, Kaibing Zhang, Yujie Dun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Semantic-Driven Global-Local Fusion Transformer for Image Super-ResolutionabstractImage Super-Resolution (SR) has seen remarkable progress with the emergence of transformer-based architectures. However, due to the high computational cost, many existing transformer-based SR methods limit their attention to local windows, which hinders their ability to model long-range dependencies and global structures. To address these challenges, we propose a novel SR framework named Semantic-Driven Global-Local Fusion Transformer (SGLFT). The proposed model enhances the receptive field by combining a Hybrid Window Transformer (HWT) and a Scalable Transformer Module (STM) to jointly capture local textures and global context. To further strengthen the semantic consistency of reconstruction, we introduce a Semantic Extraction Module (SEM) that distills high-level semantic priors from the input. These semantic cues are adaptively integrated with visual features through an Adaptive Feature Fusion Semantic Integration Module (AFFSIM). Extensive experiments on standard benchmarks demonstrate the effectiveness of SGLFT in producing visually faithful and structurally consistent SR results. The code will be available at https://github.com/kbzhang0505/SGLFT. Kaibing Zhang, Zhouwei Cheng, Xin He 0029, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Image Process. | 1 |
| 2025 | Multi-Scale Retinex Unfolding Network for Low-Light Image EnhancementabstractRetinex theory-based low-light image enhancement methods have received increasing attention and achieved tremendous advancements. However, there still exist two seldom-explored issues: 1) The above methods only formally simulate the Retinex decomposition, resulting in lacking explicit interpretability. 2) They usually are performed in single-scale space, leading to suboptimal enhancement results. In this paper, we propose an interpretable Multi-scale Retinex Unfolding Network (MRUNet) for low-light image enhancement, which can tackle both of the aforementioned issues simultaneously. Specifically, we formulate low-light image enhancement as a multi-scale Retinex optimization problem and design an iteration minimization solution to solve it. The optimization solution is further unfolded to fabricate MRUNet, which is empowered with clear physical significance and multi-scale prior knowledge in favor of image enhancement. However, it will aggravate model size and efficiency when exploiting multiple proximal mapping networks to extract multi-scale prior from multi-scale inputs. To surmount the issue, we propose a Scale-Aware Proximal mapping Module (SAPM), which efficiently collect multi-scale prior knowledge via the weight sharing strategy. In SAPM, we tailor a scale-aware transformer to model the specific scale-similarity among different scales. Extensive experiments manifest that MRUNet surpasses other Retinex-based low-light image enhancement methods on multiple benchmarks. Huake Wang, Xingsong Hou, Jutao Li, Yadi Yan, Wenke Sun, Kaibing Zhang, Xiangyong Cao |
IEEE Trans. Multim. | 7 |
| 2025 | CS-VITON: a realistic virtual try-on network based on clothing region alignment and SPM
Jinguang Chen, Kaibing Zhang |
Vis. Comput. | 5 |
| 2024 | Focusing on Significant Guidance: Preliminary Knowledge Guided Distillation
Qizhi Cao, Kaibing Zhang, Dinghua Xue, Zhouqiang Zhang |
PRCV (3) | 2 |
| 2024 | SCAMS: Semantic Category-Aware Multi-scale Network for Video Quality Assessment
Longgang Ren, Kaibing Zhang |
PRCV (10) | 2 |
| 2024 | LSGRNet: Local Spatial Latent Geometric Relation Learning Network for 3D point cloud semantic segmentation
Liguo Luo, Xiaogai Chen, Kaibing Zhang |
Comput. Graph. | 4 |
| 2024 | MFCT: Multi-Frequency Cascade Transformers for no-reference SR-IQA
Kaibing Zhang, Longgang Ren |
Comput. Vis. Image Underst. | 2 |
| 2024 | Discriminative transfer regression for low-rank and sparse subspace learning
Weihua Ou, Kaibing Zhang, Zhihui Lai 0001 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | Robust manifold discriminative distribution adaptation for transfer subspace learning
Weihua Ou, Kaibing Zhang |
Expert Syst. Appl. | 3 |
| 2024 | Coordinate Attention Guided Dual-Teacher Adaptive Knowledge Distillation for image classification
Dongtong Ma, Kaibing Zhang, Qizhi Cao, Jie Li 0001, Xinbo Gao 0001 |
Expert Syst. Appl. | 2 |
| 2024 | Low Communication-Cost PSI Protocol for Unbalanced Two-Party Private SetsabstractTwo‐party private set intersection (PSI) plays a pivotal role in secure two‐party computation protocols. The communication cost in a PSI protocol is normally influenced by the sizes of the participating parties. However, for parties with unbalanced sets, the communication costs of existing protocols mainly depend on the size of the larger set, leading to high communication cost. In this paper, we propose a low communication‐cost PSI protocol designed specifically for unbalanced two‐party private sets, aiming to enhance the efficiency of communication. For each item in the smaller set, the receiver queries whether it belongs to the larger set, such that the communication cost depends solely on the smaller set. The queries are implemented by private information retrieval which is constructed with trapdoor hash function. Our investigation indicates that in each instance of invoking the trapdoor hash function, the receiver is required to transmit both a hash key and an encoding key to the sender, thus incurring significant communication cost. In order to address this concern, we propose the utilization of a seed hash key, a seed encoding key, and a Latin square. By employing these components, the sender can autonomously generate all the necessary hash keys and encoding keys, obviating the multiple transmissions of such keys. The proposed protocol is provably secure against a semihonest adversary under the Decisional Diffie–Hellman assumption. Through implementation demonstration, we showcase that when the sizes of the two sets are 2 8 and 2 14 , the communication cost of our protocol is only 3.3% of the state‐of‐the‐art protocol and under 100 Kbps bandwidth, we achieve 1.46x speedup compared to the state‐of‐the‐art protocol. Our source code is available on GitHub: https://github.com/TAN-OpenLab/Unbanlanced-PSI . Jingyu Ning, Zhenhua Tan, Kaibing Zhang, Weizhong Ye |
IET Inf. Secur. | 3 |
| 2024 | Multi-temporal scale aggregation refinement graph convolutional network for skeleton-based action recognitionabstractAbstract Skeleton‐based human action recognition is gaining significant attention and finding widespread application in various fields, such as virtual reality and human‐computer interaction systems. Recent studies have highlighted the effectiveness of graph convolutional network (GCN) based methods in this task, leading to a remarkable improvement in prediction accuracy. However, most GCN‐based methods overlook the varying contributions of self, centripetal and centrifugal subsets. Besides, only a single‐scale temporal feature is adopted, and the multi‐temporal scale information is ignored. To this end, firstly, in order to differentiate the importance of different skeleton subsets, we develop a refinement graph convolution, which can adaptively learn a weight for each subset feature. Secondly, a multi‐temporal scale aggregation module is proposed to extract more discriminative temporal dynamic information. Furthermore, a multi‐temporal scale aggregation refinement graph convolutional network (MTSA‐RGCN) is proposed, and four‐stream structure is also adopted in this paper, which can comprehensively model complementary features and eventually achieves a significant performance boost. In the empirical experiments, the performance of our approach has been greatly improved on both NTU‐RGB+D 60 and NTU‐RGB+D 120 datasets, compared to other state‐of‐the‐art methods. Xuanfeng Li, Kaibing Zhang |
Comput. Animat. Virtual Worlds | 5 |
| 2024 | Division gets better: Learning brightness-aware and detail-sensitive representations for low-light image enhancement
Huake Wang, Xiaoyang Yan, Xingsong Hou, Yujie Dun, Kaibing Zhang |
Knowl. Based Syst. | 6 |
| 2024 | SS-CRE: A Continual Relation Extraction Method Through SimCSE-BERT and Static Relation PrototypesabstractAbstract Continual relation extraction aims to learn new relations from a continuous stream of data while avoiding forgetting old relations. Existing methods typically use the BERT encoder to obtain semantic embeddings, ignoring the fact that the vector representations suffer from anisotropy and uneven distribution. Furthermore, the relation prototypes are usually computed by memory samples directly, resulting in the model being overly sensitive to memory samples. To solve these problems, we propose a new continual relation extraction method. Firstly, we modified the basic structure of the sample encoder to generate uniformly distributed semantic embeddings using the supervised SimCSE-BERT to obtain richer sample information. Secondly, we introduced static relation prototypes and dynamically adjust their proportion with dynamic relation prototypes to adapt to the feature space. Lastly, through experimental analysis on the widely used FewRel and TACRED datasets, the results demonstrate that the proposed method effectively enhances semantic embeddings and relation prototypes, resulting in a further alleviation of catastrophic forgetting in the model. The code will be soon released at https://github.com/SuyueW/SS-CRE . Jinguang Chen, Suyue Wang, Kaibing Zhang |
Neural Process. Lett. | 5 |
| 2024 | CFGPFSR: A Generative Method Combining Facial and GAN Priors for Face Super-ResolutionabstractAbstract In recent years, facial prior has been widely applied to enhance the quality of super-resolution (SR) facial images in face super-resolution (FSR) methods based on deep learning. However, most of the existing facial prior-based FSR methods have insufficient attention to local texture details, which can cause the generated SR facial images with overly smooth and unrealistic texture details, and show obvious artifacts under large magnification. With the help of GAN prior, recent advances can produce excellent results in terms of fidelity and realness. A generative framework for FSR is proposed in this work, which combines GAN and facial prior, termed CFGPFSR. Firstly, we pre-train a face StyleGAN2 and a face parsing network (FPN) that can generate decent parsing maps, in which the proposed CFGPFSR exploits rich and varied priors encapsulated in the face StyleGAN2 (GAN prior) and face parsing maps extracted from the FPN (facial prior) for FSR. Moreover, we introduce the Channel-Split Spatial Feature Transform (CS-SFT) method to further improve FSR performance. GAN and facial priors are introduced into the FSR process through the designed CS-SFT layers so that SR facial images obtain a promising balance between fidelity and realness. Unlike GAN inversion methods which necessitate costly image optimization at runtime, the proposed CFGPFSR can jointly recover facial details by only utilizing one forward pass. Experimental results on synthetic and real images indicate that the proposed CFGPFSR obtains remarkable performance in 16 × SR task, and some of its metrics such as peak signal to noise ratio (PSNR) and structural similarity (SSIM) are higher than that of the comparison methods. Meanwhile, it shows impressive results in reconstructing high-quality facial images. Weihua Ou, Kaibing Zhang |
Neural Process. Lett. | 4 |
| 2024 | Coupled discriminative manifold alignment for low-resolution face recognition
Kaibing Zhang, Jie Li 0001, Xinbo Gao 0001 |
Pattern Recognit. | 1 |
| 2024 | Multi-scale non-local attention network for image super-resolution
Kaibing Zhang, Yanting Hu, Xin He 0029, Xinbo Gao 0001 |
Signal Process. | 2 |
| 2024 | High-Performance Feature Extraction Network for Point Cloud Semantic SegmentationabstractThe key to point cloud semantic segmentation lies in the efficient extraction of features from the point cloud data. However, previous research has often suffered from the ineffective capture of fine-grained spatial features of points or issues with ambiguous regional feature representation. To address this problem, We propose a method for point cloud surface construction to extract fine local geometric topology. We then embed the surface topology into each feature aggregation process to enrich feature representation, and propose a novel feature slice extraction method to capture significant local geometric features and contextual information. Furthermore, to enhance the performance of the Transformer network, we employ neighborhood grouping and double convolution operations at the initial network layer to aggregate the raw features of the point cloud. Numerous comparative experiments prove the effectiveness of the method in this letter, and we achieve state-of-the-art performance with mIoU of 74.7% on ScanNet V2 and 73.7% on S3DIS Area5. Youcheng Liang, Xiaogai Chen, Kaibing Zhang |
IEEE Signal Process. Lett. | 4 |
| 2024 | An Adaptive Region Proposal Network With Progressive Attention Propagation for Tiny Person Detection From UAV ImagesabstractTwo-stage detectors, which consist of the multi-scale feature representations and the prediction of region proposal boxes, have been recognized as an effective paradigm for tiny object detection in Unmanned Aerial Vehicle (UAV) images. Although most previous methods primarily concentrated on developing efficient feature fusion strategies within the feature pyramid network (FPN), few studies elaborated on improving the performance of region proposal network (RPN). Conventional RPNs exhibit two key weaknesses in the majority of existing two-stage object detection approaches. Firstly, the quality of proposal boxes generated by the RPN is heavily reliant on rich feature representations extracted from the FPN backbone. Secondly, the fixed number of generated proposal boxes limits adaptability to the distribution of tiny person objects. To mitigate the aforementioned problems, in this paper we propose a novel adaptive region proposal network (ARPN) to improve the quality of the proposal boxes and generate particularly compact yet accurate proposal boxes. On one hand, a progressive attention mechanism is devised to make the ARPN focus more on prospective object regions, where a series of multi-scale front attention modules (FAM) are applied to coarsely filter out most of irrelevant background areas and a group of top-to-bottom back attention modules (BAM) aid the ARPN to finely pinpoint tiny objects of interest in a coarse-to-fine manner. On the other hand, a mini-density map, which is inspired by the philosophy of crowd counting, is elaborately designed to adaptively determine the number of region proposal boxes. This approach significantly reduces redundancy while maintaining high-quality proposal boxes. Extensive experiments verify the superiority of proposed ARPN and show obvious improvement over other competitors in terms of two performance indicators of average precision (AP) and average recall (AR). The code will be available at https://github.com/kbzhang0505/ARPN. Youjiang Yu, Kaibing Zhang, Nannan Wang 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Hierarchical Kernel Interaction Network for Remote Sensing Object CountingabstractDifferent from object counting in surveillance scenes, remote sensing object counting encounters knotty challenges due to its tiny scale and cluttered background. However, existing remote sensing counting methods pursue favorable performance by sacrificing resolution to obtain semantic information, resulting in the loss of significant features of tiny-scale objects. To surmount the above issue, we propose a novel hierarchical kernel interaction network, dubbed HKINet, for remote sensing object counting. HKINet is comprised of several hierarchical kernel interaction modules (HKIMs) to simultaneously preserve high-resolution features and extract deep-layer semantic information. Specifically speaking, HKIM hierarchically performs multiresolution convolutions to avoid information loss in low resolution. Moreover, a scale interaction block (SIB) is used to combine multiresolution features for semantic information interaction. Finally, hierarchical resolutions are fused to output the prediction density map. To validate the superiority of our proposed HKINet, we conduct extensive experiments on four remote sensing object counting datasets, e.g., RSOC, CARPK, PUCPR+, and DroneCrowd datasets, and experimental results demonstrate HKINet outperforms other state-of-the-art remote sensing counting methods in terms of mean absolute error (MAE) and root mean squared error (RMSE). Huake Wang, Jinjiang Wei, Xingsong Hou, Hengfeng Wu, Kaibing Zhang |
IEEE Trans. Geosci. Remote. Sens. | 7 |
| 2024 | Semi-Supervised Domain Adaptation via Joint Transductive and Inductive Subspace LearningabstractMost existing shallow semi-supervised domain adaptation (SSDA) algorithms are based mainly on the framework adopting the maximum mean discrepancy (MMD) criterion, which is unstable and easily becomes stuck in a poor local minimum. Moreover, existing SSDA methods typically assume that the influence of the source domain is equivalent to that of the target domain, which is unreasonable and severely limits their performance. To address such drawbacks, we propose a novel SSDA framework derived from simple least squares regression (LSR) in a joint transductive and inductive learning paradigm, named transferable LSR (TLSR). Specifically, TLSR first learns domain-shared features using transfer component analysis (TCA) in a transductive paradigm. Then, TLSR augments the TCA features into the raw sample feature, formulating them into a block-diagonal matrix and training them in an inductive learning paradigm. This joint transductive and inductive learning paradigm helps alleviate the negative impacts of the MMD criterion of TCA but preserves the useful learned domain-shared knowledge. Moreover, the proposed block-diagonal input structure helps to separate the learned projections into independent domain-specific parts. Owing to the block-diagonal input structure, the influence of each domain can be reweighted, leading to significant improvements in performance. The experimental results demonstrate that the proposed TLSR outperforms the other shallow state-of-the-art competitors in 68 out of 90 cross-domain tasks. The source code of TLSR is available at:https://github.com/Evelhz/TLSR. Kaibing Zhang, Guofa Wang, Shaoyi Du |
IEEE Trans. Multim. | 3 |
| 2024 | Soft-edge-guided significant coordinate attention network for scene text image super-resolution
Chenchen Xi, Kaibing Zhang, Yanting Hu, Jinguang Chen |
Vis. Comput. | 2 |
| 2023 | Uncertainty awareness with adaptive propagation for multi-view stereo
Jinguang Chen, Zonghua Yu, Kaibing Zhang |
Appl. Intell. | 4 |
| 2023 | TADSRNet: A triple-attention dual-scale residual network for super-resolution image quality assessment
Xing Quan, Kaibing Zhang, Yanting Hu, Jinguang Chen |
Appl. Intell. | 2 |
| 2023 | Learning cascade regression for super-resolution image quality assessment
Xing Quan, Kaibing Zhang, Danni Zhu, Yanting Hu, Jinguang Chen |
Appl. Intell. | 2 |
| 2023 | Salient double reconstruction-based discriminative projective dictionary pair learning for crowd counting
Kaibing Zhang, Huake Wang, Minqi Li |
Appl. Intell. | 3 |
| 2023 | Corrigendum to "Unsupervised Simple Siamese Representation Learning for Blind Super-Resolution" [Eng. Appl. Artif. Intell. 114 (2022) 105092]
Pengfei Yin, Hua Huo, Kaibing Zhang |
Eng. Appl. Artif. Intell. | 6 |
| 2023 | Discriminative sparse least square regression for semi-supervised learning
Zhihui Lai 0001, Weihua Ou, Kaibing Zhang, Hua Huo |
Inf. Sci. | 4 |
| 2023 | Multi-scale information distillation network for efficient image super-resolution
Yanting Hu, Yuanfei Huang, Kaibing Zhang |
Knowl. Based Syst. | 3 |
| 2023 | Context Attention Fusion Network for crowd counting
Kaibing Zhang, Huake Wang, Minqi Li |
Knowl. Based Syst. | 3 |
| 2023 | Multi-distribution fitting for multi-view stereo
Jinguang Chen, Zonghua Yu, Kaibing Zhang |
Mach. Vis. Appl. | 4 |
| 2023 | Dynamic classifier approximation for unsupervised domain adaptation
Kaiming Shi, Danmei Niu, Hua Huo, Kaibing Zhang |
Signal Process. | 5 |
| 2023 | Be an Excellent Student: Review, Preview, and CorrectionabstractIn the letter, we propose a novel yet effective knowledge distillation scheme which mimics an all-round learning process of an excellent student from the teacher, i.e, knowledge review, knowledge preview, and knowledge correction, to acquire more informative and complementary knowledge. In the newly proposed method, to better leverage comprehensive feature knowledge from the teacher model, we propose Knowledge Review and Knowledge Preview Distillation to amalgamate multi-level features from different intermediate layers in both forward and backward pathways and fully distill them through hierarchical context loss, which greatly improves the student's feature learning efficiency. Moreover, we further present a Response Correction Mechanism to reinforce the prediction of student, which can more fully excavate the student's own knowledge, effectively alleviating the negative influence caused by the knowledge gap between the teacher and the student. We verify the effectiveness of our method with various networks on the CIFAR-100 datasets and the proposed method achieves competitive results compared with other state-of-the-art competitors. The code will be available athttps://github.com/kbzhang0505/RPC. Qizhi Cao, Kaibing Zhang, Xin He 0029, Junge Shen |
IEEE Signal Process. Lett. | 2 |
| 2023 | Multi-Branch and Progressive Network for Low-Light Image EnhancementabstractLow-light images incur several complicated degradation factors such as poor brightness, low contrast, color degradation, and noise. Most previous deep learning-based approaches, however, only learn the mapping relationship of single channel between the input low-light images and the expected normal-light images, which is insufficient enough to deal with low-light images captured under uncertain imaging environment. Moreover, too deeper network architecture is not conducive to recover low-light images due to extremely low values in pixels. To surmount aforementioned issues, in this paper we propose a novel multi-branch and progressive network (MBPNet) for low-light image enhancement. To be more specific, the proposed MBPNet is comprised of four different branches which build the mapping relationship at different scales. The followed fusion is performed on the outputs obtained from four different branches for the final enhanced image. Furthermore, to better handle the difficulty of delivering structural information of low-light images with low values in pixels, a progressive enhancement strategy is applied in the proposed method, where four convolutional long short-term memory networks (LSTM) are embedded in four branches and an recurrent network architecture is developed to iteratively perform the enhancement process. In addition, a joint loss function consisting of the pixel loss, the multi-scale perceptual loss, the adversarial loss, the gradient loss, and the color loss is framed to optimize the model parameters. To evaluate the effectiveness of proposed MBPNet, three popularly used benchmark databases are used for both quantitative and qualitative assessments. The experimental results confirm that the proposed MBPNet obviously outperforms other state-of-the-art approaches in terms of quantitative and qualitative results. The code will be available at https://github.com/kbzhang0505/MBPNet. Kaibing Zhang, Jie Li 0001, Xinbo Gao 0001, Minqi Li |
IEEE Trans. Image Process. | 1 |
| 2022 | Transfer Subspace Learning based on Double Relaxed Regression for Image Classification
Hua Huo, Chunlei Yang, Kaibing Zhang |
Appl. Intell. | 5 |
| 2022 | Learning graph-constrained cascade regressors for single image super-resolution
Jianqiang Yan, Kaibing Zhang, Zenggang Xiong |
Appl. Intell. | 2 |
| 2022 | Joint channel-spatial attention network for super-resolution image quality assessment
Tingyue Zhang, Kaibing Zhang, Zenggang Xiong |
Appl. Intell. | 2 |
| 2022 | Unsupervised simple Siamese representation learning for blind super-resolution
Pengfeng Yin, Di Wu 0058, Hua Huo, Kaibing Zhang |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | Robust sparse manifold discriminant analysis
Kaibing Zhang, Qingtao Wu, Mingchuan Zhang |
Multim. Tools Appl. | 3 |
| 2022 | Multi-level landmark-guided deep network for face super-resolution
Cheng Zhuang, Minqi Li, Kaibing Zhang |
Neural Networks | 3 |
| 2022 | C$^{2}$MT: A Credible and Class-Aware Multi-Task Transformer for SR-IQAabstractIn this letter a novel credible and class-aware multi-task transformer abbreviated as C$^{2}$MT for SRIQA, is proposed. In the proposed C$^{2}$MT, a quality-aware task for the quality prediction and the other class-aware task for the classification of SR algorithms are jointly framed to mine mutual information between the quality of SR images and the class of SR algorithms for more discriminative perceptual representation. In the class-aware task, we develop a supervised contrastive learning strategy to learn embedding perceived features related to a class-specific SR algorithm. While in the other quality-aware task, we employ a novel credible pseudo quality label generation strategy to actively adjust the quality labels by ranking the pair-wise consistency between the predicted quality scores and subjective perceptual scores but keep the image-level quality labels unchanged. The developed supervised contrastive learning and the variant of active learning strategies benefit learning a more consistent quality predictor for SR images. Experiment results indicate that our proposed C$^{2}$MT achieves state-of-the-art results on five popular SRIQA benchmark databases. Kaibing Zhang, Zhenxing Niu |
IEEE Signal Process. Lett. | 2 |
| 2022 | Locality-Adaptive Structured Dictionary Learning for Cross-Domain RecognitionabstractDictionary learning has achieved remarkable success on a wide range of machine learning-based applications. In this paper, a locality-adaptive structured dictionary learning (LASDL) algorithm for cross-domain recognition is proposed. In the LASDL, a projective structured double reconstruction strategy is developed to train class-oriented sub-dictionaries from the specific classes of cross-domain samples. The strategy benefits to make full advantage of the discriminative information of cross-domain data and bridge the distribution divergence between two different domains. Meanwhile, an adaptive geometrical structure preserving function is designed to not only preserve the local manifold structures spanned by the representation coefficient spaces of the source and target domains, but also impose a constraint that the coefficients should keep closer to their class centers, which is propitious to reduce the distribution divergence and make the representation more accurate. With the structured linear coding technique, the final cross-domain recognition can be efficiently performed by determining class-specific reconstruction error. The optimization of the proposed LASDL model can be efficiently solved by simple least square method and alternating direction method of multipliers (ADMM) algorithm. Extensive experimental results validate the superiority of the proposed algorithm in contrast to other state-of-the-art predecessors. Kaibing Zhang, Jie Li 0001, Xinbo Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Pseudo-label growth dictionary pair learning for crowd counting
Huake Wang, Kaibing Zhang, Zenggang Xiong |
Appl. Intell. | 4 |
| 2021 | PTANet: Triple Attention Network for point cloud semantic segmentation
Haozhe Cheng, Mao-Xin Luo, Kaibing Zhang |
Eng. Appl. Artif. Intell. | 5 |
| 2021 | Face hallucination based on cluster consistent dictionary learningabstractAbstract Face hallucination is a super‐resolution technique specially designed to reconstruct high‐resolution faces from low‐resolution faces. Most state‐of‐the‐art algorithms leverage position‐patch prior knowledge of human faces to better super‐resolve face images. However, most of them assume the training face dataset is sufficiently large, well cropped or aligned. This paper, proposes a novel example‐based face hallucination method, based on cluster consistent dictionary learning with the assumption that human faces have similar facial structures. In this method, the paired face image patches are firstly labelled as face areas including eyes, nose, mouth and other parts, as well as non‐face areas without requiring the training face images cropped and aligned. Then, the training patches are clustered according their labels and textures. The cluster consistent dictionary is learned to represent the low‐resolution patches and the high‐resolution patches. Finally, the high‐resolution patches of the input low‐resolution face image can be efficiently generated by using the adjusted anchored neighbourhood regression. As utilizing the labelled facial parts prior knowledge, the proposed method represents more details in the reconstruction. Experimental results demonstrate that the authors' algorithm outperforms many state‐of‐the‐art techniques for face hallucination under different datasets. Minqi Li, Xiangjian He, Kin-Man Lam 0001, Kaibing Zhang, Junfeng Jing |
IET Image Process. | 4 |
| 2021 | Learning stacking regression for no-reference super-resolution image quality assessment
Kaibing Zhang, Danni Zhu, Jie Li 0001, Xinbo Gao 0001, Fei Gao 0006 |
Signal Process. | 1 |
| 2020 | Learning stacking regressors for single image super-resolution
Kaibing Zhang, Minqi Li, Junfeng Jing, Zenggang Xiong |
Appl. Intell. | 1 |
| 2020 | Discriminative sparse embedding based on adaptive graph for dimension reduction
Kaiming Shi, Kaibing Zhang, Weihua Ou, Lin Wang 0039 |
Eng. Appl. Artif. Intell. | 3 |
| 2020 | Fabric defect detection using saliency of multi-scale local steering kernelabstractFabric defect detection (FDD) plays an important role in the quality control in textile industry. In this study, the authors propose an efficient FDD method by using the saliency analysis of multi‐scale local steering kernel (LSK). In the proposed method, a given RGB fabric image is first converted into the Commission International Eclairage (CIE) L*a*b colour space and then the LSK in each colour channel is computed by the singular value decomposition and the centre surrounding definition. Next, the matrix cosine similarity is employed to measure the similarity between different LSK features for generating the desired defective maps. Finally, a multi‐scale averaging fusion scheme is applied to integrate the obtained defective maps at different scales for the final defective map. The experimental results indicate that the proposed method achieves the state‐of‐the‐art performance on FDD compared to the other competitors. Kaibing Zhang, Yadi Yan, Junfeng Jing, Zenggang Xiong |
IET Image Process. | 1 |
| 2020 | Fast non-rigid points registration with cluster correspondences projection
Minqi Li, Jing Xin, Kaibing Zhang, Junfeng Jing |
Signal Process. | 4 |
| 2020 | Structured optimal graph based sparse feature extraction for semi-supervised learning
Zhihui Lai 0001, Weihua Ou, Kaibing Zhang, Ruijuan Zheng |
Signal Process. | 4 |
| 2019 | Learning a Cascade Regression for No-Reference Super-Resolution Image Quality AssessmentabstractNo-reference super-resolution image quality assessment (NRSRIQA) technique has been recognized an effective way to evaluate the quality of SR images and the performance of SR algorithms. In this paper, we propose a novel NR-SRIQA method by learning a two-layer regression model to establish the mapping relationship between the multiple natural statistical features and visual perceptual scores. First, we exploit three types of statistical features to quantify the degradation of SR images. Next, a cascade two-layer regression model, which integrates AdaBoost Decision Tree Regression and ridge regression, is trained to predict the quality of SR images in a coarse-to-fine manner. The experimental results demonstrate that the proposed method is superior to other previous SR quality evaluation approaches and shows better consistency with visual perception quality. Kaibing Zhang, Danni Zhu, Junfeng Jing, Xinbo Gao 0001 |
ICIP | 1 |
| 2019 | Learning recurrent residual regressors for single image super-resolution
Kaibing Zhang, Zhen Wang 0037, Jie Li 0001, Xinbo Gao 0001, Zenggang Xiong |
Signal Process. | 1 |
| 2018 | Learning local dictionaries and similarity structures for single image super-resolution
Kaibing Zhang, Jie Li 0001, Xiuping Liu, Xinbo Gao 0001 |
Signal Process. | 1 |
| 2017 | Fast single image super-resolution using sparse Gaussian process regression
Xinbo Gao 0001, Kaibing Zhang, Jie Li 0001 |
Signal Process. | 3 |
| 2017 | Single Image Super-Resolution Using Gaussian Process Regression With Dictionary-Based Sampling and Student-t LikelihoodabstractGaussian process regression (GPR) is an effective statistical learning method for modeling non-linear mapping from an observed space to an expected latent space. When applying it to example learning-based super-resolution (SR), two outstanding issues remain. One is its high computational complexity restricts SR application when a large data set is available for learning task. The other is that the commonly used Gaussian likelihood in GPR is incompatible with the true observation model for SR reconstruction. To alleviate the above two issues, we propose a GPR-based SR method by using dictionary-based sampling (DbS) and student-t likelihood. Considering that dictionary atoms effectively span the original training sample space, we adopt a DbS strategy by combining all the neighborhood samples of each atom into a compact representative training subset so as to reduce the computational complexity. Based on statistical tests, we statistically validate that student-t likelihood is more suitable to build the observation model for the SR problem. Extensive experimental results show that the proposed method outperforms other competitors and produces more pleasing details in texture regions. Xinbo Gao 0001, Kaibing Zhang, Jie Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2017 | Coarse-to-Fine Learning for Single-Image Super-ResolutionabstractThis paper develops a coarse-to-fine framework for single-image super-resolution (SR) reconstruction. The coarse-to-fine approach achieves high-quality SR recovery based on the complementary properties of both example learning-and reconstruction-based algorithms: example learning-based SR approaches are useful for generating plausible details from external exemplars but poor at suppressing aliasing artifacts, while reconstruction-based SR methods are propitious for preserving sharp edges yet fail to generate fine details. In the coarse stage of the method, we use a set of simple yet effective mapping functions, learned via correlative neighbor regression of grouped low-resolution (LR) to high-resolution (HR) dictionary atoms, to synthesize an initial SR estimate with particularly low computational cost. In the fine stage, we devise an effective regularization term that seamlessly integrates the properties of local structural regularity, nonlocal self-similarity, and collaborative representation over relevant atoms in a learned HR dictionary, to further improve the visual quality of the initial SR estimation obtained in the coarse stage. The experimental results indicate that our method outperforms other state-of-the-art methods for producing high-quality images despite that both the initial SR estimation and the followed enhancement are cheap to implement. Kaibing Zhang, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001, Jie Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2016 | Image super-resolution using non-local Gaussian process regression
Xinbo Gao 0001, Kaibing Zhang, Jie Li 0001 |
Neurocomputing | 3 |
| 2016 | Single image super-resolution using regularization of non-local steering kernel regression
Kaibing Zhang, Xinbo Gao 0001, Jie Li 0001, Hongxing Xia |
Signal Process. | 1 |
| 2016 | Single-Image Super-Resolution Using Active-Sampling Gaussian Process RegressionabstractAs well known, Gaussian process regression (GPR) has been successfully applied to example learning-based image super-resolution (SR). Despite its effectiveness, the applicability of a GPR model is limited by its remarkably computational cost when a large number of examples are available to a learning task. For this purpose, we alleviate this problem of the GPR-based SR and propose a novel example learning-based SR method, called active-sampling GPR (AGPR). The newly proposed approach employs an active learning strategy to heuristically select more informative samples for training the regression parameters of the GPR model, which shows significant improvement on computational efficiency while keeping higher quality of reconstructed image. Finally, we suggest an accelerating scheme to further reduce the time complexity of the proposed AGPR-based SR by using a pre-learned projection matrix. We objectively and subjectively demonstrate that the proposed method is superior to other competitors for producing much sharper edges and finer details. Xinbo Gao 0001, Kaibing Zhang, Jie Li 0001 |
IEEE Trans. Image Process. | 3 |
| 2016 | Similarity Constraints-Based Structured Output Regression Machine: An Approach to Image Super-ResolutionabstractFor regression-based single-image super-resolution (SR) problem, the key is to establish a mapping relation between high-resolution (HR) and low-resolution (LR) image patches for obtaining a visually pleasing quality image. Most existing approaches typically solve it by dividing the model into several single-output regression problems, which obviously ignores the circumstance that a pixel within an HR patch affects other spatially adjacent pixels during the training process, and thus tends to generate serious ringing artifacts in resultant HR image as well as increase computational burden. To alleviate these problems, we propose to use structured output regression machine (SORM) to simultaneously model the inherent spatial relations between the HR and LR patches, which is propitious to preserve sharp edges. In addition, to further improve the quality of reconstructed HR images, a nonlocal (NL) self-similarity prior in natural images is introduced to formulate as a regularization term to further enhance the SORM-based SR results. To offer a computation-effective SORM method, we use a relative small nonsupport vector samples to establish the accurate regression model and an accelerating algorithm for NL self-similarity calculation. Extensive SR experiments on various images indicate that the proposed method can achieve more promising performance than the other state-of-the-art SR methods in terms of both visual quality and computational cost. Cheng Deng 0002, Jie Xu 0012, Kaibing Zhang, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2015 | Learning Multiple Linear Mappings for Efficient Single Image Super-ResolutionabstractExample learning-based superresolution (SR) algorithms show promise for restoring a high-resolution (HR) image from a single low-resolution (LR) input. The most popular approaches, however, are either time- or space-intensive, which limits their practical applications in many resource-limited settings. In this paper, we propose a novel computationally efficient single image SR method that learns multiple linear mappings (MLM) to directly transform LR feature subspaces into HR subspaces. In particular, we first partition the large nonlinear feature space of LR images into a cluster of linear subspaces. Multiple LR subdictionaries are then learned, followed by inferring the corresponding HR subdictionaries based on the assumption that the LR-HR features share the same representation coefficients. We establish MLM from the input LR features to the desired HR outputs in order to achieve fast yet stable SR recovery. Furthermore, in order to suppress displeasing artifacts generated by the MLM-based method, we apply a fast nonlocal means algorithm to construct a simple yet effective similarity-based regularization term for SR enhancement. Experimental results indicate that our approach is both quantitatively and qualitatively superior to other application-oriented SR methods, while maintaining relatively low time and space complexity. Kaibing Zhang, Dacheng Tao, Xinbo Gao 0001, Xuelong Li 0001, Zenggang Xiong |
IEEE Trans. Image Process. | 1 |
| 2014 | Secure Multimedia Big Data Sharing in Social Networks Using Fingerprinting and Encryption in the JPEG2000 Compressed DomainabstractWith the advent of social networks and cloud computing, the amount of multimedia data produced and communicated within social networks is rapidly increasing. In the mean time, social networking platform based on cloud computing has made multimedia big data sharing in social network easier and more efficient. The growth of social multimedia, as demonstrated by social networking sites such as Facebook and YouTube, combined with advances in multimedia content analysis, underscores potential risks for malicious use such as illegal copying, piracy, plagiarism, and misappropriation. Therefore, secure multimedia sharing and traitor tracing issues have become critical and urgent in social network. In this paper, we propose a scheme for implementing the Tree-Structured Harr (TSH) transform in a homomorphic encrypted domain for fingerprinting using social network analysis with the purpose of protecting media distribution in social networks. The motivation is to map hierarchical community structure of social network into tree structure of TSH transform for JPEG2000 coding, encryption and fingerprinting. Firstly, the fingerprint code is produced using social network analysis. Secondly, the encrypted content is decomposed by the TSH transform. Thirdly, the content is fingerprinted in the TSH transform domain. At last, the encrypted and fingerprinted contents are delivered to users via hybrid multicast-unicast. The use of fingerprinting along with encryption can provide a double-layer of protection to media sharing in social networks. Theory analysis and experimental results show the effectiveness of the proposed scheme. Conghuan Ye, Zenggang Xiong, Yaoming Ding, Jiping Li, Guangwei Wang, Kaibing Zhang |
TrustCom | 7 |
| 2014 | A Unified Learning Framework for Single Image Super-ResolutionabstractIt has been widely acknowledged that learning- and reconstruction-based super-resolution (SR) methods are effective to generate a high-resolution (HR) image from a single low-resolution (LR) input. However, learning-based methods are prone to introduce unexpected details into resultant HR images. Although reconstruction-based methods do not generate obvious artifacts, they tend to blur fine details and end up with unnatural results. In this paper, we propose a new SR framework that seamlessly integrates learning- and reconstruction-based methods for single image SR to: 1) avoid unexpected artifacts introduced by learning-based SR and 2) restore the missing high-frequency details smoothed by reconstruction-based SR. This integrated framework learns a single dictionary from the LR input instead of from external images to hallucinate details, embeds nonlocal means filter in the reconstruction-based SR to enhance edges and suppress artifacts, and gradually magnifies the LR input to the desired high-quality SR result. We demonstrate both visually and quantitatively that the proposed framework produces better results than previous methods from the literature. Jifei Yu, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001, Kaibing Zhang |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2013 | Image super-resolution via non-local steering kernel regression regularizationabstractIn this paper, we employ the non-local steering kernel regression to construct an effective regularization term for the single image super-resolution problem. The proposed method seamlessly integrates the properties of local structural regularity and non-local self-similarity existing in natural images, and solves a least squares minimization problem for obtaining the desired high-resolution image. Extensive experimental results on both simulated and real low-resolution images demonstrate that the proposed method can restore compelling results with sharp edges and fine textures. Kaibing Zhang, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001 |
ICIP | 1 |
| 2013 | Single Image Super-Resolution With Multiscale Similarity LearningabstractExample learning-based image super-resolution (SR) is recognized as an effective way to produce a high-resolution (HR) image with the help of an external training set. The effectiveness of learning-based SR methods, however, depends highly upon the consistency between the supporting training set and low-resolution (LR) images to be handled. To reduce the adverse effect brought by incompatible high-frequency details in the training set, we propose a single image SR approach by learning multiscale self-similarities from an LR image itself. The proposed SR approach is based upon an observation that small patches in natural images tend to redundantly repeat themselves many times both within the same scale and across different scales. To synthesize the missing details, we establish the HR-LR patch pairs using the initial LR input and its down-sampled version to capture the similarities across different scales and utilize the neighbor embedding algorithm to estimate the relationship between the LR and HR image pairs. To fully exploit the similarities across various scales inside the input LR image, we accumulate the previous resultant images as training examples for the subsequent reconstruction processes and adopt a gradual magnification scheme to upscale the LR input to the desired size step by step. In addition, to preserve sharper edges and suppress aliasing artifacts, we further apply the nonlocal means method to learn the similarity within the same scale and formulate a nonlocal prior regularization term to well pose SR estimation under a reconstruction-based SR framework. Experimental results demonstrate that the proposed method can produce compelling SR recovery both quantitatively and perceptually in comparison with other state-of-the-art baselines. Kaibing Zhang, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2012 | Multi-scale dictionary for single image super-resolutionabstractReconstruction- and example-based super-resolution (SR) methods are promising for restoring a high-resolution (HR) image from low-resolution (LR) image(s). Under large magnification, reconstruction-based methods usually fail to hallucinate visual details while example-based methods sometimes introduce unexpected details. Given a generic LR image, to reconstruct a photo-realistic SR image and to suppress artifacts in the reconstructed SR image, we introduce a multi-scale dictionary to a novel SR method that simultaneously integrates local and non-local priors. The local prior suppresses artifacts by using steering kernel regression to predict the target pixel from a small local area. The non-local prior enriches visual details by taking a weighted average of a large neighborhood as an estimate of the target pixel. Essentially, these two priors are complementary to each other. Experimental results demonstrate that the proposed method can produce high quality SR recovery both quantitatively and perceptually. Kaibing Zhang, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001 |
CVPR | 1 |
| 2012 | A Novel JFE Scheme for Social Multimedia Distribution in Compressed Domain Using SVD and CA
Conghuan Ye, Fuhao Zou, Zhengding Lu, Zenggang Xiong, Kaibing Zhang |
IWDW | 6 |
| 2012 | Video super-resolution with 3D adaptive normalized convolution
Kaibing Zhang, Guangwu Mu, Yuan Yuan 0001, Xinbo Gao 0001, Dacheng Tao |
Neurocomputing | 1 |
| 2012 | Joint Learning for Single-Image Super-Resolution via a Coupled ConstraintabstractThe neighbor-embedding (NE) algorithm for single-image super-resolution (SR) reconstruction assumes that the feature spaces of low-resolution (LR) and high-resolution (HR) patches are locally isometric. However, this is not true for SR because of one-to-many mappings between LR and HR patches. To overcome or at least to reduce the problem for NE-based SR reconstruction, we apply a joint learning technique to train two projection matrices simultaneously and to map the original LR and HR feature spaces onto a unified feature subspace. Subsequently, the k -nearest neighbor selection of the input LR image patches is conducted in the unified feature subspace to estimate the reconstruction weights. To handle a large number of samples, joint learning locally exploits a coupled constraint by linking the LR-HR counterparts together with the K-nearest grouping patch pairs. In order to refine further the initial SR estimate, we impose a global reconstruction constraint on the SR outcome based on the maximum a posteriori framework. Preliminary experiments suggest that the proposed algorithm outperforms NE-related baselines. Xinbo Gao 0001, Kaibing Zhang, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Image Super-Resolution With Sparse Neighbor EmbeddingabstractUntil now, neighbor-embedding-based (NE) algorithms for super-resolution (SR) have carried out two independent processes to synthesize high-resolution (HR) image patches. In the first process, neighbor search is performed using the Euclidean distance metric, and in the second process, the optimal weights are determined by solving a constrained least squares problem. However, the separate processes are not optimal. In this paper, we propose a sparse neighbor selection scheme for SR reconstruction. We first predetermine a larger number of neighbors as potential candidates and develop an extended Robust-SL0 algorithm to simultaneously find the neighbors and to solve the reconstruction weights. Recognizing that the k-nearest neighbor (k-NN) for reconstruction should have similar local geometric structures based on clustering, we employ a local statistical feature, namely histograms of oriented gradients (HoG) of low-resolution (LR) image patches, to perform such clustering. By conveying local structural information of HoG in the synthesis stage, the k-NN of each LR input patch is adaptively chosen from their associated subset, which significantly improves the speed of synthesizing the HR image while preserving the quality of reconstruction. Experimental results suggest that the proposed method can achieve competitive SR quality compared with other state-of-the-art baselines. Xinbo Gao 0001, Kaibing Zhang, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 2 |
| 2012 | Single Image Super-Resolution With Non-Local Means and Steering Kernel RegressionabstractImage super-resolution (SR) reconstruction is essentially an ill-posed problem, so it is important to design an effective prior. For this purpose, we propose a novel image SR method by learning both non-local and local regularization priors from a given low-resolution image. The non-local prior takes advantage of the redundancy of similar patches in natural images, while the local prior assumes that a target pixel can be estimated by a weighted average of its neighbors. Based on the above considerations, we utilize the non-local means filter to learn a non-local prior and the steering kernel regression to learn a local prior. By assembling the two complementary regularization terms, we propose a maximum a posteriori probability framework for SR recovery. Thorough experimental results suggest that the proposed SR method can reconstruct higher quality results both quantitatively and perceptually. Kaibing Zhang, Xinbo Gao 0001, Dacheng Tao, Xuelong Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2011 | Single image super resolution with high resolution dictionaryabstractImage super resolution (SR) is a technique to estimate or synthesize a high resolution (HR) image from one or several low resolution (LR) images. This paper proposes a novel framework for single image super resolution based on sparse representation with high resolution dictionary. Unlike the previous methods, the training set is constructed from the HR images instead of HR-LR image pairs. Due to this property, there is no need to retrain a new dictionary when the zooming factor changed. Given a testing LR image, the patch-based representation coefficients and the desired image are estimated alternately through the use of dynamic group sparsity, the fidelity term and the non-local means regularization. Experimental results demonstrate the effectiveness of the proposed algorithm. Guangwu Mu, Xinbo Gao 0001, Kaibing Zhang, Xuelong Li 0001, Dacheng Tao |
ICIP | 3 |
| 2011 | Zernike-Moment-Based Image Super ResolutionabstractMultiframe super-resolution (SR) reconstruction aims to produce a high-resolution (HR) image using a set of low-resolution (LR) images. In the process of reconstruction, fuzzy registration usually plays a critical role. It mainly focuses on the correlation between pixels of the candidate and the reference images to reconstruct each pixel by averaging all its neighboring pixels. Therefore, the fuzzy-registration-based SR performs well and has been widely applied in practice. However, if some objects appear or disappear among LR images or different angle rotations exist among them, the correlation between corresponding pixels becomes weak. Thus, it will be difficult to use LR images effectively in the process of SR reconstruction. Moreover, if the LR images are noised, the reconstruction quality will be affected seriously. To address or at least reduce these problems, this paper presents a novel SR method based on the Zernike moment, to make the most of possible details in each LR image for high-quality SR reconstruction. Experimental results show that the proposed method outperforms existing methods in terms of robustness and visual effects. Xinbo Gao 0001, Xuelong Li 0001, Dacheng Tao, Kaibing Zhang |
IEEE Trans. Image Process. | 5 |