EDBT 2026 Demo / reviewers in the wild / expert
Baojiang Zhong
dblp:30/6793
· DBLP profile ↗
61ranked-venue papers
7as first author
43since 2021 · last 2026
0000-0002-9899-524XORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 44 · 5 first-author · 31 since 2021Artificial intelligence and machine learning · 18 · 2 first-author · 13 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DEALSD: A deep edge assisted line segment detector
Zhongyi Sha, Baojiang Zhong, Xueyuan Chen, Zikai Wang 0007 |
Expert Syst. Appl. | 2 |
| 2026 | Deep bilateral learning for image interpolation
Jiahuan Ji, Kai-Kuang Ma, Baojiang Zhong, Fuhui Zhou, Qihui Wu 0001 |
Knowl. Based Syst. | 3 |
| 2026 | ARPose: Anatomical relation-driven token pruning for human pose estimation
Xiaodi Sun, Baojiang Zhong, Minghao Piao, Kai-Kuang Ma |
Multim. Syst. | 2 |
| 2025 | Perception-Enhanced Network for Accurate Human Pose EstimationabstractHuman pose estimation in computer vision is particularly challenging with images containing multiple individuals. Existing methods often integrate spatial and channel attention by simply adding them up through a cascade or parallel connection. The features extracted in this way could lead to less accurate key point predictions, especially in cases where the limbs of different people are obstructed or tangled with each other. To tackle this crucial issue, we develop a novel network that enhances key point detection by combining the spatial and channel attention in a more effective manner. Specifically, our network features a lightweight perception-enhanced module (PEM) that adaptively fuses spatial and channel features through a Hadamard product, thereby refining the overall feature representation. Moreover, by exploiting the initial feature map as a guide to generate the pixel attention, we further boost key point prediction accuracy. Extensive experimental results show that our developed network can clearly outperform the current state-of-the-art methods. Xiaodi Sun, Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 2 |
| 2025 | Deformable Attention-Based Edge-Aware Network for Single Image Super-ResolutionabstractAccurately reconstructing object edges is a key challenge in single image super-resolution (SISR), as it greatly influences our visual perception of image quality. To address this fundamental issue, we propose a novel SISR approach named the deformable attention-based edge-aware (DAE) network. The DAE network features a deformable attention block that dynamically adjusts attention weights to align with edge structures, thereby improving edge awareness and reconstruction quality. Moreover, our network incorporates a multi-patterns window block that captures fine-grained edge details and enhances information flow across network layers. This combination results in visually superior SISR outputs. Extensive experiments and comparisons with the state-of-the-art methods demonstrate that our DAE network excels in both quantitative metrics and visual quality. Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 2 |
| 2025 | Lightweight Single Image Super-Resolution With High-Continuity AttentionabstractWindow attention has become a popular choice in single image super-resolution (SISR) network design due to its efficient computation. However, its self-attention is restricted to fixed-size windows, leading to a lack of cross-window interaction. To address this, the benchmark SwinIR model adopts a shifted window strategy to capture long-range dependencies. However, we observe that its attention still suffers from discontinuities at window boundaries, resulting in inferior SISR performance. To address this issue, we propose a newscale-dual attention(SDA) module, consisting of three parallel branches that integratewindowattention andpoolingattention via three complementary scales. This enables hierarchical local-global interactions, yielding high-continuity attention maps. To validate the effectiveness of our proposed SDA, we develop a lightweightscale-dual attention network(SDAN) with approximately 878K parameters for SISR. Extensive experiments demonstrate that our SDAN achieves superior performance, outperforming state-of-the-art methods in both accuracy and efficiency. Baojiang Zhong, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2025 | BAN: A Boundary-Aware Network for Accurate Colorectal Polyp SegmentationabstractColonoscopy images exhibit multi-frequency features, with polyp boundaries residing in a mid-frequency range, which are critical for accurate polyp segmentation. However, current deep learning models tend to prioritize low-frequency features, leading to reduced segmentation performance. To address this challenge, we propose a novelboundary-aware network(BAN) that integrates trainable Gabor filters into the polyp segmentation process through a dedicated module calledGabor-driven feature extraction(GFE). By developing and using atrajectory-directed frequency learningapproach, Gabor filters are trained along adamping sinusoidalpath, dynamically optimizing their frequency parameters within a proper mid-frequency range. This enhances boundary feature representation and significantly improves polyp segmentation accuracy. Extensive experiments demonstrate that our BAN outperforms existing state-of-the-art methods. Nengxiang Zhang, Baojiang Zhong, Minghao Piao, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2025 | Cognitive Contour Detection of Sparse-Structured Objects in the Alpha-Shape Scale SpaceabstractIn this paper, we introduce cognitive contour, a novel image attribute that encapsulates the global shape perceived from sparsely distributed, identical or similar objects-such as drone swarms or flocks of geese-collectively termed sparse-structured objects. Unlike traditional contour analysis that delineates the boundaries of individual objects, cognitive contours reflect a gestalt-inspired perception of the overall structure formed by the ensemble, capturing higher-level visual organization. Detecting cognitive contours is challenging due to the sparsity and multiplicity of constituent elements. To tackle this, we propose a scale-space method that integrates alpha shapes into a scale-space framework. An alpha-shape scale space is constructed for the sparse-structured object, and the optimal scale is adaptively selected to extract cognitively meaningful contours with appropriate structural detail. Extensive experiments validate the effectiveness and robustness of the proposed method, enhancing visual inference and offering flexibility across diverse image-based applications. Code and data are available at: https://github.com/CookiC/Sparse. Yuxiang Shen, Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2025 | Crucial-Diff: A Unified Diffusion Model for Crucial Image and Annotation Synthesis in Data-Scarce ScenariosabstractThe scarcity of data in various scenarios, such as medical, industry and autonomous driving, leads to model overfitting and dataset imbalance, thus hindering effective detection and segmentation performance. Existing studies employ the generative models to synthesize more training samples to mitigate data scarcity. However, these synthetic samples are repetitive or simplistic and fail to provide "crucial information" that targets the downstream model's weaknesses. Additionally, these methods typically require separate training for different objects, leading to computational inefficiencies. To address these issues, we propose Crucial-Diff, a domain-agnostic framework designed to synthesize crucial samples. Our method integrates two key modules. The Scene Agnostic Feature Extractor (SAFE) utilizes a unified feature extractor to capture target information. The Weakness Aware Sample Miner (WASM) generates hard-to-detect samples using feedback from the detection results of downstream model, which is then fused with the output of SAFE module. Together, our Crucial-Diff framework generates diverse, high-quality training data, achieving a pixel-level AP of 83.63% and an F1-MAX of 78.12% on MVTec. On polyp dataset, Crucial-Diff reaches an mIoU of 81.64% and an mDice of 87.69%. Code is publicly available at https://github.com/JJessicaYao/Crucial-diff. Siyue Yao, Mingjie Sun, Eng Gee Lim, Ran Yi 0002, Baojiang Zhong, Moncef Gabbouj |
IEEE Trans. Image Process. | 5 |
| 2025 | APSNR: Artifact Peak Signal-to-Noise Ratio for Image Quality AssessmentabstractAn image processing pipeline typically involves key operations like compression, denoising, and resizing, along with enhancements such as sharpening, histogram equalization, and low-light compensation. Within this pipeline, image artifacts are often introduced, which could severely degrade perceptual quality and mislead downstream vision tasks. Yet, current image quality assessment (IQA) models fail to distinguish between harmful artifacts and beneficial enhancements, as they generally apply a rigid fidelity criterion that penalizes all deviations from the reference image. We address this gap with the artifact peak signal-to-noise ratio (APSNR), a new IQA metric that adopts a selective fidelity criterion-allowing legitimate enhancements while penalizing only spurious artifacts. Specifically, APSNR detects artifacts by identifying pixels that violate an "artifact-free" intensity mapping between the processed and reference images, and then computes PSNR exclusively within the artifact-corrupted regions. Extensive experiments demonstrate that our APSNR consistently correlates with human perception of artifacts while remaining robust to enhancements. This enables a more nuanced evaluation of image processing algorithms and provides a principled tool for benchmarking artifact suppression. Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2024 | Prediction-Correction Line Segment DetectionabstractEdge-drawing methods have gained increasing popularity in line segment detection due to their notable efficiency. However, existing algorithms commonly impose a pre-determined threshold on the gradient magnitude of the input image to control false positives, which could lead to the detection of line segments with insufficient completeness. To address this fundamental problem, we propose a novel method called the prediction-correction line segment detector (PCLSD). The PCLSD initiates with a prediction stage utilizing a Canny-based approach to generate line segment predictions. In the subsequent correction stage, each predicted line segment undergoes refinement. Specifically, a directional routing method is employed to extend and refit the line segment, improving the accuracy of its orientation, position, and completeness. The corrected line segment is then validated to ensure confidence. Experimental results demonstrate the superior performance of the proposed PCLSD compared to current state-of-the-art methods. Zhongyi Sha, Baojiang Zhong |
ICASSP | 2 |
| 2024 | Ellipse Detection Based On Structure-Preserving Anisotropic Edge ExtractionabstractExisting methods for ellipse detection popularly adopt the edge-linking strategy—i.e., first combining the elliptical arcs extracted from an edge map into groups and then fitting each group of arcs to an ellipse. However, such methods generally use the Canny operator to extract edges, which tends to corrupt salient elliptical shapes in texture regions, thus preventing the effective detection of ellipses. To overcome the difficulty, we propose a novel ellipse detector based on our developed structure-preserving anisotropic edge extraction (SPAEE) approach, which can remove redundant textures and preserve continuous structural edges, thus improving the performance of ellipse detection. In addition, an adaptive validation strategy is proposed to further enhance the detection quality. To our best knowledge, this is the first attempt to detect ellipses with an anisotropic edge extraction process for preserving structural edges. Experimental results have shown that the mean F-score on five benchmark datasets has increased from 0.61 to 0.67 (about 10% performance gain). Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 2 |
| 2024 | Corner Detection Based on a Rotation-Invariant and Noise-Insensitive Curvature MeasurementabstractCorner detection is extensively applied across various computer vision tasks. Current corner detectors typically assume that the distance between every two nearby pixels is constant. However, this assumption is invalid in real-world scenarios. As a result, the pixel-based curvature measurements designed and used in these corner detectors may suffer instability under rotation transformation and noise interference. To tackle this issue, a novel curvature measurement is proposed in this paper, which exploits the length of the subpixelized chord to estimate the discrete curvatures of digital curves. The proposed curvature model is invariant to rotation transformation and insensitive to image noise interference. Based on this curvature measurement, a new corner detector is further developed. Experimental results demonstrate that our proposed corner detector outperforms existing state-of-the-art methods. Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 2 |
| 2024 | Ellipse Detection Based on Contrast-Guided Arc EnhancementabstractA common practice of existing ellipse detection methods is to produce an ellipse by grouping elliptic arcs. Thus, the quality (i.e., continuity) of elliptic arcs will have a significant impact on the performance of ellipse detection. In this paper, a novel contrast-guided ellipse detection method is proposed for generating elliptic arcs with improved quality. After applying an edge detector to the input image, all the nearby pixels along each edge are evaluated for possibly re-classifying them to the opposite class; that is, non-edge pixels might be re-classified as edge pixels, and vice versa. With the proposed contrast-guided arc enhancement, the produced elliptic arcs tend to have better quality, which is essential to detect ellipses more accurately. To further assure the detected ellipses, a novel ellipse validation process is developed and used to discard those ellipses with low confidence. Extensive experimental results have shown that the proposed method can deliver superior performance in nearly all test cases. Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 2 |
| 2024 | Contrast-Guided Wireframe ParsingabstractExisting deep wireframe parsing methods typically focus on the semantic saliency of scene structural lines without paying particular attention to their visual saliency. As a result, these methods often face the challenge of multiple responses to proximate line segments or erroneous responses to non-structural elements like shadows. To address this fundamental issue, a novel Contrast Guidance Module (CGM) is proposed. In the CGM, a low-level image attribute, i.e., the image contrast, is leveraged to measure the visual saliency of structural lines. The CGM augments feature maps in CNN networks, subtly balancing the interpretation of line segments with their contextual significance in the image. This approach not only refines detection accuracy but also enriches the understanding of spatial geometry. Extensive experiments conducted on benchmark datasets have shown that our proposed CGM consistently outperforms the current state-of-the-art methods. Xueyuan Chen, Baojiang Zhong |
ICIP | 2 |
| 2024 | ATU-NET: An Adaptive Transformation-Based U-NET for Medical Image SegmentationabstractBoth the U-Net and its variants, which produce the state-of-the-art performance in the field of medical image segmentation, are founded on an encoder-decoder architecture. However, this architecture generally processes the input image in the spatial domain only, overlooking potential insights that could be gained from other transform domains. For that, an adaptive transformation-based U-Net (ATU-Net) is proposed in this paper. Our ATU-Net is based on a novel network architecture called the adaptive transformation-encoder-decoder (ATU), which adaptively transforms the input image into a more suitable domain by training the transformation kernel for processing to make it easier for the encoder and decoder to extract key features of the image. Extensive experimental results obtained on benchmark datasets have shown that our proposed ATU-Net can deliver superior performance to the existing methods. Qianyu Du, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 2 |
| 2024 | An Image Decomposition-Guided Network for Image InterpolationabstractA novel image decomposition-guided network (IDGN) for image interpolation is proposed in this paper by incorporating the fundamentals of subband image decomposition into the design of our deep-learning network. In our work, a filter bank consisting of a Gaussian filter and a differenceof-Gaussian filter is designed for decomposing the low-resolution input image into multiple subbands of the same resolution without downsampling. These subbands are inherited with different low-frequency and high-frequency information and are ready to be interpolated individually in our developed network. For training our IDGN, the decomposed low-resolution subbands need to be paired up with their corresponding ground-truth high-resolution subbands. Since our human visual system is sensitive to high-frequency signals, a perception-regulated (PR) loss function is proposed to guide our IDGN by putting more emphasis on the high-frequency subbands during the training process. Extensive experimental results have shown that our IDGN can achieve superior performance when compared with a number of state-of-the-art image interpolation methods. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma, Fuhui Zhou, Qihui Wu 0001 |
ICIP | 2 |
| 2024 | DALSM: A Direction-Aware Line Segment Matching MethodabstractMatching line segments between a pair of images depicting the same scene is popularly achieved through the utilization of image feature point correspondences. However, existing methods of this type often exhibit inferior performance as they typically neglect other important image attributes. To address this fundamental problem, a novel direction-aware line segment matching (DALSM) method is proposed in this paper. Specifically, when establishing a potential match between a pair of line segments, two direction-aware image attributes are incorporated: the intersection angle between the two line segments and the gradient direction of each line segment. By integrating these direction-aware image attributes with feature point correspondences, the accuracy of the matching results is significantly improved. To further enhance matching performance, a prediction-correction scheme is also developed and exploited. In the prediction stage, a set of loose geometric constraints is used to filter out low-confidence line match candidates. Subsequently, in the correction stage, the aforementioned direction-aware image attributes are employed to assess similarities among line match candidates. Experimental results on benchmark datasets demonstrate that the proposed DALSM outperforms current state-of-the-art methods clearly. Baojiang Zhong |
ICIP | 2 |
| 2024 | A Robust Airport Detection Method Based on Environment-Insensitive Saliency Analysis
Hongtao Chen, Baojiang Zhong, Kai-Kuang Ma |
PRICAI (5) | 2 |
| 2024 | Antecedent hash modality learning and representation for enhanced wafer map defect pattern recognition
Minghao Piao, Cheng Hao Jin, Baojiang Zhong |
Expert Syst. Appl. | 3 |
| 2024 | MPG-LSD: A high-quality line segment detector based on multi-scale perceptual grouping
Baojiang Zhong, Xueyuan Chen, Hangjia Zheng |
Pattern Recognit. | 2 |
| 2024 | A Channel-Wise Multi-Scale Network for Single Image Super-ResolutionabstractExisting multi-scale feature extraction methods extract image features using various convolution window sizes conducted on the spatial dimension of the feature maps. However, such an approach inevitably encounters redundant convolution operations. To address this concern, we propose to extract multiscale features on the channel dimension rather than on the spatial dimension. To demonstrate, a channel-wise multi-scale network (CMSN) is proposed for conducting single image super-resolution (SISR). In our CMSN, a sequence of channel-wise multi-scale blocks (CMSBs) is designed to extract multi-scale features at increasing levels by performing convolutions with different channel numbers (i.e., scales). To fuse the image features generated from different levels in our CMSN, a hybrid attention-aware feature fusion block (HAFFB) is proposed. Extensive experimental results have clearly shown the superiority of our CMSN to that of several state-of-the-art SISR methods on delivering superior high-resolution images, both objectively and subjectively. This reveals the potential of channel-wise, versus spatial-wise, on the effectiveness of multi-scale feature extraction. Jiahuan Ji, Baojiang Zhong, Qihui Wu 0001, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2024 | Contrast-Guided Line Segment DetectionabstractDue to the effects of quantization error and image noise, detecting ‘meaningful’ line segments from an image with high continuity is a challenging task. To pursue this goal, a novel line segment detector, called thecontrast-guided line segment detector(CGLSD), is proposed in this paper. Our basic idea is to integrate a low-level image attribute, i.e.,edge contrast, into the line segment detection process for improving line continuity. After applying an edge detector to the input image, the edge contrast is exploited to guide the growth of aline-support regionfor each line segment individually. This is achieved by evaluating edge pixels as well as those non-edge pixels that are nearby the edges. As a result, some of the non-edge pixels are re-considered as ‘edge’ pixels and included for establishing the support region. Reversely, certain edge pixels might be treated as ‘non-edge’ pixels instead and excluded from the region. Since each support region is supposed to yield onlyoneline segment, each formed support region needs to have arefinementby removing those edge pixels that do not belong to it. Lastly, the support region is required to pass through avalidationcheck that might lead to a complete discard of the line segment due to its low confidence. Extensive experiments are conducted and compared with multiple state-of-the-arts on two datasets, including the one from us with manually-annotated ground truth. The results have shown that the proposed CGLSD can deliver superior performance in nearly all test cases. Zikai Wang 0007, Baojiang Zhong, Dongxu Han, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2024 | Anisotropic Scale-Invariant Ellipse DetectionabstractDetecting ellipses poses a challenging low-level task indispensable to many image analysis applications. Existing ellipse detection methods commonly encounter two fundamental issues. First, inferior detection accuracy could be incurred on a small ellipse than that on a large one; this introduces the scale issue. Second, inferior detection accuracy could be yielded along the minor axis than along the major one of the same ellipse; this leads to the anisotropy issue. To address these issues simultaneously, a novel anisotropic scale-invariant (ASI) ellipse detection methodology is proposed. Our basic idea is to perform ellipse detection in a transformed image space referred to as the ellipse normalization (EN) space, in which the desired ellipse from the original image is 'normalized' to the unit circle. With the establishment of the EN-space, an analytical ellipse fitting scheme and a set of distance measures are developed. Theoretical justifications are then derived to prove that both our ellipse fitting scheme and distance measures are invariant to anisotropic scaling, and thus each ellipse can be detected with the same accuracy regardless of its size and ellipticity. By incorporating these components into two recent state-of-the-art algorithms, two ASI ellipse detectors are finally developed and exploited to verify the effectiveness of our proposed methodology. Zikai Wang 0007, Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2023 | A Multi-scale Method for Cell Segmentation in Fluorescence Microscopy Images
Yating Fang, Baojiang Zhong |
ICANN (2) | 2 |
| 2023 | A Content-Based Multi-Scale Network for Single Image Super-ResolutionabstractA novel content-based multi-scale network (CMNet) is proposed in this paper for conducting single image super-resolution (SISR). Its core lies in a content-based multi-scale image representation (CMIR), which is motivated by the fact that the contents of real-world images normally have different scales. Thus, it is expected that individual treatments of these contents would yield superior SISR performance. In our CMIR, the difference curvature (DCurv) is first exploited to generate a primal sketch of the input image. Then, a filter bank is designed and used to obtain a set of coefficient matrices, and each matrix reflects the characteristics of the image content at the corresponding scale. Based on these coefficient matrices, the CMIR of the input image is formed. To conduct SISR, each scale of CMIR is processed individually in our CMNet, and the produced multi-scale outputs are then integrated to arrive at the final SISR image with a higher image quality. Extensive experiments have demonstrated that our developed CMNet can deliver superior performance compared with a number of state-of-the-art SISR methods. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Mu |
ICASSP | 2 |
| 2023 | Line Segment Matching Based on Intersection-Enhanced Point CorrespondencesabstractLine segment matching between two images of the same scene is popularly performed with the assistant of feature point correspondences. However, each existing method only relies on one specific kind of feature points to establish the correspondences, and could yield inferior performance when insufficient high-quality point correspondences are supplied. To address this fundamental problem, a novel line segment matching method is proposed. In our method, ASIFT points are exploited to establish an initial set of correspondences. Then, line intersections are matched between the two input images to establish an additional set of correspondences, i.e., intersection correspondences. Finally, an enhanced set of point correspondences is generated via a union of the ASIFT and intersection correspondences. Based on the yielded intersection-enhanced point correspondences, it will be shown that the performance of line segment matching can be greatly improved. Baojiang Zhong |
ICASSP | 2 |
| 2023 | A Multi-Scale Cell Segmentation Method for Detecting Hematological DisordersabstractCell segmentation, conducted on a microscopic image that contains blood cells, plays a crucial role for the detection of various hematological disorders. Existing methods often yield inferior performance in the presence of elongated and irregularly-shaped cells, as well as to those adjacent cells with partial overlapping among themselves. To address these issues, a novel multi-scale cell segmentation (MCS) method is proposed in this paper that involves three scales, denoted by coarse, medium, and fine, for demonstrating the effectiveness and efficacy of the proposed multiscale approach. It has been shown in our work that noise and insignificant cell structures can be effectively suppressed at the coarse scale. Consequently, those elongated and irregularly-shaped cells are more accurately identified. Furthermore, our developed coarse-to-fine scale tracing algorithm, performed to trace each identified cell, improves segmentation accuracy. To solve the problem of cell overlapping, an effective cell identification technique is developed for extracting cell contours for each scale. Extensive experimental results obtained from three benchmark datasets have shown that the proposed MCS outperforms the state-of-the-art methods by a large margin. Yating Fang, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 2 |
| 2023 | Learning Multi-Scale Features for Jpeg Image Artifacts RemovalabstractA recently proposed quantization-table convolutional network (QCN) has proven as a state-of-the-art JPEG image artifacts removal method. To reduce computational complexity, the QCN learns image features from the down-sampled version of the input image. Consequently, the performance might be compromised as some salient features that can only be learned from the original input image with full resolution will be lost. To solve this problem, a novel multi-scale feature extraction block (MFEB) is proposed in this paper, which contains a coarse-scale branch and a fine-scale branch for learning salient image features from the down-sampled input image and the original full resolutions, respectively. To avoid introducing too much additional computational complexity due to the fine-scale branch, in our MFEB, this branch uses only one layer, while the coarse-scale branch exploits multiple layers. The feature sets obtained from the two branches are then fused. With the MFEB, a novel multi-scale artifacts removal network (MARN) is then developed to remove JPEG image artifacts. Extensive experiments have clearly shown that our MARN can deliver superior performance to that of a number of state-of-the-art methods. Jiahuan Ji, Baojiang Zhong, Weigang Song, Kai-Kuang Ma |
ICIP | 2 |
| 2023 | Window Attention with Multiple Patterns for Single Image Super-ResolutionabstractThe non-local attention (NLA) is a generic method-ology that has proven its effectiveness in solving various image processing and computer vision tasks. Unfortunately, the NLA over the entire input image requires extremely high computational burden. A potential solution is to use the window attention instead, which restricts the NLA computation within a square window. However, this will sacrifice the capacity of capturing long-range dependencies (LRDs). To overcome the difficulty, a novel attention module termed the window attention with multiple patterns (WAMP) is proposed. In our WAMP, three window patterns are jointly exploited, including the square, horizontal, and vertical patterns, by which the attention receptive field can be expanded to the entire image for measuring LRDs with favorably low computational complexity. To cope with image features at different granularity levels, our WAMP is further performed with multiple window granularities. Moreover, by blending our WAMP with shift convolutions, a novel architectural element is also proposed, which incorporates non-local attention with local image information for boosting the model performance. To verify the effectiveness of our proposed attention module, a WAMP network is finally developed for conducting single image super-resolution (SISR). Extensive experimental results clearly demonstrate the superiority of our WAMP network over the state-of-the-art SISR methods. Xianwei Xiao, Baojiang Zhong |
ICTAI | 2 |
| 2023 | Multi-Scale Non-Local Sparse Attention for Single Image Super-ResolutionabstractThe non-local attention (NLA) has demonstrated its success in deep learning to solve various image processing and computer vision tasks. However, the NLA over the entire input image requires extremely high computational complexity and inevitably introduces irrelevant information since all the feature points are involved in calculating attention map. To address these problems, a novel attention module, called the multi-scale non-local sparse attention (MNSA), is proposed in this paper. In our MNSA, attention calculation is constrained within non-overlapping windows, and then only the most relevant feature points are selected to compute an attention map. The resulting sparse attention prevents the model from attending to irrelevant information and noise while reducing the computational complex-ity from quadratic to linear with respect to the input image size. To obtain receptive fields at different scales, our MNSA is further performed by exploiting different sizes of windows. Moreover, a novel local feature extraction (LFE) is proposed to extract the local structural information of natural images. To verify the effectiveness of our proposed attention module, a MNSA network is finally developed for conducting single image super-resolution (SISR). Extensive experimental results have clearly shown that our MNSA network can deliver superior performance over a number of state-of-the-art SISR methods. Xianwei Xiao, Baojiang Zhong |
IJCNN | 2 |
| 2023 | MSED: A Robust Ellipse Detector with Multi-scale Merging and Validation
Baojiang Zhong |
PRCV (11) | 2 |
| 2023 | On the generalized k-cosine arithmetic-mean curvature for multi-scale corner detection
Baojiang Zhong, Kai-Kuang Ma |
Expert Syst. Appl. | 2 |
| 2023 | Corner Detection via Scale-Space Behavior-Guided Trajectory TracingabstractExistingcurvature scale-space(CSS) methods detect corners by tracing the CSS trajectories from a determined high scale toward the lowest one. For those images with sophisticated details, such approach could often yield unsatisfactory corner detection results; i.e., miss-detectedtruecorners (false negatives) androundcorners (false positives). In this letter, these two fundamental problems are investigated. To tackle them, a novel CSS-based corner detector is proposed by incorporating our mathematically derived scale-space properties of the planar curves and corner points into the developed trajectory tracing algorithm, called thescale-space behavior-guided trajectory tracing(SBTT). In view of lacking a benchmark dataset with the ground truth, another contribution from our work is on the establishment of an augmented test image dataset, containing 147 test images with manually-labelled ground truth and their augmented images up to 62,328 images in total. Based on the ground truth, four commonly-used metrics are exploited to conduct corner detection performance evaluation. The obtained simulation results show that our proposed corner detector yields the highest F-score, when compared with that of nine state-of-the-art methods. Baojiang Zhong, Jianyu Yang 0002, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2022 | A Lightweight Local-Global Attention Network for Single Image Super-Resolution
Zijiang Song, Baojiang Zhong |
ACCV (6) | 2 |
| 2022 | MA-NET: Multi-Scale Attention-Aware Network for Optical Flow EstimationabstractExisting methods for optical flow estimation can perform well on images with small offsets of large objects. However, they could fail when there are a number of small and fast-moving objects. In particular, current deep learning methods often impose a down-sampling of the input data for a reduction of computational complexity. This practice inevitably incurs a loss of image details and causes small objects be ignored. To solve the problem, a multi-scale attention-aware network (MA-Net) is proposed, in which coarse-scale and fine-scale features are extracted in parallel and then attention is paid to them for producing the optical flow estimation. Extensive experiments conducted on the Sintel and KITTI2015 datasets show that the proposed MA-Net can capture fast-moving small objects with high accuracy and thus deliver superior performance over a number of state-of-the-art methods. Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 2 |
| 2022 | Reference-Based Jpeg Image Artifacts RemovalabstractSince information loss incurred in lossy image compression is irreversible, it is extremely challenging to achieve a satisfactory effect of artifacts removal by using the compressed image itself only. To solve the problem, a novel reference-based artifacts removal network (RARN) is proposed in this paper, which exploits a high-quality reference image to provide useful information for facilitating the removal of artifacts and the reconstruction of details. In our RARN, a feature extraction module is first established, which takes both the compressed and reference images as inputs and produces multi-scale feature pairs as outputs. Then, a feature transfer non-local block (FTNB) is developed to match the feature pairs and transfer relevant features from the reference image to the compressed one in the feature space. Finally, image information is recovered from the multi-scale outputs of FTNB by using a reconstruction module. Extensive experimental results clearly show that our proposed RARN can deliver superior performance over a number of state-of-the-art methods. Weigang Song, Jiahuan Ji, Baojiang Zhong |
ICIP | 3 |
| 2022 | Anisotropic Edge Detection in Catadioptric ImagesabstractCatadioptric images are produced in omnidirectional vision systems and can be expressed on Riemannian manifolds. The existing edge detectors are operated either in Euclidean space, or on Riemannian manifolds with isotropic image filtering. In this paper, a new type of edge detection is proposed—it is operated on Riemannian manifolds with anisotropic image filtering. For that, an anisotropic image filtering kernel on Riemannian manifolds is derived by solving the anisotropic heat equation embedded with Riemannian metric. With this kernel, a novel anisotropic edge detector is then developed. Compared to an edge detector operated in Euclidean space, our edge detector is more suitable for catadioptric images, since their geometric structure information will be taken into account in the detection process. Compared to existing edge detectors customized for catadioptric images, the new edge detector has a higher efficiency in preserving image edges and thus can produce more true positives. Enzhuang Zheng, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 2 |
| 2022 | Object-Scale Adaptive Optical Flow Estimation Network
Baojiang Zhong, Kai-Kuang Ma |
PRICAI (3) | 2 |
| 2022 | Nested Multi-Axis Learning Network for Single Image Super-Resolution
Xianwei Xiao, Baojiang Zhong |
PRICAI (3) | 2 |
| 2022 | Corner Detection Based on a Dynamic Measure of Cornerity
Baojiang Zhong |
PRICAI (3) | 2 |
| 2022 | A Direction-Decoupled Non-Local Attention Network for Single Image Super-ResolutionabstractThenon-local attentionmechanism has often been exploited in deep learning to capturelong-range dependencies(LRDs) from the same image for enhancing the performance of various image processing methods. However, the initially proposed non-local attention process inevitably yields extremely-high computation complexity, sinceallthe feature points are involved in computing the LRDs. To address this concern, a recently proposedcriss-cross network(CCNet), which has arecurrent criss-cross attention(RCCA) module, is used to compute the LRDs by involving only a small set of feature points for significantly reducing computation. Motivated by the RCCA, a noveldirection-decoupled non-local attention(DNA) module is proposed in this paper that is able to further reduce the computation complexity of RCCA by half approximately. To verify the performance of our new non-local attention module, a DNA network is developed for conducting single image super-resolution (SISR). Extensive experimental results have clearly demonstrated the superiority of using our DNA network for SISR when compared with that of state-of-the-art methods. Zijiang Song, Baojiang Zhong, Jiahuan Ji, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2021 | Single Image Super-Resolution Using Asynchronous Multi-Scale NetworkabstractAn existing multi-scale residual network (MSRN) has demonstrated its success on conducting the single image super-resolution (SISR) task. The MSRN consists of a number of multi-scale residual blocks (MSRBs), and each MSRB performs convolutions by exploiting two different sizes of windows for conducting multi-scale feature extraction. The smaller window is used to extract image features at a low scale, while the larger one is used for a high scale. To significantly reduce the number of parameters involved in the MSRB, a new feature extraction module, called the asynchronous multi-scale block (AMB), is proposed in this paper. It is based on the fact that the larger window used in the MSRB can be replaced by two smaller windows without affecting the original MSRB's function. Consequently, by replacing each MSRB with our AMB, an asynchronous multi-scale network (AMNet) is then constructed, which can yield a significant reduction on computational complexity. This means that more AMBs can be used in our AMNet to deliver superior SISR performance, while maintaining the same or comparable computational complexity to that of the MSRN. To consolidate all image features generated from all scales, a new fusion scheme, called the adaptive feature fusion block (AFFB), is proposed that weights the extracted features according to their importance for further increasing SISR's performance. Extensive experimental results have clearly shown the superiority of our proposed AMNet when compared with multiple state-of-the-arts. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2020 | Single Image Super-Resolution Via A Progressive Mixture ModelabstractIn this paper, a progressive mixture model (PMM) for single image super-resolution is proposed. Our model consists of an offline training stage and an online reconstruction stage, and both stages are conducted progressively by exploiting a uniform iterative scheme. In the training stage, the training dataset is clustered into finer and finer groups, and a set of mixture models are sequentially learned. In the reconstruction stage, residuals of the low-resolution (LR) input image are estimated by using the trained mixture models at increasing levels, and then they are progressively added to the LR image for producing the high-resolution (HR) image. Extensive experimental simulation results have clearly shown that the proposed model consistently delivers highly accurate and visually pleasant HR images, compared to that of the state-of the-art image super-resolution methods. Run Su, Baojiang Zhong, Jiahuan Ji, Kai-Kuang Ma |
ICIP | 2 |
| 2020 | Image Interpolation Using Multi-Scale Attention-Aware Inception NetworkabstractA new multi-scale deep learning (MDL) framework is proposed and exploited for conducting image interpolation in this paper. The core of the framework is a seeding network that needs to be designed for the targeted task. For image interpolation, a novel attention-aware inception network (AIN) is developed as the seeding network; it has two key stages: 1) feature extraction based on the low-resolution input image; and 2) feature-to-image mapping to enlarge image's size or resolution. Note that the designed seeding network, AIN, needs to be trained with a matched training dataset at each scale. For that, multi-scale image patches are generated using our proposed pyramid cut, which outperforms the conventional image pyramid method by completely avoiding aliasing issue. After training, the trained AINs are then combined for processing the input image in the testing stage. Extensive experimental simulation results obtained from seven image datasets (comprising 359 images in total) have clearly shown that the proposed MAIN consistently delivers highly accurate interpolated images. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2020 | Color Image Demosaicing Using Progressive Collaborative RepresentationabstractIn this paper, a progressive collaborative representation (PCR) framework is proposed that is able to incorporate any existing color image demosaicing method for further boosting its demosaicing performance. Our PCR consists of two phases: (i) offline training and (ii) online refinement. In phase (i), multiple training-and-refining stages will be performed. In each stage, a new dictionary will be established through the learning of a large number of feature-patch pairs, extracted from the demosaicked images of the current stage and their corresponding original full-color images. After training, a projection matrix will be generated and exploited to refine the current demosaicked image. The updated image with improved image quality will be used as the input for the next training-and-refining stage and performed the same processing likewise. At the end of phase (i), all the projection matrices generated as above-mentioned will be exploited in phase (ii) to conduct online demosaicked image refinement of the test image. Extensive simulations conducted on two commonly-used test datasets (i.e., the IMAX and Kodak) for evaluating the demosaicing algorithms have clearly demonstrated that our proposed PCR framework is able to constantly boost the performance of any image demosaicing method we experimented, in terms of the objective and subjective performance evaluations. Zhangkai Ni, Kai-Kuang Ma, Huanqiang Zeng, Baojiang Zhong |
IEEE Trans. Image Process. | 4 |
| 2019 | Multi-Scale Defense of Adversarial ImagesabstractDeep learning has achieved great success in image classification. However, recent researches show that existing deep learning-based classifiers remain weak for recognizing adversarial images. In this paper, an effective multi-scale defense method is proposed to solve the problem. In our method, an input image is first evolved by Gaussian kernels of different intensities to generate a multi-scale representation of the image. These evolved images are then fed into a classifier trained with a multi-scale strategy to yield multi-scale confidences. Finally, an average confidence is exploited to generate classification result. Furthermore, by monitoring the change of confidence values during the image evolution process, our method is able to achieve an indication of the attacking risk. Experimental results show that our method performs favorably against a number of state-of-the-art methods. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 2 |
| 2019 | Shape Matching Based on Rectangularized Curvature Scale-Space MapsabstractThe curvature scale-space (CSS) technique is a shape descriptor in the MPEG-7 standard. To match a shape under query, a common approach is to exploit a simplified CSS map, which is produced by representing each arch-shaped contour of the original CSS map by a vertical line segment with the same height. However, we consider this approach is oversimplified, since some very useful information presented in the original CSS map has been totally ignored. To solve the problem, a new CSS map approximation is proposed in this paper, which is a rectangularized bin circumscribing each arch-shaped contour of the CSS map so that both the height and width characteristics are simultaneously captured and preserved. For conducting shape matching, an efficient method is further proposed, which has a significantly more succinct structure than the existing CSS matching technique used in MPEG-7. Simulation results have shown that the proposed method can yield superior performance. Baojiang Zhong, Kai-Kuang Ma |
ICIP | 2 |
| 2019 | Confusion-Aware Convolutional Neural Network for Image Classification
Liguang Yan, Baojiang Zhong, Kai-Kuang Ma |
ICONIP (1) | 2 |
| 2019 | Shape Description and Retrieval in a Fused Scale Space
Baojiang Zhong, Jianyu Yang 0002 |
ICONIP (2) | 2 |
| 2019 | Predictor-corrector image interpolation
Baojiang Zhong, Kai-Kuang Ma, Zhifang Lu |
J. Vis. Commun. Image Represent. | 1 |
| 2017 | Extract feature curves on noisy triangular meshes
Hao Liu 0029, Baojiang Zhong, Tao Li 0007, Jun Wang 0039 |
Graph. Model. | 3 |
| 2015 | A location-aware scale-space method for salient object detectionabstractMany existing saliency detection methods made an assumption that the salient object is on the center of the image and incorporated such center-biased assumption in the design of their algorithms. Obviously, this is not always proper to set, especially for those imageries acquired by unmanned monitoring system or device (e.g., surveillance camera), in which the salient object could appear in any location within the image. Consequently, the resulted saliency detection performance could be greatly degraded. In this paper, an existing hypercomplex Fourier transform (HFT) based saliency detection algorithm is investigated and modified for improving the saliency detection performance. In details, we remove its prior assumption on `center bias' and exploit a location-aware strategy to identify the optimal saliency map across multiple scales of the image. Extensive simulation results have justified that the proposed location-aware HFT-based approach clearly outperforms existing five state-of-the-art algorithms on saliency detection. Dan Xiang, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 2 |
| 2013 | Curvature scale-space of open curves: Theory and shape representationabstractThe problem of extending the curvature scale-space (CSS) technique to represent open curves is addressed. Various approaches for dealing with the endpoint problem of open curves are considered, and one is selected which allows us to handle the evolution of the open curves as a special case of the evolution of closed curves. The convergence theory of evolved open curves is established, and the CSS shape representation is investigated. Baojiang Zhong, Kai-Kuang Ma, Jiwen Yang |
ICIP | 1 |
| 2013 | Vertical corner line detection on buildings in quasi-Manhattan worldabstractAn efficient algorithm is proposed for detecting vertical corner lines on buildings in quasi-Manhattan world scene. The vertical corner lines are useful geometric information of buildings in urban scene, whose importance is similar to that of corners on planar curves for various applications in computer vision. The algorithm employs a bottom-up and step-by-step pipeline processing procedure. First, straight-line segments are extracted as low-level features from the input building image by using a fast and accurate line segment detector. Secondly, the extracted straight-line segments are clustered to groups and associated to different vanishing points, which are computed by a J-Linkage estimator and act as the mid-level features. Finally, high-level features, vertical corner lines, are detected with several geometric constraints in quasi-Manhattan world. Experimental results demonstrate the efficiency of the new algorithm. Baojiang Zhong, Jiwen Yang |
ICIP | 1 |
| 2013 | Classification by nearness in complementary subspaces
Menglong Yang, Yiguang Liu, Baojiang Zhong |
Pattern Anal. Appl. | 3 |
| 2012 | A scale-space technique for polygonal approximation of planar curvesabstractA novel technique is proposed for polygonal approximation of planar curves under a given maximum of approximation error and with a given initial vertex. Different to the existing techniques in this field, which usually accept a fixed error, the proposed technique uses a flexible acceptable error to obtain the approximate polygon. This idea is based on a scale-space concept in computer vision. It can not only ensure the error between the original and approximated curves is no lager than the maximal acceptable error, but also make the description of the detail information of the curve more precise without significant loss. Experiments are conducted to compare the new technique with the existing error-fixed techniques. Baojiang Zhong |
ICIP | 2 |
| 2010 | On the Convergence of Planar Curves Under SmoothingabstractCurve smoothing has two important applications in computer vision and image processing: 1) the curvature scale-space (CSS) technique for shape analysis, and 2) the Gaussian filter for noise suppression. In this paper, we study how planar curves converge as they are smoothed with increasing scales. First, two types of convergence behavior are clarified. The coined term shrinkage refers to the reduction of arc-length of a smoothed planar curve, which describes the convergence of the curve latitudinally; and another coined term collapse refers to the movement of each point to its limiting position, which describes the convergence of the curve longitudinally. A systematic study on the shrinkage and collapse of three categories of curve models is then presented. The corner models helps to reveal how the local structures of planar curves collapse and what the smoothed curves may converge to. The sawtooth models allows us to gain insights regarding how noise is suppressed from noisy planar curves by the Gaussian filter. Our investigation on the closed curves shows that each curve collapses to a point at its center of mass. However, different curves may yield different limiting shapes at the infinity scale. Finally, based upon the derived results the performance of the CSS technique in corner detection and shape representation is analyzed, and a fast implementation method of the Gaussian filter for noise suppression is proposed. Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 1 |
| 2009 | Scale-Space Behavior of Planar-Curve CornersabstractThe curvature scale-space (CSS) technique is suitable for extracting curvature features from objects with noisy boundaries. To detect corner points in a multiscale framework, Rattarangsi and Chin investigated the scale-space behavior of planar-curve corners. Unfortunately, their investigation was based on an incorrect assumption, viz., that planar curves have no shrinkage under evolution. In the present paper, this mistake is corrected. First, it is demonstrated that a planar curve may shrink nonuniformly as it evolves across increasing scales. Then, by taking into account the shrinkage effect of evolved curves, the CSS trajectory maps of various corner models are investigated and their properties are summarized. The scale-space trajectory of a corner may either persist, vanish, merge with a neighboring trajectory, or split into several trajectories. The scale-space trajectories of adjacent corners may attract each other when the corners have the same concavity, or repel each other when the corners have opposite concavities. Finally, we present a standard curvature measure for computing the CSS maps of digital curves, with which it is shown that planar-curve corners have consistent scale-space behavior in the digital case as in the continuous case. Baojiang Zhong, Kai-Kuang Ma, Wenhe Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2007 | Direct Curvature Scale Space: Theory and Corner DetectionabstractThe Curvature Scale Space (CSS) technique is considered to be a modern tool in image processing and computer vision. Direct Curvature Scale Space (DCSS) is defined as the CSS that results from convolving the curvature of a planar curve with a Gaussian kernel directly. In this paper we present a theoretical analysis of DCSS in detecting corners on planar curves. The scale space behavior of isolated single and double corner models is investigated and a number of model properties are specified which enable us to transform a DCSS image into a tree organization and, so that corners can be detected in a multiscale sense. To overcome the sensitivity of DCSS to noise, a hybrid strategy to apply CSS and DCSS is suggested. Baojiang Zhong, Wenhe Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | A Hybrid Method for Fast Computing the Curvature Scale Space ImageabstractThe curvature scale space (CSS) technique is one of the key techniques of the MPEG-7 international standard in computer vision and image processing. It was selected as a contour shape descriptor for MPEG-7 after substantial and comprehensive testing. However, to compute a CSS image in general needs to wait a long time. This is very disadvantageous when the CSS technique is applied to an object recognition system to perform real-time recognition. In order to solve this bottleneck problem, a hybrid method for fast computing the CSS image is proposed. In the method, firstly the curve is evolved in low scale space, and after image noise is suppressed then the curvature is evolved directly. Numerical experiments show that the hybrid method can perform equally well as the existing method. It is suitable for recognizing a noisy curve of arbitrary shape at any scale or orientation. On the other hand, the hybrid method only requires 1/3 /spl sim/ 1/5 CPU time of the existing one. As a result, the CSS technique is improved significantly for real-time recognition. Baojiang Zhong, Wenhe Liao |
GMP | 1 |