Shipeng Zhu

dblp:252/0041 · DBLP profile ↗
← Back
15ranked-venue papers
5as first author
13since 2021 · last 2026
0000-0002-8296-6743ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 5 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Adaptive Hyperbolic Kernels: Modulated Embedding in de Branges-Rovnyak Spaces
abstract
Hierarchical data pervades diverse machine learning applications, including natural language processing, computer vision, and social network analysis. Hyperbolic space, characterized by its negative curvature, has demonstrated strong potential in such tasks due to its capacity to embed hierarchical structures with minimal distortion. Previous evidence indicates that the hyperbolic representation capacity can be further enhanced through kernel methods. However, existing hyperbolic kernels still suffer from mild geometric distortion or lack adaptability. This paper addresses these issues by introducing a curvature-aware de Branges–Rovnyak space, a reproducing kernel Hilbert space (RKHS) that is isometric to a Poincaré ball. We design an adjustable multiplier to select the appropriate RKHS corresponding to the hyperbolic space with any curvature adaptively. Building on this foundation, we further construct a family of adaptive hyperbolic kernels, including the novel adaptive hyperbolic radial kernel, whose learnable parameters modulate hyperbolic features in a task-aware manner. Extensive experiments on visual and language benchmarks demonstrate that our proposed kernels outperform existing hyperbolic kernels in modeling hierarchical dependencies.
Leping Si, Meimei Yang, Hui Xue 0002, Shipeng Zhu, Pengfei Fang
AAAI4
2026 InsNet: Deep indefinite spectral kernel network
Yanfang Xue, Hui Xue 0002, Shipeng Zhu
Pattern Recognit.3
2025 Triple-Directional Fusion Attention for Infrared Small Target Detection
abstract
Attention mechanism has gained popularity due to its effectiveness. However, most existing mechanisms are designed for large-sized targets currently, with limited improvement in single-frame infrared small target (SIRST) detection tasks. In this letter, we propose a novel attention mechanism to enhance the extraction capacity of deep networks for infrared small targets, termed the triple-directional fusion attention module (TFAM). This module aggregates channel-, height-, and width-dimension into three independent directional perception attention vectors, preserving both accurate channel and spatial information. Through adaptive cross-direction interaction, TFAM establishes inter-directional dependencies essential for enhancing faint target signatures in deep layers. Notably, TFAM only requires minimal complexity for modeling and offers flexibility in integration. Experiments conducted on the NUDT-SIRST and NUAA-SIRST datasets demonstrate consistent improvements.
Jun Chen 0007, Shipeng Zhu, Boyang Li 0007, Jianpeng Fan, Zaiping Lin, Wei An 0003
IEEE Geosci. Remote. Sens. Lett.4
2024 Text Image Inpainting via Global Structure-Guided Diffusion Models
abstract
Real-world text can be damaged by corrosion issues caused by environmental or human factors, which hinder the preservation of the complete styles of texts, e.g., texture and structure. These corrosion issues, such as graffiti signs and incomplete signatures, bring difficulties in understanding the texts, thereby posing significant challenges to downstream applications, e.g., scene text recognition and signature identification. Notably, current inpainting techniques often fail to adequately address this problem and have difficulties restoring accurate text images along with reasonable and consistent styles. Formulating this as an open problem of text image inpainting, this paper aims to build a benchmark to facilitate its study. In doing so, we establish two specific text inpainting datasets which contain scene text images and handwritten text images, respectively. Each of them includes images revamped by real-life and synthetic datasets, featuring pairs of original images, corrupted images, and other assistant information. On top of the datasets, we further develop a novel neural framework, Global Structure-guided Diffusion Model (GSDM), as a potential solution. Leveraging the global structure of the text as a prior, the proposed GSDM develops an efficient diffusion model to recover clean texts. The efficacy of our approach is demonstrated by thorough empirical study, including a substantial boost in both recognition accuracy and image quality. These findings not only highlight the effectiveness of our method but also underscore its potential to enhance the broader field of text image understanding and processing. Code and datasets are available at: https://github.com/blackprotoss/GSDM.
Shipeng Zhu, Pengfei Fang, Chenjie Zhu, Zuoyan Zhao, Hui Xue 0002
AAAI1
2024 PEAN: A Diffusion-Based Prior-Enhanced Attention Network for Scene Text Image Super-Resolution
abstract
Scene text image super-resolution (STISR) aims at simultaneously increasing the resolution and readability of low-resolution scene text images, thus boosting the performance of the downstream recognition task. Two factors in scene text images, visual structure and semantic information, affect the recognition performance significantly. To mitigate the effects from these factors, this paper proposes a Prior-Enhanced Attention Network (PEAN). Specifically, an attention-based modulation module is leveraged to understand scene text images by neatly perceiving the local and global dependence of images, despite the shape of the text. Meanwhile, a diffusion-based module is developed to enhance the text prior, hence offering better guidance for the SR network to generate SR images with higher semantic accuracy. Additionally, a multi-task learning paradigm is employed to optimize the network, enabling the model to generate legible SR images. As a result, PEAN establishes new SOTA results on the TextZoom benchmark. Experiments are also conducted to analyze the importance of the enhanced text prior as a means of improving the performance of the SR network. Code is available at https://github.com/jdfxzzy/PEAN.
Zuoyan Zhao, Hui Xue 0002, Pengfei Fang, Shipeng Zhu
ACM Multimedia4
2024 Reproducing the Past: A Dataset for Benchmarking Inscription Restoration
abstract
Inscriptions on ancient steles, as carriers of culture, encapsulate the humanistic thoughts and aesthetic values of our ancestors. However, these relics often deteriorate due to environmental and human factors, resulting in significant information loss. Since the advent of inscription rubbing technology over a millennium ago, archaeologists and epigraphers have devoted immense effort to manually restoring these cultural imprints, endeavoring to unlock the storied past within each rubbing. This paper approaches this challenge as a multi-modal task, aiming to establish a novel benchmark for the inscription restoration from rubbings. In doing so, we construct the Chinese Inscription Rubbing Image (CIRI) dataset, which includes a wide variety of real inscription rubbing images characterized by diverse calligraphy styles, intricate character structures, and complex degradation forms. Furthermore, we develop a synthesis approach to generate "intact-degraded'' paired data, mirroring real-world degradation faithfully. On top of the datasets, we propose a baseline framework that achieves visual consistency and textual integrity through global and local diffusion-based restoration processes and explicit incorporation of domain knowledge. Comprehensive evaluations confirm the effectiveness of our pipeline, demonstrating significant improvements in visual presentation and textual integrity. The project is available at: https://github.com/blackprotoss/CIRI.
Shipeng Zhu, Hui Xue 0002, Na Nie, Chenjie Zhu, Haiyue Liu, Pengfei Fang
ACM Multimedia1
2024 Dynamic functional connections analysis with spectral learning for brain disorder detection
Yanfang Xue, Hui Xue 0002, Pengfei Fang, Shipeng Zhu, Lishan Qiao, Yuexuan An
Artif. Intell. Medicine4
2024 Improving Scene Text Retrieval via Stylized Middle Modality
abstract
Scene text retrieval addresses the challenge of localizing and searching for all text instances within scene images based on a query text. This cross-modal task has significant applications in various domains, such as intelligent transportation systems and social media analysis. In practice, ensuring consistency of the same content between two modalities is crucial in improving retrieval accuracy. This article addresses the issue by introducing a stylized middle modality, which fuses the graphical query text with the style of the extracted text proposal. To this end, we propose a stylized middle modality learning (SM 2 L) framework. The proposed stylized middle modality enables the network to jointly enforce constraints on visual feature coherence and text semantic feature consistency in the optimization phase, thereby minimizing the modality gap in the retrieval space. This brings in two major advantages: (1) SM 2 L will pave the way to seamlessly benefit the scene text retrieval and (2) the proposed learning paradigm enables the machine to avoid adding redundant computing resources in the inference phase. Substantial experiments demonstrate that the proposed method outperforms the state-of-the-art retrieval performance considerably.
Shipeng Zhu, Pengfei Fang, Hui Xue 0002
ACM Trans. Multim. Comput. Commun. Appl.1
2023 Improving Scene Text Image Super-resolution via Dual Prior Modulation Network
abstract
Scene text image super-resolution (STISR) aims to simultaneously increase the resolution and legibility of the text images, and the resulting images will significantly affect the performance of downstream tasks. Although numerous progress has been made, existing approaches raise two crucial issues: (1) They neglect the global structure of the text, which bounds the semantic determinism of the scene text. (2) The priors, e.g., text prior or stroke prior, employed in existing works, are extracted from pre-trained text recognizers. That said, such priors suffer from the domain gap including low resolution and blurriness caused by poor imaging conditions, leading to incorrect guidance. Our work addresses these gaps and proposes a plug-and-play module dubbed Dual Prior Modulation Network (DPMN), which leverages dual image-level priors to bring performance gain over existing approaches. Specifically, two types of prior-guided refinement modules, each using the text mask or graphic recognition result of the low-quality SR image from the preceding layer, are designed to improve the structural clarity and semantic accuracy of the text, respectively. The following attention mechanism hence modulates two quality-enhanced images to attain a superior SR result. Extensive experiments validate that our method improves the image quality and boosts the performance of downstream tasks over five typical approaches on the benchmark. Substantial visualizations and ablation studies demonstrate the advantages of the proposed DPMN. Code is available at: https://github.com/jdfxzzy/DPMN.
Shipeng Zhu, Zuoyan Zhao, Pengfei Fang, Hui Xue 0002
AAAI1
2023 CosNet: A Generalized Spectral Kernel Network
abstract
Complex-valued representation exists inherently in the time-sequential data that can be derived from the integration of harmonic waves. The non-stationary spectral kernel, realizing a complex-valued feature mapping, has shown its potential to analyze the time-varying statistical characteristics of the time-sequential data, as a result of the modeling frequency parameters. However, most existing spectral kernel-based methods eliminate the imaginary part, thereby limiting the representation power of the spectral kernel. To tackle this issue, we propose a generalized spectral kernel network, namely, \underline{Co}mplex-valued \underline{s}pectral kernel \underline{Net}work (CosNet), which includes spectral kernel mapping generalization (SKMG) module and complex-valued spectral kernel embedding (CSKE) module. Concretely, the SKMG module is devised to generalize the spectral kernel mapping in the real number domain to the complex number domain, recovering the inherent complex-valued representation for the real-valued data. Then a following CSKE module is further developed to combine the complex-valued spectral kernels and neural networks to effectively capture long-range or periodic relations of the data. Along with the CosNet, we study the effect of the complex-valued spectral kernel mapping via theoretically analyzing the bound of covering number and generalization error. Extensive experiments demonstrate that CosNet performs better than the mainstream kernel methods and complex-valued neural networks.
Yanfang Xue, Pengfei Fang, Jinyue Tian, Shipeng Zhu, Hui Xue 0002
NeurIPS4
2023 On learning distribution alignment for video-based visible-infrared person re-identification
Pengfei Fang, Yaojun Hu, Shipeng Zhu, Hui Xue 0002
Comput. Vis. Image Underst.3
2023 FATE: a three-stage method for arithmetical exercise correction
Qipeng Zhu, Zhuoyan Luo, Shipeng Zhu, Zihang Xu, Hui Xue 0002
Neural Comput. Appl.3
2022 Multispectral Image Matching Method Based on Histogram of Maximum Gradient and Edge Orientation
abstract
Motivated by the problem of nonlinear intensity changes in multispectral image matching, this letter introduces a robust and efficient image-matching method. The proposed method consists of three steps. First, control-point candidates are identified that are widely distributed in the areas of effective main structure. Then, a novel feature descriptor called the histogram of maximum gradient and edge orientation (HGEO) is proposed for the purpose of multispectral image matching. Finally, a bilateral matching process is carried out to perform the matching process and remove mismatches. The proposed method is successfully applied for matching various multispectral remote sensing images, and experiments are performed with typical datasets that are widely applied in tests of multispectral image matching. According to some popular feature descriptors, the test results demonstrate that the proposed HGEO achieves better matching performance than do many currently used methods.
Quan Wu, Shipeng Zhu
IEEE Geosci. Remote. Sens. Lett.2
2020 Joint Optimization on Cache Size and File Placement in Backhaul Limited Dense Networks
abstract
Backhaul capability is usually considered as a bottleneck for system performance in dense wireless networks, caching popular file in base stations is an effective solution. In the dense deployed network, it is necessary to consider the interaction between cache size and file placement, which will lead to a rise of cache deployment cost. In order to solve this problem, this paper introduces the concept of user group and proposes a scheme with joint optimization of cache size and file placement. In this scheme, considering constraint of file delivery latency and other factors, joint optimization for minimizing the total cache size is established, with iteration of the coordinate axis descent algorithm and cache value competition algorithm. Simulation results show that, compared with the greedy algorithm, the proposed algorithm can effectively reduce the total cache size and and is very close to the optimal algorithm. Influences of user group size and preference on file placement are also discussed.
Jin Xu 0001, Shipeng Zhu, Qimei Cui, Xiaofeng Tao 0001
VTC Fall2
2019 BDGAN: Image Blind Denoising Using Generative Adversarial Networks
Shipeng Zhu, Guili Xu, Yuehua Cheng, Xiaodong Han, Zhengsheng Wang
PRCV (2)1