Bochuan Zheng

dblp:66/8175 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
20since 2021 · last 2026
0000-0003-4495-8299ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 3 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Hyperspectral remote sensing image classification based on domain-level complementarity of spatial-spectral component
Bochuan Zheng
Neural Networks3
2026 LIIA -Net: A lightweight illumination iterative adjustment network for low-light image enhancement
Chengwan You, Wenxu Shi, Guibin Hu, Bochuan Zheng
Neural Networks4
2026 Enhancing small object detection: LDNet with location awareness and detail enhancement
Piao Chen, Guibin Hu, Bochuan Zheng
Pattern Recognit. Lett.4
2026 C-GAN: Medical Image Steganography Based on Convergent GANs With Localization
abstract
Image steganography aims to hide secret message into cover image in an imperceptible and undetectable way, and only allows the informed receivers to decode stego image. Generative adversarial nets have been proved to be promising against other generative models, and some recently propose to use GANs to hide secret message to reach image steganography. However, it is still facing low embedding capacity, high detectability and poor convergence. To hand the pitfalls, we propose a novel medical image steganography method with convergent divergence measurement to achieve large capacity and undetectable hiding. Specifically, generator, extractor and discriminator are jointed into end-to-end framework where generator yields visually and detectably indistinguishable steganography image from which extractor recovers diagnose report while discriminator tries to distinguish the steganographic and original images. Then, we design Zero-centered Wasserstein distance to achieve controllable and stable training. It can be proved that the proposed method with the defined Zero-centered Wasserstein can converge to a local equilibrium with finite discriminator updates per generator updates. Besides, local regularization which can be also proved to be effective for accelerating convergence is imposed on generator to improve embedding capacity and achieve homeomorphism manifold mapping in low-dimension latent space. Extensive experiments on Open-I, LGK and COV-CTR medical dataset show that the proposed method outperforms recent state-of-the-art methods in capacity, detectability and convergence rate.
Liming Xu, Bochuan Zheng, Weisheng Li 0001
IEEE Trans. Dependable Secur. Comput.3
2025 Lightweight multi-scale global attention enhancement network for image super-resolution
Yumei Zheng, Bochuan Zheng
Image Vis. Comput.4
2025 Dual-modality visual feature flow for medical report generation
abstract
Medical report generation, a cross-modal task of generating medical text information, aiming to provide professional descriptions of medical images in clinical language. Despite some methods have made progress, there are still some limitations, including insufficient focus on lesion areas, omission of internal edge features, and difficulty in aligning cross-modal data. To address these issues, we propose Dual-Modality Visual Feature Flow (DMVF) for medical report generation. Firstly, we introduce region-level features based on grid-level features to enhance the method's ability to identify lesions and key areas. Then, we enhance two types of feature flows based on their attributes to prevent the loss of key information, respectively. Finally, we align visual mappings from different visual feature with report textual embeddings through a feature fusion module to perform cross-modal learning. Extensive experiments conducted on four benchmark datasets demonstrate that our approach outperforms the state-of-the-art methods in both natural language generation and clinical efficacy metrics.
Quan Tang 0006, Liming Xu, Yongheng Wang, Bochuan Zheng, Jiancheng Lv 0001, Weisheng Li 0001
Medical Image Anal.4
2025 Pre- to post-contrast medical image synthesis with outline-guide accelerate diffusion model
Xueying Fan, Liming Xu, Bochuan Zheng
Neural Networks3
2025 Deep Differential Lifelong Cross-modal Hashing for Stream Medical Data Retrieval
abstract
With the explosive growth of stream medical multi-modal data, it is significant to develop an efficient cross-modal retrieval algorithm to achieve effective medical data search. Within it, deep cross-modal hashing which maps cross-modal data into low-dimensional Hamming space where similarity in high-dimension space is preserved has made much progress. However, most of deep cross-modal hashing algorithms are usually facing disability of adapting to dynamic stream medical data, non-differentiable optimization, and unaligned semantic across modalities. To address these, we, in this article, propose a novel deep differential lifelong cross-modal hashing method for large-scale stream medical data retrieval. Specifically, we first design lifelong learning module to keep the learned hash code of base data unchanged and directly learn hash code of incremental data with new categories to achieve continuous retrieval of stream medical data, which effectively mitigates catastrophic forgetting, as well as significantly reduces training time and computation resource. Then, we introduce differential cross-modal hashing module to generate discriminative binary hash codes, which yields continuous and differentiable optimization and improves accuracy. Besides, we design semantic alignment module which embeds intra-modal and inter-modal losses to maintain the semantic similarity and dis-similarity among stream medical data across modalities. Extensive experiments on benchmark medical datasets show that our proposed method can retrieve dynamic stream medical cross-modal data effectively and obtain higher retrieval performance comparing with recent state-of-the-art approaches.
Liming Xu, Dengping Zhao, Bochuan Zheng
ACM Trans. Multim. Comput. Commun. Appl.5
2024 Conflict-Alleviated Gradient Descent for Adaptive Object Detection
Wenxu Shi, Bochuan Zheng
IJCAI2
2024 Alleviating the Equilibrium Challenge with Sample Virtual Labeling for Adversarial Domain Adaptation
abstract
Many domain adaptive object detection (DAOD) methods employ domain adversarial training to align features and mitigate the domain gap. In this approach, a feature extractor is trained to deceive a domain classifier, thereby aligning feature distributions. However, the domain classifier's discrimination capability can easily fall into a local optimum due to the equilibrium challenge, hindering the effective training of the feature extractor. In this work, we propose an efficient optimization strategy called Virtual-label Fooled Domain Discrimination (VFDD), which revitalizes the domain classifier during training using virtual domain labels. Such virtual label makes the separable distributions less separable, and thus leads to a more easily confused domain classifier, which in turn further drives feature alignment. Particularly, we introduce a novel concept of virtual domain label for the unaligned samples and propose the VirtualH -divergence to overcome the problem of falling into local optimum due to the equilibrium challenge. VFDD is orthogonal to most existing DAOD methods and can be integrated as a plug-and-play module to enhance these models. Theoretical insights and experimental analyses demonstrate that VFDD improves many popular baselines and surpasses recent unsupervised DAOD models.
Wenxu Shi, Bochuan Zheng
ACM Multimedia2
2024 DIFNet: Dual-Domain Information Fusion Network for Image Denoising
Zedong Wu, Wenxu Shi, Liming Xu, Zicheng Ding, Bochuan Zheng
PRCV (8)7
2024 A dynamically class-wise weighting mechanism for unsupervised cross-domain object detection under universal scenarios
Wenxu Shi, Dailun Tan, Bochuan Zheng
Knowl. Based Syst.4
2024 Confused and disentangled distribution alignment for unsupervised universal adaptive object detection
Wenxu Shi, Zedong Wu, Bochuan Zheng
Knowl. Based Syst.4
2024 Deep Lifelong Cross-Modal Hashing
abstract
Hashing methods have made significant progress in cross-modal retrieval tasks with fast query speed and low storage cost. Among them, deep learning-based hashing achieves better performance on large-scale data due to its excellent extraction and representation ability for nonlinear heterogeneous features. However, there are still two main challenges in catastrophic forgetting when data with new categories arrive continuously, and time-consuming for non-continuous hashing retrieval to retrain for updating. To this end, we, in this paper, propose a novel deep lifelong cross-modal hashing to achieve lifelong hashing retrieval instead of re-training hash function repeatedly when new data arrive. Specifically, we design lifelong learning strategy to update hash functions by directly training the incremental data instead of retraining new hash functions using all the accumulated data, which significantly reduce training time. Then, we propose lifelong hashing loss to enable original hash codes participate in lifelong learning but remain invariant, and further preserve the similarity and dis-similarity among original and incremental hash codes to maintain performance. Additionally, considering distribution heterogeneity when new data arriving continuously, we introduce enhanced-semantic similarity to supervise hash learning, and it has been proven that the similarity improves performance with detailed analysis. Experimental results on benchmark datasets show that our proposed method achieves comparative performance comparing with recent state-of-the-art cross-modal hashing methods, and it yields substantial average increments over 20% in retrieval accuracy and almost reduces over 80% training time when new data arrives continuously.
Liming Xu, Bochuan Zheng, Weisheng Li 0001, Jiancheng Lv 0001
IEEE Trans. Circuits Syst. Video Technol.3
2024 CGFTrans: Cross-Modal Global Feature Fusion Transformer for Medical Report Generation
abstract
Medical report generation, as a cross-modal automatic text generation task, can be highly significant both in research and clinical fields. The core is to generate diagnosis reports in clinical language from medical images. However, several limitations persist, including a lack of global information, inadequate cross-modal fusion capabilities, and high computational demands. To address these issues, we propose cross-modal global feature fusion Transformer (CGFTrans) to extract global information meanwhile reduce computational strain. Firstly, we introduce mesh recurrent network to capture inter-layer information at different levels to address the absence of global features. Then, we design feature fusion decoder and define 'mid-fusion' strategy to separately fuse visual and global features with medical report embeddings, which enhances the ability of the cross-modal joint learning. Finally, we integrate shifted window attention into Transformer encoder to alleviate computational pressure and capture pathological information at multiple scales. Extensive experiments conducted on three datasets demonstrate that the proposed method achieves average increments of 2.9%, 1.5%, and 0.7% in terms of the BLEU-1, METEOR and ROUGE-L metrics, respectively. Besides, it achieves average increments -22.4% and 17.3% training time and images throughput, respectively.
Liming Xu, Quan Tang 0006, Bochuan Zheng, Jiancheng Lv 0001, Weisheng Li 0001
IEEE J. Biomed. Health Informatics3
2023 LoTE-Animal: A Long Time-span Dataset for Endangered Animal Behavior Understanding
abstract
Understanding and analyzing animal behavior is increasingly essential to protect endangered animal species. However, the application of advanced computer vision techniques in this regard is minimal, which boils down to lacking large and diverse datasets for training deep models. To break the deadlock, we present LoTE-Animal, a large-scale endangered animal dataset collected over 12 years, to foster the application of deep learning in rare species conservation. The collected data contains vast variations such as ecological seasons, weather conditions, periods, viewpoints, and habitat scenes. So far, we retrieved at least 500K videos and 1.2 million images. Specifically, we selected and annotated 11 endangered animals for behavior understanding, including 10K video sequences for the action recognition task, 28K images for object detection, instance segmentation, and pose estimation tasks. In addition, we gathered 7K web images of the same species as source domain data for the domain adaptation task. We provide evaluation results of representative vision understanding approaches and cross-domain experiments. LoTE-Animal dataset would facilitate the community to research more advanced machine learning models and explore new tasks to aid endangered animal conservation. Our dataset will be released with the paper. Our dataset can be found at https://LoTE-Animal.github.io
Jin Hou, Shaoli Huang, Bochuan Zheng, Jifeng Ning
ICCV6
2023 Deep image captioning: A review of methods, trends and future challenges
Liming Xu, Quan Tang 0006, Jiancheng Lv 0001, Bochuan Zheng, Weisheng Li 0001
Neurocomputing4
2023 MFGAN: Multi-modal Feature-fusion for CT Metal Artifact Reduction Using GANs
abstract
Due to the existence of metallic implants in certain patients, the Computed Tomography (CT) images from these patients are often corrupted by undesirable metal artifacts, which causes severe problem of metal artifact. Although many methods have been proposed to reduce metal artifact, reduction is still challenging and inadequate. Some reduced results are suffering from symptom variance, second artifact, and poor subjective evaluation. To address these, we propose a novel method based on generative adversarial nets (GANs) to reduce metal artifacts. Specifically, we firstly encode interactive information (text) and imaging CT (image) to yield multi-modal feature-fusion representation, which overcomes representative ability limitation of single-modal CT images. The incorporation of interaction information constrains feature generation, which ensures symptom consistency between corrected and target CT. Then, we design an enhancement network to avoid second artifact and enhance edge as well as suppress noise. Besides, three radiology physicians are invited to evaluate the corrected CT image. Experiments show that our method gains significant improvement over other methods. Objectively, ours achieves an average increment of 7.44% PSNR and 6.12% SSIM on two medical image datasets. Subjectively, ours outperforms others in comparison in term of sharpness, resolution, invariance, and acceptability.
Liming Xu, Weisheng Li 0001, Bochuan Zheng
ACM Trans. Multim. Comput. Commun. Appl.4
2022 Multi-Manifold Deep Discriminative Cross-Modal Hashing for Medical Image Retrieval
abstract
Benefitting from the low storage cost and high retrieval efficiency, hash learning has become a widely used retrieval technology to approximate nearest neighbors. Within it, the cross-modal medical hashing has attracted an increasing attention in facilitating efficiently clinical decision. However, there are still two main challenges in weak multi-manifold structure perseveration across multiple modalities and weak discriminability of hash code. Specifically, existing cross-modal hashing methods focus on pairwise relations within two modalities, and ignore underlying multi-manifold structures across over 2 modalities. Then, there is little consideration about discriminability, i.e., any pair of hash codes should be different. In this paper, we propose a novel hashing method named multi-manifold deep discriminative cross-modal hashing (MDDCH) for large-scale medical image retrieval. The key point is multi-modal manifold similarity which integrates multiple sub-manifolds defined on heterogeneous data to preserve correlation among instances, and it can be measured by three-step connection on corresponding hetero-manifold. Then, we propose discriminative item to make each hash code encoded by hash functions be different, which improves discriminative performance of hash code. Besides, we introduce Gaussian-binary Restricted Boltzmann Machine to directly output hash codes without using any continuous relaxation. Experiments on three benchmark datasets (AIBL, Brain and SPLP) show that our proposed MDDCH achieves comparative performance to recent state-of-the-art hashing methods. Additionally, diagnostic evaluation from professional physicians shows that all the retrieved medical images describe the same object and illness as the queried image.
Liming Xu, Bochuan Zheng, Weisheng Li 0001
IEEE Trans. Image Process.3
2021 Covariance Self-Attention Dual Path UNet for Rectal Tumor Segmentation
abstract
Deep learning algorithms are recognized as the most effective method for rectal tumor segmentation. However, since the multi-scale detailed feature information of rectal tumor cannot be fully extracted and applied, the segmentation and identification results of most algorithms are not always perfect. In this work, we introduce a Covariance Self-Attention Dual Path UNet (CSA-DPUNet), that is modified on the basis of UNet network and self-attention mechanism to improve network performance in feature processing and representation. The proposed network mainly makes two improvements. First, the UNet structure with single path is extended to dual paths (DPUNet). By broadening the network connections, our network is able to learn more local features with multiple contextual scales from CT images. Second, an improved criss-cross self-attention module is incorporated into DPUNet(CSA-DPUNet), instead of correlation method, we adopt covariance operation to calculate the attention weight map of self-attention mechanism, which can adaptively enhance feature combination and characterization ability. Experiments illustrate that our network called CSA-DPUNet can obviously improves the segmentation accuracy of rectal tumors, which brings 15.31%, 7.2%, 11.8%, and 9.5% improvement in Dice coefficient, P, R, F1, respectively compared with state-of-the-art. The above characteristics make the proposed CSA-DPUNet suitable for segmenting rectal tumor in practice.
Haijun Gao, Xiangyin Zeng, Dazhi Pan, Bochuan Zheng
ICRA4
2014 A winner-take-all Lotka-Volterra recurrent neural network with only one winner in each row and each column
Bochuan Zheng
Neural Comput. Appl.1
2013 Using competitive layer model implemented by Lotka-Volterra recurrent neural networks for detecting brain activated regions from fMRI data
Bochuan Zheng, Zhang Yi 0001
Neural Comput. Appl.1
2010 A Method for MRI Segmentation of Brain Tissue
Bochuan Zheng, Zhang Yi 0001
ISNN (2)1