Bocheng Ren

dblp:357/4004 · DBLP profile ↗
← Back
15ranked-venue papers
4as first author
15since 2021 · last 2026
0000-0002-6594-1421ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Efficient, Secure, Differentially Private Deep Learning in the Two-Server Model
abstract
Existing solutions on differentially private deep learning (DPDL) either require the assumption of a trusted data server (centralized DPDL) or suffer from poor utility (local DPDL); and hence their adoptions are hampered in real-world scenarios.We present CRYPTDP, a crypto-assisted differentially private deep learning approach in the two-server model. CRYPTDP employs two non-colluding servers to collaboratively and efficiently train differentially private deep learning over the secret shares of data owners' private data while protecting the confidentiality of the data from untrusted servers. CRYPTDP is the first approach with the best of both local DPDL and centralized DPDL models, which does not resort to trusted server like local DPDL and has the utility like centralized DPDL. In particular, we also make innovations for addressing the major challenges like poor performance and security that beset CRYPTDP: We introduce a new secure computation and differential privacy friendly activation function; we propose a novel garbled-circuits-free most significant bit extraction protocol, and using the protocol we propose an efficient and secure garbled-circuits-free protocol for activation function over secret shares. Exhaustive experiments show that CRYPTDP delivers significantly better performance than the state-of-the-art local DPDL, yields higher accuracy than the state-of-the-art centralized DPDL, and can achieve two orders of magnitude faster runtime than the state-of-the-art approach.
Jun Feng 0007, Pengfei Zhang 0010, Bocheng Ren, Shunli Zhang 0003
AAAI4
2026 Stabilizing Cross-Modal Bidirectional Attribution: Few-Shot Adversarial Prompt Tuning for Robust Vision-Language Models
abstract
Large-scale pre-trained vision-language models (VLMs) like CLIP show exceptional performance and zero-shot generalization. However, their reliability may be severely undermined by a critical vulnerability to subtle adversarial perturbations. Our work reveals a critical cross-modal vulnerability: visual-only perturbations induce substantial, synchronous shifts in decision attribution maps across both image and text. This phenomenon signifies a fundamental disruption of the VLM's internal logic, as it alters both the model's perceptual focus and its decision rationale. To counter this vulnerability, we introduce Cross-modal Bidirectional Attribution guided Few-shot Adversarial Prompt Tuning (CBA-FAPT), a novel method that leverages the model's internal decision rationale as a regularizer for robust learning. Our framework's core mechanism is the alignment of a novel bidirectional attribution map. This map is a unique fusion of two components. It combines forward feature attention to capture the model's perceptual focus. It also incorporates backward decision gradients to act as a proxy for the model's decision rationale, quantifying how each feature influences the final outcome. We enforce consistency on this bidirectional map between clean and adversarial examples. This approach corrects the model's internal logic on two fronts and effectively restores its adversarial robustness. Comprehensive experiments on 11 datasets demonstrate that CBA-FAPT outperforms the state-of-the-art, establishing a superior trade-off between robust and natural accuracy.
Jun Feng 0007, Shuhong Wu, Pengfei Zhang 0010, Bocheng Ren, Shunli Zhang 0003
AAAI5
2025 SADBA: Self-Adaptive Distributed Backdoor Attack Against Federated Learning
abstract
Backdoor attacks in federated learning (FL) face challenges such as lower attack success rates and compromised main task accuracy (MA) compared to local training. Existing methods like distributed backdoor attack (DBA) mitigate these issues by modifying malicious clients’ updates and partitioning global triggers to enhance backdoor persistence and stealth. The recent full combination backdoor attack (FCBA) further improves backdoor efficiency with a full combination strategy. However, these methods are mainly applicable in small-scale FL. In large-scale FL, small trigger patterns weaken impact, and scaling them requires controlling exponentially more clients, which poses significant challenges, while simply reverting to DBA may decrease backdoor performance. To overcome these challenges, we propose the self-adaptive distributed backdoor attack (SADBA), which achieves similar performance to FCBA with a lower percentage of malicious clients (PMC). It also adapts more flexibly through an optimized model poisoning strategy and a self-adaptive data poisoning strategy. Experiments demonstrate SADBA outperforms state-of-the-art methods, achieving higher or comparable backdoor performance and MA across various datasets with limited PMC.
Jun Feng 0007, Yuzhe Lai, Bocheng Ren
AAAI4
2025 Factorization-based Attribute Residual Summary for Adaptive Edge-based Autonomous System Security
abstract
Due to the particularity of the marginal environment, edge-based autonomous systems face significant risks associated with security operations. Traffic anomaly detection in edge-based autonomous systems has become increasingly crucial for ensuring the security of these systems. Existing works lack consideration of the relationship between traffic attributes and anomaly types. In particular, existing solutions struggle with detecting anomalies that primarily manifest statistical signs in only a few attributes. To address this, we propose a nonnegative factorization-based attribute residual summary and a nonparametric statistic framework for adaptive security monitoring in edge-based autonomous systems. Specifically, the nonnegative factorization, which depends on the multiplicative update rules, is introduced to extract attribute features. Using the tensor linear representation, the attribute residual summary is built, which depicts the statistic discrepancy well even if only a few of traffic attributes are affected, to implement adaptive security monitoring for various attacks in edge-based autonomous systems. Then, a nonparametric statistic framework is developed, which achieves the real-time detection by accumulating and comparing each statistic evidence. Extensive experiments with real-world traffic trace datasets validate the adaptivity, accuracy, real-time performance, and superiority of our method, particularly in dealing with anomalies that exhibit statistical signs in only a few traffic attributes.
Jiuzhen Zeng, Laurence T. Yang, Chao Wang 0014, Bocheng Ren, Honglu Zhao
ACM Trans. Auton. Adapt. Syst.5
2025 Zero-Shot Recognition for Healthcare Social Networks via Tensor-Based Vision-Semantic Manifold Alignment
abstract
Healthcare social networks (HSNs) are pivotal in spreading healthcare knowledge, providing support to both potential patients and medical professionals, and enhancing healthcare services. However, identifying unseen data in HSN poses a significant challenge due to their intrinsic heterogeneity, dynamic characteristics, and the scarcity of labeled data. Employing semantic knowledge transfer for class-agnostic zero-shot recognition stands out as a promising and innovative solution to this problem, but the visual-semantic gap and domain shift problems considerably hinder advancements in zero-shot recognition capabilities. Previous zero-shot models often impose constraints between vision and semantics in the loss part without explicitly injecting intermodality guidance into the feature refinement process. This article yields a novel zero-shot recognition framework for HSN, named the dual tensor prototype graph network, devoted to improving the performance of recognizing unseen objects in HSN leveraging semantic knowledge. We have developed an iterative and interactive updating strategy for dual tensor prototype graphs, explicitly leveraging the distribution information from one modality to guide the prototype graph updates of another modality. We constrain the update process of the dual prototype graphs by several tailored loss functions and episodic training, alleviating the inconsistency between semantic and visual manifolds. Extensive comparative experiments conducted on two medical imaging datasets and five zero-shot benchmarks affirm the stronger generalization ability of our proposed method compared with other advanced approaches, showing the potential of addressing zero-shot problems in HSN.
Bocheng Ren, Yuanyuan Yi, Laurence T. Yang, Zecan Yang, Jun Feng 0007
IEEE Trans. Comput. Soc. Syst.1
2025 Zero-Shot Image Recognition via Learning Dual Prototype Accordance Across Meta-Domains
abstract
Zero-shot learning (ZSL) aims to recognize unseen classes by transferring semantic knowledge from seen categories. However, existing methods often struggle with the persistent semantic gap caused by limited semantic descriptors and rigid visual feature modeling. In particular, modeling pre-defined class-level attribute descriptions as ground truth hinders effective semantic-to-visual alignment to some extent. To mitigate these issues, we propose the Bilateral-guided Prototype Refinement Network (BPRN), a novel ZSL framework designed to refine dual prototypes across meta-domains of varying scales. Specifically, we first disentangle the relationships among class-level semantics and use them to generate corresponding pseudo-visual prototypes. Then, by leveraging distribution information across dual prototypes in different meta-domains, BPRN achieves bidirectional calibration between visual-to-semantic and semantic-to-visual modalities. Finally, a synthesized class-level representation derived from the refined dual prototypes is employed for inference, instead of relying on a single prototype. Extensive experiments conducted on five widely-used ZSL benchmark datasets demonstrate that BPRN consistently achieves competitive or even superior performance. Specifically, in the GZSL scenario, BPRN shows improvements of 2.1%, 7.3%, 6.1%, and 4.8% on AWA1, AWA2, SUN, and aPY, respectively, compared to existing embedding-based ZSL methods. Ablation studies and visualization analyses further validate the effectiveness of the proposed components.
Bocheng Ren, Yuanyuan Yi, Qingchen Zhang 0001, Debin Liu
IEEE Trans. Image Process.1
2025 Zero-Shot Fault Diagnosis for Smart Process Manufacturing via Tensor Prototype Alignment
abstract
Identifying unseen faults is a crux of the digital transformation of process manufacturing. The ever-changing manufacturing process requires preset models to cope with unseen problems. However, most current works focus on recognizing objects seen during the training phase. Conventional zero-shot recognition methods perform poorly when they are applied directly to these tasks due to the different scenarios and limited generalizability. This article yields a tensor-based zero-shot fault diagnosis framework, termed MetaEvolver, which is dedicated to improving fault diagnosis accuracy and unseen domain generalizability for practical process manufacturing scenarios. MetaEvolver learns to evolve the dual prototype distributions for each uncertain meta-domain from seen faults and then adapt to unseen faults. We first propose the concept of the uncertain meta-domain and then construct corresponding sample prototypes with the guidance of class-level attributes, which produce the sample-attribute alignment at the prototype level. MetaEvolver further collaboratively evolves the uncertain meta-domain dual prototypes by injecting the prototype distribution information of another modality, boosting the sample-attribute alignment at the distribution level. Building on the uncertain meta-domain strategy, MetaEvolver is prone to achieving knowledge transferring and unseen domain generalization with the optimization of several devised loss functions. Comprehensive experimental results on five process manufacturing data groups and five zero-shot benchmarks demonstrate that our MetaEvolver has great superiority and potential to tackle zero-shot fault diagnosis for smart process manufacturing.
Bocheng Ren, Laurence T. Yang, Jun Feng 0007, Xianjun Deng, Chenlu Zhu
IEEE Trans. Neural Networks Learn. Syst.1
2024 General Point Model Pretraining with Autoencoding and Autoregressive
abstract
The pre-training architectures of large language models encompass various types, including autoencoding models, autoregressive models, and encoder-decoder models. We posit that any modality can potentially benefit from a large language model, as long as it undergoes vector quantization to become discrete tokens. Inspired by the General Language Model, we propose a General Point Model (GPM) that seamlessly integrates autoencoding and autoregressive tasks in a point cloud transformer. This model is versatile, allowing fine-tuning for downstream point cloud representation tasks, as well as unconditional and conditional generation tasks. GPM enhances masked prediction in autoencoding through various forms of mask padding tasks, leading to improved performance in point cloud understanding. Additionally, GPM demonstrates highly competitive results in unconditional point cloud generation tasks, even exhibiting the potential for conditional generation tasks by modifying the input's conditional information. Compared to models like Point-BERT, MaskPoint. and PointMAE, our GPM achieves superior performance in point cloud understanding tasks. Furthermore, the integration of autoregressive and autoencoding within the same transformer underscores its versatility across different downstream tasks. Codes are available at https://github.com/gentlefress/GPM
Zhe Li 0038, Zhangyang Gao, Cheng Tan 0012, Bocheng Ren, Laurence T. Yang, Stan Z. Li
CVPR4
2024 MLIP: Enhancing Medical Visual Representation with Divergence Encoder and Knowledge-guided Contrastive Learning
abstract
The scarcity of annotated data has sparked signifi-cant interest in unsupervised pre-training methods that leverage medical reports as auxiliary signals for medi-cal visual representation learning. However, existing re-search overlooks the multi-granularity nature of medical visual representation and lacks suitable contrastive learning techniques to improve the models' generalizability across different granularities, leading to the underutilization of image-text information. To address this, we pro-pose MLIP, a novel framework leveraging domain-specific medical knowledge as guiding signals to integrate language information into the visual domain through image-text contrastive learning. Our model includes global contrastive learning with our designed divergence encoder, lo-cal token-knowledge-patch alignment contrastive learning, and knowledge-guided category-level contrastive learning with expert knowledge. Experimental evaluations reveal the efficacy of our model in enhancing transfer performance for tasks such as image classification, object detection, and semantic segmentation. Notably, MLIP surpasses state-of-the-art methods even with limited annotated data, highlighting the potential of multimodal pre-training in advancing medical representation learning.11Codes are available at https://github.com/gentlefress/MLIP
Zhe Li 0038, Laurence T. Yang, Bocheng Ren, Zhangyang Gao, Cheng Tan 0012, Stan Z. Li
CVPR3
2024 Capturing Detail Variations for Lightweight Neural Radiance Fields
abstract
Neural Radiance Fields (NeRF) has recently overhauled novel view synthesis, but it requires extensive computations for training and captures variations in detail with difficulty. In this paper, we propose a novel framework, termed CD-TDRF, to mitigate these dilemmas. CD-TDRF factorizes a density voxel grid into a core tensor and three matrices via Tucker decomposition, reducing memory usage and accelerating training. To better capture variations in complex scenes, CD-TDRF uses a fully convolutional network to extract prior information from the training images. Moreover, three learnable appearance planes are constructed to preserve information about scene details, which enhances the rendering quality significantly. Our experimental results demonstrate that CD-TDRF has achieved competitive rendering quality on three popular datasets and speeds up training compared with traditional NeRF models.
Laurence T. Yang, Bocheng Ren, Jinglin Zhao, Zhe Li 0038, Guolei Zeng
ICASSP3
2024 Sparse Bayesian Tensor Completion for Data Recovery in Intelligent IoT Systems
abstract
Intelligent Internet of Things (IoT), is an emerging paradigm that integrates lightweight intelligence algorithms to various IoT devices to provide convenient and intelligent services for modern life and production. For this purpose, data should be efficiently processed to explore the hidden information to elevate the intelligence of services. However, the IoT data are collected from a complex environment with high speed, and high noise, which inevitably brings problems about missing and imparting challenges to the progression of intelligent IoT services. To recover the missing data with higher precision and provide data cornerstones for intelligent IoT systems, a sparse Bayesian tensor completion (SBTC) method is proposed in this article. With the hierarchical sparse prior, the proposed tensor completion model can obtain the underlying low-rank structure from the incomplete tensor, thereby recovering missing data with high accuracy. For model learning, a variational Bayesian inference method is developed in the frequency domain, which improves the model’s efficiency. The model proposed is within a fully Bayesian framework, thereby endowing the model with commendable robustness. The superiority of our model is fully demonstrated by comparing other state-of-the-art methods on synthetic data, traffic data, logistics data, and visual data. In particular, on traffic data and video data, our method has improved by at least 2% and 10dB.
Honglu Zhao, Laurence T. Yang, Zecan Yang, Debin Liu, Bocheng Ren
IEEE Internet Things J.6
2024 Prototype rectification for zero-shot learning
Yuanyuan Yi, Guolei Zeng, Bocheng Ren, Laurence T. Yang, Bin Chai
Pattern Recognit.3
2024 Tensor Recurrent Neural Network With Differential Privacy
abstract
Recurrent neural network (RNN), a branch of deep learning, is a powerful model for sequential data that has outstanding performance on a wide range of important Internet of Things (IoT) tasks. This unprecedented growth of RNN model has however encountered both heterogeneous IoT data and privacy issues. Existing RNN model can not deal with heterogeneous sequential data; often the larger datasets used in training of RNN model contain sensitive information. To tackle these challenges and for the first time, this research proposes a novel differentially private tensor-based RNN (DPTRNN) that can be applied in many challenging deep learning sequence tasks for IoT systems. Specifically, to process heterogeneous sequential data, we propose a tensor-based RNN model. To guarantee privacy, we develop a tensor-based back-propagation through time algorithm with perturbation to avoid exposing the sensitive information for training the tensor-based RNN model within the framework of differential privacy. Thorough security analysis shows that the differential private tensor-based RNN efficiently protects the confidentiality of sensitive user information for IoT. Our results from extensive experiments on two challenging large video datasets suggest that our proposed scheme is practical with guarantee of data privacy preservation and acceptable accuracy loss.
Jun Feng 0007, Laurence T. Yang, Bocheng Ren, Deqing Zou, Mianxiong Dong, Shunli Zhang 0003
IEEE Trans. Computers3
2023 Enhancing Sentence Representation with Visually-supervised Multimodal Pre-training
abstract
Large-scale pre-trained language models have garnered significant attention in recent years due to their effectiveness in extracting sentence representations. However, most pre-trained models currently use transformer-based encoder with a single modality and are primarily designed for specific tasks such as natural language inference and question-answering. Unfortunately, this approach neglects the complementary information provided by multimodal data, which can enhance the effectiveness of sentence representation. To address this issue, we propose a Visually-supervised Pre-trained Multimodal Model (ViP) for sentence representation. Our model leverages diverse label-free multimodal proxy tasks to embed visual information into language, facilitating effective modality alignment and complementarity exploration. Additionally, our model utilizes a novel approach to distinguish highly similar negative and positive samples. We conduct comprehensive downstream experiments on natural language understanding and sentiment classification, demonstrating that ViP outperforms both existing unimodal and multimodal pre-trained models. Our contributions include a novel approach to multimodal pre-training and a state-of-the-art model for sentence representation that incorporates visual information.1 Our code is available at https://github.com/gentlefress/ViP
Zhe Li 0038, Laurence T. Yang, Bocheng Ren, Xianjun Deng
ACM Multimedia4
2023 Tensor-Empowered Adaptive Learning for Few-Shot Streaming Tasks
abstract
Various stream learning methods are emerging in an endless stream to provide a wealth of solutions for artificial intelligence in streaming data scenarios. However, when each data stream is oriented to a different target space, it forces stream learning approaches oriented to the same task to be no longer applicable. Due to inconsistent target spaces for different tasks, the previous approaches fail on the new streaming tasks or it is impracticable to be trained from scratch with few labeled samples at the beginning. To this end, we have proposed an adaptive learning scheme for few-shot streaming tasks with the contributions of tensor and meta-learning. This adaptive scheme is conducive to mitigating the domain shift when a new task has few labeled samples. We elaborate a novel tensor-empowered attention mechanism derived from nonlocal neural networks, which enables to capture long-range dependency and preserve the high-dimensional structure to refine the global features of streaming tasks. Furthermore, we develop a fine-grained similarity computing approach, which is prone to better characterize the difference across few-shot streaming tasks. To show the superiority of our method, we have carried out extensive experiments on three popular few-shot datasets to simulate streaming tasks and evaluate the performance of adaptation. The results show that our proposed method has achieved competitive performance for few-shot streaming tasks compared with the state-of-the-art (SOTA).
Bocheng Ren, Laurence T. Yang, Qingchen Zhang 0001, Jun Feng 0007
IEEE Trans. Neural Networks Learn. Syst.1