Pengyang Li

dblp:210/5359 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RDA-CAFL: Reputation-Dynamic and Distillation-based Asynchronous Conflict-Aware Federated Learning for UAV Networks
Pengyang Li, Xingwei Wang
CCGrid2
2026 Visual Information Facilitation Scene Text Retrieval
Mayire Ibrayim, Hailong Luo, Pengyang Li
ICPR (9)3
2026 Adaptive dynamic graph interaction network for fine-grained image-text retrieval
Pengyang Li, Mayire Ibrayim, Alkut Mardan, Peichao Jiang
Knowl. Based Syst.1
2025 Iterative Self-Training with Class-Aware Text-to-Image Synthesis for Visual Task Learning
abstract
Generative models are widely used to produce synthetic images with annotations, alleviating the burden of image collection and annotation for training deep visual models. However, challenges such as limited image diversity, noisy pseudo labels, and domain gaps between synthetic and real images often undermine their effectiveness in downstream visual tasks. This paper introduces the Iterative Self-Training with Class-Aware Text-to-Image Synthesis (IST-CATS) framework, which addresses these challenges by integrating a class-aware text-to-image synthesis (CATS) component with an iterative self-training (IST) strategy. CATS innovatively introduces a class-aware chain approach to generate detailed descriptions. These descriptions act as prompts for a diffusion model, enabling the creation of a diverse of images accompanied by distinguishable objects against the background. The generated images can be easily pseudo-labeled by an unsupervised instance segmentation method, and then noisy pseudo labels can be effectively purified by a novel feature similarity-based filtering mechanism. The generated images underpin our IST, which progressively enhances vision models and refines pseudo labels through self-training and our proposed label filtering strategy (LabFilt). LabFilt meticulously improves the quality of pseudo labels by employing class-adaptive techniques at both the pixel and object levels, ensuring refined pseudo-label accuracy. IST-CATS demonstrates superior performance in object detection and semantic segmentation compared to traditional synthetic and semi/weakly-supervised methods, effectively addressing data collection and annotation challenges.
Xiang Zhang 0018, Wanqing Zhao, Pengyang Li, Hangzai Luo, Sheng Zhong 0006, Jinye Peng 0001, Jianping Fan 0001
AAAI3
2025 DPAcc: An FPGA-based Differential Privacy Acceleration Framework
abstract
In the data-driven era, privacy protection has become a critical concern. Differential privacy is an effective technique that incorporates random noise during data processing to ensure that alterations to individual data points do not significantly affect overall outputs. However, the additional operations required by differential privacy can result in prolonged training times and degraded model performance. This work proposes DPAcc, an FPGA-based acceleration framework for differential privacy that utilizes hardware implementation to decrease training time. The designed FPGA module efficiently executes clipping and noise addition operations, significantly reducing the overhead compared to standard training. Experimental results demonstrate that DPAcc improves training efficiency across multiple models, achieving up to 2× speedup compared to standard differential privacy training methods.
Ao Dong, Pengyang Li, Yifei Tian, Xiaobai Chen, Jieming Yin
ISCAS3
2025 FlexAcc: Accelerating Batch Normalization through GPU-FPGA Integration
abstract
Convolutional neural networks are fundamental to deep learning, especially in computer vision. However, their computational demands, particularly during batch normalization, create significant inefficiencies due to excessive data movement between memory and processing units. To address this, we propose FlexAcc, a novel architecture that integrates GPUs and FPGAs to offload BN computations to FPGAs, reducing data movement and improving hardware utilization. FlexAcc accelerates end-to-end training performance by up to 1.1× across various models. This approach bridges the performance gap between convolutional and non-convolutional layers, advancing deep learning model deployment.
Haishuai Zhang, Pengyang Li, Xuehuai Shi, Xiaobai Chen, Jieming Yin
ISCAS3
2024 Stratifying heart failure patients with graph neural network and transformer using Electronic Health Records to optimize drug response prediction
abstract
OBJECTIVES: Heart failure (HF) impacts millions of patients worldwide, yet the variability in treatment responses remains a major challenge for healthcare professionals. The current treatment strategies, largely derived from population based evidence, often fail to consider the unique characteristics of individual patients, resulting in suboptimal outcomes. This study aims to develop computational models that are patient-specific in predicting treatment outcomes, by utilizing a large Electronic Health Records (EHR) database. The goal is to improve drug response predictions by identifying specific HF patient subgroups that are likely to benefit from existing HF medications. MATERIALS AND METHODS: A novel, graph-based model capable of predicting treatment responses, combining Graph Neural Network and Transformer was developed. This method differs from conventional approaches by transforming a patient's EHR data into a graph structure. By defining patient subgroups based on this representation via K-Means Clustering, we were able to enhance the performance of drug response predictions. RESULTS: Leveraging EHR data from 11 627 Mayo Clinic HF patients, our model significantly outperformed traditional models in predicting drug response using NT-proBNP as a HF biomarker across five medication categories (best RMSE of 0.0043). Four distinct patient subgroups were identified with differential characteristics and outcomes, demonstrating superior predictive capabilities over existing HF subtypes (best mean RMSE of 0.0032). DISCUSSION: These results highlight the power of graph-based modeling of EHR in improving HF treatment strategies. The stratification of patients sheds light on particular patient segments that could benefit more significantly from tailored response predictions. CONCLUSIONS: Longitudinal EHR data have the potential to enhance personalized prognostic predictions through the application of graph-based AI techniques.
Shaika Chowdhury, Yongbin Chen, Pengyang Li, Sivaraman Rajaganapathy, Andrew Wen, Xiao Ma 0019, Qiying Dai, Yue Yu 0012, Sunyang Fu, Xiaoqian Jiang, Zhe He 0001, Sunghwan Sohn, Xiaoke Liu, Suzette J. Bielinski, Alanna M. Chamberlain, James R. Cerhan, Nansu Zong
J. Am. Medical Informatics Assoc.3
2022 Transform Cold-Start Users into Warm via Fused Behaviors in Large-Scale Recommendation
abstract
Recommendation for cold-start users who have very limited data is a canonical challenge in recommender systems. Existing deep recommender systems utilize user content features and behaviors to produce personalized recommendations, yet often face significant performance degradation on cold-start users compared to existing ones due to the following challenges: (1) Cold-start users may have a quite different distribution of features from existing users. (2) The few behaviors of cold-start users are hard to be exploited. In this paper, we propose a recommender system called Cold-Transformer to alleviate these problems. Specifically, we design context-based Embedding Adaption to offset the differences in feature distribution. It transforms the embedding of cold-start users into a warm state that is more like existing ones to represent corresponding user preferences. Furthermore, to exploit the few behaviors of cold-start users and characterize the user context, we propose Label Encoding that models Fused Behaviors of positive and negative feedback simultaneously, which are relatively more sufficient. Last, to perform large-scale industrial recommendations, we keep the two-tower architecture that de-couples user and target item. Extensive experiments on public and industrial datasets show that Cold-Transformer significantly outperforms state-of-the-art methods, including those that are deep coupled and less scalable.
Pengyang Li, Quan Liu 0008, Jian Xu 0015, Bo Zheng 0007
SIGIR1
2021 Inference Fusion with Associative Semantics for Unseen Object Detection
abstract
We study the problem of object detection when training and test objects are disjoint, i.e. no training examples of the target classes are available. Existing unseen object detection approaches usually combine generic detection frameworks with a single-path unseen classifier, by aligning object regions with semantic class embeddings. In this paper, inspired from human cognitive experience, we propose a simple but effective dual-path detection model that further explores associative semantics to supplement the basic visual-semantic knowledge transfer. We use a novel target-centric multiple-association strategy to establish concept associations, to ensure that the predictor generalized to unseen domain can be learned during training. In this way, through a reasonable inference fusion mechanism, those two parallel reasoning paths can strengthen the correlation between seen and unseen objects, thus improving detection performance. Experiments show that our inductive method can significantly boost the performance by 7.42% over inductive models, and even 5.25% over transductive models on MSCOCO dataset.
Yanan Li 0002, Pengyang Li
AAAI2
2013 Dezert-Smarandache theory for multiple targets tracking in natural environment
abstract
The aim of this article was to investigate multiple targets tracking in natural environment based on Dezert–Smarandache theory (DSmT). On the basis of establishing conflict strategy and combination model, the basic framework and algorithm of fusing multi‐source information were described. The multiple targets tracking platform which embedded location and colour cues into the particle filters (PFs) was developed in the framework of DSmT. Three sets of experiments with comparisons were carried out to validate the suggested tracking approach. Results showed that the conflict strategy and DSmT combination model were available, and the introduced approach exhibited a significantly better performance for dealing with high conflict between evidences than a PF. As a result, the approach was suitable for real‐time video‐based targets tracking, and it had the ability to track interesting targets. Furthermore, the approach can easily be generalised to deal with larger number of targets and additional cues in a complicated environment.
Yingwu Fang, Pengyang Li
IET Comput. Vis.4