VLDB 2026 Research / reviewers in the wild / expert
Yongqi Sun
dblp:19/8652
· DBLP profile ↗
18ranked-venue papers
1as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 2 · 2 since 2021Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FastPFRec: A fast personalized federated recommendation with secure sharing
Zhenxing Yan, Jidong Yuan, Yongqi Sun, Zhihui Gao |
Expert Syst. Appl. | 3 |
| 2026 | HFAVL: Hard-Label Fusion Attack Toward Vision-Language ModelabstractWith the advancement of artificial intelligence, vision-language models (VLMs) that integrate text and image modalities have become central to multimodal learning and are increasingly deployed in Internet of Things (IoT) environments such as smart surveillance, autonomous driving and industrial monitoring. However, VLMs are also highly susceptible to adversarial attacks. Most previous black-box attack methods towards the VLMs rely on either the soft label or substitute models. However, most of them need detailed information on the target model, which is often unavailable for real-world threat models due to limitations on additional access and resource restrictions. In this study, we propose a surrogate-free hard-label attack towards the VLMs, which does not require a substitute model, called Hard-Label Fusion Attack towards Vision-Language models. HFAVL directly leverages the model’s final decision to jointly generate the perturbation for both text and image modalities instead of the confidence score necessary for the soft label attack. By transforming the text into a continuous embedding space that enables them to be aligned with the image in a unified manner, HFAVL applies cooperative Monte Carlo-based estimation on the multimodal gradient. Aiming to improve the efficiency of attacks, we introduce a Mutual Iterative Refinement mechanism (MIR), which gradually searches for a better adversarial example through small iterative steps that blend perturbations across both modalities. Additionally, to alleviate the imbalance between textual and visual noise during the attack, we propose a new multimodal edit distance evaluation metric to measure the quality of multimodal adversarial examples. Extensive experimental results demonstrate that HFAVL consistently outperforms existing multimodal adversarial attack methods across various vision-language pretraining (VLP) models. Notably, under the setting with recall@10, HFAVL achieves over 80% attack success rate on most benchmarks, demonstrating its efficiency and high precision in generating effective adversarial examples. Hongbo Cao, Yongqi Sun, Yifan Sui, Xisu Wang |
IEEE Internet Things J. | 2 |
| 2026 | P3D: Plug-and-play prompt-driven framework for RGB-thermal semantic segmentationabstract• A plug-and-play prompt-driven framework for RGB-thermal image semantic segmentation. • LoRA-based fine-tuning strategy for SAM series model integration. • A model-agnostic encoder to generate statistical distributed prompts for training. The semantic segmentation of RGB-thermal images is critical for applications with low-light conditions. Existing works primarily focus on feature fusion strategies and model design to enhance performance. While Visual Foundation Models (VFMs) have been introduced in previous studies to improve generalization and segmentation accuracy, they suffer from poor compatibility with other models thus requiring full model retraining. Additionally, the domain gap and modality gap between VFM pre-training datasets and RGB-thermal semantic segmentation datasets pose significant challenges to VFM adaptation for downstream tasks. To address these issues, in this paper a plug-and-play prompt driven framework P 3 D is proposed. Unlike existing VFM-based methods that require complete retraining for each specific architecture, P 3 D is designed with a model-agnostic training strategy that enables one-time training and seamless integration with various existing methods without requiring retraining. First, a dual-branch LoRA (Low-Rank Adaptation) fine-tuned (DBLF) image encoder for the RGB and thermal image branches is proposed to narrow the domain gap and modality gap when incorporating SAM series models into our task. Second, a unified prompt generation and representation (UPGR) encoder is proposed. It generates diverse prompts using semantic labels during the training stage, ensuring the generated prompts are model-agnostic and compatible with existing methods. Finally, a cross-modality spatial-channel attention (CM-SCA) decoder is developed to fuse the embeddings from two-modality images and prompts for the final prediction. Extensive experiments are conducted on three popular benchmarks. Results demonstrate that P 3 D not only improves the performance of existing models but also outperforms current state-of-the-art (SOTA) methods leveraging < 1% trainable parameters. More importantly, by simply plugging P 3 D into existing methods, we consistently achieve significant performance improvements without retraining these base models, demonstrating the practical value of our plug-and-play design. Yongqi Sun, Chenguang Dai, Hanyun Wang, Longguang Wang, Wenke Li, Anzhu Yu |
Pattern Recognit. | 1 |
| 2026 | SGAN: Shapelet-based GAN for local time series black-box attack
Kaiyu Yang, Jidong Yuan, Haoyu Yan, Yongqi Sun |
Pattern Recognit. | 5 |
| 2026 | Backdoor Detection in Federated Learning With Feature Map: A Multi-Task Learning PerspectiveabstractBackdoor attacks pose severe security challenges to federated learning systems due to their stealthy nature. Existing detection methods primarily focus on identifying anomalies by analyzing discrepancies in client model updates. However, in federated learning, the non-independent and identically distributed (non-IID) nature of client data leads to inconsistencies among local model updates, which can mask the distinguishing features of backdoor attacks and consequently degrade the performance of detection methods. Unlike benign models, which are trained solely for a single classification task, backdoored models are simultaneously optimized for both the classification (main) task and the backdoor task. Therefore, training the backdoored models can be regarded as a multi-task learning problem. Inspired by information bottleneck theory, we observe that backdoored models exhibit more stable feature representations than benign models when performing the main task. Based on this insight, we propose a novel stability metric that quantitatively captures the disparity in feature map stability between backdoored and benign models. Leveraging this metric, we develop a new backdoor detection framework for federated learning. Our method computes anomaly scores for each client and selectively aggregates models with benign characteristics, effectively defending against backdoor attacks. We validate our approach through extensive experiments on multiple benchmark datasets under non-IID settings. The results demonstrate that our method consistently achieves high detection performance across a range of backdoor scenarios and data heterogeneity levels. Yifan Sui, Yongqi Sun, Naiyue Chen, Hongbo Cao, Baomin Xu |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2025 | A Language-Assisted Semantic-Aware Disentangled Method for Link Prediction on Heterogeneous GraphsabstractLink prediction serves as a fundamental task in graph-based applications, where graph neural networks (GNNs) are extensively applied to estimate node connectivity likelihood. However, GNN-based methods for homogeneous graphs encounter the semantic mixing issue in heterogeneous graphs. Previous works leverage the disentangled-based model to separate the semantic information into different factors and conduct the message passing for link prediction. However, their models suffer from information loss and inadequate expression, which harms link prediction performance. To address these limitations, we propose a language-assisted semantic-aware disentangled method for link prediction on heterogeneous graphs. First, we employ a factor-wise attention mechanism to reduce the information loss caused by the disentangled model. Specifically, we design a factor selection strategy to select disentangled factors and combine them to utilize more semantic information. Second, a language-graph learning method is developed to enhance contextual expression by fusing the features of nodes and edge textual information. Extensive experiments show that the proposed method outperforms existing state-of-the-art baselines. Rongqiang Fang, Yongqi Sun, Jidong Yuan, Hongbo Cao, Jinkun Dong |
ACM Multimedia | 2 |
| 2025 | DMIA: A Disentangled-Based Method for Graph Convolutional Network Against Membership Inference AttackabstractAs a well-known graph embedding method, Graph Convolutional Networks (GCNs) have been widely applied to recommendation systems and social media analysis, in which privacy concerns regarding sensitive data have emerged in the public view due to the collection of personal preferences. Although the regularization methods are introduced to improve the network's security, the GCN tends to memorize individual user information in latent representations susceptible to the Membership Inference Attack (MIA). In addition, the previous works focus on improving the security while hurting the utility, or vice versa, which induces “negative transfer”. In this paper, we propose a novel disentangled-based framework to defend MIA and alleviate the issue of negative transfer in multi-task learning. First, we divide the sensitive and practical channels from the latent representations of graph nodes to minimize their linear dependency. Then, to effectively train our model, we employ the sub-computational graphs to generate local gradients for different tasks and allocate losses to them. Finally, we propose a novel mixed updating strategy to accumulate the updating information of sub-computational graphs. Extensive experiments show that the proposed method can mitigate the risk of membership inference while ensuring model accuracy. Rongqiang Fang, Yongqi Sun, Jidong Yuan, Hongbo Cao |
IEEE Trans. Dependable Secur. Comput. | 2 |
| 2025 | FedFOC: Personalized Federated Learning via Fine-Grained One-Shot ClusteringabstractThe rapid growth of industrial Internet of Things (IoT) devices presents significant opportunities for federated learning to advance secure machine learning in industrial IoT systems. However, the performance of global federated learning models deteriorates significantly due to data heterogeneity inherent in industrial IoT systems. To address this issue, we propose a novel personalized federated learning framework via fine-grained one-shot clustering (FedFOC). Our method simultaneously considers the dataset structure and the relative positions of samples to extract compact features for clustering. Specifically, we combine clients’ locally calculated sample-level distances with MinHash signatures of their label sets to transform their data distributions into 1-D vectors to reduce communication costs. These vectors are then clustered to identify the fine-grained similarities between clients’ data distributions. To capture the complex relationships among local data distributions, we perform soft clustering in a one-shot manner and quantify client cluster association strengths. Subsequently, the clustering results are utilized to improve the local models’ personalized performance. Extensive experiments conducted on four widely used public datasets under various heterogeneous scenarios demonstrate that FedFOC outperforms state-of-the-art personalized federated learning methods in the accuracy of local models in almost all scenarios partitioned by Dirichlet. In addition, we also validate the effectiveness of our approach on the industrial dataset, indicating that our method has potential value in real-world applications. Yongqi Sun, Naiyue Chen, Yifan Sui, Hongbo Cao |
IEEE Trans. Ind. Informatics | 2 |
| 2024 | Local perturbation-based black-box federated learning attack for time series classification
Shengbo Chen, Jidong Yuan, Yongqi Sun |
Future Gener. Comput. Syst. | 4 |
| 2024 | D³STN: Dynamic Delay Differential Equation Spatiotemporal Network for Traffic Flow ForecastingabstractTraffic flow forecasting is crucial for intelligent transportation systems. Currently, most models need to pay more attention to the delay (history) state to improve forecasting performance. In this paper, we propose a dynamic delay differential equation spatiotemporal network for traffic flow forecasting, named D3STN. First, our method integrates dynamic delay state optimization into delay differential equations to enhance delay state inputs and model forecasting performance. In addition, we propose a hybrid graph neural network and convolution multi-head attention mechanism. With the hybrid graph neural network, our method takes the multi-granularity correlation relationship into account to capture spatial characteristics from relevant nodes. With the convolution multi-head attention mechanism, our method balances the attention distribution between short-term and long-term attention. Empirical experiments are executed on highway traffic flow and metro flow datasets to evaluate the performance of our method. The results demonstrate that D3STN achieves significant advancements in traffic flow forecasting tasks. Compared with CorrSTN (baseline), D3STN makes improvements of 10.93%, 16.24% and 24.56% in terms of the MAE, RMSE and MAPE, respectively, on the HZME (outflow) dataset. Weiguo Zhu, Caiyuan Liu, Yongqi Sun |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Real-Time Adaptive Partition and Resource Allocation for Multi-User End-Cloud Inference Collaboration in Mobile EnvironmentabstractThe deployment of Deep Neural Networks (DNNs) requires significant computational and storage resources, which is challenging for resource-constrained end devices. To this end, collaborative deep inference is proposed, in which the DNN is divided into two parts and executed on the end device and cloud respectively. The selection of DNN partition point is the key challenge to realize end-cloud collaborative deep inference, especially in mobile environments with unstable networks. In this paper, we propose a Real-time Adaptive Partition (RAP) framework, in which a fast split point decision algorithm is proposed to realize real-time adaptive DNN model partition in the mobile network. A weighted joint optimization of DNN quantization loss, inference and transmission latency is performed. We further propose a Joint Multi-user Model Partition and Resource Allocation (JM-MPRA) algorithm under RAP framework. JM-MPRA aims to guarantee the optimized latency, accuracy and resource utilization in the multi-user scene. Experimental evaluations have demonstrated the effectiveness of RAP with JM-MPRA in improving the performance of real-time end-cloud collaborative inference in both stable and unstable mobile networks. Compared with the state-of-the-art methods, the proposed approaches can achieve up to 5.06x decrease in inference latency and bring performance improvement of 1.52% in inference accuracy. Zhen Liu 0052, Ze Kou, Yannan Wang, Yidong Li, Yongqi Sun |
IEEE Trans. Mob. Comput. | 7 |
| 2023 | A correlation information-based spatiotemporal network for traffic flow forecasting
Weiguo Zhu, Yongqi Sun, Xintong Yi |
Neural Comput. Appl. | 2 |
| 2023 | A Low-Memory Community Detection Algorithm With Hybrid Sparse Structure and Structural Information for Large-Scale NetworksabstractCommunity detection plays an essential role in the domains of social, bioinformatics, and e-commerce. The innovative structural information theory (SInfo, introduced by Li et al.) has achieved excellent performance for network analysis. Nevertheless, similar to traditional network algorithms, the SInfo algorithm will exhaust all system memory resources when processing large-scale networks based on adjacency data structures. Moreover, due to the irregular and sequential characteristics, the SInfo algorithm is challenging to parallelize. In this article, we propose a hybrid sparse data structure and design a low-memory community detection algorithm based on structural information theory (HSSInfo), which shrinks memory requirements and achieves high parallelism. Specifically, we first propose a general quantization method to quantify the information change in community transformation, with which community inner and outer connection complexity can be semantically interpreted. Second, we design a general hybrid sparse structure to store network data among CPU and GPU, which can reduce memory resource consumption for community detection on large-scale networks. Finally, we develop an HSSInfo algorithm based on the quantization method and hybrid sparse structure, in which parallelism intersection and community fusion algorithms are employed to improve parallel scalability. We execute the comparison experiments for HSSInfo and eleven baseline algorithms on nine real-world datasets. Empirically, HSSInfo can achieve orders of magnitude speedups and extend structural information theory to large-scale datasets of billions of edges. Meanwhile, compared with various community detection algorithms, HSSInfo can achieve higher or similar accuracy with up to 19x memory shrinkage ratio. Weiguo Zhu, Yongqi Sun, Rongqiang Fang, Baomin Xu |
IEEE Trans. Parallel Distributed Syst. | 2 |
| 2022 | StyleFuse: An unsupervised network based on style loss function for infrared and visible image fusion
Yongqi Sun |
Signal Process. Image Commun. | 3 |
| 2021 | SVMAC: Unsupervised 3D Human Pose Estimation from a Single Image with Single-view-multi-angle ConsistencyabstractRecovering 3D human pose from 2D joints is still a challenging problem, especially without any 3D annotation, video information, or multi-view information. In this paper, we present an unsupervised GAN-based model consisting of multiple weight-sharing generators to estimate a 3D human pose from a single image without 3D annotations. In our model, we introduce single-view-multi-angle consistency (SVMAC) to significantly improve the estimation performance. With 2Djoint locations as input, our model estimates a 3D pose and a camera simultaneously. During training, the estimated 3D pose is rotated by random angles and the estimated camera projects the rotated 3D poses back to 2D. The 2D reprojections will be fed into weight-sharing generators to estimate the corresponding 3D poses and cameras, which are then mixed to impose SVMAC constraints to self-supervise the training process. The experimental results show that our method outpetforms the state-of-the-art unsupervised methods on Human 3.6M and MPI-INF-3DHP. Moreover, qualitative results on MPII and LSP show that our method can generalize well to unknown data. Yicheng Deng, Yongqi Sun |
3DV | 4 |
| 2021 | An effective dynamic spatiotemporal framework with external features information for traffic prediction
Jichen Wang, Weiguo Zhu, Yongqi Sun, Chunzi Tian |
Appl. Intell. | 3 |
| 2020 | EPMDA: Edge Perturbation Based Method for miRNA-Disease Association PredictionabstractIn the recent few years, plenty of research has shown that microRNA (miRNA) is likely to be involved in the formation of many human diseases. So effectively predicting potential associations between miRNAs and diseases helps to understand the development and treatment of diseases. In this study, an edge perturbation based method is proposed for predicting potential miRNA-disease association (EPMDA). Different from the previous studies, we design an feature vector to describe each edge of a graph by structural Hamiltonian information. Moreover, the extracted features are used to train a multi-layer perception model to predict the candidate disease-miRNA associations. The experimental results on the HMDD dataset show that EPMDA achieves the AUC value of 0.9818 through 5-fold cross-validation, which improves the AUC values by approximately 3.5 percent compared to the latest method DeepMDA. For the leave-one-disease-out cross-validation, EPMDA achieves the AUC value of 0.9371, which improves the AUC values by approximately 7.4 percent compared to DeepMDA. In the case study, we verify the prediction performance of EPMDA on three human diseases. As a result, there are 42, 46, and 41 of the top 50 predicted miRNAs for these three diseases which are confirmed by the published experimental discoveries, respectively. Yadong Dong, Yongqi Sun, Weiguo Zhu |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2015 | Wheel and star-critical Ramsey numbers for quadrilateral
Yongqi Sun, Stanislaw P. Radziszowski |
Discret. Appl. Math. | 2 |