Hongfeng Chai

dblp:157/2952 · DBLP profile ↗
← Back
31ranked-venue papers
2as first author
28since 2021 · last 2027
0000-0002-8577-4771ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 14 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Computer networks · 3 · 3 since 2021Systems, architecture and hardware · 2 · 2 since 2021Security and privacy · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author
YearPublicationVenuePosition
2027 LETAGNN: Label-enhanced time-aware graph neural network for phishing detection on Ethereum
Changhao Wu, Jitao Wang, Wenke Zhu, Weili Han, Hongfeng Chai
Expert Syst. Appl.6
2026 Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents
abstract
Zeping Li, Hongru Wang, Yiwen Zhao, Guanhua Chen, Yixia Li, Keyang Chen, Yixin Cao, Guangnan Ye, Hongfeng Chai, Zhenfei Yin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zeping Li, Hongru Wang 0003, Guanhua Chen 0001, Yixia Li, Keyang Chen, Yixin Cao 0002, Guangnan Ye, Hongfeng Chai, Zhenfei Yin
ACL (1)9
2026 PII-Bench: Evaluating Query-Aware Privacy Protection Systems
abstract
The widespread adoption of Large Language Models (LLMs) has raised significant privacy concerns regarding the exposure of personally identifiable information (PII) in user prompts.To address this challenge, we propose a queryunrelated PII masking strategy and introduce PII-Bench, the first comprehensive evaluation framework for assessing privacy protection systems.PII-Bench comprises 2,842 test samples across 7 PII types with 55 fine-grained subcategories, featuring diverse scenarios from singlesubject descriptions to complex multi-party interactions.Each sample is carefully crafted with a user query, context description, and standard answer indicating query-relevant PII.Our empirical evaluation reveals that while current models perform adequately in basic PII detection, they show significant limitations in determining PII query relevance.Even advanced LLMs struggle with this task, particularly in handling complex multi-subject scenarios, indicating substantial room for improvement in achieving intelligent PII masking.
Zhouhong Gu, Haokai Hong, Weili Han, Hongfeng Chai
ACL (1)5
2026 GQLBench: A Large-Scale Cross-Domain, Cross-Dialect Benchmark for NL2GQL
abstract
Despite growing interest in NL2GQL, benchmarking progress has been constrained by the lack of resources that are simultaneously largescale, cross-domain, and cross-dialect.To address this gap, we present GQLBench, a new benchmark built through an automated and scalable framework that integrates NL2SQL-to-NL2GQL conversion with graph-native data generation.GQLBench supports executionbased evaluation on both Cypher and ISO GQL, covering hundreds of graph databases and over 20k natural language questions for each dialect.By combining converted data from mature NL2SQL resources with synthetic graphspecific queries, it captures both schema diversity from real-world relational sources and graph-native reasoning challenges, including long paths and cycles.Beyond overall performance comparison, GQLBench also enables fine-grained evaluation across dialects, graph patterns, and query complexity.Experiments on advanced LLMs show that even strong proprietary models struggle on GQLBench, with gemini-3-flash achieving only 35.40% average execution accuracy across the two dialects.Our data and code are available at https://github.com/qxssadf/GQLBench.
Yanning Su, Guangnan Ye, Hongfeng Chai
ACL (1)6
2026 TabLoft: Tabular Data Generation Based on LLM with Ordered Features
Luyu Chen, Changhao Wu, Guangnan Ye, Hongfeng Chai
ICDE6
2026 CATS : An enhanced framework in LLM-based tabular data synthesis by correlation augmentation
abstract
As AI technology advances, sectors such as finance and healthcare are increasingly adopting AI tools. However, due to privacy concerns that arise with the use of AI and the high cost of real data collection, generating realistic tabular data to replace original data has become an increasingly popular solution to address these limitations. Although the output of existing tabular data generation algorithms seems to match the distribution of original data, they often fail to preserve the correlations present in the original data. This oversight can lead to significant negative impacts on downstream tasks. This paper proposes a tabular data synthesis enhancement framework based on Large Language Model (LLM), i.e., C orrelation A ugmentation T abular data S ynthesis ( CATS ), which emphasizes improving the quality of generated tabular data while preserving the correlations between the features of the original data. We explore the technical details of the CATS framework and demonstrate that CATS achieves state-of-the-art performance compared with baseline methods across seven benchmark datasets. Additionally, it enhances the performance of existing strong LLM-based tabular data generators by an average improvement of over 4.7% across all datasets.
Luyu Chen, Mingxuan Jiang, Ziyue Dai, Sen Liu 0002, Hongfeng Chai
Expert Syst. Appl.5
2026 ETTracker: A fund tracking framework for anti-money laundering on Ethereum
Changhao Wu, Kai Wang 0062, Weili Han, Hongfeng Chai
Expert Syst. Appl.5
2026 Dynamic graph learning for integrating temporal relationships in stock prediction
Ziyue Dai, Qianru Zeng, Nianwang Lin, Hongjie Xia, Keyu Zhao, Sen Liu 0002, Guangnan Ye, Jie Wu 0003, Hongfeng Chai
Expert Syst. Appl.11
2026 MasterKey: A multi-target backdoor attack in federated learning
Haohe Jia, Hongbin Zhu, Guangnan Ye, Hongfeng Chai
Knowl. Based Syst.5
2025 FALI: Fusion Adapter for Multiple LoRA Models Inference
Kexuan Chang, Nianwang Lin, Zhixin Li 0003, Sen Liu 0002, Hongfeng Chai
ICA3PP (8)6
2025 Personalized Graph Transformer for Federated Graph Learning
abstract
Federated graph learning (FGL) empowers distributed training of subgraphs across multiple institutions, overcoming the challenges of data silos and inter-institutional data sharing. Existing federated subgraph methods achieve collaborative training by training local models on clients and uploading them to a server for parameter aggregation. However, most existing methods do not take into account for the heterogeneity of client subgraph data, leading to suboptimal results. To deal with this challenging problem, we propose a personalized FGL method named FedPGT. FedPGT employs a single-layer transformer to capture long-term dependencies between client nodes, and pools client subgraphs into an averaged node representation for subgraph similarity calculation. We validate the effectiveness of FedPGT through extensive experiments on two subgraph scenarios and five datasets. The results demonstrate that FedPGT significantly outperforms the baseline methods and mitigates the adverse impact of heterogeneity.
Haohe Jia, Hongbin Zhu, Hongfeng Chai
ICASSP4
2025 Enhancing Federated Knowledge Distillation in Heterogeneous and Non-IID Scenarios
abstract
Federated Learning (FL) allows multiple participants to train models together while keeping their data private. Some FL frameworks use Knowledge Distillation to address model heterogenity, but many struggle in non-IID and heterogeneous environments, making it hard for clients to learn from each other. In this work, we show that the entropy of the softmax-averaged logits from clients reflects the model’s convergence. Based on this, we propose a new loss function, Sharpened Symmetric KL Divergence Loss (SSKL), which combines KL and Reverse KL Divergence with Label Sharpening to reduce the impact of non-IID data. Experiments demonstrate that our approach improves performance and reduces accuracy decline in non-IID and heterogeneous settings.
Wenjie Lv, Xingjun Ma, Guangnan Ye, Hongfeng Chai
ICASSP7
2025 FedCAda: Adaptive Client-Side Optimization for Accelerated and Stable Federated Learning
abstract
Federated learning (FL) enables collaborative model training across distributed clients while preserving data privacy. However, achieving both acceleration and stability, particularly on the client side, remains a challenge. In this paper, we introduce FedCAda, an adaptive algorithm that leverages an Adam-like approach to adjust first and second moment estimates on the client side while aggregating adaptive parameters on the server side. This design aims to accelerate convergence without compromising stability and performance. We also explore several algorithms with different adjustment functions and find that stronger constraints on adaptive parameters are necessary in the early stages of FL, when information from other clients is limited. Experiments on public datasets demonstrate that FedCAda surpasses state-of-the-art methods in adaptability, convergence, and stability, advancing adaptive algorithms for FL.
Liuzhi Zhou, Kun Zhai, Xingjun Ma, Guangnan Ye, Hongfeng Chai
ICASSP8
2025 FSPFL: Mitigating Communication Gap in Personalization Federated Learning Through Flexible Sparsity Allocation
abstract
Federated learning has emerged as a promising distributed learning paradigm that enables model training across decentralized devices while preserving data privacy. However, two critical challenges hinder its practical deployment: model performance degradation due to client drift in high data heterogeneity scenarios and communication gap due to varying client communication capabilities. In this paper, we provide a comprehensive analysis of these challenges and propose Flexible Sparse Personalized Federated Learning (FSPFL), a novel framework that jointly optimizes model personalization and communication efficiency. FSPFL adaptively allocates local model sparsity while incorporating personalization mechanisms to trade off model performance and communication efficiency. Extensive experiments demonstrate that FSPFL significantly mitigates the communication gap, and outperforms existing methods. Our results show that FSPFL improves communication efficiency by up to$4.1 \times$than the baselines while maintaining similar model accuracy in heterogeneous scenarios where clients have diverse data distributions and communication capabilities.
Liuzhi Zhou, Nianwang Lin, Sen Liu 0002, Guangnan Ye, Yimin Yu, Hongfeng Chai
IWQoS8
2025 Distributed DRL-Based Integrated Sensing, Communication, and Computation in Cooperative UAV-Enabled Intelligent Transportation Systems
abstract
The integration of sensing, communication, and computation (ISCC) is a critical technology that will support various emerging wireless services in future 6G networks. The unmanned aerial vehicles (UAVs) equipped with edge servers can be used as an aerial service platform in intelligent transportation systems (ITSs) to offer ISCC services to vehicles. This article studies an aerial UAV network comprising a central UAV and secondary UAVs to realize sensing of the global ITS environment and data fusion computation through collaborative UAVs. To enhance the service performance of ISCC, we maximize the success rate of ISCC services and the energy efficiency of UAVs by jointly optimizing bandwidth allocation, power allocation, and computing capacity control while ensuring the sensing and data processing latency requirements. Leveraging the network architecture and collaboration requirements of UAVs, we propose the multi-UAV collaborative Air-ISCC (MCAI) algorithm based on the asynchronous advantage actor-critic algorithm, which obtains the optimal ISCC service policy by co-training a deep reinforcement learning model with multiple UAVs. Sufficient experimental results show that MCAI enhances energy efficiency by 10.51% to 80.12% compared with the baselines. Moreover, MCAI exhibits good scalability, strengthening its feasibility in real scenarios.
Peng Hou 0003, Yi Huang 0020, Hongbin Zhu, Zhihui Lu 0002, Shih-Chia Huang, Yang Yang 0001, Hongfeng Chai
IEEE Internet Things J.7
2025 PerFedGT: A personalized federated graph transformer for scale-heterogeneous graph data
Haohe Jia, Hongbin Zhu, Hongfeng Chai
Inf. Process. Manag.6
2025 LacGCL: Lightweight message masking with linear attention and cross-view interaction graph contrastive learning for recommendation
Haohe Jia, Hongbin Zhu, Hongfeng Chai
Inf. Process. Manag.5
2025 MultiHGPT: Multi-task heterogeneous graph prompt tuning
Yixin Cao 0002, Zeping Li, Guangnan Ye, Hongfeng Chai
Inf. Process. Manag.6
2024 CI-STHPAN: Pre-trained Attention Network for Stock Selection with Channel-Independent Spatio-Temporal Hypergraph
abstract
Quantitative stock selection is one of the most challenging FinTech tasks due to the non-stationary dynamics and complex market dependencies. Existing studies rely on channel mixing methods, exacerbating the issue of distribution shift in financial time series. Additionally, complex model structures they build make it difficult to handle very long sequences. Furthermore, most of them are based on predefined stock relationships thus making it difficult to capture the dynamic and highly volatile stock markets. To address the above issues, in this paper, we propose Channel-Independent based Spatio-Temporal Hypergraph Pre-trained Attention Networks (CI-STHPAN), a two-stage framework for stock selection, involving Transformer and HGAT based stock time series self-supervised pre-training and stock-ranking based downstream task fine-tuning. We calculate the similarity of stock time series of different channel in dynamic intervals based on Dynamic Time Warping (DTW), and further construct channel-independent stock dynamic hypergraph based on the similarity. Experiments with NASDAQ and NYSE markets data over five years show that our framework outperforms SOTA approaches in terms of investment return ratio (IRR) and Sharpe ratio (SR). Additionally, we find that even without introducing graph information, self-supervised learning based on the vanilla Transformer Encoder also surpasses SOTA results. Notable improvements are gained on the NYSE market. It is mainly attributed to the improvement of fine-tuning approach on Information Coefficient (IC) and Information Ratio based IC (ICIR), indicating that the fine-tuning method enhances the accuracy and stability of the model prediction.
Hongjie Xia, Huijie Ao, Guangnan Ye, Hongfeng Chai
AAAI7
2024 Universal adversarial backdoor attacks to fool vertical federated learning
Peng Chen 0030, Xin Du 0002, Zhihui Lu 0002, Hongfeng Chai
Comput. Secur.4
2024 DPTVAE: Data-driven prior-based tabular variational autoencoder for credit data synthesizing
Yandan Tan, Hongbin Zhu, Hongfeng Chai
Expert Syst. Appl.4
2024 Distributed DRL-Based Intelligent Over-the-Air Computation in Unmanned Aerial Vehicle Swarm-Assisted Intelligent Transportation System
abstract
Unmanned aerial vehicle (UAV)-based edge computing has been widely applied in intelligent transportation systems (ITSs) owing to its ease of deployment and high mobility. In this article, we study intelligent over-the-air computation (AirComp) in UAV swarm-assisted ITS. To develop a holistic service framework for UAV swarm, we consider the heterogeneity of Internet of Things Devices (IoTDs) and UAVs. We model the 3-D deployment of UAVs, service configuration, bandwidth allocation, the control of computing capacity, and transmission power as a joint optimization problem. To tackle this complex problem, we first propose a dual time-scale architecture based on deep reinforcement learning (DRL). This architecture enables UAVs to achieve seamless coverage of IoTDs on larger time scales, while collaborative UAVs dynamically provide services on smaller time scales. Next, we propose an intelligent AirComp algorithm D2IAC based on distributed DRL to obtain the optimal UAV deployment and dynamic service policies on different time scales. The D2IAC algorithm consists of three subalgorithms, i.e., TD3-based UAV deployment (TBUD), UAV services configuration (USC), and REINFORCE-based dynamic service (RBDS). Sufficient experimental results show that the proposed algorithm can achieve 3-D deployment of UAVs with coverage improvement from 9% to 36% compared to clustering, center layout, and random algorithms. Regarding dynamic services, compared with the deep deterministic policy gradient algorithm, greedy, fixed, and random strategies, the service durations of UAV swarm are improved by 32.95%–93.72% and the resource utilization is improved by 36.19%–49.61%.
Peng Hou 0003, Yi Huang 0020, Hongbin Zhu, Zhihui Lu 0002, Shih-Chia Huang, Yang Yang 0001, Hongfeng Chai
IEEE Internet Things J.7
2024 Towards transferable adversarial attacks on vision transformers for image classification
Xu Guo 0004, Peng Chen 0030, Zhihui Lu 0002, Hongfeng Chai, Xin Du 0002
J. Syst. Archit.4
2024 TabSAL: Synthesizing Tabular data with Small agent Assisted Language models
Run Qian, Yandan Tan, Zhixin Li 0003, Luyu Chen, Sen Liu 0002, Jie Wu 0003, Hongfeng Chai
Knowl. Based Syst.8
2023 Tab-Attention: Self-Attention-Based Stacked Generalization for Imbalanced Credit Default Prediction
abstract
Accurately credit default prediction faces challenges due to imbalanced data and low correlation between features and labels. Existing default prediction studies on the basis of gradient boosting decision trees (GBDT), deep learning techniques, and feature selection strategies can have varying degrees of success depending on the specific task. Motivated by this, we propose Tab-Attention, a novel self-attention-based stacked generalization method for credit default prediction. This approach ensembles the potential proprietary knowledge contributions from multi-view feature spaces, to cope with low feature correlation and imbalance. We organize multi-view feature spaces according to the latent linear or nonlinear strengths between features and labels. Meanwhile, the f1 score assists the model in imbalance training to find the optimal state for identifying minority default samples. Our Tab-Attention achieves superior Recall1 and f11 of default intention recognition than existing GBDT-based models and advanced deep learning by about 32.92% and 16.05% on average, respectively, while maintaining outstanding overall performance and prediction performance for non-default samples. The proposed method could ensemble essential knowledge through the self-attention mechanism, which is of great significance for a more robust future prediction system.
Yandan Tan, Hongbin Zhu, Hongfeng Chai
ECAI4
2023 A Practical Clean-Label Backdoor Attack with Limited Information in Vertical Federated Learning
abstract
Vertical Federated Learning (VFL) facilitates collaboration on model training among multiple parties, each owning partitioned features of the distributed dataset. Although backdoor attacks have been found as one of the main threats to FL security, research on backdoor attacks in VFL is still in the infant stage. Existing methods for VFL backdoor attacks rely on predicting sample pseudo-labels using approaches such as label inference, which require substantial additional information not readily available in practical FL scenarios. To evaluate the practical vulnerability of VFL to backdoor attacks, we present a target-efficient clean backdoor (TECB) attack for VFL. The TECB approach consists of two phases – i) Clean Backdoor Poisoning (CBP) and Target Gradient Alignment (TGA). In the CBP phase, the adversary trains a backdoor trigger and poisons the model during VFL training. The poisoned model is further fine-tuned in the TGA phase to enhance its efficacy in complex multi-classification tasks. Compared to the existing methods, the proposed TECB achieves a highly effective backdoor attack with very limited information about the target class samples, which is more practical in typical VFL settings. Experimental results verify the superior performance of TECB, achieving above 97% attack success rate (ASR) on three widely used datasets (CIFAR10, CIFAR100, and CINIC-10) with only 0.1% of target labels known, which outperforms the state-of-the-art attack methods. This study uncovers the potential backdoor risks in VFL, enabling the development of secure VFL applications in areas like finance, healthcare, and beyond. Source code is available at: https://github.com/13thDayOLunarMay/TECB-attack
Peng Chen 0030, Jirui Yang, Junxiong Lin, Zhihui Lu 0002, Qiang Duan 0002, Hongfeng Chai
ICDM6
2022 MCHPT: A Weakly Supervise Based Merchant Pre-trained Model
Zehua Zeng, Xiaohan She, Xuetao Qiu, Hongfeng Chai, Yanming Yang
ICONIP (4)4
2021 StarFL: Hybrid Federated Learning Architecture for Smart Urban Computing
abstract
From facial recognition to autonomous driving, Artificial Intelligence (AI) will transform the way we live and work over the next couple of decades. Existing AI approaches for urban computing suffer from various challenges, including dealing with synchronization and processing of vast amount of data generated from the edge devices, as well as the privacy and security of individual users, including their bio-metrics, locations, and itineraries. Traditional centralized-based approaches require data in each organization be uploaded to the central database, which may be prohibited by data protection acts, such as GDPR and CCPA. To decouple model training from the need to store the data in the cloud, a new training paradigm called Federated Learning (FL) is proposed. FL enables multiple devices to collaboratively learn a shared model while keeping the training data on devices locally, which can significantly mitigate privacy leakage risk. However, under urban computing scenarios, data are often communication-heavy, high-frequent, and asynchronized, posing new challenges to FL implementation. To handle these challenges, we propose a new hybrid federated learning architecture called StarFL. By combining with Trusted Execution Environment (TEE), Secure Multi-Party Computation (MPC), and (Beidou) satellites, StarFL enables safe key distribution, encryption, and decryption, and provides a verification mechanism for each participant to ensure the security of the local data. In addition, StarFL can provide accurate timestamp matching to facilitate synchronization of multiple clients. All these improvements make StarFL more applicable to the security-sensitive scenarios for the next generation of urban computing.
Anbu Huang, Yang Liu 0165, Tianjian Chen, Yongkai Zhou, Hongfeng Chai, Qiang Yang 0001
ACM Trans. Intell. Syst. Technol.6
2018 Fraud detection within bankcard enrollment on mobile device based payment using machine learning
abstract
The rapid growth of mobile Internet technologies has induced a dramatic increase in mobile payments as well as concomitant mobile transaction fraud. As the first step of mobile transactions, bankcard enrollment on mobile devices has become the primary target of fraud attempts. Although no immediate financial loss is incurred after a fraud attempt, subsequent fraudulent transactions can be quickly executed and could easily deceive the fraud detection systems if the fraud attempt succeeds at the bankcard enrollment step. In recent years, financial institutions and service providers have implemented rule-based expert systems and adopted short message service (SMS) user authentication to address this problem. However, the above solution is inadequate to face the challenges of data loss and social engineering. In this study, we introduce several traditional machine learning algorithms and finally choose the improved gradient boosting decision tree (GBDT) algorithm software library for use in a real system, namely, XGBoost. We further expand multiple features based on analysis of the enrollment behavior and plan to add historical transactions in future studies. Subsequently, we use a real card enrollment dataset covering the year 2017, provided by a worldwide payment processor. The results and framework are adopted and absorbed into a new design for a mobile payment fraud detection system within the Chinese payment processor.
Hao Zhou 0035, Hongfeng Chai, Mao-lin Qiu
Frontiers Inf. Technol. Electron. Eng.2
2016 UStore: An optimized storage system for enterprise data warehouses at UnionPay
abstract
UnionPay's inter-bank transaction settlement platform (ITSP) generates a huge amount of bankcard transaction data everyday, recording different bankcard activities. In order to unleash the business value of these data, UnionPay has built a customized data warehouse based on Hadoop to manage and query the massive data imported from ITSP. However, the original system suffers from low storage utilization due to various types of data redundancy. Such data redundancy is caused by the long-term evolution of the system architecture. It dramatically wastes storage space, degrades query performance and leads to data inconsistency problem. In order to address these issues, we have developed UStore, an optimized storage system to reduce most data redundancies and improve query performance. In this paper, we present the design and implementation of UStore in detail. We test the performance of UStore on UnionPay's real data and the results show significant improvements in both storage utilization and query performance. To date, UStore has been deployed to process over 15 years' bankcard transaction data (over 3PB in plain text format) in UnionPay.
Hongfeng Chai, Hao Liu 0026, Xibo Zhou, Yanjun Xu, Jinzhi Hua, Dongjie He, Weihuai Liu
IEEE BigData1
2014 TEEI - A Mobile Security Infrastructure for TEE Integration
abstract
Mobile security becomes a hot topic recently, especially in mobile payment and privacy data fields. Traditional solution can't keep a good balance between convenience and security. Against this background, a dual OS security solution named Trusted Execution Environment (TEE) is proposed and implemented by many institutions and companies. However, it raised TEE fragmentation and control problem. Addressing this issue, a mobile security infrastructure named Trusted Execution Environment Integration (TEEI) is presented to integrate multiple different TEEs. By using Trusted Virtual Machine (TVM) tech-nology, TEEI allows multiple TEEs running on one secure world on one mobile device at the same time and isolates them safely. Furthermore, a Virtual Network protocol is proposed to enable communication and cooperation among TEEs which includes TEE on TVM and TEE on SE. At last, a SOA-like Internal Trusted Service (ITS) framework is given to facilitate the development and maintenance of TEEs.
Hongfeng Chai, Zhijun Lu, Qingyang Meng, Xiubang Zhang
TrustCom1