EDBT 2026 Demo / reviewers in the wild / expert
Peiyi Han
dblp:189/7657
· DBLP profile ↗
19ranked-venue papers
0as first author
14since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 8 since 2021Systems, architecture and hardware · 3 · 3 since 2021Security and privacy · 3 · 1 since 2021Computer networks · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Free Energy-Driven Reinforcement Learning with Adaptive Advantage Shaping for Unsupervised Reasoning in LLMsabstractYiming Huang, Zhenbo Shi, Xin-Cheng Wen, Jichuan Zeng, Cuiyun Gao, Peiyi Han, Chuanyi Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yiming Huang 0001, Zhenbo Shi, Xin-Cheng Wen, Jichuan Zeng, Cuiyun Gao 0001, Peiyi Han, Chuanyi Liu |
ACL (1) | 6 |
| 2026 | AFT-Tab: Adversarial Fine-Tuning for Tabular Data Synthesis with Long Text ColumnsabstractYuhao Zhang, Liang Yan, Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu |
ACL (1) | 6 |
| 2025 | CRED-SQL: Enhancing Real-World Large Scale Database Text-to-SQL Parsing Through Cluster Retrieval and Execution DescriptionabstractRecent advances in large language models (LLMs) have significantly improved the accuracy of Text-to-SQL systems. However, a critical challenge remains: the semantic mismatch between natural language questions (NLQs) and their corresponding SQL queries. This issue is exacerbated in large-scale databases, where semantically similar attributes hinder schema linking and semantic drift during SQL generation, ultimately reducing model accuracy. To address these challenges, we introduce CRED-SQL, a framework designed for large-scale databases that integrates Cluster Retrieval and Execution Description. CRED-SQL first performs cluster-based large-scale schema retrieval to pinpoint the tables and columns most relevant to a given NLQ, alleviating schema mismatch. It then introduces an intermediate natural language representation—Execution Description Language (EDL)—to bridge the gap between NLQs and SQL. This reformulation decomposes the task into two stages: Text-to-EDL and EDL-to-SQL, leveraging LLMs’ strong general reasoning capabilities while reducing semantic deviation. Extensive experiments on two large-scale, cross-domain benchmarks—SpiderUnion and BirdUnion—demonstrate that CRED-SQL achieves new state-of-the-art (SOTA) performance, validating its effectiveness and scalability. Our code is available at https://github.com/smduan/CRED-SQL.git Shaoming Duan, Chuanyi Liu, Peiyi Han, Zewu Peng |
ECAI | 6 |
| 2025 | Uncertainty-Aware Probabilistic Risk Quantification of SOTIF for Autonomous VehiclesabstractEnsuring the Safety of the Intended Functionality (SOTIF) for autonomous vehicles (AVs) is critical. Effective risk assessment helps AVs make decisions and avoid risks. However, existing methods face challenges due to environmental uncertainties, insufficient multi-dimensional risk quantification, and limited predictive accuracy. To address this challenge, we propose an uncertainty-aware probabilistic risk assessment framework that quantifies the risk of AVs violating safety constraints and calculates the expected average severity of such violations in uncertain environments. We first establish a general SOTIF risk model to characterize the static risk of the AV and surrounding traffic participants. Following this, we introduce a method for predicting dynamic uncertainty risks, resulting in probabilistic risk quantification. This framework accounts for multi-dimensional uncertainties and enhances safety under dynamic conditions. Extensive evaluations across typical traffic scenarios-including highways, intersections, and roundabouts-demonstrate that our method outperforms typical algorithms like Time Headway (THW) and Time-toCollision (TTC). Empirical studies in extreme scenarios further validate the framework's ability to reduce risks and improve system generalization. The related code is available at: https://github.com/idslab-autosec/risk_uncertainty. Botao Yao, Shuohan Huang, Chuanyi Liu, Peiyi Han, Shaoming Duan |
ICRA | 4 |
| 2025 | MQA-SQL: Mitigating Question Ambiguity in Text-to-SQL with Multi-model Collaboration and Multi-variant Query Rephrasing
Yiming Huang 0001, Jiyu Guo, Jichuan Zeng, Cuiyun Gao 0001, Peiyi Han, Chuanyi Liu |
NLPCC (2) | 5 |
| 2025 | SPS-SQL: Enhancing Text-to-SQL generation on small-scale LLMs with pre-synthesized queries
Qichen Wan, Chuanyi Liu, Shaoming Duan, Peiyi Han, Yong Xu 0001 |
Pattern Recognit. Lett. | 5 |
| 2024 | FEDKA: Federated Knowledge Augmentation for Multi-Center Medical Image Segmentation on non-IID DataabstractFederated learning (FL) allows decentralized medical institutions to collaboratively learn a shared global model without breaching data privacy. However, in the context of medical image segmentation, data distributions across centers may vary a lot due to the diverse imaging protocols, vendors and partial annotation, which usually hampers the optimization convergence and the performance of FL. In this paper, we propose a novel approach called federated knowledge augmentation (FedKA) to address the non-IID (non-independent and identically distributed) problem in medical image segmentation within FL. FedKA first designs a pixel-wise knowledge augmentation method to preserve the knowledge of globally labeled regions for the local model during training, and augments each local feature statistical knowledge based on a mixture of Gaussian distribution. Our experiments on public datasets show the superiority of FedKA over the state-of-the-art methods in test performance. Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu |
ICASSP | 5 |
| 2024 | Generative data augmentation with differential privacy for non-IID problem in decentralized clinical machine learning
Tianyu He, Peiyi Han, Shaoming Duan, Wentai Wu, Chuanyi Liu, Jianrun Han |
Future Gener. Comput. Syst. | 2 |
| 2023 | A Framework of Large-Scale Peer-to-Peer Learning System
Peiyi Han, Wenjian Luo, Shaocong Xue, Kesheng Chen, Linqi Song |
ICONIP (2) | 2 |
| 2023 | FIGAT: Accurately Classify Individual Crime Risks With Multi-Information FusionabstractCrime prediction plays a vital role in public security. Existing studies infer crime locations or crime groups without considering individual GPS trajectory data. They ignore joint influence on crime patterns coming from the internal relationship between criminals, locations, and time. In this study, we propose Fusion Information Graph Attention Networks (FIGAT), which classifies individuals into high and low risks with personal movement time series and location trajectories. To solve the independence of individual crime behavior and the fusion information loss problem, FIGAT proposes Multi-dimension Fusion Information Graph to combine semantic correlation features with conventional person basic features, time features, and location features. FIGAT constructs a multi-relation graph attention layer, which utilizes the semantic relationship and node information to accurately classify individuals into high and low risks. We evaluate FIGAT with 14,625,884 GPS trajectories from 1038 individuals collected by a real-world public safety department. The results demonstrate that FIGAT improves F1 score by 41%, 32%, and 23% compared with legacy machine learning, RNN-based deep learning, and graph neural network SOTA methods, respectively. T-SNE results and ablation experiments further prove the effectiveness of FIGAT. Peiyi Han, Shaoming Duan, Chuanyi Liu |
IEEE Trans. Serv. Comput. | 2 |
| 2022 | On-manifold adversarial attack based on latent space substitute model
Chunkai Zhang, Xiaofeng Luo, Peiyi Han |
Comput. Secur. | 3 |
| 2022 | Fed-DR-Filter: Using global data representation to reduce the impact of noisy labels on the performance of federated learning
Shaoming Duan, Chuanyi Liu, Zhengsheng Cao, Xiaopeng Jin, Peiyi Han |
Future Gener. Comput. Syst. | 5 |
| 2021 | Reconstruct Anomaly to Normal: Adversarially Learned and Latent Vector-Constrained Autoencoder for Time-Series Anomaly Detection
Chunkai Zhang, Wei Zuo, Shaocong Li, Xuan Wang 0002, Peiyi Han, Chuanyi Liu |
PRICAI (2) | 5 |
| 2021 | Log Sequence Anomaly Detection Based on Local Information Extraction and Globally Sparse Transformer ModelabstractAnomaly detection for log sequences is a necessary task for system intelligent operation and fault diagnosis. In a log sequence, adjacent logs have the property of local correlation, while long-distance logs have remote dependencies. It is helpful to fully mine these information during modeling for improving the performance of anomaly detection. Meanwhile, there are some redundant information or noise in the log sequence, which has no contribution to the detection, and may even bring negative impact. The existing methods for log sequence anomaly detection do not take the above problems into account when constructing models. In this paper, we propose LSADNET, an unsupervised log sequence anomaly detection network based on local information extraction and globally sparse Transformer model. LSADNET applies multi-layer convolution to capture the local correlation between adjacent logs, and utilizes Transformer to learn the global dependency among long-distance logs. Meanwhile, we propose a globally sparse Transformer model to improve the self-attention mechanism, which can help to retain important information adaptively and eliminate the irrelevant information in the log sequence. In addition, according to the co-occurrence mode of log templates, we put forward the calculation formula of log template transfer value, and apply it to log vectorization. Through sufficient experiments on two public datasets, it is confirmed that LSADNET has better performance than the state-of-art methods. Chunkai Zhang, Xinyu Wang 0028, Hongye Zhang, Peiyi Han |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2020 | Pixel re-representations for better classification of images
Junqian Wang, Peiyi Han, Chuanyi Liu, Yong Xu 0001 |
Pattern Recognit. Lett. | 3 |
| 2020 | Scene text reading based cloud compliance access
Hezhong Pan, Chuanyi Liu, Shaoming Duan, Peiyi Han, Binxing Fang |
World Wide Web | 4 |
| 2019 | CloudDLP: Transparent and Automatic Data Sanitization for Browser-Based Cloud StorageabstractBecause cloud storage services have been broadly used in enterprises for online sharing and collaboration, sensitive information in images or documents may be easily leaked outside the trust enterprise on-premises due to such cloud services. Existing solutions to this problem have not fully explored the tradeoffs among application performance, service scalability, and user data privacy. Therefore, we propose CloudDLP, a generic approach for enterprises to automatically sanitize sensitive data in images and documents in browser-based cloud storage. To the best of our knowledge, CloudDLP is the first system that automatically and transparently detects and sanitizes both sensitive images and textual documents without compromising user experience or application functionality on browser-based cloud storage. To prevent sensitive information escaping from on-premises, CloudDLP utilizes deep learning methods to detect sensitive information in both images and textual documents. We have evaluated the proposed method on a number of typical cloud applications. Our experimental results show that it can achieve transparent and automatic data sanitization on the cloud storage services with relatively low overheads, while preserving most application functionalities. Chuanyi Liu, Peiyi Han, Yingfei Dong, Hezhong Pan, Shaoming Duan, Binxing Fang |
ICCCN | 2 |
| 2019 | Fingerprinting SDN Applications via Encrypted Control Traffic
Jiahao Cao 0001, Zijie Yang, Kun Sun 0001, Qi Li 0002, Peiyi Han |
RAID | 6 |
| 2017 | Query Recovery Attacks on Searchable Encryption Based on Partial Knowledge
Chuanyi Liu, Yingfei Dong, Hezhong Pan, Peiyi Han, Binxing Fang |
SecureComm | 5 |