Shaoming Duan

dblp:249/7599 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-7546-9562ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 AFT-Tab: Adversarial Fine-Tuning for Tabular Data Synthesis with Long Text Columns
abstract
Yuhao Zhang, Liang Yan, Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu
ACL (1)3
2025 CRED-SQL: Enhancing Real-World Large Scale Database Text-to-SQL Parsing Through Cluster Retrieval and Execution Description
abstract
Recent advances in large language models (LLMs) have significantly improved the accuracy of Text-to-SQL systems. However, a critical challenge remains: the semantic mismatch between natural language questions (NLQs) and their corresponding SQL queries. This issue is exacerbated in large-scale databases, where semantically similar attributes hinder schema linking and semantic drift during SQL generation, ultimately reducing model accuracy. To address these challenges, we introduce CRED-SQL, a framework designed for large-scale databases that integrates Cluster Retrieval and Execution Description. CRED-SQL first performs cluster-based large-scale schema retrieval to pinpoint the tables and columns most relevant to a given NLQ, alleviating schema mismatch. It then introduces an intermediate natural language representation—Execution Description Language (EDL)—to bridge the gap between NLQs and SQL. This reformulation decomposes the task into two stages: Text-to-EDL and EDL-to-SQL, leveraging LLMs’ strong general reasoning capabilities while reducing semantic deviation. Extensive experiments on two large-scale, cross-domain benchmarks—SpiderUnion and BirdUnion—demonstrate that CRED-SQL achieves new state-of-the-art (SOTA) performance, validating its effectiveness and scalability. Our code is available at https://github.com/smduan/CRED-SQL.git
Shaoming Duan, Chuanyi Liu, Peiyi Han, Zewu Peng
ECAI1
2025 Uncertainty-Aware Probabilistic Risk Quantification of SOTIF for Autonomous Vehicles
abstract
Ensuring the Safety of the Intended Functionality (SOTIF) for autonomous vehicles (AVs) is critical. Effective risk assessment helps AVs make decisions and avoid risks. However, existing methods face challenges due to environmental uncertainties, insufficient multi-dimensional risk quantification, and limited predictive accuracy. To address this challenge, we propose an uncertainty-aware probabilistic risk assessment framework that quantifies the risk of AVs violating safety constraints and calculates the expected average severity of such violations in uncertain environments. We first establish a general SOTIF risk model to characterize the static risk of the AV and surrounding traffic participants. Following this, we introduce a method for predicting dynamic uncertainty risks, resulting in probabilistic risk quantification. This framework accounts for multi-dimensional uncertainties and enhances safety under dynamic conditions. Extensive evaluations across typical traffic scenarios-including highways, intersections, and roundabouts-demonstrate that our method outperforms typical algorithms like Time Headway (THW) and Time-toCollision (TTC). Empirical studies in extreme scenarios further validate the framework's ability to reduce risks and improve system generalization. The related code is available at: https://github.com/idslab-autosec/risk_uncertainty.
Botao Yao, Shuohan Huang, Chuanyi Liu, Peiyi Han, Shaoming Duan
ICRA6
2025 SPS-SQL: Enhancing Text-to-SQL generation on small-scale LLMs with pre-synthesized queries
Qichen Wan, Chuanyi Liu, Shaoming Duan, Peiyi Han, Yong Xu 0001
Pattern Recognit. Lett.4
2024 FEDKA: Federated Knowledge Augmentation for Multi-Center Medical Image Segmentation on non-IID Data
abstract
Federated learning (FL) allows decentralized medical institutions to collaboratively learn a shared global model without breaching data privacy. However, in the context of medical image segmentation, data distributions across centers may vary a lot due to the diverse imaging protocols, vendors and partial annotation, which usually hampers the optimization convergence and the performance of FL. In this paper, we propose a novel approach called federated knowledge augmentation (FedKA) to address the non-IID (non-independent and identically distributed) problem in medical image segmentation within FL. FedKA first designs a pixel-wise knowledge augmentation method to preserve the knowledge of globally labeled regions for the local model during training, and augments each local feature statistical knowledge based on a mixture of Gaussian distribution. Our experiments on public datasets show the superiority of FedKA over the state-of-the-art methods in test performance.
Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu
ICASSP2
2024 Generative data augmentation with differential privacy for non-IID problem in decentralized clinical machine learning
Tianyu He, Peiyi Han, Shaoming Duan, Wentai Wu, Chuanyi Liu, Jianrun Han
Future Gener. Comput. Syst.3
2023 FIGAT: Accurately Classify Individual Crime Risks With Multi-Information Fusion
abstract
Crime prediction plays a vital role in public security. Existing studies infer crime locations or crime groups without considering individual GPS trajectory data. They ignore joint influence on crime patterns coming from the internal relationship between criminals, locations, and time. In this study, we propose Fusion Information Graph Attention Networks (FIGAT), which classifies individuals into high and low risks with personal movement time series and location trajectories. To solve the independence of individual crime behavior and the fusion information loss problem, FIGAT proposes Multi-dimension Fusion Information Graph to combine semantic correlation features with conventional person basic features, time features, and location features. FIGAT constructs a multi-relation graph attention layer, which utilizes the semantic relationship and node information to accurately classify individuals into high and low risks. We evaluate FIGAT with 14,625,884 GPS trajectories from 1038 individuals collected by a real-world public safety department. The results demonstrate that FIGAT improves F1 score by 41%, 32%, and 23% compared with legacy machine learning, RNN-based deep learning, and graph neural network SOTA methods, respectively. T-SNE results and ablation experiments further prove the effectiveness of FIGAT.
Peiyi Han, Shaoming Duan, Chuanyi Liu
IEEE Trans. Serv. Comput.3
2022 Fed-DR-Filter: Using global data representation to reduce the impact of noisy labels on the performance of federated learning
Shaoming Duan, Chuanyi Liu, Zhengsheng Cao, Xiaopeng Jin, Peiyi Han
Future Gener. Comput. Syst.1
2022 Bi-TCCS: Trustworthy Cloud Collaboration Service Scheme Based on Bilateral Social Feedback
abstract
As a complementary technology to traditional network security, trust computing scheme has been playing an increasingly important role in providing cloud service. However, many organizations constantly face trust computing challenges; moreover, establishing a highly trustworthy cloud ecosystem can be costly and time-consuming. In this article, we originally propose the conceptual model and formal definitions for a trustworthy collaboration service ecosystem, and construct a Bi-trustworthy cloud collaboration service (Bi-TCCS), which is a scheme based on an innovative bilateral social feedback (referred to as “bi-feedback”) scheme. First, a trust-aware collaboration service model is proposed based on cloud service brokerages (CSBs), which can provide intermediation and aggregation capabilities to enable organizations to deploy their services across a collaborative cloud environment. Then, we propose a bi-feedback scheme based on the inherent social relationship among three network communities, which are composed of three types of network entities: cloud users, CSBs, and cloud service providers. The proposed scheme is effective and reliable against garnished and bad-mouthing attacks resulting from the traditional social feedback scheme. Moreover, we innovatively adopt an aggregating method for overall trust based on deviation analysis. This method can minimize errors and overcome the limitations of traditional schemes, where trust attributes are weighted manually. Theoretical analysis and experiments verify the effectiveness ofBi-TCCS. Compared with existing approaches, the service successful ratio ofBi-TCCSincreased by 12 percent under highly dishonest cloud environment. These results also indicate thatBi-TCCSis more adaptable both in the random walk and cheating profiles, which represents a substantial improvement in tracking the dynamic behavior of cloud services.
Chuanyi Liu, Xiaoyong Li 0003, Mingliang Sun, Yali Gao 0004, Jie Yuan 0001, Shaoming Duan
IEEE Trans. Cloud Comput.6
2020 Scene text reading based cloud compliance access
Hezhong Pan, Chuanyi Liu, Shaoming Duan, Peiyi Han, Binxing Fang
World Wide Web3
2019 CloudDLP: Transparent and Automatic Data Sanitization for Browser-Based Cloud Storage
abstract
Because cloud storage services have been broadly used in enterprises for online sharing and collaboration, sensitive information in images or documents may be easily leaked outside the trust enterprise on-premises due to such cloud services. Existing solutions to this problem have not fully explored the tradeoffs among application performance, service scalability, and user data privacy. Therefore, we propose CloudDLP, a generic approach for enterprises to automatically sanitize sensitive data in images and documents in browser-based cloud storage. To the best of our knowledge, CloudDLP is the first system that automatically and transparently detects and sanitizes both sensitive images and textual documents without compromising user experience or application functionality on browser-based cloud storage. To prevent sensitive information escaping from on-premises, CloudDLP utilizes deep learning methods to detect sensitive information in both images and textual documents. We have evaluated the proposed method on a number of typical cloud applications. Our experimental results show that it can achieve transparent and automatic data sanitization on the cloud storage services with relatively low overheads, while preserving most application functionalities.
Chuanyi Liu, Peiyi Han, Yingfei Dong, Hezhong Pan, Shaoming Duan, Binxing Fang
ICCCN5
2019 Combining dissimilarity measures for image classification
Chuanyi Liu, Junqian Wang, Shaoming Duan, Yong Xu 0001
Pattern Recognit. Lett.3