EDBT 2026 Demo / reviewers in the wild / expert
Shaoming Duan
dblp:249/7599
· DBLP profile ↗
12ranked-venue papers
2as first author
9since 2021 · last 2026
0000-0002-7546-9562ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Systems, architecture and hardware · 4 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AFT-Tab: Adversarial Fine-Tuning for Tabular Data Synthesis with Long Text ColumnsabstractYuhao Zhang, Liang Yan, Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu |
ACL (1) | 3 |
| 2025 | CRED-SQL: Enhancing Real-World Large Scale Database Text-to-SQL Parsing Through Cluster Retrieval and Execution DescriptionabstractRecent advances in large language models (LLMs) have significantly improved the accuracy of Text-to-SQL systems. However, a critical challenge remains: the semantic mismatch between natural language questions (NLQs) and their corresponding SQL queries. This issue is exacerbated in large-scale databases, where semantically similar attributes hinder schema linking and semantic drift during SQL generation, ultimately reducing model accuracy. To address these challenges, we introduce CRED-SQL, a framework designed for large-scale databases that integrates Cluster Retrieval and Execution Description. CRED-SQL first performs cluster-based large-scale schema retrieval to pinpoint the tables and columns most relevant to a given NLQ, alleviating schema mismatch. It then introduces an intermediate natural language representation—Execution Description Language (EDL)—to bridge the gap between NLQs and SQL. This reformulation decomposes the task into two stages: Text-to-EDL and EDL-to-SQL, leveraging LLMs’ strong general reasoning capabilities while reducing semantic deviation. Extensive experiments on two large-scale, cross-domain benchmarks—SpiderUnion and BirdUnion—demonstrate that CRED-SQL achieves new state-of-the-art (SOTA) performance, validating its effectiveness and scalability. Our code is available at https://github.com/smduan/CRED-SQL.git Shaoming Duan, Chuanyi Liu, Peiyi Han, Zewu Peng |
ECAI | 1 |
| 2025 | Uncertainty-Aware Probabilistic Risk Quantification of SOTIF for Autonomous VehiclesabstractEnsuring the Safety of the Intended Functionality (SOTIF) for autonomous vehicles (AVs) is critical. Effective risk assessment helps AVs make decisions and avoid risks. However, existing methods face challenges due to environmental uncertainties, insufficient multi-dimensional risk quantification, and limited predictive accuracy. To address this challenge, we propose an uncertainty-aware probabilistic risk assessment framework that quantifies the risk of AVs violating safety constraints and calculates the expected average severity of such violations in uncertain environments. We first establish a general SOTIF risk model to characterize the static risk of the AV and surrounding traffic participants. Following this, we introduce a method for predicting dynamic uncertainty risks, resulting in probabilistic risk quantification. This framework accounts for multi-dimensional uncertainties and enhances safety under dynamic conditions. Extensive evaluations across typical traffic scenarios-including highways, intersections, and roundabouts-demonstrate that our method outperforms typical algorithms like Time Headway (THW) and Time-toCollision (TTC). Empirical studies in extreme scenarios further validate the framework's ability to reduce risks and improve system generalization. The related code is available at: https://github.com/idslab-autosec/risk_uncertainty. Botao Yao, Shuohan Huang, Chuanyi Liu, Peiyi Han, Shaoming Duan |
ICRA | 6 |
| 2025 | SPS-SQL: Enhancing Text-to-SQL generation on small-scale LLMs with pre-synthesized queries
Qichen Wan, Chuanyi Liu, Shaoming Duan, Peiyi Han, Yong Xu 0001 |
Pattern Recognit. Lett. | 4 |
| 2024 | FEDKA: Federated Knowledge Augmentation for Multi-Center Medical Image Segmentation on non-IID DataabstractFederated learning (FL) allows decentralized medical institutions to collaboratively learn a shared global model without breaching data privacy. However, in the context of medical image segmentation, data distributions across centers may vary a lot due to the diverse imaging protocols, vendors and partial annotation, which usually hampers the optimization convergence and the performance of FL. In this paper, we propose a novel approach called federated knowledge augmentation (FedKA) to address the non-IID (non-independent and identically distributed) problem in medical image segmentation within FL. FedKA first designs a pixel-wise knowledge augmentation method to preserve the knowledge of globally labeled regions for the local model during training, and augments each local feature statistical knowledge based on a mixture of Gaussian distribution. Our experiments on public datasets show the superiority of FedKA over the state-of-the-art methods in test performance. Shaoming Duan, Xinyu Zha, Jinhang Su, Peiyi Han, Chuanyi Liu |
ICASSP | 2 |
| 2024 | Generative data augmentation with differential privacy for non-IID problem in decentralized clinical machine learning
Tianyu He, Peiyi Han, Shaoming Duan, Wentai Wu, Chuanyi Liu, Jianrun Han |
Future Gener. Comput. Syst. | 3 |
| 2023 | FIGAT: Accurately Classify Individual Crime Risks With Multi-Information FusionabstractCrime prediction plays a vital role in public security. Existing studies infer crime locations or crime groups without considering individual GPS trajectory data. They ignore joint influence on crime patterns coming from the internal relationship between criminals, locations, and time. In this study, we propose Fusion Information Graph Attention Networks (FIGAT), which classifies individuals into high and low risks with personal movement time series and location trajectories. To solve the independence of individual crime behavior and the fusion information loss problem, FIGAT proposes Multi-dimension Fusion Information Graph to combine semantic correlation features with conventional person basic features, time features, and location features. FIGAT constructs a multi-relation graph attention layer, which utilizes the semantic relationship and node information to accurately classify individuals into high and low risks. We evaluate FIGAT with 14,625,884 GPS trajectories from 1038 individuals collected by a real-world public safety department. The results demonstrate that FIGAT improves F1 score by 41%, 32%, and 23% compared with legacy machine learning, RNN-based deep learning, and graph neural network SOTA methods, respectively. T-SNE results and ablation experiments further prove the effectiveness of FIGAT. Peiyi Han, Shaoming Duan, Chuanyi Liu |
IEEE Trans. Serv. Comput. | 3 |
| 2022 | Fed-DR-Filter: Using global data representation to reduce the impact of noisy labels on the performance of federated learning
Shaoming Duan, Chuanyi Liu, Zhengsheng Cao, Xiaopeng Jin, Peiyi Han |
Future Gener. Comput. Syst. | 1 |
| 2022 | Bi-TCCS: Trustworthy Cloud Collaboration Service Scheme Based on Bilateral Social FeedbackabstractAs a complementary technology to traditional network security, trust computing scheme has been playing an increasingly important role in providing cloud service. However, many organizations constantly face trust computing challenges; moreover, establishing a highly trustworthy cloud ecosystem can be costly and time-consuming. In this article, we originally propose the conceptual model and formal definitions for a trustworthy collaboration service ecosystem, and construct a Bi-trustworthy cloud collaboration service (Bi-TCCS), which is a scheme based on an innovative bilateral social feedback (referred to as “bi-feedback”) scheme. First, a trust-aware collaboration service model is proposed based on cloud service brokerages (CSBs), which can provide intermediation and aggregation capabilities to enable organizations to deploy their services across a collaborative cloud environment. Then, we propose a bi-feedback scheme based on the inherent social relationship among three network communities, which are composed of three types of network entities: cloud users, CSBs, and cloud service providers. The proposed scheme is effective and reliable against garnished and bad-mouthing attacks resulting from the traditional social feedback scheme. Moreover, we innovatively adopt an aggregating method for overall trust based on deviation analysis. This method can minimize errors and overcome the limitations of traditional schemes, where trust attributes are weighted manually. Theoretical analysis and experiments verify the effectiveness ofBi-TCCS. Compared with existing approaches, the service successful ratio ofBi-TCCSincreased by 12 percent under highly dishonest cloud environment. These results also indicate thatBi-TCCSis more adaptable both in the random walk and cheating profiles, which represents a substantial improvement in tracking the dynamic behavior of cloud services. Chuanyi Liu, Xiaoyong Li 0003, Mingliang Sun, Yali Gao 0004, Jie Yuan 0001, Shaoming Duan |
IEEE Trans. Cloud Comput. | 6 |
| 2020 | Scene text reading based cloud compliance access
Hezhong Pan, Chuanyi Liu, Shaoming Duan, Peiyi Han, Binxing Fang |
World Wide Web | 3 |
| 2019 | CloudDLP: Transparent and Automatic Data Sanitization for Browser-Based Cloud StorageabstractBecause cloud storage services have been broadly used in enterprises for online sharing and collaboration, sensitive information in images or documents may be easily leaked outside the trust enterprise on-premises due to such cloud services. Existing solutions to this problem have not fully explored the tradeoffs among application performance, service scalability, and user data privacy. Therefore, we propose CloudDLP, a generic approach for enterprises to automatically sanitize sensitive data in images and documents in browser-based cloud storage. To the best of our knowledge, CloudDLP is the first system that automatically and transparently detects and sanitizes both sensitive images and textual documents without compromising user experience or application functionality on browser-based cloud storage. To prevent sensitive information escaping from on-premises, CloudDLP utilizes deep learning methods to detect sensitive information in both images and textual documents. We have evaluated the proposed method on a number of typical cloud applications. Our experimental results show that it can achieve transparent and automatic data sanitization on the cloud storage services with relatively low overheads, while preserving most application functionalities. Chuanyi Liu, Peiyi Han, Yingfei Dong, Hezhong Pan, Shaoming Duan, Binxing Fang |
ICCCN | 5 |
| 2019 | Combining dissimilarity measures for image classification
Chuanyi Liu, Junqian Wang, Shaoming Duan, Yong Xu 0001 |
Pattern Recognit. Lett. | 3 |