EDBT 2026 Demo / reviewers in the wild / expert
Shuyuan Zheng
dblp:243/3290
· DBLP profile ↗
10ranked-venue papers
3as first author
9since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Legal Fact Prediction: The Missing Piece in Legal Judgment PredictionabstractJunkai Liu, Yujie Tong, Hui Huang, Bowen Zheng, Yiran Hu, Peicheng Wu, Chuan Xiao, Makoto Onizuka, Muyun Yang, Shuyuan Zheng. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yujie Tong, Hui Huang 0021, Yiran Hu, Peicheng Wu, Chuan Xiao 0001, Makoto Onizuka, Muyun Yang, Shuyuan Zheng |
EMNLP | 10 |
| 2025 | Text-IRSTD: Leveraging Semantic Text to Promote Infrared Small Target Detection in Complex Scenes
Feng Huang 0007, Shuyuan Zheng, Zhaobing Qiu, Huanxian Liu, Huanxin Bai, Liqiong Chen |
ICCV | 2 |
| 2025 | Reinforcement Learning with Argument-Structured Reward for Court Decision Abstractive SummarizationabstractCourt decision summarization is challenging due to the significant length and structural complexity of legal documents, which makes existing reinforcement learning (RL)-based abstractive summarization methods less effective. We propose an RL-based abstractive legal summarization model with a novel argument-structured reward mechanism. It leverages Issue–Reason–Conclusion components to compute fine-grained sub-rewards and aggregate them into a final reward. This design provides more reliable learning signals for model optimization. Experiments on the Indian Supreme Court dataset with Longformer-Encoder-Decoder (LED) and Llama demonstrate consistent improvements over non-argument-structured methods. To the best of our knowledge, this is the first work to incorporate argumentative structures into RL-based summarization, offering a novel direction for improving legal summarization. Yuntao Kong, Ye Xiong, Shuyuan Zheng, Ken Satoh |
JURIX | 3 |
| 2025 | Language-Bias-Resilient Visual Question Answering via Adaptive Multi-Margin Collaborative DebiasingabstractLanguage bias in Visual Question Answering (VQA) arises when models exploit spurious statistical correlations between question templates and answers, particularly in out-of-distribution scenarios, thereby neglecting essential visual cues and compromising genuine multimodal reasoning. Despite numerous efforts to enhance the robustness of VQA models, a principled understanding of how such bias originates and influences model behavior remains underdeveloped. In this paper, we address this gap through a comprehensive empirical and theoretical analysis, revealing that modality-specific gradient imbalances, which originate from the inherent heterogeneity of multimodal data, lead to skewed feature fusion and biased classifier weights. To alleviate these issues, we propose a novel Multi-Margin Collaborative Debiasing (MMCD) framework that adaptively integrates frequency-, confidence-, and difficulty-aware angular margins with a dynamic difficulty-aware contrastive learning mechanism, to dynamically reshape decision boundaries. Extensive experiments across multiple challenging VQA benchmarks confirm the consistent superiority of our proposed MMCD over state-of-the-art baselines in combating language bias. Huanjia Zhu, Shuyuan Zheng, Yishu Liu 0001, Sudong Cai, Bingzhi Chen |
NeurIPS | 2 |
| 2024 | ICPR 2024 Competition on Resource-Limited Infrared Small Target Detection Challenge: Methods and Results
Boyang Li 0007, Xinyi Ying, Ruojing Li, Yongxian Liu, Yangsi Shi, Xin Zhang 0170, Mingyuan Hu, Yukai Zhang, Dongli Tang, Qiang Ling 0002, Zaiping Lin, Weidong Sheng, Chenxu Peng, Huoren Yang, Lingjie Liu, Zelin Shi, Yunpeng Liu 0001, Chuang Yu 0003, Jinmiao Zhao, Heng Xiang, Tianyu Li 0005, Minghang Zhou, Chenxi Lan, Dongyu Xi, Chaofan Qiao, Yupeng Gao, Yongxu Liu 0006, Deping Chen, Xiaopeng Song, Jiuping Yang, Zhaobing Qiu, Rixiang Ni, Changhai Luo, Shuyuan Zheng, Baojin Huang, Xiaoqi Zhou, Qingshan Guo, Dangxuan Wu, Haodong Zeng, Qiang Fu 0017, Yimian Dai, Renke Kou, Jian Song 0007, Changfeng Feng, Zihao Xiong, Mengxuan Xiao, Yingxu Liu, Quanyi Zhao |
ICPR (34) | 50 |
| 2024 | Privacy-Enhanced Database Synthesis for Benchmark PublishingabstractBenchmarking is crucial for evaluating a DBMS, yet existing benchmarks often fail to reflect the varied nature of user workloads. As a result, there is increasing momentum toward creating databases that incorporate real-world user data to more accurately mirror business environments. However, privacy concerns deter users from directly sharing their data, underscoring the importance of creating synthesized databases for benchmarking that also prioritize privacy protection. Differential privacy (DP)-based data synthesis has become a key method for safeguarding privacy when sharing data, but the focus has largely been on minimizing errors in aggregate queries or downstream ML tasks, with less attention given to benchmarking factors like query runtime performance. This paper delves into differentially private database synthesis specifically for benchmark publishing scenarios, aiming to produce a synthetic database whose benchmarking factors closely resemble those of the original data. Introducing PrivBench , an innovative synthesis framework based on sum-product networks (SPNs), we support the synthesis of high-quality benchmark databases that maintain fidelity in both data distribution and query runtime performance while preserving privacy. We validate that PrivBench can ensure database-level DP even when generating multi-relation databases with complex reference relationships. Our extensive experiments show that PrivBench efficiently synthesizes data that maintains privacy and excels in both data distribution similarity and query runtime similarity. Yunqing Ge, Jianbin Qin, Shuyuan Zheng, Yongrui Zhong, Bo Tang 0016, Yu-Xuan Qiu, Rui Mao 0001, Ye Yuan 0001, Makoto Onizuka, Chuan Xiao 0001 |
Proc. VLDB Endow. | 3 |
| 2023 | Secure Shapley Value for Cross-Silo Federated LearningabstractThe Shapley value (SV) is a fair and principled metric for contribution evaluation in cross-silo federated learning (cross-silo FL), wherein organizations, i.e., clients, collaboratively train prediction models with the coordination of a parameter server. However, existing SV calculation methods for FL assume that the server can access the raw FL models and public test data. This may not be a valid assumption in practice considering the emerging privacy attacks on FL models and the fact that test data might be clients' private assets. Hence, we investigate the problem of secure SV calculation for cross-silo FL. We first propose HESV , a one-server solution based solely on homomorphic encryption (HE) for privacy protection, which has limitations in efficiency. To overcome these limitations, we propose SecSV , an efficient two-server protocol with the following novel features. First, SecSV utilizes a hybrid privacy protection scheme to avoid ciphertext-ciphertext multiplications between test data and models, which are extremely expensive under HE. Second, an efficient secure matrix multiplication method is proposed for SecSV. Third, SecSV strategically identifies and skips some test samples without significantly affecting the evaluation accuracy. Our experiments demonstrate that SecSV is 7.2--36.6× as fast as HESV, with a limited loss in the accuracy of calculated SVs. Shuyuan Zheng, Yang Cao 0011, Masatoshi Yoshikawa |
Proc. VLDB Endow. | 1 |
| 2022 | FL-Market: Trading Private Models in Federated LearningabstractAcquiring a sufficient amount of training data is a significant bottleneck for machine learning (ML) based data analytics. Recently, commoditizing ML models has been proposed as an economical and moderate solution to ML-oriented data acquisition. However, existing model marketplaces assume that the broker can access data owners’ private training data, which may not be realistic in practice. In this paper, to promote trustworthy data acquisition for ML tasks, we propose FL-Market, a locally private model marketplace that protects privacy against not only model buyers but also an untrusted broker. FL-Market decouples ML from the need to centrally gather training data on the broker’s side using federated learning, a privacy-preserving ML paradigm in which data owners collaboratively train an ML model by uploading local gradients (to be aggregated into a global gradient for model updating). Then, FL-Market enables data owners to locally perturb their gradients by local differential privacy and thus further prevents privacy risks. To drive FL-Market, we propose a deep learning-empowered auction mechanism for intelligently deciding the local gradients’ perturbation levels and an optimal aggregation mechanism for aggregating the perturbed gradients. Our auction and aggregation mechanisms can jointly maximize the global gradient’s accuracy, which optimizes model buyers’ utility. Our experiments verify the effectiveness of the proposed mechanisms. Shuyuan Zheng, Yang Cao 0011, Masatoshi Yoshikawa, Huizhong Li, Qiang Yan 0001 |
IEEE Big Data | 1 |
| 2022 | Diversity-Oriented Route Planning for Tourists
Wei Kun Kong, Shuyuan Zheng, Minh Le Nguyen 0001, Qiang Ma 0001 |
DEXA (2) | 2 |
| 2020 | Money Cannot Buy Everything: Trading Mobile Data with Controllable Privacy LossabstractAs personal data has been the new oil of the digital era, there is a growing trend perceiving personal data as a commodity. Existing studies have built theories on how to map the privacy loss to an arbitrage-free price. They assumed that a data buyer could purchase arbitrarily accurate results as long as she could compensate data owners for their privacy loss. However, it may not be a viable business model under strict privacy regulations, such as GDPR and CCPA, and data owners' emerging privacy concerns. In this paper, we study how to empower data owners with the control of privacy loss when continuously trading their personal mobile data. Concretely, we propose a framework for trading infinite streaming mobile data which enables each data owner to bound her privacy loss in a w-length sliding window. Introducing such upper bounds of privacy loss makes the existing trading frameworks invalid and raises new technical challenges in terms of budget allocation and arbitrage-free pricing. To address these problems, we propose a modularized trading framework with instances that allows data owners to personalize their privacy loss while the price is still arbitrage-free. Finally, we conduct experiments to verify the effectiveness of the proposed trading protocols. Shuyuan Zheng, Yang Cao 0011, Masatoshi Yoshikawa |
MDM | 1 |