EDBT 2026 Demo / reviewers in the wild / expert
Peiran Wang
dblp:308/9920
· DBLP profile ↗
17ranked-venue papers
3as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 4 · 4 since 2021Computer networks · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FineSteer: A Unified Framework for Fine-Grained Inference-Time Steering in Large Language ModelsabstractLarge language models (LLMs) often exhibit undesirable behaviors, such as safety violations and hallucinations.Although inference-time steering offers a cost-effective way to adjust model behavior without updating its parameters, existing methods often fail to be simultaneously effective, utilitypreserving, and training-efficient due to their rigid, one-size-fits-all designs and limited adaptability.In this work, we present FineSteer, a novel steering framework that decomposes inference-time steering into two complementary stages-conditional steering and fine-grained vector synthesis-allowing finegrained control over when and how to steer internal representations.In the first stage, we introduce a Subspace-guided Conditional Steering (SCS) mechanism that preserves model utility by avoiding unnecessary steering.In the second stage, we propose a Mixture-of-Steering-Experts (MoSE) mechanism that captures the multimodal nature of desired steering behaviors and generates query-specific steering vectors for improved effectiveness.Through tailored designs in both SCS and MoSE, FineSteer maintains robust performance on general queries while adaptively optimizing steering vectors for targeted inputs in a training-efficient manner.Extensive experiments on safety and truthfulness benchmarks show that FineSteer outperforms the state-of-the-art methods in overall performance (e.g., A 7.6% improvement on TruthfulQA over Llama-3), achieving stronger steering performance with minimal utility loss.The code is available at https://github.com/YukinoAsuna/FineSteer. Zixuan Weng, Jinghuai Zhang, Kunlin Cai, Ying Li 0095, Peiran Wang, Yuan Tian 0001 |
ACL (1) | 5 |
| 2025 | Towards Adversarially Robust Dataset Distillation by Curvature RegularizationabstractDataset distillation (DD) allows datasets to be distilled to fractions of their original size while preserving the rich distributional information so that models trained on the distilled datasets can achieve a comparable accuracy while saving significant computational loads. Recent research in this area has been focusing on improving the accuracy of models trained on distilled datasets. In this paper, we aim to explore a new perspective of DD. We study how to embed adversarial robustness in distilled datasets, so that models trained on these datasets maintain the high accuracy and meanwhile acquire better adversarial robustness. We propose a new method that achieves this goal by incorporating curvature regularization into the distillation process with much less computational overhead than standard adversarial training. Extensive empirical experiments suggest that our method not only outperforms standard adversarial training on both accuracy and robustness with less computation overhead but is also capable of generating robust distilled datasets that can withstand various adversarial attacks. Eric Xue 0002, Yijiang Li, Haoyang Liu 0001, Peiran Wang, Haohan Wang |
AAAI | 4 |
| 2025 | Dataset Distillation via the Wasserstein MetricabstractDataset Distillation (DD) aims to generate a compact synthetic dataset that enables models to achieve performance comparable to training on the full large dataset, significantly reducing computational costs. Drawing from optimal transport theory, we introduce WMDD (Wasserstein Metric-based Dataset Distillation), a straightforward yet powerful method that employs the Wasserstein metric to enhance distribution matching. We compute the Wasserstein barycenter of features from a pretrained classifier to capture essential characteristics of the original data distribution. By optimizing synthetic data to align with this barycenter in feature space and leveraging per-class BatchNorm statistics to preserve intra-class variations, WMDD maintains the efficiency of distribution matching approaches while achieving state-of-the-art results across various high-resolution datasets. Our extensive experiments demonstrate WMDD's effectiveness and adaptability, highlighting its potential for advancing machine learning applications at scale. Haoyang Liu 0001, Yijiang Li, Tiancheng Xing, Peiran Wang, Vibhu Dalal, Luwei Li, Jingrui He, Haohan Wang |
ICCV | 4 |
| 2025 | Real-Time Video Analytics for Urban Safety: Deployment over Edge and End DevicesabstractThis paper introduces PAVE (Pedestrian Awareness Via Edge analytics), a scalable real-time video analytics system that uses street cameras to enhance pedestrian safety while preserving their privacy. PAVE processes live camera streams on an edge server to track pedestrians and vehicles in real-time, predict vehicles' trajectories, and identify danger zones where pedestrians are present. The coordinates of these zones are sent to pedestrians' mobile devices via a custom iOS app, which locally determines if they are at risk without sharing any data with the edge server, hence preserving privacy. Moreover, anonymized metadata, including real-time location and speed/direction of pedestrians and vehicles, are visualized on a public map. PAVE's effectiveness was validated through deployment on the NSF COSMOS testbed, processing live video from cameras in diverse urban environments. Live field tests show that PAVE can alert at-risk pedestrians ~0.9 s before a vehicle reaches them. Through extensive profiling, we show that optimizing memory/compute configuration per pipeline stage can reduce latency by up to 10× compared to the default operating system configurations. Mahshid Ghasemi, Yongjie Fu, Peiran Wang, Mehmet Kerem Türkcan, Jhonatan Tavori, Sofia Kleisarchaki, Thomas Calmant, Levent Gürgen, Zoran Kostic, Xuan Di, Gil Zussman, Javad Ghaderi |
SEC | 4 |
| 2025 | Demo: Real-Time Video Analytics for Urban Safety, Deployment over Edge and End DevicesabstractWe showcase the workflow of PAVE (Pedestrian Awareness Via Edge analytics), a scalable system for real-time video analytics that leverages street cameras to improve pedestrians' safety while maintaining their privacy. PAVE distributes computation across edge servers and end-user mobile devices. Cameras' live streams are processed at the edge to forecast vehicles' trajectories and detect danger zones. Pedestrians' mobile devices then locally determine if the user is inside a danger zone and trigger timely alerts via a custom iOS app. In addition, anonymized metadata, such as pedestrian and vehicle positions, speeds, and directions, are aggregated and displayed on a public map for broader situational awareness. We evaluated PAVE's performance through implementation on the NSF COSMOS testbed's edge server while processing real-time video stream from cameras in diverse urban environments. Live field tests at an intersection in New York City show that PAVE can alert at-risk pedestrians about 0.9 s before a vehicle reaches them. With low-latency cameras, this lead time extends to around 1.6 s which is within the 1–2 s window pedestrians typically need to react. Mahshid Ghasemi, Yongjie Fu, Peiran Wang, Mehmet Kerem Türkcan, Jhonatan Tavori, Sofia Kleisarchaki, Thomas Calmant, Levent Gürgen, Zoran Kostic, Xuan Di, Gil Zussman, Javad Ghaderi |
SEC | 4 |
| 2025 | SAP: Privacy-Preserving Fine-Tuning on Language Models with Split-and-Privatize FrameworkabstractPre-trained Language Models (PLM) have enabled a cost-effective approach to handling various downstream applications via Parameter-Efficient-Fine-Tuning (PEFT) techniques. In this context, service providers have introduced a popular fine-tuning-based product service known as Model-as-a-Service (MaaS). This service offers users access to extensive PLMs and training resources. With MaaS, users can fine-tune, deploy, and utilize their customized models seamlessly, leveraging a one-stop platform that allows them to work with their private datasets efficiently. However, this service paradigm has recently been exposed to the possibility of leaking user private data. To this end, we identify the data privacy leakage risks in MaaS-based PEFT and propose a Split-and-Privatize (SAP) framework, mitigating the privacy leakage by integrating split learning and differential privacy into MaaS PEFT. Furthermore, we propose Contributing-Token-Identification (CTI), a novel method to balance model utility degradation and privacy leakage. As a result, the proposed framework is comprehensively evaluated, demonstrating a 65% improvement in empirical privacy with only a 1% degradation in model performance on the Stanford Sentiment Treebank dataset, outperforming existing state-of-the-art baselines. Xicong Shen, Yi Liu 0057, Peiran Wang, Huiqi Liu, Jue Hong, Bing Duan, Zirui Huang, Yunlong Mao, Sheng Zhong 0002 |
IJCAI | 4 |
| 2025 | FedUFD: Personalized Edge Computing Using Federated Uncertainty-Driven Feature DistillationabstractRecently, federated learning (FL) has been considered a promising and well-suited technique for edge computing applications, such as intelligent traffic control, autonomous driving, and mobile crowdsensing. However, since each edge device may perform individual-specific tasks, they often have heterogeneous data distributions that impact the performance of collaborative training models. Personalized FL (PFL) has then received considerable attention to tackle this problem. Many existing PFL works often employ knowledge distillation to mitigate the negative effects of data heterogeneity. Nevertheless, these works often neglect the fact that the knowledge transferred from the teacher models is not completely correct, which limits the personalization performance of edge devices. In this work, we leverage the knowledge contained in global features to explore the potential of global models and propose a novel uncertainty-driven feature distillation framework called FedUFD. Specifically, we design an uncertainty estimation module in local models, by estimating the uncertainty of the personalized feature distribution, FedUFD can measure the difficulty of learning different personalized features, and then combine the global features to distill the corresponding personalized features. Extensive experiments show that FedUFD outperforms fourteen state-of-the-art PFL frameworks in edge computing, beating the best-performing traditional and personalized baselines by up to 45.45% and 3.55%, respectively. Zerui Shao, Beibei Li 0002, Zhibo Wang 0001, Yanbing Yang 0001, Peiran Wang, Jun Luo 0001 |
INFOCOM | 5 |
| 2025 | Fair-MoE: Medical Fairness-Oriented Mixture of Experts in Vision-Language Models
Peiran Wang, Linjie Tong, Jian Wu 0001, Zuozhu Liu |
MICCAI (5) | 1 |
| 2025 | CVE-Bench: Benchmarking LLM-based Software Engineering Agent's Ability to Repair Real-World CVE VulnerabilitiesabstractPeiran Wang, Xiaogeng Liu, Chaowei Xiao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Peiran Wang, Xiaogeng Liu, Chaowei Xiao |
NAACL (Long Papers) | 1 |
| 2025 | FedLoRE: Communication-Efficient and Personalized Edge Intelligence Framework via Federated Low-Rank EstimationabstractFederated learning (FL) has recently garnered significant attention in edge intelligence. However, FL faces two major challenges: First, statistical heterogeneity can adversely impact the performance of the global model on each client. Second, the model transmission between server and clients leads to substantial communication overhead. Previous works often suffer from the trade-off issue between these seemingly competing goals, yet we show that it is possible to address both challenges simultaneously. We propose a novel communication-efficient personalized FL framework for edge intelligence that estimates the low-rank component of the training model gradient and stores the residual component at each client. The low-rank components obtained across communication rounds have high similarity, and sharing these components with the server can significantly reduce communication overhead. Specifically, we highlight the importance of previously neglected residual components in tackling statistical heterogeneity, and retaining them locally for training model updates can effectively improve the personalization performance. Moreover, we provide a theoretical analysis of the convergence guarantee of our framework. Extensive experimental results demonstrate that our framework outperforms state-of-the-art approaches, achieving up to 89.18% reduction in communication overhead and 91.00% reduction in computation overhead while maintaining comparable personalization accuracy compared to previous works. Zerui Shao, Beibei Li 0002, Peiran Wang, Yi Zhang 0018, Kim-Kwang Raymond Choo |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2024 | Moderator: Moderating Text-to-Image Diffusion Models through Fine-grained Context-based PoliciesabstractWe present Moderator, a policy-based model management system that allows administrators to specify fine-grained content moderation policies and modify the weights of a text-to-image (TTI) model to make it significantly more challenging for users to produce images that violate the policies. In contrast to existing general-purpose model editing techniques, which unlearn concepts without considering the associated contexts, Moderator allows admins to specify what content should be moderated, under which context, how it should be moderated, and why moderation is necessary. Given a set of policies, Moderator first prompts the original model to generate images that need to be moderated, then uses these self-generated images to reverse fine-tune the model to compute task vectors for moderation and finally negates the original model with the task vectors to decrease its performance in generating moderated content. We evaluated Moderator with 14 participants to play the role of admins and found they could quickly learn and author policies to pass unit tests in approximately 2.29 policy iterations. Our experiment with 32 stable diffusion users suggested that Moderator can prevent 65% of users from generating moderated content under 15 attempts and require the remaining users an average of 8.3 times more attempts to generate undesired content. Peiran Wang, Qiyu Li 0001, Longxuan Yu, Ang Li 0005, Haojian Jin |
CCS | 1 |
| 2024 | Distributed Boosting: An Enhancing Method on Dataset DistillationabstractDataset Distillation (DD) is a technique for synthesizing smaller, compressed datasets from large original datasets while retaining essential information to maintain efficacy. Efficient DD is a current research focus among scholars. Squeeze, Recover and Relabel (SRe2L) and Adversarial Prediction Matching (APM) are two advanced and efficient DD methods, yet their performance is moderate with lower volumes of distilled data. This paper proposes an ingenious improvement method, Distributed Boosting (DB), capable of significantly enhancing the performance of these two algorithms at low distillation volumes, leading to DB-SRe2L and DB-APM. Specifically, DB is divided into three stages: Distribute & Encapsulate, Distill, and Integrate & Mix-relabel. DB-SRe2L, compared to SRe2L, demonstrates performance improvements of 25.2%, 26.9%, and 26.2% on full 224×224 ImageNet-1k at Images Per Class (IPC) 10, CIFAR-10 at IPC 10, and CIFAR-10 at IPC 50, respectively. Meanwhile, DB-APM, in comparison to APM, exhibits performance enhancements of 21.2% and 20.9% on CIFAR-10 at IPC 10, CIFAR-100 at IPC 1, respectively. Additionally, we provide a theoretical proof of convergence for DB. To the best of our knowledge, DB is the first method suitable for distributed parallel computing scenarios. Xuechao Chen, Wenchao Meng, Peiran Wang, Qihang Zhou |
CIKM | 3 |
| 2023 | TinyG: Accurate IP Geolocation Using a Tiny Number of ProbersabstractIP geolocation is essential for various applications. However, the reliability of IP geolocation databases has been proven to be inadequate. In recent years, the growing number of public probers has offered the potential for more accurate geolocation results through active measurement. The conventional practice is to probe the target IP address using all available probers and feed the measurement results to the active geolocation method. However, this practice is cost-inefficient and may trigger the anti-flood mechanism. Moreover, public probers typically impose user-level limits on the frequency and quantity of measurements. Therefore, it is important to reduce the average number of probers (ANP) selected for successfully probing each target. Researchers have discovered that geolocation accuracy primarily depends on the minimum delay between probers and the target. Inspired by that, we propose TinyG, a prober selection algorithm designed to reduce the ANP needed to find probers within a sufficiently small delay from the target. TinyG divides the probing process into multiple rounds and leverages previous measurement results to guide the selection of probers for subsequent rounds. Experimental results show that when the associated minimum delay is within 2 ms, various active geolocation methods can provide credible geolocation results. TinyG outperforms other algorithms in reducing the ANP needed to obtain credible results. Compared to using more than 1,300 probers, TinyG can achieve an ANP of 6.7 with only a 6% coverage loss of credible results. Hui Wang 0011, Jilong Wang 0001, Peiran Wang |
CNSM | 4 |
| 2023 | Top AS Router Geolocation in Databases: Performance and TechniquesabstractAutonomous systems (ASes) at the top of the global transit hierarchy play core roles in the Internet. Thus, accurately geolocating their routers is crucial for drawing credible conclusions on topics such as network resilience and traffic censorship. IP geolocation databases (DBs) are frequently used for this purpose. However, the accuracy of DBs on top AS routers has not been fully evaluated due to the lack of a comprehensive ground truth dataset. In this study, we address this gap by constructing a geolocation ground truth dataset that contains more than 12,000 router interfaces in the top 10 ASes, utilizing delay measurements from carefully selected Looking Glass vantage points. We evaluate the coverage and accuracy of 6 DBs, including 4 widely used ones and 2 new ones. Our evaluation shows that most DBs exhibit poor accuracy when geolocating top AS routers. We conduct an in-depth analysis to uncover the primary techniques behind DBs. Our investigation reveals the reasons behind the poor performance of certain DBs. Moreover, we discover that the best-performing DB heavily relies on hostnames, and all DBs perform poorly in geolocating routers without hostnames. Our dataset, which is the largest ground truth dataset of top AS routers to the best of our knowledge, will be publicly accessible to the Internet research community. Hui Wang 0011, Jilong Wang 0001, Peiran Wang |
GLOBECOM | 4 |
| 2023 | Defending Byzantine attacks in ensemble federated learning: A reputation-based phishing approach
Beibei Li 0002, Peiran Wang, Zerui Shao, Ao Liu 0005, Yukun Jiang 0001 |
Future Gener. Comput. Syst. | 2 |
| 2021 | FedVANET: Efficient Federated Learning with Non-IID Data for Vehicular Ad Hoc NetworksabstractThe vehicular ad hoc networks (VANETs) play a significant role in intelligent transportation systems (ITS). In recent years, federated learning (FL) has been widely used in VANETs to preserve the privacy-sensitive data, such as vehicle locations, drivers' driving patterns, on-board camera data, etc. However, conventional FL faces the challenges of non-independent and identically distributed (Non-IID) data and high communication overheads in VANETs. To address these challenges, we propose a novel FL framework for VANETs, named FedVANET, where a hierarchical inner-cluster FL model and a weighted inter-cluster cycling update algorithm are, respectively, developed. Extensive experiments demonstrate the high efficiency of the FedVANET in inner-cluster communications, effectiveness in handling Non-IID data, and robustness in dynamic VANET topologies. Beibei Li 0002, Yukun Jiang 0001, Weina Niu, Peiran Wang |
GLOBECOM | 5 |
| 2021 | FLPhish: Reputation-based Phishing Byzantine Defense in Ensemble Federated LearningabstractThe increasing demand for privacy protection facilitates growing interests in Federated Learning (FL). Nevertheless, most of existing FL schemes are susceptible to malicious participating clients compromised by Byzantine attacks, which remains a challenging issue. In this paper, we propose a novel Byzantine-robust FL scheme, coined FLPhish. Specifically, we first design a ensemble learning-based FL architecture, named Ensemble Federated Learning (Ensemble FL). Second, a phishing mechanism is crafted for the FL architecture to detect abnormal client behaviors. Third, a reputation mechanism is developed to further identify malicious participating clients compromised by Byzantine attackers. We evaluate the performance of FLPhish by considering various fractions of Byzantine clients and various imbalance degrees of the data distribution. Extensive experiments demonstrate the high effectiveness of the proposed FLPhish scheme in resisting Byzantine attacks in Ensemble FL. Beibei Li 0002, Peiran Wang, Hanyuan Huang, Shang Ma, Yukun Jiang 0001 |
ISCC | 2 |