Oubo Ma

dblp:326/3872 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2025
0000-0002-6572-972XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Security and privacy · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2025 PRSA: Prompt Stealing Attacks against Real-World Prompt Services
Yong Yang 0017, Changjiang Li, Qingming Li, Oubo Ma, Zonghui Wang, Yandong Gao, Wenzhi Chen, Shouling Ji
USENIX Security Symposium4
2024 SUB-PLAY: Adversarial Policies against Partially Observed Multi-Agent Reinforcement Learning Systems
abstract
Recent advancements in multi-agent reinforcement learning (MARL) have opened up vast application prospects, such as swarm control of drones, collaborative manipulation by robotic arms, and multi-target encirclement. However, potential security threats during the MARL deployment need more attention and thorough investigation. Recent research reveals that attackers can rapidly exploit the victim's vulnerabilities, generating adversarial policies that result in the failure of specific tasks. For instance, reducing the winning rate of a superhuman-level Go AI to around 20%. Existing studies predominantly focus on two-player competitive environments, assuming attackers possess complete global state observation.
Oubo Ma, Yuwen Pu, Linkang Du, Ruo Wang, Xiaolei Liu 0001, Yingcai Wu, Shouling Ji
CCS1
2024 Is Mamba Compatible with Trajectory Optimization in Offline Reinforcement Learning?
abstract
Transformer-based trajectory optimization methods have demonstrated exceptional performance in offline Reinforcement Learning (offline RL). Yet, it poses challenges due to substantial parameter size and limited scalability, which is particularly critical in sequential decision-making scenarios where resources are constrained such as in robots and drones with limited computational power. Mamba, a promising new linear-time sequence model, offers performance on par with transformers while delivering substantially fewer parameters on long sequences. As it remains unclear whether Mamba is compatible with trajectory optimization, this work aims to conduct comprehensive experiments to explore the potential of Decision Mamba (dubbed DeMa) in offline RL from the aspect of data structures and essential components with the following insights: (1) Long sequences impose a significant computational burden without contributing to performance improvements since DeMa's focus on sequences diminishes approximately exponentially. Consequently, we introduce a Transformer-like DeMa as opposed to an RNN-like DeMa. (2) For the components of DeMa, we identify the hidden attention mechanism as a critical factor in its success, which can also work well with other residual structures and does not require position embedding. Extensive evaluations demonstrate that our specially designed DeMa is compatible with trajectory optimization and surpasses previous methods, outperforming Decision Transformer (DT) with higher performance while using 30\% fewer parameters in Atari, and exceeding DT with only a quarter of the parameters in MuJoCo.
Oubo Ma, Xingxing Liang, Shengchao Hu, Mengzhu Wang, Shouling Ji, Jincai Huang 0001, Li Shen 0008
NeurIPS2
2023 Text Laundering: Mitigating Malicious Features Through Knowledge Distillation of Large Foundation Models
Chenghui Shi, Oubo Ma, Youliang Tian, Shouling Ji
Inscrypt (2)3
2023 RLID-V: Reinforcement Learning-Based Information Dissemination Policy Generation in VANETs
abstract
Ciphertext policy attribute-based encryption (CP-ABE) is popularly used to implement secure and accurate access control of disseminated information in vehicular ad hoc networks (VANETs). Nevertheless, how to improve the policy generation of CP-ABE for accurate information dissemination in the dynamic VANETs remains a challenge, as there are several access control policies rising from moving vehicles and road side units (RSUs) with different sensing boarder regarding to a specific event, such as moving vehicles and road side units (RSUs). To solve this problem, this paper proposes a reinforcement learning-based information dissemination policy generation scheme in VANETs, named RLID-V. The scheme firstly combines multiple attribute-based access control policies and resolves policy conflicts between vehicles and RSUs. Then, a manual feedback policy construction method is designed by applying decision tree to the collected feedback from all receivers. Finally, we employ reinforcement learning to dynamically update the confidence weights of different policy sources. The experiments are conducted in two classic VANETs scenarios, traffic guidance and accident warning, demonstrating that RLID-V achieves better performance in the accuracy and effectiveness of information dissemination compared with three existing schemes. Otherwise, RLID-V outperforms the compared schemes in robustness with 20% error feedback and takes a negligible cost of less than 1% of the overall delay overhead for policy generation.
Yingjie Xia, Xuejiao Liu 0002, Jing Ou, Oubo Ma
IEEE Trans. Intell. Transp. Syst.4
2022 HDRS: A Hybrid Reputation System With Dynamic Update Interval for Detecting Malicious Vehicles in VANETs
abstract
The reputation-based scheme is a promising solution to prevent malicious behaviors in Vehicular Ad-hoc Networks (VANETs). However, traditional centralized reputation schemes are not suited for distributed networks, while decentralized reputation schemes are vulnerable to malicious vehicles spreading false messages. Most of these schemes assume that the behavior of vehicles can be accurately measured as reputation from the communication, ignoring that malicious vehicles may behave intelligently to avoid being detected. In this paper, we propose a hybrid reputation system (HDRS) which allows vehicles and roadside units (RSU) to complete reputation evaluations separately and provide references to each other. HDRS utilizes a reliability evaluation module to filter out unreliable calculation results and reference records. Furthermore, HDRS includes a dynamic adjustment mechanism for the reputation update interval, employing Analytic Hierarchy Process (AHP) and reliability evaluation results to resist intelligent attacks. Simulation results illustrate that HDRS can maintain a high detection rate and low false-positive rate for detecting malicious vehicles in different environments. Compared with existing schemes, HDRS increases the detection rates of collusion and intelligent attacks by 30% and 16%, respectively.
Xuejiao Liu 0002, Oubo Ma, Wei Chen 0147, Yingjie Xia
IEEE Trans. Intell. Transp. Syst.2