EDBT 2026 Demo / reviewers in the wild / expert
Hao Wu 0067
dblp:72/4250-67
· DBLP profile ↗
36ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0002-0980-9805ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Computer networks · 11 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 8 since 2021Security and privacy · 6 · 1 first-author · 6 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021Software engineering, systems software and programming languages · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Diverse Human Driving Vehicle Simulation in Background Traffic for Autonomous Driving TestsabstractRealistic background traffic is critical to the simulation platforms for autonomous driving (AD) testing. Given that most vehicles in reality are driven by human beings, introducing human driving (HD) vehicles to the background traffic is necessary to be able to discover more problems of the tested AD vehicle in the simulation stage. However, existing methods rely on ad-hoc rules or data-driven training to mimic partial human driver behaviors, which are not comprehensive and lack transparency. In this work, we design a smart human driving vehicle simulator HDSim which is empowered by cognitively inspired modeling and AI models. HDSim enables diverse, realistic, and scalable HD traffic simulation on AD testing platforms like CARLA in a non-intrusive manner. There are two novel components in HDSim. First, we introduce a driver model to guide the generation of diverse human driving styles by using different combinations of latent cognitive factors in a hierarchy. Second, we design a Perception-Mediated Behavior Influence (PMBI) mechanism to use LLM-assisted perceptual transformations to indirectly fuse driving actions with driving styles. Experiments show that HDSim traffic can help simulation platforms like CARLA to reveal 68% more failures of tested AD vehicles, and the explainability of reported accidents is also improved. Wendi Li, Hao Wu 0067, Bing Mao 0001, Fengyuan Xu, Sheng Zhong 0002 |
AAAI | 2 |
| 2026 | Integrated Load-Balanced Scheduling for Human-Vehicle Collaborative Urban Sanitation
Lingzi Zhao, Huali Lu, Hao Wu 0067, Shucheng Li, Longye Li, Wenlong Liao, Feng Lyu 0001 |
ICDCS | 3 |
| 2026 | KAT: Knowledge-Context Augmentation for Evolving LLM-Based Telecom Troubleshooting
Feng Lyu 0001, Hao Wu 0067, Shucheng Li, Fan Wu 0014, Fengyuan Xu |
INFOCOM | 3 |
| 2026 | Seeing the Whole Through the Parts: Discovering Objects through Semantic Part Mining in Weak Supervision
Shucheng Li, Weixuan Xu, Hao Wu 0067, Fengyuan Xu, Fan Wu 0014, Feng Lyu 0001 |
SIGIR | 4 |
| 2026 | SynDiSC: High-Quality Tabular Data Synthesis with Distributional and Semantic ConsistencyabstractSynthesizing high-quality tabular data is essential for privacy-preserving data analysis. However, this task remains challenging due to two key factors: (1) distribution complexity : imbalanced and skewed data make it challenging to learn the data distribution accurately; and (2) semantic coherence : implicit relationships and logical dependencies among fields must be preserved to ensure valid and meaningful synthetic samples. To address these issues, we propose SynDiSC, a high-quality tabular data synthesis approach that enforces both distributional and semantic consistency. It comprises three core designs: (1) a distribution-aware encoding that effectively handles heterogeneous data types and complex distributions; (2) a multi-dimensional semantic conditioning that leverages multi-dimensional conditional dependencies to enforce semantic validity during generation; and (3) a conditional consistency controller that guides the generator to produce diverse samples satisfying multiple conditional constraints while mitigating mode collapse. Extensive experiments on datasets from various application domains demonstrate that SynDiSC significantly improves data quality, conditional controllability, and downstream task performance compared to state-of-the-art methods. Our code and data samples are open-sourced in the GitHub repository. https://github.com/Knightz9/SynDiSC. Fan Wu 0014, Haoye Pan, Hao Wu 0067, Shucheng Li, Feng Lyu 0001 |
SIGIR | 3 |
| 2026 | "Say What You Mean": Natural Language Access Control With Large Language Models for Internet of Things
Ye Cheng, Minghui Xu 0001, Yue Zhang 0025, Kun Li 0026, Hao Wu 0067, Yechao Zhang, Shao-Yong Guo 0001, Wangjie Qiu, Dongxiao Yu, Xiuzhen Cheng |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2026 | H2O: Heterogeneity-Aware Hierarchical Orchestration for Memory-Efficient On-Device LLM InferenceabstractOn-device Large Language Model (LLM) inference enables private, personalized AI but faces memory constraints. Despite memory optimization efforts, scaling laws continue to increase model sizes and memory pressure. In this paper, we revisit the core memory bottlenecks in on-device LLM inference and conduct a comprehensive analysis of mainstream optimization techniques. We uncover several overlooked inefficiencies: (1) model weights, not KV caches, dominate memory usage; (2) weight sparsity remains underutilized; (3) OS-level memory behaviors cause redundancy; and (4) naive weight loading leads to excessive memory residency. To address these challenges, we propose H2O, a heterogeneity-aware hierarchical orchestration framework for memory-efficient on-device LLM inference. H2O introduces three key techniques, including hierarchical weight orchestration to reduce redundant memory retention, zero copy I/O–compute parallelism for safe and efficient memory reuse, and heterogeneity-aware inference planning to adapt to diverse mobile hardware constraints. Extensive experimental results show that H2O reduces peak memory usage by up to 60%, eliminates out-of-memory (OOM) failures for 7B–13B models, and improves inference latency by 34%–94% under tight memory budgets. We open-source our implementation at: https://github.com/ccfeiker/H2O. Feng Lyu 0001, Hao Wu 0067, Zhanxi Li, Shucheng Li, Fengyuan Xu |
IEEE Trans. Mob. Comput. | 3 |
| 2026 | SlimFit-Gens: Toward Low Bandwidth One-on-One Video Calls on COTS SmartphonesabstractMobile video calls play an essential role in our daily lives. However, in bandwidth-limited scenarios (e.g., inadequate cellular coverage, congested satellite links, and metered connections), users often experience poor quality of experience (QoE) during video calls. While recent advances in deep learning have demonstrated significant improvements in video compression over traditional methods, existing approaches are ill-suited for bidirectional video streaming on smartphones. The primary challenge lies in simultaneously achieving high video quality, computational and bandwidth efficiency, and practical usability on constrained mobile devices. In this work, we present SlimFit-Gens, the first practical video calling system for smartphones capable of delivering real-time 480p video at as low as 30 kbps. SlimFit-Gens addresses the challenge with joint algorithm and system-level optimizations. The core technique is a fine-grained model personalization design tailored for mobile video calling, enabling high-fidelity video generation at low model complexity. SlimFit-Gens achieves effective personalized adaptation through a novel two-stage personalization mechanism working upon an optimized model architecture. It also incorporates a privacy preserving, resource-efficient system design, featuring TEE-based (e.g., Confidential VM/NVIDIA Confidential Computing) fine-tuning on the server side and heterogeneity-aware inference on the device side. We implement SlimFit-Gens on four commercial off-the-shelf (COTS) smartphones with different system-on-chip (SoC) configurations and conduct extensive evaluations. Compared to prior work, SlimFit-Gens simultaneously improves generation quality with a 0.09-0.12 reduction in LPIPS and system efficiency through a 1.6-1.8× increase in video frame rate. Jingzhou Zhu, Lizhi Sun, Peiwen Dong, Wendi Li, Yixin Xu 0003, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
IEEE Trans. Mob. Comput. | 8 |
| 2026 | Game in Motion: Heterogeneous Task Offloading in Dynamic Vehicular Edge ComputingabstractThe rise of vehicular edge computing (VEC) enables vehicles to offload resource-intensive tasks to roadside units (RSUs), improving efficiency and reducing latency. However, high mobility, dynamic resource availability, and heterogeneous quality-of-experience (QoE) requirements make task offloading and coordination highly challenging. In this paper, we investigate the heterogeneous task offloading problem in dynamic VEC environments by proposing a two-stage optimization framework, TOVEC. Our TOVEC decouples the spatio-temporally coupled decision space into two tractable subproblems and solves it with a two-stage design. In the first stage, we employ a TD3-based deep reinforcement learning algorithm to handle RSU-channel access decisions under dynamic network conditions. In the second stage, we formulate the interaction between vehicular users (VUs) and RSUs as a Stackelberg game, enabling joint task scheduling and dynamic pricing that balances VUs’ QoE optimization and RSUs’ revenue maximization. We theoretically prove the existence of a Stackelberg equilibrium and validate our TOVEC using real-world vehicular traces. The experimental results show that our method has improved the QoE (measured by task delay and energy cost) of VU and the benefits of RSU by 6.8%-41.3% and 2.2%-117.4%, respectively, compared with the baseline schemes. Jie Zhao 0041, Feng Lyu 0001, Hao Wu 0067, Fan Wu 0014, Shucheng Li |
IEEE Trans. Netw. | 3 |
| 2025 | StreamMind: Unlocking Full Frame Rate Streaming Video Dialogue through Event-Gated Cognition
Hao Wu 0067, Yifan Yang 0004, Shiqi Jiang 0002, Qianxi Zhang, Donglin Bai, Zhibo Chen 0001, Ting Cao 0003 |
ICCV | 2 |
| 2025 | HiPOD: Hierarchical Pruning for Low-Distinction Multi-Scale Object Detection on Edge DevicesabstractThe deployment of high-accuracy low-distinction and multi-scale object detection models on resource-constrained edge devices is essential for ubiquitous intelligent applications, from autonomous obstacle avoidance to anomaly object recognition. However, these models' computational burden, energy consumption, and memory footprint pose significant challenges for distributed and pervasive systems. In this paper, we propose HiPOD, a hierarchical pruning framework designed to prune low-distinction and multi-scale object detection models by jointly learning layer-wise and path-wise pruning strategies. HiPOD balances the accuracy-efficiency trade-off through two novel components: Layer-Adaptive Ratio Learning, named AdaLR, which leverages network structural characteristics and a reinforcement learning-based action-feedback mechanism to adaptively generate balanced layer-wise pruning ratios; and Genetic Path Optimization, named GenPath, which employs crossover and mutation operations to optimize inter-layer kernel pruning paths, preserving critical semantic and spatial information. Extensive experiments on three benchmark datasets demonstrate the effectiveness of HiPOD, with only a 3 % drop in mAP50 and a 2 % drop in mAP50:95 compared to the full models. The ablation studies and impact analysis further validate the effectiveness of each module and highlight the robustness of our framework. Furthermore, evaluations on an edge device demonstrate the practicality of the proposed solution for powerline inspection. Jieyu Zhou, Feng Lyu 0001, Mingliu Liu, Hao Wu 0067, Fan Wu 0014, Yaoxue Zhang |
ICPADS | 4 |
| 2025 | Uncovering Prompt Elements: Cloning System Prompts from Behavioral TracesabstractWe introduce prompt cloning, a new black-box attack that reconstructs functionally equivalent system prompts rather than extracts original system prompts. Unlike prompt stealing, prompt cloning exploits the insight that system prompts leave persistent behavioral traces in outputs, even under strong alignment and prompt-level defenses. Our method decomposes system behavior into semantically interpretable elements, selectively elicits them through carefully designed queries, and aggregates representative traces to synthesize high-fidelity cloned prompts. Extensive evaluations show that cloned prompts replicate functional behavior with up to 85% semantic similarity, outperforming base LLMs by up to 8%, and even exceeding original system prompts when transferred to different back-end models. We also conduct a large-scale study on GitHub repositories, revealing that single-prompt architectures remain widespread in open-source LLM applications, reinforcing the real-world relevance of our threat model. Our findings reveal that prompt cloning enables unauthorized replication of confidential LLM behavior and underscore the urgent need for defenses that go beyond hiding prompt text. Hao Wu 0067, Ligeng Chen, Bing Mao 0001 |
ASE | 3 |
| 2025 | Training Data Attribution: Was Your Model Secretly Trained On Data Created By Mine?abstractThe emergence of text-to-image models has recently sparked significant interest, but the attendant is a looming shadow of potential infringement by violating user terms. Specifically, an adversary may exploit data created by a commercial model to train their own without proper authorization. To address such risk, it is crucial to investigate the attribution of a suspicious model's training data by determining whether its training data originates, wholly or partially, from a specific source model. To trace the generated data, existing methods need to apply additional watermarks during either the training or inference phases of the source model. However, these methods are impractical for pre-trained models that have been released, especially when model owners lack security expertise. To tackle this challenge, we propose an injection-free training data attribution method for text-to-image models. It can identify whether a model's training data stems from a certain source model without adding additional watermarks on the source model. The rationale of our method lies in the inherent memorization characteristic of text-to-image models. The memorization of training data is inherited through the data generated by the source model to the model trained on that data, making the source model and the infringing model exhibit consistent behaviors on specific samples. Therefore, from instance-level, we develop detection-based and generation-based strategies to uncover these distinct samples and using them as inherent watermarks to verify if a suspicious model originates from the source model. Besides, we also propose a statistical-level attribution method, utilizing the shadow model technique to train an attribution discriminator. Experiments demonstrate that the attribution accuracy and AUC scores of our methods are over 80% even when the infringing model only uses a small proportion of generated data. Hao Wu 0067, Lingcui Zhang, Fengyuan Xu, Jin Cao 0001, Fenghua Li 0001, Ben Niu 0001 |
KDD (2) | 2 |
| 2025 | When LLMs Copy to Think: Uncovering Copy-Guided Attacks in Reasoning LLMsabstractLarge Language Models (LLMs) have become integral to automated code analysis, enabling tasks such as vulnerability detection and code comprehension. However, their integration introduces novel attack surfaces. In this paper, we identify and investigate a new class of prompt-based attacks, termed Copy-Guided Attacks (CGA), which exploit the inherent copying tendencies of reasoning-capable LLMs. By injecting carefully crafted triggers into external code snippets, adversaries can induce the model to replicate malicious content during inference. This behavior enables two classes of vulnerabilities: inference length manipulation, where the model generates abnormally short or excessively long reasoning traces; and inference result manipulation, where the model produces misleading or incorrect conclusions. We formalize CGA as an optimization problem and propose a gradient-based approach to synthesize effective triggers. Empirical evaluation on state-of-the-art reasoning LLMs shows that CGA reliably induces infinite loops, premature termination, false refusals, and semantic distortions in code analysis tasks. While highly effective in targeted settings, we observe challenges in generalizing CGA across diverse prompts due to computational constraints, posing an open question for future research. Our findings expose a critical yet underexplored vulnerability in LLM-powered development pipelines and call for urgent advances in prompt-level defense mechanisms. Yue Li 0002, Xiao Li 0082, Hao Wu 0067, Yue Zhang 0025, Fengyuan Xu, Xiuzhen Cheng, Sheng Zhong 0002 |
MASS | 3 |
| 2025 | PriCAF: Privacy-Preserving Contribution Assessment in Federated Learning Before Model Training
Yixin Xu 0003, Hao Wu 0067, Jingzhou Zhu, Fengyuan Xu, Sheng Zhong 0002 |
ACM Multimedia | 2 |
| 2025 | Auto-UIT: Automating UAV Inspection Trajectory by Recognizing Pylon Structure from 3D Point CloudabstractUAV-assisted inspection is critical for modern power grid maintenance, enhancing efficiency and safety in remote areas. However, automatically designing UAV inspection trajectories is challenging due to the cluttered inspection environments, small inspection targets, and pervasive obstacles. We propose Auto-UIT, a novel method for generating inspection trajectories in noisy, sparse, and complex 3D point cloud. Auto-UIT has three core techniques: (1) A local structure-enhanced pylon segmentation, which accurately segments pylons, power lines, and surroundings in noisy point cloud for effective inspection target identification and trajectory planning. (2) A 3D fingerprint-based pylon type recognition that compensates for point cloud sparsity to complete missing inspection targets based on the pylon type. (3) An adaptive trajectory generation that samples positions in response to diverse pylon orientations and pervasive environmental obstacles, ensuring UAV operational safety. Our experiments on a real-world dataset across four distinct areas demonstrate that Auto-UIT outperforms existing baseline methods in all three tasks. Furthermore, a four-month deployment in a power grid inspection system—covering a 270 km2 primary mountainous area—yielded an expert first-review acceptance rate of 91.86% for the generated trajectories, and reduced design time by an average of 88.19% compared to manual methods, significantly improving inspection efficiency. Feng Lyu 0001, Lijuan He, Mingliu Liu, Sijing Duan, Hao Wu 0067, Jieyu Zhou, Yi Ding 0011, Zaixun Ling |
MobiCom | 5 |
| 2025 | Demo: UAV Trajectory Generation from Sparse and Noisy 3D Point CloudsabstractUAV-assisted inspection is critical for modern power grid maintenance, enhancing efficiency and safety in remote areas. However, automatically designing UAV inspection trajectories is challenging due to the cluttered inspection environments, small inspection targets, and pervasive obstacles. We propose a novel method for generating inspection trajectories in noisy, sparse, and complex 3D point cloud. It has three core techniques: (1) A local structure-enhanced pylon segmentation, which accurately segments pylons, power lines, and surroundings in noisy point cloud for effective inspection target identification and trajectory planning. (2) A 3D fingerprint-based pylon type recognition that compensates for point cloud sparsity to complete missing inspection targets based on the pylon type. (3) An adaptive trajectory generation that samples positions in response to diverse pylon orientations and pervasive environmental obstacles, ensuring UAV operational safety. A four-month deployment in a power grid inspection system—covering a 270 km2 primary mountainous area—yielded an expert first-review acceptance rate of 91.86% for the generated trajectories, and reduced design time by an average of 88.19% compared to manual methods, significantly improving inspection efficiency. Demo video and dataset are available at https://ljhe006.github.io/autouit/. Lijuan He, Feng Lyu 0001, Mingliu Liu, Hao Wu 0067, Sijing Duan, Jieyu Zhou, Yi Ding 0011, Zaixun Ling |
MobiCom | 4 |
| 2025 | Demo: Task Cooperation for Urban Unmanned Sanitation VehiclesabstractUnmanned sanitation vehicles (USVs) promise cleaner cities, yet efficiently coordinating multiple USVs in large urban areas remains challenging due to constraints such as limited waste capacity and battery life. In this demo, we present MRTC, a multi-robot task cooperation system. First, Dynamic Task Assignment employs an Actor-Critic policy within a Markov decision framework to allocate cleaning tasks and decide the required number of USVs. Second, Single-USV Path Planning refines each route via a fast two-layer iterative search. Over an eight-month real-world deployment in three urban testbeds, our MRTC system markedly improved cleaning efficiency while lowering operating costs. Operating over a combined 10,775 km of routes per month, the system achieved average monthly savings of 20,575 kWh of energy and 2,744 labour hours. A demonstration video is available at https://llq978.github.io/Demo/. Lingzi Zhao, Feng Lyu 0001, Hao Wu 0067, Huaqing Wu, Huali Lu, Shucheng Li, Wenlong Liao, Sheng Zhong 0002 |
MobiCom | 3 |
| 2025 | Make a Feint to the East While Attacking in the West: Blinding LLM-Based Code Auditors with Flashboom AttacksabstractLLM-based vulnerability auditors (e.g., GitHub Copilot) represent a significant advancement in automated code analysis, offering precise detection of security vulnerabilities. This paper explores the potential to circumvent LLM-based vulnerability auditors by diverting their focus, decided by the LLM attention mechanism, away from real vulnerable code segments. In these LLM-based vulnerability auditors, the attention mechanism is supposed to focus on potentially vulnerable code sections to identify security issues. Our approach introduces high-attention code snippets (code fragments designed to draw focus) into the codebase under review. By strategically diverting the model's focus away from actual vulnerabilities, this technique effectively “blinds” the LLM, resulting in missed detections. To scale this approach, we present Crazy-Ivan11Source code, dataset and attack results are available at https://github.com/oxygen-hunter/Flashboom., an automated system that identifies and seamlessly integrates high-attention code snippets, shifting focus away from genuine vulnerabilities to decoy functions. Through systematic function-level prioritization and refinement, Crazy-Ivan optimizes the blinding effect, producing the Flashboom that can reduce the model's capacity to detect true security risks. Our evaluation underscores the effectiveness of Flashboom, achieving blinding success rates of up to 96.3% on CodeLlama and 83.05% on Gemma, with notable cross-model transferability and applicability across multiple programming languages. In a case study with GitHub Copilot, Flashboom led the tool to overlook a critical blockchain vulnerability, underscoring the security implications of such attention-diverting attacks and the risks inherent in relying solely on LLM-based automated auditing systems. We have reported our findings to the respective LLM-based code auditor vendors, who have acknowledged the issues and are currently working on fixes. Xiao Li 0082, Yue Li 0002, Hao Wu 0067, Yue Zhang 0025, Kaidi Xu, Xiuzhen Cheng, Sheng Zhong 0002, Fengyuan Xu |
SP | 3 |
| 2025 | UTRDCL: Stealthy DCL-Based Obfuscation and Its Attacks and Defenses in AndroidabstractDynamic Class Loading (DCL) is a legitimate technique extensively used by Android developers to incorporate additional functionalities into applications at runtime. However, adversaries can exploit DCL as a stealthy obfuscation technique to dynamically load malicious code and evade detection. While prior studies have analyzed typical DCL-based obfuscation and the attacks it enables—such as identifying payloads on storage, inspecting DCL-related APIs, or profiling dynamic behaviors—existing solutions remain insufficient against increasingly evasive DCL threats. In this paper, we propose UTRDCL, a novel stealthy obfuscation technique that leverages system APIs instead of conventional DCL-related APIs, and employs an automated footprint cleanup strategy to minimize runtime traces. Based on UTRDCL, we construct three real-world attack instances by embedding it into existing malware and benign applications, demonstrating how it can be used to evade detection. To counter such threats, we design and implement a lightweight defense mechanism by patching a previously overlooked vulnerability in the Android system that UTRDCL exploits in this specific context. This system-level mitigation closes the attack surface leveraged by UTRDCL, offering a more fundamental defense than behavioral detection. Extensive experiments show that attacks leveraging UTRDCL can evade 11 state-of-the-art malware detectors from open-source, academic, and commercial sources. We further validate our defense mechanism on real devices, demonstrating its effectiveness in preventing UTRDCL-based attacks without introducing noticeable overhead. Our proof-of-concept of UTRDCL and its defense is publicly available1. Hao Wu 0067, Sheng Zhong 0002, Fengyuan Xu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2024 | MuSR: Multi-Scale 3D Scenes Reconstruction based on Monocular VideoabstractThree-dimensional (3D) scene reconstruction, particularly from monocular videos, is a significant challenge in large-scale scenarios due to difficulty handling varying object sizes and high computational resource needs. This paper introduces MuSR, a novel multi-scale reconstruction method addressing these issues. MuSR features a dynamic multi-resolution spatial structure that adaptively adjusts voxel resolution for objects of different sizes to improve reconstruction quality. MuSR also employs a block-based sparse 3D data structure and hardware resource management strategy to reduce GPU memory usage while maintaining efficient reconstruction. Evaluated on ScanNet and 7-Scenes datasets, as well as real-world scenes, MuSR outperforms state-of-the-art methods in terms of efficiency, completeness, and geometric shape reconstruction, proving its applicability in practical multi-scale 3D reconstructions. Hao Wu 0067, Peiwen Dong, Yixin Xu 0003, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 2 |
| 2024 | VIDAR: Data Quality Improvement for Monocular 3D Reconstruction through In-situ Visual Interactionabstract3D reconstruction based on monocular videos has attracted wide attention, and existing reconstruction methods usually work in a reconstruction-after-scanning manner. However, these methods suffer from insufficient data collection problems due to the lack of effective guidance for users during the scanning process, which affects reconstruction quality. We propose VIDAR, which visually guides users with the streaming incremental reconstructed mesh in data collection for monocular 3D reconstruction. We propose an incremental mesh extraction algorithm to achieve lossless fusion of streaming incremental mesh data via slice-style management for guidance quality. We also design an incremental mesh rendering algorithm to achieve precise memory reallocation by updating the buffer in a fill-in-the-blank pattern for guidance efficiency. Besides, we introduce several optimizations on data transmission and human-computer interaction to improve the overall system performance. The experiment results on real-world scenes show that VIDAR efficiently delivers high-quality visual guidance and outperforms the non-interactive data collection methods for scene reconstruction. Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ICRA | 4 |
| 2024 | CoAst: Validation-Free Contribution Assessment for Federated Learning based on Cross-Round ValuationabstractIn the federated learning (FL) process, since the data held by each participant is different, it is necessary to figure out which participant has a higher contribution to the model performance. Effective contribution assessment can help motivate data owners to participate in the FL training. Research works in this field can be divided into two directions based on whether a validation dataset is required. Validation-based methods need to use representative validation data to measure the model accuracy, which is difficult to obtain in practical FL scenarios. Existing validation-free methods assess the contribution based on the parameters and gradients of local models and the global model in a single training round, which is easily compromised by the stochasticity of model training. In this work, we propose CoAst, a practical method to assess the FL participants' contribution without access to any validation data. The core idea of CoAst involves two aspects: one is to only count the most important part of model parameters through a weights quantization, and the other is a cross-round valuation based on the similarity between the current local parameters and the global parameter updates in several subsequent communication rounds. Extensive experiments show that CoAst has comparable assessment reliability to existing validation-based methods and outperforms existing validation-free methods. Hao Wu 0067, Shucheng Li, Fengyuan Xu, Sheng Zhong 0002 |
ACM Multimedia | 1 |
| 2024 | TTFL: Towards Trustworthy Federated Learning with Arm Confidential ComputingabstractFederated learning (FL), as a distributed training paradigm, has drawn great attention from both academia and industry. Recently, privacy and security concerns have been raised for FL. Despite many efforts to protect privacy and security, an FL framework that can systematically provide privacy and security guarantees is lacking. In this work, we present TTFL, a trustworthy FL framework in practice to defend the security and privacy issues based on Arm Confidential Compute Architecture (CCA). TTFL has two core designs. (1) It achieves a high-availability privacy protection based on flexible Trusted Execution Environments (TEEs). It leverages the resource-rich and conveniently accessed features of the latest TEE on Arm CCA, combined with our TEE secure interconnection design, to enable the whole FL process performed in distributed TEEs, which efficiently protects parameter confidentiality and protocol integrity. (2) It achieves effective security protection by proposing an effective poisoning-resisted secure aggregation scheme and protecting it within TEE. The new proposed secure aggregation combines the advantages of existing defenses and is placed in the flexible TEE to ensure a secure, effective, and non-bypassable aggregation procedure. We implement a prototype of TTFL and evaluate it regarding security, privacy, and system performance. Evaluation results show that TTFL can comprehensively and efficiently address the main privacy and security threats in FL. For instance, compared with previous work, it improves the model accuracy by 1.9% and reduces the attack success rate by 79.7% on the CIFAR-10 dataset with only about 19.8% training time overhead. Lizhi Sun, Jingzhou Zhu, Boyu Chang, Yixin Xu 0003, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
TrustCom | 6 |
| 2024 | TIM: Enabling Large-Scale White-Box Testing on In-App Deep Learning ModelsabstractIntelligent Applications (iApps), equipped with in-App deep learning (DL) models, are emerging to provide reliable DL inference services. However, in-App DL models are typically compiled into inference-only versions to enhance system performance, thereby impeding the evaluation of DL models. Specifically, the assessment of in-App models currently relies on black-box testing methods rather than direct white-box testing approaches. In this work, we propose TIM, an automated tool designed for conducting large-scale white-box testing of in-App models. Taking an iApp as input, TIM can lift the black-box (i.e., inference-only) in-App DL model into a backpropagation-enabled one and package it together, allowing comprehensive DL model testing or security issues detection. TIM proposes two reconstruction techniques to convert the inference-only model to a backpropagation-enabled version and reconstruct the DL-related IO processing code. In our experiments, we utilize TIM to extract 100 unique commercial in-App models and convert the models to white-box models, enabling backpropagation functionality. Experimental results show that TIM’s reconstruction techniques exhibit high accuracy. We open-source our prototype and part of the experimental data on the websitehttps://zenodo.org/record/7548141. Hao Wu 0067, Yuhang Gong, Xiaopeng Ke, Hanzhong Liang, Fengyuan Xu, Yunxin Liu 0001, Sheng Zhong 0002 |
IEEE Trans. Inf. Forensics Secur. | 1 |
| 2024 | Multi-Label and Evolvable Dataset Preparation for Web-Based Object DetectionabstractIn this article, we focus on the emerging field of web-based object detection, which has gained considerable attention due to its ability to utilize large amounts of web data for training, thus eliminating the need for labor-intensive manual annotations. However, the noisy and ever-evolving nature of web data poses challenges in preparing high-quality datasets for web-based object detection. To address these challenges, we propose a fully automatic dataset preparation method in this article. Our proposed method incorporates a hierarchical clustering module that assigns multiple precise labels to each image. This module is based on our observation that web image data exhibits different distributions at varying granularities. Furthermore, an evolutionary relabeling module ensures the adaptability of both the prepared dataset and trained detection models to the ever-evolving web data. Extensive experiments demonstrate that our method outperforms other web-based methods, and achieves a comparable performance to those manually labeled benchmark datasets. Shucheng Li, Jingzhou Zhu, Boyu Chang, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ACM Trans. Knowl. Discov. Data | 4 |
| 2023 | GAPter: Gray-Box Data Protector for Deep Learning Inference Services at User SideabstractThe widespread deployment of Deep Learning Inference Services (DLISes) has raised people’s concerns about their data privacy being breached. Although data privacy enhancement has recently attracted a lot of attention, existing solutions all require the cooperation of service providers. Users lose control of their data when making data privacy enhancement decisions. However, it is difficult to enable the user-side control of data abuse prevention because users do not have any programming skills, deep learning knowledge, or rich computing resources. In this work, we propose a fully-automatic userside data privacy enhancement solution, GAPter, for DLISes. Given such a DLIS, GAPter can adaptively fuzz the service for a suitable enhancement strategy, with no cooperation between the DLIS provider and the user. We have implemented and comprehensively evaluated GAPter. The experimental results show that GAPter can find good balance points between privacy enhancement and user data utility. Hao Wu 0067, Xiaopeng Ke, Siyi He, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 1 |
| 2023 | SIEGE: Self-Supervised Incremental Deep Graph Learning for Ethereum Phishing Scam DetectionabstractThe phishing scams pose a serious threat to the ecosystem of Ethereum which is one of the largest blockchains in the world. Such a type of cyberattack recently has caused losses of millions of dollars. In this paper, we propose a Self-supervised IncrEmental deep Graph lEarning (SIEGE) model, for the phishing scam detection problem on Ethereum. To overcome the data scalability challenge, we propose splitting the original Ethereum transaction data and constructing transaction graphs for each split. Confronted with the minimal labeled data available, we resort to graph-based self-supervised learning. We design a spatial pretext task to learn high-quality node embeddings inside a single graph split, as well as an incremental learning paradigm and a temporal pretext task to facilitate information flow between different graph splits. To evaluate the effectiveness of SIEGE, we gather a real-world dataset consisting of six-month Ethereum transaction records. The results demonstrate that our model consistently outperforms baseline approaches in both transductive and inductive settings. Shucheng Li, Runchuan Wang, Hao Wu 0067, Sheng Zhong 0002, Fengyuan Xu |
ACM Multimedia | 3 |
| 2023 | Dataset Preparation for Arbitrary Object Detection: An Automatic Approach based on Web Information in EnglishabstractAutomatic dataset preparation can help users avoid labor-intensive and costly manual data annotations. The difficulty in preparing a high-quality dataset for object detection involves three key aspects: relevance, naturality, and balance, which are not addressed by existing works. In this paper, we leverage information from the web, and propose a fully-automatic dataset preparation mechanism without any human annotation, which can automatically prepare a high-quality training dataset for the detection task with English text terms describing target objects. It contains three key designs, i.e., keyword expansion, data de-noising, and data balancing. Our experiments demonstrate that the object detectors trained with auto-prepared data are comparable to those trained with benchmark datasets and outperform other baselines. We also demonstrate the effectiveness of our approach in several more challenging real-world object categories that are not included in the benchmark datasets. Shucheng Li, Boyu Chang, Hao Wu 0067, Sheng Zhong 0002, Fengyuan Xu |
SIGIR | 4 |
| 2023 | LEAP: TrustZone Based Developer-Friendly TEE for Intelligent Mobile AppsabstractARM TrustZone is widely deployed on commercial-off-the-shelf mobile devices for secure execution. However, many Apps cannot enjoy this feature because it brings many constraints to App developers. Previous works have been proposed to build a secure execution environment for developers on top of TrustZone. Unfortunately, these works are still not a fully-fledged solution for mobile Apps, especially for the emerging intelligent Apps. To this end, we propose LEAP, which is a lightweight developer-friendly TEE solution for mobile Apps. LEAP enables isolated codes to execute in parallel and access peripheral (e.g., mobile GPUs) with ease, flexibly manages system resources upon different workloads, and offers the auto DevOps tool to help developers prepare the codes running on it. We implement the LEAP prototype on the off-the-shelf ARM platform and conduct extensive experiments on it. The experimental results show that Apps can be adapted to run with LEAP easily and efficiently. Compared to the state-of-the-art work along this research line, LEAP can achieve an average 3.57× speedup in supporting intelligent Apps using mobile GPU acceleration. Lizhi Sun, Shuocheng Wang, Hao Wu 0067, Yuhang Gong, Fengyuan Xu, Yunxin Liu 0001, Sheng Zhong 0002 |
IEEE Trans. Mob. Comput. | 3 |
| 2022 | Towards Practical and Efficient Long Video SummaryabstractRecently, video summarization (VS) techniques are widely used to alleviate huge processing pressure brought by numerous long videos. However, it is hard to summarize long videos efficiently since processing hundreds of frames is still time-consuming. In this paper, we find that the Kernel Temporal Segmentation (KTS) method designed for detecting the shot boundaries in SOTA VS methods is time-consuming while handling long videos. To address this issue, we propose the Distribution-based KTS (D-KTS) by fully considering the characteristic of shot length distribution. Furthermore, we propose the Hash-based Adaptive Frame Selection (HAFS) to improve the system performance by fully taking advantage of the temporal locality of long videos. Our experiments present that the proposed D-KTS is 92.70% faster and takes up 90.08% less memory than the baseline KTS method on average. Xiaopeng Ke, Boyu Chang, Hao Wu 0067, Fengyuan Xu, Sheng Zhong 0002 |
ICASSP | 3 |
| 2022 | AVMiner: Expansible and Semantic-Preserving Anti-Virus Labels Mining MethodabstractWith the increase in the variety and quantity of malware, there is an urgent need to speed up the diagnosis and analysis of malware. Extracting the malware family-related tokens from AV (Anti-Virus) labels, provided by online antivirus engines, paves the way for pre-diagnosing the malware. Automatically extracting vital information from AV labels will greatly enhance the detection ability of security enterprises and equip the research ability of security analysts. Recent works like AVCLASS and AVCLASS2 try to extract the attributes of malware from AV labels and establish the taxonomy based on expert knowledge. However, due to the uncertain trend of complicated malicious behaviors, the system needs the following abilities to face the challenge: preserving vital semantics, being expansible, and being free from expert knowledge. In this work, we present AVMiner, an expansible malware tagging system that can mine the most vital tokens from AV labels. AVMiner adopts natural language processing techniques and clustering methods to generate a sequence of tokens without expert knowledge ranked by importance. AVMiner can self-update when new samples come. Finally, we evaluate AVMiner on over 8,000 samples from well-known datasets with manually labeled ground truth, which outperforms previous works. Ligeng Chen, Zhongling He, Hao Wu 0067, Yuhang Gong, Bing Mao 0001 |
TrustCom | 3 |
| 2022 | DIComP: Lightweight Data-Driven Inference of Binary Compiler Provenance with High AccuracyabstractBinary analysis is pervasively utilized to assess software security and test vulnerabilities without accessing source codes. The analysis validity is heavily influenced by the inferring ability of information related to the code compilation. Among the compilation information, compiler type and optimization level, as the key factors determining how binaries look like, are still difficult to be inferred efficiently with existing tools. In this paper, we conduct a thorough empirical study on the binary's appearance under various compilation settings and propose a lightweight binary analysis tool based on the simplest machine learning method, called DIComP to infer the compiler and optimization level via most relevant features according to the observation. Our comprehensive evaluations demonstrate that DIComP can fully recognize the compiler provenance, and it is effective in inferring the optimization levels with up to 90% accuracy. Also, it is efficient to infer thousands of binaries at a millisecond level with our lightweight machine learning model (1MB). Ligeng Chen, Zhongling He, Hao Wu 0067, Fengyuan Xu, Bing Mao 0001 |
SANER | 3 |
| 2021 | PECAM: privacy-enhanced video streaming and analytics via securely-reversible transformationabstractAs Video Streaming and Analytics (VSA) systems become increasingly popular, serious privacy concerns have risen on exposing too much unnecessary private information to the VSA providers. Yet, it is challenging to protect privacy while still preserving desired VSA features, i.e., effective analytics, forensic support, resource efficiency, and real-time execution. In this paper, we present a VSA privacy enhancement system (PECAM), which addresses the above challenge with no change in the VSA back-end. PECAM leverages a novel Generative Adversarial Network to perform the privacy-enhanced securely-reversible video transformation. PECAM also incorporates a couple of system optimizations into its VSA workflow to reduce network bandwidth usage and enable real-time processing on cameras. We implement our PECAM prototype on commodity hardware and evaluate its performance via both security study and extensive experiments. Results demonstrate that PECAM can effectively enhance the visual privacy of VSA in the presence of an adversary, and its transformed videos, when taken as input for various VSA back-end tasks, maintain a 96% accuracy of corresponding original videos. Additionally, it performs 12.3× and 1.8× better than baseline methods in terms of the computing cost and network bandwidth usage, respectively. Hao Wu 0067, Xuejin Tian, Minghao Li 0003, Yunxin Liu 0001, Ganesh Ananthanarayanan, Fengyuan Xu, Sheng Zhong 0002 |
MobiCom | 1 |
| 2021 | DAPter: Preventing User Data Abuse in Deep Learning Inference ServicesabstractThe data abuse issue has risen along with the widespread development of the deep learning inference service (DLIS). Specifically, mobile users worry about their input data being labeled to secretly train new deep learning models that are unrelated to the DLIS they subscribe to. This unique issue, unlike the privacy problem, is about the rights of data owners in the context of deep learning. However, preventing data abuse is demanding when considering the usability and generality in the mobile scenario. In this work, we propose, to our best knowledge, the first data abuse prevention mechanism called DAPter. DAPter is a user-side DLIS-input converter, which removes unnecessary information with respect to the targeted DLIS. The converted input data by DAPter maintains good inference accuracy and is difficult to be labeled manually or automatically for the new model training. DAPter’s conversion is empowered by our lightweight generative model trained with a novel loss function to minimize abusable information in the input data. Furthermore, adapting DAPter requires no change in the existing DLIS backend and models. We conduct comprehensive experiments with our DAPter prototype on mobile devices and demonstrate that DAPter can substantially raise the bar of the data abuse difficulty with little impact on the service quality and overhead. Hao Wu 0067, Xuejin Tian, Yuhang Gong, Minghao Li 0003, Fengyuan Xu |
WWW | 1 |
| 2020 | EMO: real-time emotion recognition from single-eye images for resource-constrained eyewear devicesabstractReal-time user emotion recognition is highly desirable for many applications on eyewear devices like smart glasses. However, it is very challenging to enable this capability on such devices due to tightly constrained image contents (only eye-area images available from the on-device eye-tracking camera) and computing resources of the embedded system. In this paper, we propose and develop a novel system called EMO that can recognize, on top of a resource-limited eyewear device, real-time emotions of the user who wears it. Unlike most existing solutions that require whole-face images to recognize emotions, EMO only utilizes the single-eye-area images captured by the eye-tracking camera of the eyewear. To achieve this, we design a customized deep-learning network to effectively extract emotional features from input single-eye images and a personalized feature classifier to accurately identify a user's emotions. EMO also exploits the temporal locality and feature similarity among consecutive video frames of the eye-tracking camera to further reduce the recognition latency and system resource usage. We implement EMO on two hardware platforms and conduct comprehensive experimental evaluations. Our results demonstrate that EMO can continuously recognize seven-type emotions at 12.8 frames per second with a mean accuracy of 72.2%, significantly outperforming the state-of-the-art approach, and consume much fewer system resources. Hao Wu 0067, Xuejin Tian, Edward Sun, Yunxin Liu 0001, Fengyuan Xu, Sheng Zhong 0002 |
MobiSys | 1 |