EDBT 2026 Demo / reviewers in the wild / expert
Yang Bai 0011
dblp:39/6825-11
· DBLP profile ↗
44ranked-venue papers
13as first author
35since 2021 · last 2026
0000-0002-2475-4232ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 7 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 4 first-author · 11 since 2021Security and privacy · 4 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 3 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Note2Chat: Improving LLMs for Multi-Turn Clinical History Taking Using Medical NotesabstractEffective clinical history taking is a foundational yet underexplored component of clinical reasoning. While large language models (LLMs) have shown promise on static benchmarks, they often fall short in dynamic, multi-turn diagnostic settings that require iterative questioning and hypothesis refinement. To address this gap, we propose Note2Chat, a note-driven framework that trains LLMs to conduct structured history taking and diagnosis by learning from widely available medical notes. Instead of relying on scarce and sensitive dialogue data, we convert real-world medical notes into high-quality doctor-patient dialogues using a decision tree-guided generation and refinement pipeline. We then propose a three-stage fine-tuning strategy combining supervised learning, simulated data augmentation, and preference learning. Furthermore, we propose a novel single-turn reasoning paradigm that reframes history taking as a sequence of single-turn reasoning problems. This design enhances interpretability and enables local supervision, dynamic adaptation, and greater sample efficiency. Experimental results show that our method substantially improves clinical reasoning, achieving gains of +16.9 F1 and +21.0 Top-1 diagnostic accuracy over GPT-4o. Yang Zhou 0017, Zhenting Sheng, Mingrui Tan, Yuting Song, Jun Zhou 0014, Yu Heng Kwan, Lian Leng Low, Yang Bai 0011, Yong Liu 0026 |
AAAI | 8 |
| 2026 | ConRF: Zero-shot stylization of 3D scenes with conditioned radiation fieldsabstract• We propose a novel method that leverages CLIP for zero-shot 3D scene artistic style transfer by a single condition (i.e. image or text). • We introduce a mapping network to alleviate the ambiguity in CLIP features related to style. • We present a 3D selection volume that allows for localized style manipulation within 3D scenes, expanding the possibilities in scene stylization and manipulation. Most of the existing works on arbitrary 3D NeRF style transfer required retraining on each single style condition. This work aims to achieve zero-shot controlled stylization in 3D scenes utilizing text or visual input as conditioning factors. We introduce ConRF, a novel method of zero-shot stylization. Specifically, due to the ambiguity of CLIP features, we employ a conversion process that maps the CLIP feature space to the style space of a pre-trained VGG network and then refine the CLIP multi-modal knowledge into a style transfer neural radiation field. Additionally, we use a 3D volumetric representation to perform local style transfer. By combining these operations, ConRF offers the capability to utilize either text or images as references, resulting in the generation of sequences with novel views enhanced by global or local stylization. Our experiment demonstrates that ConRF outperforms other existing methods for 3D scene and single-text stylization in terms of visual quality. Code is available: https://xingy038.github.io/ConRF/ . Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Fan Wan, Yawen Huang, Yang Long 0001, Yefeng Zheng 0001 |
Pattern Recognit. | 2 |
| 2026 | Boosting adversarial transferability of vision-language pre-trained models via optimal transport
Simeng Qin, Sensen Gao, Dongchen Han, Xiaojun Jia, Yang Bai 0011, Jindong Gu, Xiaochun Cao |
Pattern Recognit. | 6 |
| 2025 | VQA4CIR: Boosting Composed Image Retrieval with Visual Question AnsweringabstractAlbeit progress has been made in Composed Image Retrieval (CIR), we empirically find that a certain percentage of failure retrieval results are not consistent with their relative captions. To address this issue, this work provides a Visual Question Answering (VQA) perspective to boost the performance of CIR. The resulting VQA4CIR is a post-processing approach and can be directly plugged into existing CIR methods. Given the top-C retrieved images by a CIR method, VQA4CIR aims to decrease the adverse effect of the failure retrieval results being inconsistent with the relative caption. To find the retrieved images inconsistent with the relative caption, we resort to the "QA generation → VQA" self-verification pipeline. For QA generation, we suggest fine-tuning LLM (e.g., LLaMA) to generate several pairs of questions and answers from each relative caption. We then fine-tune LVLM (e.g., LLaVA) to obtain the VQA model. By feeding the retrieved image and question to the VQA model, one can find the images inconsistent with relative caption when the answer by VQA is inconsistent with the answer in the QA pair. Consequently, the CIR performance can be boosted by modifying the ranks of inconsistently retrieved images. Experimental results show that our proposed method outperforms state-of-the-art CIR methods on the CIRR and Fashion-IQ datasets. Chun-Mei Feng 0001, Yang Bai 0011, Tao Luo 0014, Zhen Li 0026, Salman Khan 0001, Wangmeng Zuo, Rick Siow Mong Goh, Yong Liu 0026 |
AAAI | 2 |
| 2025 | Protecting Your Video Content: Disrupting Automated Video-based LLM AnnotationsabstractRecently, video-based large language models (video-based LLMs) have achieved impressive performance across various video comprehension tasks. However, this rapid advancement raises significant privacy and security concerns, particularly regarding the unauthorized use of personal video data in automated annotation by video-based LLMs. These unauthorized annotated video-text pairs can then be used to improve the performance of downstream tasks, such as text-to-video generation. To safeguard personal videos from unauthorized use, we propose two series of protective video watermarks with imperceptible adversarial perturbations, named Ramblings and Mutes. Concretely, Ramblings aim to mislead video-based LLMs into generating inaccurate captions for the videos, thereby degrading the quality of video annotations through inconsistencies between video content and captions. Mutes, on the other hand, are designed to prompt video-based LLMs to produce exceptionally brief captions, lacking descriptive detail. Extensive experiments demonstrate that our video watermarking methods effectively protect video data by significantly reducing video annotation performance across various video-based LLMs, showcasing both stealthiness and robustness in protecting personal video content. Our code is available at https://github.com/ttthhl/Protecting_Your_Video_Content. Haitong Liu, Kuofeng Gao, Yang Bai 0011, Jinmin Li, Jinxiao Shan, Tao Dai 0001, Shutao Xia |
CVPR | 3 |
| 2025 | Using Homomorphic Proxy Re-Encryption to Enhance Security and Privacy of Federated Learning-Based Intelligent Connected VehiclesabstractIntelligent connected vehicles (ICVs) are one of the fast‐growing directions that plays a significant role in the area of autonomous driving. To realize collaborative computation among ICVs, federated learning (FL) or federated‐based large language model (FedLLM) as a promising distributed approach has been used to support various collaborative application computations in ICVs scenarios, for example, analyzing vehicle driving information to realize trajectory prediction, voice‐activated controls, conversational AI assistants. Unfortunately, recent research reveals that FL systems are still faced with privacy challenges from honest‐but‐curious server, honest‐but‐curious distributed participants, or the collusion between participants and the server. These threats can lead to the leakage of sensitive private data, such as location information and driving conditions. Homomorphic encryption (HE) is one of the typical mitigation that has few effects on the model accuracy and has been studied before. However, single‐key HE cannot resist collusion between participants and the server, multikey HE is not suitable for ICVs scenarios. In this work, we proposed a novel approach that combines FL with homomorphic proxy re‐encryption (PRE) which is based on participants’ ID information. By doing so, the FL‐based ICVs can be able to successfully defend against privacy threats. In addition, we analyze the security and performance of our method, and the theoretical analysis and the experiment results show that our defense framework with ID‐based homomorphic PRE can achieve a high‐security level and efficient computation. We anticipate that our approach can serve as a fundamental point to support the extensive research on FedLLMs privacy‐preserving. Yang Bai 0011, Yutang Rao, Juan Wang 0017, Gaojie Xing, Xiaoshu Yuan |
IET Inf. Secur. | 1 |
| 2025 | Video Compression Optimization and Rate Control for Cyberspace ApplicationabstractVideo traffic has become the principal part of data resources in the current cyberspace which brings many challenges such as security, stability and scalability of streaming transmission. Moreover, how to ensure high visual quality while obtaining a significant bit-rate reduction has always been the focus of the industry. By constructing a source distortion temporal propagation (SDTP) model, this paper proposes a temporal dependent RDO (TDRDO) algorithm to resolve the global RDO problem in the temporal domain. Besides, a fuzzy logic based rate control (FLRC) algorithm is proposed to robustly regulate encoding bit-rates. The two algorithms have previously been adopted by Audio Video Coding Standard Workgroup of China and integrated into the second generation (AVS2). Experimental results prove the excellence of the proposed algorithms, for significantly improving the AVS2 video coding performance and providing AVS2 with superb efficiency to compete with HEVC/H.265 in modern video compression. Yimin Zhou 0002, Chengzong Peng, Jie Luo 0005, Juelin Liu, Siqi Yang 0009, Juan Wang 0017, Yang Bai 0011 |
Int. J. Pattern Recognit. Artif. Intell. | 7 |
| 2025 | MOVE: Effective and Harmless Ownership Verification via Embedded External FeaturesabstractCurrently, deep neural networks (DNNs) are widely adopted in different applications. Despite its commercial values, training a well-performing DNN is resource-consuming. Accordingly, the well-trained model is valuable intellectual property for its owner. However, recent studies revealed the threats of model stealing, where the adversaries can obtain a function-similar copy of the victim model, even when they can only query the model. In this paper, we propose an effective and harmless model ownership verification (MOVE) to defend against different types of model stealing simultaneously, without introducing new security risks. In general, we conduct the ownership verification by verifying whether a suspicious model contains the knowledge of defender-specified external features. Specifically, we embed the external features by modifying a few training samples with style transfer. We then train a meta-classifier to determine whether a model is stolen from the victim. This approach is inspired by the understanding that the stolen models should contain the knowledge of features learned by the victim model. In particular, we develop our MOVE method under both glass-boxand closed-box settings and analyze its theoretical foundation to provide comprehensive model protection. Extensive experiments on benchmark datasets verify the effectiveness of our method and its resistance to potential adaptive attacks. Yiming Li 0004, Linghui Zhu, Xiaojun Jia, Yang Bai 0011, Yong Jiang 0001, Shutao Xia, Xiaochun Cao, Kui Ren 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Laser: Efficient Language-Guided Segmentation in Neural Radiance FieldsabstractIn this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance. Xingyu Miao, Haoran Duan 0001, Yang Bai 0011, Tejal Shah, Jun Song 0003, Yang Long 0001, Rajiv Ranjan 0001, Ling Shao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Uncertainty-Aware Medical Diagnostic Phrase Identification and GroundingabstractMedical phrase grounding is crucial for identifying relevant regions in medical images based on phrase queries, facilitating accurate image analysis and diagnosis. However, current methods rely on manual extraction of key phrases from medical reports, reducing efficiency and increasing the workload for clinicians. Additionally, the lack of model confidence estimation limits clinical trust and usability. In this paper, we introduce a novel task-Medical Report Grounding (MRG)-which aims to directly identify diagnostic phrases and their corresponding grounding boxes from medical reports in an end-to-end manner. To address this challenge, we propose uMedGround, a a robust and reliable framework that leverages a multimodal large language model to predict diagnostic phrases by embedding a unique token, < $\mathtt {BOX}$BOX >, into the vocabulary to enhance detection capabilities. A vision encoder-decoder processes the embedded token and input image to generate grounding boxes. Critically, uMedGround incorporates an uncertainty-aware prediction model, significantly improving the robustness and reliability of grounding predictions. Experimental results demonstrate that uMedGround outperforms state-of-the-art medical phrase grounding methods and fine-tuned large visual-language models, validating its effectiveness and reliability. This study represents a pioneering exploration of the MRG task, marking the first-ever endeavor in this domain. Additionally, we demonstrate the applicability of uMedGround in medical visual question answering and class-based localization tasks, where it highlights visual evidence aligned with key diagnostic phrases, supporting clinicians in interpreting various types of textual inputs, including free-text reports, visual question answering queries, and class labels. Ke Zou, Yang Bai 0011, Bo Liu 0113, Zhihao Chen 0004, Yang Zhou 0017, Xuedong Yuan, Meng Wang 0038, Xiaojing Shen, Xiaochun Cao, Huazhu Fu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Backdoor Attack and Defense on Deep Learning: A SurveyabstractDeep learning, as an important branch of machine learning, has been widely applied in computer vision, natural language processing, speech recognition, and more. However, recent studies have revealed that deep learning systems are vulnerable to backdoor attacks. Backdoor attackers inject a hidden backdoor into the deep learning model, such that the predictions of the infected model will be maliciously changed if the hidden backdoor is activated by input with a backdoor trigger while behaving normally on any benign sample. This kind of attack can potentially result in severe consequences in the real world. Therefore, research on defending against backdoor attacks has emerged rapidly. In this article, we have provided a comprehensive survey of backdoor attacks, detections, and defenses previously demonstrated on deep learning. We have investigated widely used model architectures, benchmark datasets, and metrics in backdoor research and have classified attacks, detections and defenses based on different criteria. Furthermore, we have analyzed some limitations in existing methods and, based on this, pointed out several promising future research directions. Through this survey, beginners can gain a preliminary understanding of backdoor attacks and defenses. Furthermore, we anticipate that this work will provide new perspectives and inspire extra research into the backdoor attack and defense methods in deep learning. Yang Bai 0011, Gaojie Xing, Zhihong Rao, Chuan Ma 0001, Shiping Wang, Xiaolei Liu 0001, Yimin Zhou 0002, Jiajia Tang, Kaijun Huang, Jiale Kang |
IEEE Trans. Comput. Soc. Syst. | 1 |
| 2024 | Sentence-level Prompts Benefit Composed Image RetrievalabstractComposed image retrieval (CIR) is the task of retrieving specific images by using a query that involves both a reference image and a relative caption. Most existing CIR models adopt the late-fusion strategy to combine visual and language features. Besides, several approaches have also been suggested to generate a pseudo-word token from the reference image, which is further integrated into the relative caption for CIR. However, these pseudo-word-based prompting methods have limitations when target image encompasses complex changes on reference image, e.g., object removal and attribute modification. In this work, we demonstrate that learning an appropriate sentence-level prompt for the relative caption (SPRC) is sufficient for achieving effective composed image retrieval. Instead of relying on pseudo- word-based prompts, we propose to leverage pretrained V-L models, e.g., BLIP-2, to generate sentence-level prompts. By concatenating the learned sentence-level prompt with the relative caption, one can readily use existing text-based image retrieval models to enhance CIR performance. Furthermore, we introduce both image-text contrastive loss and text prompt alignment loss to enforce the learning of suitable sentence-level prompts. Experiments show that our proposed method performs favorably against the state-of-the-art CIR methods on the Fashion-IQ and CIRR datasets. Yang Bai 0011, Xinxing Xu, Yong Liu 0026, Salman Khan 0001, Fahad Shahbaz Khan, Wangmeng Zuo, Rick Siow Mong Goh, Chun-Mei Feng 0001 |
ICLR | 1 |
| 2024 | Inducing High Energy-Latency of Large Vision-Language Models with Verbose ImagesabstractLarge vision-language models (VLMs) such as GPT-4 have achieved exceptional performance across various multi-modal tasks. However, the deployment of VLMs necessitates substantial energy consumption and computational resources. Once attackers maliciously induce high energy consumption and latency time (energy-latency cost) during inference of VLMs, it will exhaust computational resources. In this paper, we explore this attack surface about availability of VLMs and aim to induce high energy-latency cost during inference of VLMs. We find that high energy-latency cost during inference of VLMs can be manipulated by maximizing the length of generated sequences. To this end, we propose verbose images, with the goal of crafting an imperceptible perturbation to induce VLMs to generate long sentences during inference. Concretely, we design three loss objectives. First, a loss is proposed to delay the occurrence of end-of-sequence (EOS) token, where EOS token is a signal for VLMs to stop generating further tokens. Moreover, an uncertainty loss and a token diversity loss are proposed to increase the uncertainty over each generated token and the diversity among all tokens of the whole generated sequence, respectively, which can break output dependency at token-level and sequence-level. Furthermore, a temporal weight adjustment algorithm is proposed, which can effectively balance these losses. Extensive experiments demonstrate that our verbose images can increase the length of generated sequences by 7.87× and 8.56× compared to original images on MS-COCO and ImageNet datasets, which presents potential challenges for various applications. Kuofeng Gao, Yang Bai 0011, Jindong Gu, Shutao Xia, Philip Torr 0001, Zhifeng Li 0001, Wei Liu 0005 |
ICLR | 2 |
| 2024 | UrFound: Towards Universal Retinal Foundation Models via Knowledge-Guided Masked Modeling
Kai Yu 0009, Yang Zhou 0017, Yang Bai 0011, Zhi Da Soh, Xinxing Xu, Rick Siow Mong Goh, Ching Yu Cheng, Yong Liu 0026 |
MICCAI (12) | 3 |
| 2024 | ISPPFL: An incentive scheme based privacy-preserving federated learning for avatar in metaverse
Yang Bai 0011, Gaojie Xing, Zhihong Rao, Chengzong Peng, Yutang Rao, Chuan Ma 0001, Yimin Zhou 0002 |
Comput. Networks | 1 |
| 2024 | CTNeRF: Cross-time Transformer for dynamic neural radiance field from monocular video
Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Fan Wan, Yawen Huang, Yang Long 0001, Yefeng Zheng 0001 |
Pattern Recognit. | 2 |
| 2024 | An Overview of Advanced Deep Graph Node ClusteringabstractGraph data have become increasingly important, and graph node clustering has emerged as a fundamental task in data analysis. In recent years, graph node clustering has gradually moved from traditional shallow methods to deep neural networks due to the powerful representation capabilities of deep learning. In this article, we review some representatives of the latest graph node clustering methods, which are classified into three categories depending on their principles. Extensive experiments are conducted on real-world graph datasets to evaluate the performance of these methods. Four mainstream evaluation performance metrics are used, including clustering accuracy, normalized mutual information, adjusted rand index, and F1-score. Based on the experimental results, several potential research challenges and directions in the field of deep graph node clustering are pointed out. This work is expected to facilitate researchers interested in this field to provide some insights and further promote the development of deep graph node clustering. Shiping Wang, Jinbin Yang, Yang Bai 0011, William Zhu 0001 |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2024 | DS-Depth: Dynamic and Static Depth Estimation via a Fusion Cost VolumeabstractSelf-supervised monocular depth estimation methods typically rely on the reprojection error to capture geometric relationships between successive frames in static environments. However, this assumption does not hold in dynamic objects in scenarios, leading to errors during the view synthesis stage, such as feature mismatch and occlusion, which can significantly reduce the accuracy of the generated depth maps. To address this problem, we propose a novel dynamic cost volume that exploits residual optical flow to describe moving objects, improving incorrectly occluded regions in static cost volumes used in previous work. Nevertheless, the dynamic cost volume inevitably generates extra occlusions and noise, thus we alleviate this by designing a fusion module that makes static and dynamic cost volumes compensate for each other. In other words, occlusion from the static volume is refined by the dynamic volume, and incorrect information from the dynamic volume is eliminated by the static volume. Furthermore, we propose a pyramid distillation loss to reduce photometric error inaccuracy at low resolutions and an adaptive photometric error loss to alleviate the flow direction of the large gradient in the occlusion regions. We conducted extensive experiments on the KITTI and Cityscapes datasets, and the results demonstrate that our model outperforms previously published baselines for self-supervised monocular depth estimation. Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Yawen Huang, Fan Wan, Xinxing Xu, Yang Long 0001, Yefeng Zheng 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Fast Propagation Is Better: Accelerating Single-Step Adversarial Training via Sampling SubnetworksabstractAdversarial training has shown promise in building robust models against adversarial examples. A major drawback of adversarial training is the computational overhead introduced by the generation of adversarial examples. To overcome this limitation, adversarial training based on single-step attacks has been explored. Previous work improves the single-step adversarial training from different perspectives, e.g., sample initialization, loss regularization, and training strategy. Almost all of them treat the underlying model as a black box. In this work, we propose to exploit the interior building blocks of the model to improve efficiency. Specifically, we propose to dynamically sample lightweight subnetworks as a surrogate model during training. By doing this, both the forward and backward passes can be accelerated for efficient adversarial training. Besides, we provide theoretical analysis to show the model robustness can be improved by the single-step adversarial training with sampled subnetworks. Furthermore, we propose a novel sampling strategy where the sampling varies from layer to layer and from iteration to iteration. Compared with previous methods, our method not only reduces the training cost but also achieves better model robustness. Evaluations on a series of popular datasets demonstrate the effectiveness of the proposed FB-Better. Our code has been released at https://github.com/jiaxiaojunQAQ/FP-Better. Xiaojun Jia, Jianshu Li, Jindong Gu, Yang Bai 0011, Xiaochun Cao |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2023 | Backdoor Defense via Adaptively Splitting Poisoned DatasetabstractBackdoor defenses have been studied to alleviate the threat of deep neural networks (DNNs) being backdoor attacked and thus maliciously altered. Since DNNs usually adopt some external training data from an untrusted third party, a robust backdoor defense strategy during the training stage is of importance. We argue that the core of training-time defense is to select poisoned samples and to handle them properly. In this work, we summarize the training-time defenses from a unified framework as splitting the poisoned dataset into two data pools. Under our framework, we propose an adaptively splitting dataset-based defense (ASD). Concretely, we apply loss-guided split and meta-learning-inspired split to dynamically update two data pools. With the split clean data pool and polluted data pool, ASD successfully defends against backdoor attacks during training. Extensive experiments on multiple benchmark datasets and DNN models against six state-of-the-art backdoor attacks demonstrate the superiority of our ASD. Our code is available at https://github.com/KuofengGao/ASD. Kuofeng Gao, Yang Bai 0011, Jindong Gu, Yong Yang 0001, Shutao Xia |
CVPR | 2 |
| 2023 | Towards Few-shot Image Captioning with Cycle-based Compositional Semantic Enhancement FrameworkabstractMany efforts paid attention to the multi-modal task, of which image captioning is a classic work. Especially the Clip model improves the performance of image captioning; meantime, its few-shot and zero-shot problems have become a significant research project. In this work, aiming at the image captioning task, we design the new few-shot and zero-shot settings different from popular directions. The direction focuses on the impact of the exited dataset for captioning model ability. According to analysis, we discover the frequency of the word combination can directly influence the performance of the captioning model. Based on this, we define the new few-shot and zero-shot settings. In terms of this, a Cycle-based captioning framework based on data augmentation is proposed to overcome this problem, of which the novelty switcher module is the critical component. Finally, experiments demonstrate that our framework can achieve state-of-the-art performance on both traditional, few-shot and zero-shot settings. Peng Zhang 0058, Yang Bai 0011, Jie Su 0001, Yan Huang 0008, Yang Long 0001 |
IJCNN | 2 |
| 2023 | Query efficient black-box adversarial attack on deep neural networks
Yang Bai 0011, Yisen Wang 0001, Yuyuan Zeng, Yong Jiang 0001, Shutao Xia |
Pattern Recognit. | 1 |
| 2023 | BlockExplorer: Exploring Blockchain Big Data Via Parallel ProcessingabstractToday's blockchain systems store detailed runtime information in the format of transactions and blocks, which are valuable not only to understand the finance of blockchain-based ecosystems but also to audit the security of on-chain applications. However, exploring this blockchain “big data” is challenging due to data heterogeneity and the huge amount. Existing blockchain exploration techniques are either incomplete or inefficient, making them inapt in time-sensitive applications. This paper presents ${\sf BlockExplorer}$ , an efficient and flexible blockchain exploration system for Ethereum. ${\sf BlockExplorer}$ builds on a master-slave architecture, where the master partitions all blocks into multiple non-overlapped sets and each slave simultaneously processes Ethereum big data based on a set of blocks. ${\sf BlockExplorer}$ implements a transaction-based partitioning approach to address load balance among slaves, and a code instrumentation approach to acquire complete Ethereum big data. The evaluation shows that ${\sf BlockExplorer}$ accelerates the data acquisition performance of the state-of-the-art by 4.1×, while the workload difference among slaves is up to 18%. To demonstrate the application of ${\sf BlockExplorer}$ , we develop three apps upon ${\sf BlockExplorer}$ to detect real-life attacks against Ethereum and show that our apps can detect attacks in a large range of blocks (e.g., ten million) within a short time (e.g., multiple hours). Jingwei Li 0001, Yuxing Tang, Xiapu Luo, Zheyuan He, Zihao Li 0001, Yang Bai 0011, Ting Chen 0002, Yuzhe Tang, Zhe Liu 0001, Xiaosong Zhang 0001 |
IEEE Trans. Computers | 8 |
| 2023 | Interpretable Graph Convolutional Network for Multi-View Semi-Supervised LearningabstractAs real-world data become increasingly heterogeneous, multi-view semi-supervised learning has garnered widespread attention. Although existing studies have made efforts towards this and achieved decent performance, they are restricted to shallow models and how to mine deeper information from multiple views remains to be investigated. As a recently emerged neural network, Graph Convolutional Network (GCN) exploits graph structure to propagate label signals and has achieved encouraging performance, and it has been widely employed in various fields. Nonetheless, research on solving multi-view learning problems via GCN is limited and lacks interpretability. To address this gap, in this paper we propose a framework termed Interpretable Multi-view Graph Convolutional Network (IMvGCN11Code is available athttps://github.com/ZhihaoWu99/IMvGCN.). We first combine the reconstruction error and Laplacian embedding to formulate a multi-view learning problem that explores the original space from feature and topology perspectives. In light of a series of derivations, we establish a potential connection between GCN and multi-view learning, which holds significance for both domains. Furthermore, we propose an orthogonal normalization method to guarantee the mathematical connection, which solves the intractable problem of orthogonal constraints in deep learning. In addition, the proposed framework is applied to the multi-view semi-supervised learning task. Comprehensive experiments demonstrate the superiority of our proposed method over other state-of-the-art methods. Zhihao Wu 0003, Xincan Lin, Zhenghong Lin, Zhaoliang Chen, Yang Bai 0011, Shiping Wang |
IEEE Trans. Multim. | 5 |
| 2023 | TokenAware: Accurate and Efficient Bookkeeping Recognition for Token Smart ContractsabstractTokens have become an essential part of blockchain ecosystem, so recognizing token transfer behaviors is crucial for applications depending on blockchain. Unfortunately, existing solutions cannot recognize token transfer behaviors accurately and efficiently because of their incomplete patterns and inefficient designs. This work proposes TokenAware , a novel online system for recognizing token transfer behaviors. To improve accuracy, TokenAware infers token transfer behaviors from modifications of internal bookkeeping of a token smart contract for recording the information of token holders (e.g., their addresses and shares). However, recognizing bookkeeping is challenging, because smart contract bytecode does not contain type information. TokenAware overcomes the challenge by first learning the instruction sequences for locating basic types and then deriving the instruction sequences for locating sophisticated types that are composed of basic types. To improve efficiency, TokenAware introduces four optimizations. We conduct extensive experiments to evaluate TokenAware with real blockchain data. Results show that TokenAware can automatically identify new types of bookkeeping and recognize 107,202 tokens with 98.7% precision. TokenAware with optimizations merely incurs 4% overhead, which is 1/345 of the overhead led by the counterpart with no optimization. Moreover, we develop an application based on TokenAware to demonstrate how it facilitates malicious behavior detection. Zheyuan He, Shuwei Song, Yang Bai 0011, Xiapu Luo, Ting Chen 0002, Hongwei Li 0001, Xiaodong Lin 0001, Xiaosong Zhang 0001 |
ACM Trans. Softw. Eng. Methodol. | 3 |
| 2022 | Action Quality Assessment with Temporal Parsing Transformer
Yang Bai 0011, Desen Zhou, Songyang Zhang 0001, Jian Wang 0066, Errui Ding, Yu Guan 0001, Yang Long 0001, Jingdong Wang 0001 |
ECCV (4) | 1 |
| 2022 | Watermark Vaccine: Adversarial Attacks to Prevent Watermark Removal
Jian Liu 0012, Yang Bai 0011, Jindong Gu, Xiaojun Jia, Xiaochun Cao |
ECCV (14) | 3 |
| 2022 | Imitated Detectors: Stealing Knowledge of Black-box Object DetectorsabstractDeep neural networks have shown great potential in many practical applications, yet their knowledge is at the risk of being stolen via exposed services (\eg APIs). In contrast to the commonly-studied classification model extraction, there exist no studies on the more challenging object detection task due to the sufficiency and efficiency of problem domain data collection. In this paper, we for the first time reveal that black-box victim object detectors can be easily replicated without knowing the model structure and training data. In particular, we treat it as black-box knowledge distillation and propose a teacher-student framework named Imitated Detector to transfer the knowledge of the victim model to the imitated model. To accelerate the problem domain data construction, we extend the problem domain dataset by generating synthetic images, where we apply the text-image generation process and provide short text inputs consisting of object categories and natural scenes; to promote the feedback information, we aim to fully mine the latent knowledge of the victim model by introducing an iterative adversarial attack strategy, where we feed victim models with transferable adversarial examples making victim provide diversified predictions with more information. Extensive experiments on multiple datasets in different settings demonstrate that our approach achieves the highest model extraction accuracy and outperforms other model stealing methods by large margins in the problem domain dataset. Our codes can be found at \urlhttps://github.com/LiangSiyuan21/Imitated-Detectors. Siyuan Liang 0004, Aishan Liu, Longkang Li, Yang Bai 0011, Xiaochun Cao |
ACM Multimedia | 5 |
| 2022 | Untargeted Backdoor Watermark: Towards Harmless and Stealthy Dataset Copyright ProtectionabstractDeep neural networks (DNNs) have demonstrated their superiority in practice. Arguably, the rapid development of DNNs is largely benefited from high-quality (open-sourced) datasets, based on which researchers and developers can easily evaluate and improve their learning methods. Since the data collection is usually time-consuming or even expensive, how to protect their copyrights is of great significance and worth further exploration. In this paper, we revisit dataset ownership verification. We find that existing verification methods introduced new security risks in DNNs trained on the protected dataset, due to the targeted nature of poison-only backdoor watermarks. To alleviate this problem, in this work, we explore the untargeted backdoor watermarking scheme, where the abnormal model behaviors are not deterministic. Specifically, we introduce two dispersibilities and prove their correlation, based on which we design the untargeted backdoor watermark under both poisoned-label and clean-label settings. We also discuss how to use the proposed untargeted backdoor watermark for dataset ownership verification. Experiments on benchmark datasets verify the effectiveness of our methods and their resistance to existing backdoor defenses. Yiming Li 0004, Yang Bai 0011, Yong Jiang 0001, Yong Yang 0001, Shutao Xia, Bo Li 0026 |
NeurIPS | 2 |
| 2021 | GANMIA: GAN-based Black-box Membership Inference AttackabstractMembership inference attacks (MIAs) against machine learning systems have drawn tremendous attention from information security researchers. By MIA, an adversary can speculate whether an individual data record is a member of the training set or not. Existing black-box MIA assumes that much information about the training data is available. Specifically, the attacker assumes that (s)he has the ability to query the target model without limitations or can access a sufficient dataset whose distribution is the same as the training data set. However, in a realistic scenario, MIAs usually come up with the limited number and the imbalanced proportion of target training datasets which cause significant challenges for MIAs. To launch an MIA in the realistic scenario, in this paper, we present a novel method called GANMIA, which generates synthetic data to augment the training samples of the shadow model for the black-box MIA by a Generative Adversarial Network (GAN). GANMIA firstly augments synthesized samples and then uses the generated samples to train the given shadow model to increase the training efficiency, and additionally improve the MIA’s performance. The experimental results show that the accuracy of the black-box MIA increases by 23% with the help of our synthetic data. Yang Bai 0011, Degang Chen 0003, Ting Chen 0002, Mingyu Fan |
ICC | 1 |
| 2021 | Improving Adversarial Robustness via Channel-wise Activation Suppressing
Yang Bai 0011, Yuyuan Zeng, Yong Jiang 0001, Shutao Xia, Xingjun Ma, Yisen Wang 0001 |
ICLR | 1 |
| 2021 | D2Defend: Dual-Domain based Defense against Adversarial ExamplesabstractConvolutional neural networks (CNNs) have recently been widely applied in computer vision tasks, yet they are seriously vulnerable to imperceptible adversarial perturbations. Such phenomena have caused great attention on the adversary topic. Existing adversarial defense methods mainly focus on improving the robustness of models (e.g., adversarial training) or removing adversarial perturbations (e.g., input-transformation based methods) directly, while rarely considering the accurate recovery of image structures of the input, which also play a vital role in making predictions for CNNs. To this end, we propose a Dual-Domain based Defense (D2Defend) method by recovering low-frequency and high-frequency image structures in both spatial and transform domains, while removing adversarial perturbations simultaneously. Unlike the existing input-transformation based methods, our method can decompose the input image into edge feature and texture feature layers, accompanied with bilateral filtering and short-time fourier transform (STFT) filtering. Experimental results demonstrate the effectiveness of our method against various adversarial attacks, and show the superiority of our method over other adversarial defense methods especially at strong adversarial strength. Tao Dai 0001, Yang Bai 0011, Shutao Xia |
IJCNN | 4 |
| 2021 | Discriminative Latent Semantic Graph for Video CaptioningabstractVideo captioning aims to automatically generate natural language sentences that can describe the visual contents of a given video. Existing generative models like encoder-decoder frameworks cannot explicitly explore the object-level interactions and frame-level information from complex spatio-temporal data to generate semantic-rich captions. Our main contribution is to identify three key problems in a joint framework for future video summarization tasks. 1) Enhanced Object Proposal: we propose a novel Conditional Graph that can fuse spatio-temporal information into latent object proposal. 2) Visual Knowledge: Latent Proposal Aggregation is proposed to dynamically extract visual words with higher semantic levels. 3) Sentence Validation: A novel Discriminative Language Validator is proposed to verify generated captions so that key semantic concepts can be effectively preserved. Our experiments on two public datasets (MVSD and MSR-VTT) manifest significant improvements over state-of-the-art approaches on all metrics, especially for BLEU-4 and CIDEr. Our code is available at https://github.com/baiyang4/D-LSG-Video-Caption. Yang Bai 0011, Junyan Wang 0001, Yang Long 0001, Bingzhang Hu, Yang Song 0001, Maurice Pagnucco, Yu Guan 0001 |
ACM Multimedia | 1 |
| 2021 | Clustering Effect of Adversarial Robust ModelsabstractAdversarial robustness has received increasing attention along with the study of adversarial examples. So far, existing works show that robust models not only obtain robustness against various adversarial attacks but also boost the performance in some downstream tasks. However, the underlying mechanism of adversarial robustness is still not clear. In this paper, we interpret adversarial robustness from the perspective of linear components, and find that there exist some statistical properties for comprehensively robust models. Specifically, robust models show obvious hierarchical clustering effect on their linearized sub-networks, when removing or replacing all non-linear components (e.g., batch normalization, maximum pooling, or activation layers). Based on these observations, we propose a novel understanding of adversarial robustness and apply it on more tasks including domain adaption and robustness boosting. Experimental evaluations demonstrate the rationality and superiority of our proposed clustering strategy. Our code is available at https://github.com/bymavis/AdvWeightNeurIPS2021. Yang Bai 0011, Yong Jiang 0001, Shutao Xia, Yisen Wang 0001 |
NeurIPS | 1 |
| 2021 | A Defense Framework for Privacy Risks in Remote Machine Learning ServiceabstractIn recent years, machine learning approaches have been widely adopted for many applications, including classification. Machine learning models deal with collective sensitive data usually trained in a remote public cloud server, for instance, machine learning as a service (MLaaS) system. In this scene, users upload their local data and utilize the computation capability to train models, or users directly access models trained by MLaaS. Unfortunately, recent works reveal that the curious server (that trains the model with users’ sensitive local data and is curious to know the information about individuals) and the malicious MLaaS user (who abused to query from the MLaaS system) will cause privacy risks. The adversarial method as one of typical mitigation has been studied by several recent works. However, most of them focus on the privacy-preserving against the malicious user; in other words, they commonly consider the data owner and the model provider as one role. Under this assumption, the privacy leakage risks from the curious server are neglected. Differential privacy methods can defend against privacy threats from both the curious sever and the malicious MLaaS user by directly adding noise to the training data. Nonetheless, the differential privacy method will decrease the classification accuracy of the target model heavily. In this work, we propose a generic privacy-preserving framework based on the adversarial method to defend both the curious server and the malicious MLaaS user. The framework can adapt with several adversarial algorithms to generate adversarial examples directly with data owners’ original data. By doing so, sensitive information about the original data is hidden. Then, we explore the constraint conditions of this framework which help us to find the balance between privacy protection and the model utility. The experiments’ results show that our defense framework with the AdvGAN method is effective against MIA and our defense framework with the FGSM method can protect the sensitive data from direct content exposed attacks. In addition, our method can achieve better privacy and utility balance compared to the existing method. Yang Bai 0011, Mingchuang Xie, Mingyu Fan |
Secur. Commun. Networks | 1 |
| 2020 | A Black-Box Attack on Neural Networks Based on Swarm Evolutionary Algorithm
Xiaolei Liu 0001, Kangyi Ding, Yang Bai 0011, Weina Niu |
ACISP | 4 |
| 2020 | Improving Query Efficiency of Black-Box Adversarial Attack
Yang Bai 0011, Yuyuan Zeng, Yong Jiang 0001, Yisen Wang 0001, Shutao Xia, Weiwei Guo |
ECCV (25) | 1 |
| 2020 | Self-Adaptive Feature FoolabstractRecently, deep neural networks (DNNs) are shown to be susceptible to data-agnostic quasi-imperceptible noises called Universal Adversarial Perturbations (UAPs). Moreover, the techniques to craft UAPs can be categorized into data-driven and data-free. However, data-free techniques craft UAPs without utilizing any data samples and therefore result in weaker attack capacity. In this paper, we propose a novel method to craft UAPs in the absence of data, via adaptively perturbing mid-layer outputs of the CNN. Based on our proposed self-adaptive attention mechanism, we explore the effects of feature correlation of the internal representations on generating UAPs for the first time. Experimental evaluation demonstrates that UAPs crafted by our Self-Adaptive Feature Fool (SAFF) approach achieve state-of-the-art performance in data-free scenarios. Yang Bai 0011, Shutao Xia, Yong Jiang 0001 |
ICASSP | 2 |
| 2020 | Query Twice: Dual Mixture Attention Meta Learning for Video SummarizationabstractVideo summarization aims to select representative frames to retain high-level information, which is usually solved by predicting the segment-wise importance score via a softmax function. However, softmax function suffers in retaining high-rank representations for complex visual or sequential information, which is known as the Softmax Bottleneck problem. In this paper, we propose a novel framework named Dual Mixture Attention (DMASum) model with Meta Learning for video summarization that tackles the softmax bottleneck problem, where the Mixture of Attention layer (MoA) effectively increases the model capacity by employing twice self-query attention that can capture the second-order changes in addition to the initial query-key attention, and a novel Single Frame Meta Learning rule is then introduced to achieve more generalization to small datasets with limited training sources. Furthermore, the DMASum significantly exploits both visual and sequential attention that connects local key-frame and global attention in an accumulative way. We adopt the new evaluation protocol on two public datasets, SumMe, and TVSum. Both qualitative and quantitative experiments manifest significant improvements over the state-of-the-art methods. Junyan Wang 0001, Yang Bai 0011, Yang Long 0001, Bingzhang Hu, Zhenhua Chai, Yu Guan 0001, Xiaolin Wei |
ACM Multimedia | 2 |
| 2019 | Improved Forward-Backward Propagation to Generate Adversarial Examples
Yuying Hao, Tuanhui Li, Yang Bai 0011, Li Li 0013, Yong Jiang 0001, Xuanye Cheng |
ICANN (3) | 3 |
| 2019 | Hilbert-Based Generative Defense for Adversarial ExamplesabstractAdversarial perturbations of clean images are usually imperceptible for human eyes, but can confidently fool deep neural networks (DNNs) to make incorrect predictions. Such vulnerability of DNNs raises serious security concerns about their practicability in security-sensitive applications. To defend against such adversarial perturbations, recently developed PixelDefend purifies a perturbed image based on PixelCNN in a raster scan order (row/column by row/column). However, such scan mode insufficiently exploits the correlations between pixels, which further limits its robustness performance. Therefore, we propose a more advanced Hilbert curve scan order to model the pixel dependencies in this paper. Hilbert curve could well preserve local consistency when mapping from 2-D image to 1-D vector, thus the local features in neighboring pixels can be more effectively modeled. Moreover, the defensive power can be further improved via ensembles of Hilbert curve with different orientations. Experimental results demonstrate the superiority of our method over the state-of-the-art defenses against various adversarial attacks. Yang Bai 0011, Yisen Wang 0001, Tao Dai 0001, Shutao Xia, Yong Jiang 0001 |
ICCV | 1 |
| 2019 | KnightKing: a fast distributed graph random walk engineabstractRandom walk on graphs has recently gained immense popularity as a tool for graph data analytics and machine learning. Currently, random walk algorithms are developed as individual implementations and suffer significant performance and scalability problems, especially with the dynamic nature of sophisticated walk strategies. Kang Chen 0001, Xiaosong Ma, Yang Bai 0011, Yong Jiang 0001 |
SOSP | 5 |
| 2015 | Test Generation for Embedded Executables via Concolic Execution in a Real EnvironmentabstractTraditional software testing methods are not effective for testing embedded software thoroughly due to the fact that generating effective test inputs to cover all code is extremely difficult. In this work, we propose an automatic method to generate test inputs for embedded executables which is based on concolic execution. The core idea of our method is to divide concolic execution into symbolic execution on hosts, and concrete execution on targets, so considerable development work can be saved. Our method overcomes the limitations of the software and hardware abilities of embedded systems by restricting heavy-weight work on resourceful hosts. One feature of our method is that it targets executables, so the source of tested software is not needed. Another feature is that tested programs run in a real environment rather than in a simulator, so accurate run-time information can be acquired. Symbolic execution and concrete execution are coordinated by cross-debugging functions. Then we implement our method on Wind River VxWorks. Experiments show that our method achieves high code coverage with acceptable speed. Ting Chen 0002, Xiaosong Zhang 0001, Xiao-li Ji, Cong Zhu, Yang Bai 0011 |
IEEE Trans. Reliab. | 5 |
| 2014 | Conpy: Concolic Execution Engine for Python Applications
Ting Chen 0002, Xiaosong Zhang 0001, Rui-dong Chen, Yang Bai 0011 |
ICA3PP (2) | 5 |