VLDB 2026 Research / reviewers in the wild / expert
Yazhe Wang
dblp:88/9436
· DBLP profile ↗
28ranked-venue papers
6as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Security and privacy · 8 · 1 first-authorDatabases, data management, data science and information retrieval · 6 · 5 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 6 · 1 first-author · 5 since 2021Computer networks · 5 · 4 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 2 since 2021Systems, architecture and hardware · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Automated Detection of Pre-training Text in Black-box LLMsabstractDetecting whether a given text is a member in the pre-training data of Large Language Models (LLMs) is crucial for ensuring data privacy and copyright protection. Most existing methods rely on the LLM's hidden information (e.g., model parameters or token probabilities), making them ineffective in the black-box setting, where only input and output texts are accessible. Although some methods have been proposed for the black-box setting, they rely on massive manual efforts such as designing complicated questions or instructions. To address these issues, we propose VeilProbe, the first framework for automatically detecting LLMs' pre-training texts in a black-box setting without human intervention. VeilProbe utilizes a sequence-to-sequence mapping model to infer the latent mapping feature between the input text and the corresponding output suffix generated by the LLM. Then it performs the key token perturbations to obtain more distinguishable membership features. Additionally, considering real-world scenarios where the ground-truth training text samples are limited, a prototype-based membership classifier is introduced to alleviate the overfitting issue. Extensive evaluations on three widely used datasets demonstrate that our framework is effective and superior in the black-box setting. Ruihan Hu, Yuming Shang, Jiankun Peng, Wei Luo 0016, Yazhe Wang, Xi Zhang 0008 |
IJCAI | 5 |
| 2025 | DADN: A Dynamic Anomaly Detection Network for Multivariate Time Series Data of the Industrial Internet of ThingsabstractIndustrial Internet of Things (IIoT) faces significant security challenges such as data privacy and vulnerabilities. Unsupervised anomaly detection aims to identify abnormal patterns by monitoring multivariate time series data of IIoT without anomaly annotation. Previous deep-learning-based methods have high-computation cost, which hinders their deployment in edge devices. In this article, we propose a dynamic anomaly detection network (DADN), which introduces a dynamic anomaly detection mechanism to enable efficient inference. Specifically, a bilateral early-exit mechanism is designed so that each sample can dynamically exit at a certain layer during the forward process to support the anomaly judgement, and the layer where sample exits is adaptively determined at the inference stage. Experimental results show that DADN significantly reduces computational costs and enhances F1 scores in industrial anomaly-detection benchmarks, as shown by a 58.22% decrease in GFLOPS on the SWAT dataset, outperforming previous representative method (anomaly transformer). Yusheng Kong, Lei Ren 0001, Guoliang Kang, Yazhe Wang, Jinhu Lü 0001 |
IEEE Trans. Ind. Informatics | 5 |
| 2024 | NAPGuard: Towards Detecting Naturalistic Adversarial PatchesabstractRecently, the emergence of naturalistic adversarial patch (NAP), which possesses a deceptive appearance and various representations, underscores the necessity of developing robust detection strategies. However, existing approaches fail to differentiate the deep-seated natures in adversarial patches, i.e., aggressiveness and naturalness, leading to unsatisfactory precision and generalization against NAPs. To tackle this issue, we propose NAP-Guard to provide strong detection capability against NAPs via the elaborated critical feature modulation framework. For improving precision, we propose the aggressive feature aligned learning to enhance the model's capability in capturing accurate aggressive patterns. Considering the challenge of inaccurate model learning caused by deceptive appearance, we align the aggressive features by the proposed pattern alignment loss during training. Since the model could learn more accurate aggressive patterns, it is able to detect deceptive patches more precisely. To enhance generalization, we design the natural feature suppressed inference to universally mitigate the disturbance from different NAPs. Since various representations arise in diverse disturbing forms to hinder generalization, we suppress the natural features in a unified approach via the feature shield module. Therefore, the models could recognize NAPs within less disturbance and activate the generalized detection ability. Extensive experiments show that our method surpasses state-of-the-art methods by large margins in detecting NAPs (improve 60.24% [email protected] on average).11Our code is available at https://github.com/wsynuiag/NAPGaurd. Siyang Wu, Jiakai Wang, Jiejie Zhao, Yazhe Wang, Xianglong Liu 0001 |
CVPR | 4 |
| 2023 | Auto-Tuning with Reinforcement Learning for Permissioned Blockchain SystemsabstractIn a permissioned blockchain, performance dictates its development, which is substantially influenced by its parameters. However, research on auto-tuning for better performance has somewhat stagnated because of the difficulty posed by distributed parameters; thus, it is possible only with difficulty to propose an effective auto-tuning optimization scheme. To alleviate this issue, we lay a solid basis for our research by first exploring the relationship between parameters and performance in Hyperledger Fabric, a permissioned blockchain, and we propose Athena, a Fabric-based auto-tuning system that can automatically provide parameter configurations for optimal performance. The key of Athena is designing a new Permissioned Blockchain Multi-Agent Deep Deterministic Policy Gradient (PB-MADDPG) to realize heterogeneous parameter-tuning optimization of different types of nodes in Fabric. Moreover, we select parameters with the most significant impact on accelerating recommendation. In its application to Fabric, a typical permissioned blockchain system, with 12 peers and 7 orderers, Athena achieves a throughput improvement of 470.45% and a latency reduction of 75.66% over the default configuration. Compared with the most advanced tuning schemes (CDBTune, Qtune, and ResTune), our method is competitive in terms of throughput and latency. Yazhe Wang, Shuai Ma 0001, Chao Liu 0020, Dongdong Huo, Yu Wang 0243, Zhen Xu 0009 |
Proc. VLDB Endow. | 2 |
| 2022 | FIK: Find Important Knobs for Permessioned Blockchain Hyperledger FabricabstractPermissioned blockchain technology has a significant impact on many industries due to its security and scalability, which has many configurable knobs to control all aspects of it. Too many knobs lead to an exponential increase in the configurable knob combination, so it is vital to find a subset of knobs that affect the performance. This article proposes FIK (Find Important Knobs), which uses (1) Latin hypercube sampling technology to create a data set of knobs and performance. The data set includes more than 10000 samples of three different network architectures of the permissioned blockchain network and three different chaincode. (2) The relationship between knobs and performance is modeled based on the Xgboost method. The performance of this model is better than other selected machine learning models. (3) The SHAP-based method helps decision-makers explain the contribution of the knobs of the permissioned blockchain to performance. Nanfang Li, Yazhe Wang, Zong-Rong Li, Dongdong Huo |
CSCWD | 3 |
| 2022 | A highly reliable cross-domain identity authentication protocol based on blockchain in edge computing environmentabstractAiming at the urgent need for a cross-domain authentication scheme with low computational cost, low latency, and high security in the edge computing environment, this paper proposes a blockchain-based cross-domain authentication protocol. In this protocol, the IoT devices in the edge computing environment themselves control the generation, encrypted storage, registration, use or cancellation of their identities, and the authenticity and credibility of these identities can be guaranteed through a trusted identity checking mechanism. The public information of these identities is shared among edge nodes of different domains using the consensus mechanism of blockchain. IoT devices can perform cross-domain authentication at the edge. The authentication process not only makes use of the advantages of low latency of edge computing, but also minimizes the disclosure of detailed identity information. Security analysis shows that the protocol can resist common attacks such as DDoS Attacks and Impersonation Attacks. Compared with several previous cross-domain authentication schemes, our protocol is more secure. Besides, this protocol has the advantage of low computational cost, suitable for low-performance IoT devices. Penghui Lv, Yu Wang 0243, Yazhe Wang, Chao Liu 0020, Qihui Zhou, Zhen Xu 0009 |
CSCWD | 3 |
| 2022 | An atomic member addition mechanism for permissioned blockchain based on autonomous rollbackabstractAs a distributed ledger technology, the addition of new members in permissioned blockchain is usually composed of several steps among distributed nodes. The addition can not be considered successful until all of the steps are completed. In other words, these steps are an atomic operation. However, there is no solution for the atomic operation in existing permissioned blockchain, leading to an inconsistent state when the addition of new members is partially completed. To implement the atomic member addition in permissioned blockchain, we propose a method targeting at the atomic addition of new members based on distributed and autonomous rollback. After member addition starts, distributed nodes of existing members detect the new node and decide whether to rollback or not, instead of getting commands from the coordinator. After deciding to rollback, a new configuration block is added to achieve rollback of the uncompleted member addition. In order for the new configuration block to pass the policy validation of orderers, we set a rollback mode for orderers. The evaluation results show that our method can actually implement atomic member addition and has little impact on performance. Qihui Zhou, Xianglin Dang, Yazhe Wang, Zhen Xu 0009, Penghui Lv |
MSN | 3 |
| 2021 | CallChain: Identity Authentication Based on Blockchain for Telephony NetworksabstractTelephony networks serve as a reliable communication channel, and phone numbers are usually used to authenticate users' identities in telephony networks. However, modern telephony networks does not provide end-to-end authentication for phone numbers, making it difficult for a receiver to confirm a caller's real identity. Moreover, due to the lack of effective authentication solutions for users, phone numbers are usually exploited to commit identity theft, credit card fraud, or telecom harassment by phone number spoofing attacks. In this paper, we propose a blockchain-based identity authentication system for telephony networks, CallChain. CallChain issues verifiable digital identities to users which bound a user to his/her phone number. Based on the digital identity, CallChain provides end-to-end identity authentication between the caller and receiver in the data channel without modifying the core technologies in telephony networks. In the proposed system, users first register decentralized identities (DIDs) and store them in the blockchain distributed ledger. Then Call Center issues verifiable PhoneNumber Credentials to users. And finally users conduct end-to-end identity authentication based on these PhoneNumber Credentials before the call is established. We implement a prototype as an App for Android-based phones based on an open source blockchain platform Hyperledger Indy. Moreover, we detect that CallChain only adds a worst-case 2.1 seconds call establishment overhead and can resist DDoS attacks effectively. Yazhe Wang, Yu Wang 0243, Guochao Dong, Chao Liu 0020 |
CSCWD | 2 |
| 2021 | A Novel Trojan Attack against Co-learning Based ASR DNN SystemabstractASR (Automatic Speech Recognition) technology is a key technology for human-computer interaction. Especially the DNN models of wake-up-word speech recognition, which enables the smart device to recognize wake-up words spoken by users when they are in the sleep or lock screen state, allowing the device to directly enter the wait command state, and start the first step of voice interaction. ASR technology no wonder provides great convenience for people's daily life, but it's security problem has always been a hot topic of further research. The rapid development of ASR models has made these models very vulnerable which seriously affect their performance in real scenarios. This paper proposes a new backdoor attack method TNN (Trojan Neural Network) for deep learning models in voice wake-up scenarios. The attacker leaves backdoor in the deep learning model of the ASR. Besides specific wakeup words, attackers can also use other words to force the device to wake up and get the IoT smart device's highest privileges. And when the smart device is in the awake state, any audio containing a particular vocabulary will be recognized as a specific command and executed. In this papaer, we propose a novel method to attack ASR model, and without affecting the performance of clean samples, by applying the attack method, the success rate in compulsory recognition of specific words can be up to 100%. The experimental results prove that our attack method is very effective and poses a great security problem of the voice-interactive IoT device. Dongdong Huo, Chao Liu 0020, Yazhe Wang, Yu Wang 0243, Zhen Xu 0009 |
CSCWD | 6 |
| 2021 | Achieve Better Endorsement Balance on Blockchain SystemsabstractAt the beginning of the transaction flow in blockchain system Hyperedger Fabric, client have to select endorsement peer nodes to send proposals. Fabric provides two ways for client: static configuration and random selection using service discovery. However, when a large number of transaction requests arrive, the assignment of tasks to endorsement peer nodes may become unfair. The result of the above situation may lead to the high utilization of computing resources and the working pressure of Fabric nodes, which will eventually affect the transaction performance and throughput. In this paper, we propose a calculation method to accurately measure peer node endorsement load and a dynamic stochastic load minimization algorithm, DSLM. Using DSLM to dynamically determine endorsed peer nodes based on real-time work load will make the distribute of endorsement tasks more balanced. In addition, we implement our approach on Hyperledger Fabric v1.4 LTS and carry out a series of evaluation experiments. By observing the experimental results, the proposed method can effectively balance and reduce the overall workload, and improve the performance of the blockchain system in the case of high concurrency. Chao Liu 0020, Yazhe Wang, Yu Wang 0243, Dongdong Huo |
CSCWD | 3 |
| 2021 | A Self-Sovereign Decentralized Identity Platform Based on BlockchainabstractTraditional accounts and passwords mechanism usually needs to register multiple accounts for adopting different scenarios, which makes it difficult to manage passwords and handle privacy leakage issues. On the other hand, users have no real control over their identities under centralized trusted authorities or insecure third-party operators. In this paper, we propose a self-sovereign decentralized identity platform called SSIChain. SSIChain uses distributed infrastructure blockchain to replace authorities and third-party operators. In the proposed platform, the DID and a credential represent an user's digital identity. We define DID chaincode and credential protocol to cover the full lifecycle of a user's digital identity. We build a complete consortium blockchain and application environment, and implement a prototype as an App for Android-based phones. Moreover, we detect that SSIChain only takes at most 2.1 seconds for authentication with high TPS, which can meet users' demands for identity management well. Chao Liu 0020, Yu Wang 0243, Yazhe Wang |
ISCC | 4 |
| 2021 | Potential Risk Detection System of Hyperledger Fabric Smart Contract based on Static AnalysisabstractAs an indispensable part of Hyperledger Fabric application system, smart contracts are mostly developed in general-purpose programming languages such as Golang. However, these smart contracts often have potential risks that cause serious problems such as failed transactions or sensitive information leakage on the ledger. Although there are already some detection tools for potential risks, e.g., Chaincode Scanner and Chaincode Analyzer, the accuracy and coverage of them are limited. In response to the above problems, this paper summarizes 16 potential risks in smart contracts, including three type risks: Non-determinism Risk, Logical Security Risk, and Private Data Security Risk. In order to detect them, we propose a new static analysis method based on Abstract Syntax Tree Analysis, Package Dependency Analysis, and Functional Dependency Analysis. Based on the new method, a detection system is designed that can detect 16 potential risks in smart contracts developed using Golang language more accurately and provide development suggestions to eliminate these risks. Penghui Lv, Yu Wang 0243, Yazhe Wang, Qihui Zhou |
ISCC | 3 |
| 2021 | Reviewing IoT Security via Logic Bugs in IoT Platforms and SystemsabstractIn recent years, Internet-of-Things (IoT) platforms and systems have been rapidly emerging. Although IoT is a new technology, new does not mean simpler (than existing networked systems). Contrarily, the complexity (of IoT platforms and systems) is actually being increased in terms of the interactions between the physical world and cyberspace. The increased complexity indeed results in new vulnerabilities. This article seeks to provide a review of the recently discovered logic bugs that are specific to IoT platforms and systems and discuss the lessons we learned from these bugs. In particular, 20 logic bugs and one weakness falling into seven categories of vulnerabilities are reviewed in this survey. Wei Zhou 0026, Chen Cao 0004, Dongdong Huo, Lan Zhang 0008, Le Guan, Yan Jia 0009, Yaowen Zheng, Yuqing Zhang 0001, Limin Sun 0001, Yazhe Wang, Peng Liu 0005 |
IEEE Internet Things J. | 12 |
| 2021 | Commercial hypervisor-based task sandboxing mechanisms are unsecured? But we can fix it!abstractCyber–Physical–Social Systems are frequently prescribed for providing valuable information on personalized services . The foundation of these services is big data which must be trustily collected and efficiently processed. Though High Performance Computing and Communication technique makes great contributions to addressing the issue of data processing, its effectiveness still relies on the veracity of data generated from Internet of Things (IoT) devices. Nevertheless, IoT devices, as basic production facilities to ensure data’s security, are unable to deploy expensive security extensions. Consequently, it causes the implementation of the task sandboxing, the fundamental security mechanism in Real-Time Operating Systems (RTOSs), much simpler and more vulnerable. In this paper, we take ARM Mbed uVisor as an example system, utilizing hypervisor-based task sandboxing mechanisms, and presents three new findings: First, we discover vulnerabilities against Mbed task sandboxing, which can be exploited to compromise system-maintained data structure to manipulate any tasks’ data. Second, we present LIPS (Lightweight Intra-Mode Privilege Separation), building a special protection domain to isolate particular system-maintained data structures. Finally, thorough evaluation and experimental tests show the efficiency of LIPS to defeat these attacks, with small runtime overheads and good portability. Dongdong Huo, Chen Cao 0004, Peng Liu 0005, Yazhe Wang, Zhen Xu 0009 |
J. Syst. Archit. | 4 |
| 2020 | A Machine Learning-Assisted Compartmentalization Scheme for Bare-Metal Systems
Dongdong Huo, Chao Liu 0020, Yu Wang 0243, Yazhe Wang, Peng Liu 0005, Zhen Xu 0009 |
ICICS | 6 |
| 2020 | Prihook: Differentiated context-aware hook placement for different owners' smartphonesabstractA context-aware hook is a piece of code. It checks context-aware user privacy policy before some sensitive operations happen. We propose Prihook to address specific context-aware user privacy concerns through putting specific context-aware hooks. We design User Privacy Preference Table (UPPT) to help a user express his privacy concerns and propose a mapping from the words in the UPPT lexicon to the methods in the Potential Method Set. With this mapping, Prihook is able to (a) select a specific set of methods; and (b) generate and place hooks automatically. Hence, the hook placement in Prihook is personalized. We test Prihook separately on 6 typical UPPTs representing 6 kinds of resource-sensitive UPPTs, and no user privacy violation is found. The experimental results show that the hooks placed by PriHook have small runtime overhead. Chen Tian 0004, Yazhe Wang, Peng Liu 0005, Yu Wang 0243, Ruirui Dai, Anyuan Zhou, Zhen Xu 0009 |
TrustCom | 2 |
| 2019 | Medical Protocol Security: DICOM Vulnerability Mining Based on Fuzzing TechnologyabstractDICOM is an international standard for medical images and related information, and is a medical image format that can be used for data exchange. The agreement is widely used in medical fields such as radiology and cardiovascular imaging. However, since DICOM libraries have less security considerations in protocol implementation, they have a large number of security risks. Aiming at the security issue of DICOM libraries, the paper conducts research on vulnerability mining technology for DICOM open source libraries, proposes a vulnerability mining framework based on Fuzzing technology, and implements a prototype system named DICOM-Fuzzer, which includes initialization, test case generation, automatic test, exception monitoring and other modules. Finally, the open source library DCMTK was selected for testing, and it was found that data overflow would occur when the content of the received file was greater than 7080 lines. Found that there is a vulnerability that causes the PACS system to refuse service. In conclusion, the DICOM protocol does have risks, and its information security needs to be further improved. Zhiqiang Wang 0006, Quanqi Li, Yazhe Wang, Qixu Liu |
CCS | 3 |
| 2018 | Using IM-Visor to stop untrusted IME apps from stealing sensitive keystrokesabstractThird-party IME (Input Method Editor) apps are often the preference means of interaction for Android users’ input. In this paper, we first discuss the insecurity of IME apps, including the Potentially Harmful Apps (PHAs) and malicious IME apps, which may leak users’ sensitive keystrokes. The current defense system, such as I-BOX, is vulnerable to the prefix substitution attack and the colluding attack due to the post-IME nature. We provide a deeper understanding that all the designs with the post-IME nature are subject to the prefix-substitution and colluding attacks. To remedy the above post-IME system’s flaws, we propose a new idea, pre-IME, which guarantees that “Is this touch event a sensitive keystroke?” analysis will always access user touch events prior to the execution of any IME app code. We design an innovative TrustZone-based framework named IM-Visor which has the pre-IME nature. Specifically, IM-Visor creates the isolation environment named STIE as soon as a user intends to type on a soft keyboard, then the STIE intercepts,Android event sub translates and analyzes the user’s touch input. If the input is sensitive, the translation of keystrokes will be delivered to user apps through a trusted path. Otherwise, IM-Visor replays non-sensitive keystroke touch events for IME apps or replays non-keystroke touch events for other apps. A prototype of IM-Visor has been implemented and tested with several most popular IMEs. The experimental results show that IM-Visor has small runtime overheads. Chen Tian 0004, Yazhe Wang, Peng Liu 0005, Qihui Zhou |
Cybersecur. | 2 |
| 2017 | IM-Visor: A Pre-IME Guard to Prevent IME Apps from Stealing Sensitive Keystrokes Using TrustZoneabstractThird-party IME (Input Method Editor) apps are often the preference means of interaction for Android users' input. In this paper, we first discuss the insecurity of IME apps, including the Potentially Harmful Apps (PHA) and malicious IME apps, which may leak users' sensitive keystrokes. The current defense system, such as I-BOX, is vulnerable to the prefix-substitution attack and the colluding attack due to the post-IME nature. We provide a deeper understanding that all the designs with the post-IME nature are subject to the prefix-substitution and colluding attacks. To remedy the above post-IME system's flaws, we propose a new idea, pre-IME, which guarantees that "Is this touch event a sensitive keystroke?" analysis will always access user touch events prior to the execution of any IME app code. We designed an innovative TrustZone-based framework named IM-Visor which has the pre-IME nature. Specifically, IM-Visor creates the isolation environment named STIE as soon as a user intends to type on a soft keyboard, then the STIE intercepts, translates and analyzes the user's touch input. If the input is sensitive, the translation of keystrokes will be delivered to user apps through a trusted path. Otherwise, IM-Visor replays non-sensitive keystroke touch events for IME apps or replays non-keystroke touch events for other apps. A prototype of IM-Visor has been implemented and tested with several most popular IMEs. The experimental results show that IM-Visor has small runtime overheads. Chen Tian 0004, Yazhe Wang, Peng Liu 0005, Qihui Zhou, Zhen Xu 0009 |
DSN | 2 |
| 2017 | SecHome: A Secure Large-Scale Smart Home System Using Hierarchical Identity Based Encryption
Yazhe Wang, Yuan Zhang 0004 |
ICICS | 2 |
| 2017 | PMViewer: A Crowdsourcing Approach to Fine-Grained Urban PM2.5 Monitoring in ChinaabstractIn recent years, the PM2.5 (particulate matter with a mean aerodynamic diameter of 2.5 micrometers or less) pollution has become a very serious problem in China. Currently, there are three types of monitoring approaches: government-led monitoring, Wireless Sensor Networks (WSN) approaches and Participatory Urban Sensing (PUS). There are three limitations in these state-of-the-art research approaches: a) the coarse-grained limitation of government-led monitoring, b) the deployment and maintenance cost of WSN approaches, c) the "Black Hole" problem and the "Black Time Window" problem of PMTI-based approaches in PUS. How to overcome these three main limitations is the biggest challenge. To address these limitations, we need a new way to collect PM2.5 data. Nowadays, IoT (Internet of Things) smart devices sold to various customers could steadily and directly collect the PM2.5 data in vast urban areas, but how to obtain wide-spread real PM2.5 data from tens of thousands of smart devices is another challenge. While no existing work has addressed these two challenges, how to address them is an open problem. In this paper, we propose PMViewer, a novel PUS approach to address these two challenges. PMViewer's data are collected from tens of thousands of smart devices called AirBox through a crowdsourcing approach. We aim to offer a fine-grained spatial-temporal resolution for the public to monitor the urban PM2.5 pollution near their locations. PMViewer scrawls data from AirBox's vendor server and parses the data to generate a map view to display real-time urban PM2.5 measurements. In this study, we design, implement and evaluate PMViewer. Evaluation results show that PMViewer efficiently and economically addressed these two challenges described above. Yazhe Wang, Peng Liu 0005, Lvgen Luo, Xinwang Zhuo |
MASS | 2 |
| 2017 | Visual Analysis of Android Malware Behavior Profile Based on PMCG_droid : A Pruned Lightweight APP Call Graph
Gui Peng, Yazhe Wang, Minghui Tian, Jianxing Hu, Liming Wang 0001 |
SecureComm | 4 |
| 2015 | Preserving privacy in social networks against connection fingerprint attacksabstractExisting works on identity privacy protection on social networks make the assumption that all the user identities in a social network are private and ignore the fact that in many real-world social networks, there exists a considerable amount of users such as celebrities, media users, and organization users whose identities are public. In this paper, we demonstrate that the presence of public users can cause serious damage to the identity privacy of other ordinary users. Motivated attackers can utilize the connection information of a user to some known public users to perform re-identification attacks, namely connection fingerprint (CFP) attacks. We propose two k-anonymization algorithms to protect a social network against the CFP attacks. One algorithm is based on adding dummy vertices. It can resist powerful attackers with the connection information of a user with the public users within n hops (n ≥ 1) and protect the centrality utility of public users. The other algorithm is based on edge modification. It is only able to resist attackers with the connection information of a user with the public users within 1 hop but preserves a rich spectrum of network utility. We perform comprehensive experiments on real-world networks and demonstrate that our algorithms are very efficient in terms of the running time and are able to generate k-anonymized networks with good utility. Yazhe Wang, Baihua Zheng |
ICDE | 1 |
| 2015 | Should We Use the Sample? Analyzing Datasets Sampled from Twitter's Stream APIabstractResearchers have begun studying content obtained from microblogging services such as Twitter to address a variety of technological, social, and commercial research questions. The large number of Twitter users and even larger volume of tweets often make it impractical to collect and maintain a complete record of activity; therefore, most research and some commercial software applications rely on samples, often relatively small samples, of Twitter data. For the most part, sample sizes have been based on availability and practical considerations. Relatively little attention has been paid to how well these samples represent the underlying stream of Twitter data. To fill this gap, this article performs a comparative analysis on samples obtained from two of Twitter’s streaming APIs with a more complete Twitter dataset to gain an in-depth understanding of the nature of Twitter data samples and their potential for use in various data mining tasks. Yazhe Wang, Jamie Callan, Baihua Zheng |
ACM Trans. Web | 1 |
| 2014 | On macro and micro exploration of hashtag diffusion in TwitterabstractThis exploratory work studies hashtag diffusion in Twitter. The analysis is conducted from two aspects. From the macro perspective, we study general properties of hashtag diffusion, and classify hashtags into three main classes based on their temporal dynamics referred as “single spike”, “multi-spikes”, and “fluctuation”, and find that each of these classes has some unique characteristics. From the micro perspective, we investigate individual diffusion.We adopt Edelman's “topology of influence” theory to identify four type of users with different influence levels in diffusion based on their dynamic retweet behaviors. The results of our study are useful for gaining more insights of information diffusion in Twitter. Yazhe Wang, Baihua Zheng |
ASONAM | 1 |
| 2014 | UAuth: A Strong Authentication Method from Personal Devices to Multi-accounts
Yazhe Wang |
SecureComm (1) | 1 |
| 2014 | High utility K-anonymization for social network publishing
Yazhe Wang, Long Xie, Baihua Zheng, Ken C. K. Lee |
Knowl. Inf. Syst. | 1 |
| 2011 | Utility-Oriented K-Anonymization on Social Networks
Yazhe Wang, Long Xie, Baihua Zheng, Ken C. K. Lee |
DASFAA (1) | 1 |