VLDB 2026 Research / reviewers in the wild / expert
Haizhou Wang 0001
dblp:46/8126-1
· DBLP profile ↗
50ranked-venue papers
3as first author
45since 2021 · last 2026
0000-0003-1197-5906ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 1 first-author · 23 since 2021Databases, data management, data science and information retrieval · 9 · 1 first-author · 8 since 2021Security and privacy · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 6 since 2021Computer networks · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AMSW: Adaptive Multi-view Semantic Weighting for Chinese Harmful Meme Detection
Henghua Que, Haizhou Wang 0001 |
ICIC (24) | 3 |
| 2026 | Cross-Modal Alignment Attacks: Inducing Vision Token Suppression in Large Vision-Language Models via Coherent False Contexts
Hanle Wang, Haizhou Wang 0001 |
ICIC (26) | 2 |
| 2026 | EvoJail: Evolutionary diverse jailbreak prompt generation for large language models
Rui Tang 0020, Kaiyu Xu, Pengsen Cheng, Hao Ren 0001, Haizhou Wang 0001, Shuyu Jiang |
Inf. Process. Manag. | 5 |
| 2026 | Chinese Implicit Offensive Speech Detection Based on Knowledge Graph and Fuzzy SemanticabstractRecently, publishers of offensive comments are increasingly employing strategies such as metaphors, abbreviations, and homophones to obscure the aggressive nature of their comments. These strategies pose a significant challenge for existing detection models. At present, many studies mainly focused on the detection of explicit offensive speech, and there were few studies on implicit offensive speech. Our research aims to analyze implicit offensive speech on Chinese social platforms and achieve high detection performance. Firstly, we have collected data from one of the largest Chinese social networking platforms, Weibo, and constructed the first Chinese implicit offensive speech dataset, which contains 54,714 comments. Subsequently, we introduce Enhanced-BERT-Mate-Ambiguity (EBMA), a novel fuzzy semantic interpretation framework that leverages BERT and knowledge graphs. Specifically, this model detects implicit offensive speech by extracting semantic, emotional, metaphorical, and ambiguity features. Finally, extensive experiments were conducted, including comparison tests, robustness tests, and ablation studies, to validate our approach. We tested our model against state-of-the-art models in the field, and an accuracy of 95.83% and an F1-score of 95.52% confirmed its best performance. The performance of our model is visually illustrated through visualization. Moreover, we provide an analysis of error cases to explore the limitations of our model. Tengda Guo, Chengping Zheng, Lianxin Lin, Zhijian Tu, Haizhou Wang 0001, Lei Zhang 0103 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 6 |
| 2026 | Enhancing the Security of Large Character Set CAPTCHAs Using Transferable Adversarial ExamplesabstractThe large character set CAPTCHA is an important extension of the traditional text-based CAPTCHA with larger alphabet languages to defend against automated attack programs. However, the state-of-the-art deep learning attacks have cracked such CAPTCHA. Existing defenses against such threats increase the complexity of CAPTCHA, thus decreasing usability. We propose ACG (Adversarial Large Character Set CAPTCHA Generation), a framework with two modules: aFine-grained Generation Module, combining three novel strategies to prevent attackers from recognizing characters, and anEnsemble Generation Moduleto generate global perturbations in CAPTCHAs. It not only strengthens defense against recognition attacks but also improves robustness against diverse detection architectures through adversarial perturbations. Additionally, we develop a toolkit, Adv-Eval, consisting of CAPTCHA datasets from 10 of the most popular Chinese CAPTCHA schemes and benchmarking various attacks. We conduct extensive experiments using Adv-Eval to demonstrate ACG's efficacy, especially manifesting a significant decrease in the average success rate of diverse attacks from 51.52% to 2.56%. To the best of our knowledge, ACG is the first framework to defend large character set CAPTCHAs against detection attacks using transferable adversarial examples. Guoheng Sun, Yucheng Fu, Juntian Huang, Ruimei Zhang, Haizhou Wang 0001 |
IEEE Trans. Dependable Secur. Comput. | 6 |
| 2025 | PPNA: Enabling Privacy-Preserving and Efficient Social Network AlignmentabstractSocial network alignment has made significant progress in social network analysis, with representative applications such as cross-domain recommendation and community detection. However, existing approaches require institutions to share raw user data, raising significant privacy concerns. To address this issue, we propose a Privacy-Preserving Network Alignment (PPNA) scheme that eliminates the need for raw data sharing. In concrete, PPNA leverages homomorphic encryption to enable computation over the ciphertext domain without decryption. It ensures provable data privacy. PPNA also presents a secure multiparty computation protocol to eliminate reliance on trusted third-party servers, which is often impractical in real-world scenarios. Furthermore, its well-designed iterative update mechanism is well-suited for iterative alignment algorithms. Comprehensive experimental results have demonstrated that PPNA improves performance compared to the scenario where raw data sharing is unfeasible due to privacy concerns. It achieves an average F1-score increase of 1.65 times and up to 2.28 times. The performance gain is more pronounced in decentralized settings, highlighting PPNA’s practicality in real-world scenarios when multi-institution collaboration is imperative. Rui Tang 0020, Hao Ren 0001, Haizhou Wang 0001, Xingshu Chen, Meng Li 0006, Hongwei Li 0001 |
GLOBECOM | 4 |
| 2025 | When Vision Becomes a Threat: Adversarial Prompt Injection via Visual Embedding Manipulation
Yajing Ma, Junfeng Hao, Haizhou Wang 0001 |
PRICAI (4) | 3 |
| 2025 | MFGAF: Multi-faceted Granular Analysis Framework for LLM-Generated Text DetectionabstractLarge Language Models (LLMs) have become increasingly proficient in generating human-like text, yet their widespread deployment raises significant concerns, including disseminating fake information, privacy violations, and academic dishonesty. Detecting LLM-generated text is vital for mitigating these risks and is often framed as a binary classification task. Although zero-shot textual analysis methods have gained popularity due to their generalizability, they face key challenges: (1) limited ability to capture holistic and granular consistency and (2) insufficiently comprehensive textual analysis, particularly regarding syntax and lexical patterns. To address these problems, we propose a novel Multi-faceted Granular Analysis Framework (MFGAF), which leverages Rewriting Concordance and Completive Concordance to detect LLM-generated text through multi-granular textual dissection. Specifically, MFGAF is designed with two perspectives, Rewriting and Completion, to comprehensively capture global and local LLM-generated features. Additionally, for each perspective, a Multi-Granular Textual Dissection mechanism is constructed to thoroughly analyze LLM-generated text. Finally, MFGAF leverages LLMs to achieve adaptive integration of text analysis results and reflect on the correctness of these results. Our method achieves an average F1 improvement of 3.06% compared to the best baselines, demonstrating its robustness and effectiveness through extensive evaluations. Haizhou Wang 0001 |
SMC | 2 |
| 2025 | ReZG: Retrieval-augmented zero-shot counter narrative generation for hate speech
Shuyu Jiang, Wenyi Tang, Xingshu Chen, Rui Tang 0020, Haizhou Wang 0001, Wenxian Wang |
Neurocomputing | 5 |
| 2025 | A Toxic Euphemism Detection framework for online social network based on Semantic Contrastive Learning and dual channel knowledge augmentation
Haizhou Wang 0001, Wenxian Wang, Shuyu Jiang, Rui Tang 0020, Xingshu Chen |
Inf. Process. Manag. | 2 |
| 2025 | A metadata-aware detection model for fake restaurant reviews based on multimodal fusion
Yifei Jian, Xiaoda Wang, Xingshu Chen, Xiao Lan, Wenxian Wang, Haizhou Wang 0001 |
Neural Comput. Appl. | 8 |
| 2025 | A Novel Retrospective-Reading Model for Detecting Chinese Sarcasm Comments of Online Social NetworkabstractThrough the use of sarcastic sentences on social media, people can express their strong emotions. Therefore, the detection of sarcasm in social media has received more and more attention over the past years. Classifying a sentence as sarcastic or nonsarcastic heavily relies on the contextual information of the sentence. However, only focusing on the features of target text is the main solution of most existing research. Moreover, the scale of publicly available Chinese sarcasm dataset is very small and does not contain the contextual information. To address the issues mentioned above, we build a Chinese sarcasm dataset from Bilibili, which is one of the most widely used social network platforms in China and has a significant number of sarcastic comments and contextual information. As far as we know, our dataset is the first publicly available large-scale Chinese sarcasm dataset including contextual information. Additionally, we have proposed a novel retrospective reading method for detecting sarcasm that leverages contextual information to improve model's performance. The experimental results show the effectiveness of the proposed model and the significance of contextual information for Chinese sarcasm detection: achieving the highest F-score of 0.6942, outperforming existing state-of-the-art (SOTA) approaches. The study presented in this article offers approaches and ideas for future Chinese sarcasm detection studies. Lei Zhang 0103, Qinfeng Mao, Yanbing Yang 0001, Dong Li 0051, Haizhou Wang 0001 |
IEEE Trans. Comput. Soc. Syst. | 6 |
| 2025 | Detecting Lifecycle-Related Concurrency Bugs in ROS Programs via Coverage-Guided FuzzingabstractRobot Operating System (ROS) is very popular in robotic software development. To ease the process management of ROS programs, ROS provides a speciallifecyclemechanism that can conveniently manage the state of each running process, which often involves resource allocation, initialization, and release; and this mechanism has been widely used in real-world ROS programs. However, due to code concurrency of ROS programs, a lifecycle-related function is inevitably concurrently executed with other functions, introducing the security risk of dangerous concurrency bugs involving null-pointer dereference and use after free. Due to the non-determinism of thread scheduling, these concurrency bugs are difficult to find and reproduce. In this paper, we design and implement a new coverage-guided fuzzing framework named ROCF, which can effectively detect and reproduce lifecycle-related concurrency bugs in ROS programs, with two novel techniques. First, we propose alifecycle-aware fuzzing approachthat useslifecycle pair sequenceas a new coverage metric to effectively describe lifecycle-related thread interleavings, for input-mutation guidance of ROS concurrency fuzzing. Second, we propose aheuristic-based reproducing methodthat identifies minimal input sequences that can stably and efficiently reproduce the found concurrency bugs, with strategical input pruning and delay injection. We evaluate ROCF on eight popular robotic programs in ROS2, and it finds 32 new and real concurrency bugs, all of which have been confirmed by ROS developers, and 19 have been assigned CVE IDs. Si-Miao Gao, Pengcheng Wang 0003, Jia-Ju Bai, Jia-Wei Yu, Haizhou Wang 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | TSMGAN-II: Generative Adversarial Network Based on Two-Stage Mask Transformer and Information Interaction for Speech Enhancement
Lianxin Lin, Yaowen Li, Haizhou Wang 0001 |
ICIC (4) | 3 |
| 2024 | SPC-GAN-Attack: Attacking Slide Puzzle CAPTCHAs by Human-Like Sliding Trajectories Based on Generative Adversarial Network
Enbo Yu, Qianqian Qiao, Zhenyu Mao, Haizhou Wang 0001 |
SecureComm (1) | 6 |
| 2024 | A deep semantic-aware approach for Cantonese rumor detection in social networks with graph convolutional network
Yifei Jian, Liang Ke, Yunxiang Qiu, Xingshu Chen, Yunya Song, Haizhou Wang 0001 |
Expert Syst. Appl. | 7 |
| 2024 | A novel cross-domain adaptation framework for unsupervised criminal jargon detection via pre-trained contextual embedding of darknet corpus
Liang Ke, Shui Yu 0001, Xingshu Chen, Haizhou Wang 0001 |
Expert Syst. Appl. | 6 |
| 2024 | Detecting Offensive Language Based on Graph Attention Networks and Fusion FeaturesabstractThe pervasiveness of offensive language on social networks has caused adverse effects on society, such as abusive behavior online. It is urgent to detect offensive language and curb its spread. In the popular datasets, the distribution of users and tweets is imbalanced, which limits the generalization ability of the model. In addition, existing research shows that methods with community information extracted from the social graphs effectively improve the performance of offensive language detection. However, the existing models deal with social graphs independently, which seriously affects the effectiveness of detection models. In this article, we release a new dataset with users and social relationships. To encode community information, we construct the social graphs based on the user historical behavior information and social relationships. Moreover, we propose a model based on graph attention networks (GATs) and fusion features for offensive language detection (GF-OLD). Specifically, the community information is directly captured by the GAT module, and the text embeddings are taken from the last hidden layer of bidirectional encoder representation from transformer (BERT). Attention mechanisms and position encoding are used to fuse these features. Our method outperforms baselines with the F1-score of 89.94%. The results show that our model effectively learns the potential information of social graphs and text, and user historical behavior information is more suitable for user attribute in the social graphs. Zhenxiong Miao, Xingshu Chen, Haizhou Wang 0001, Rui Tang 0020, Tiemai Huang, Wenyi Tang |
IEEE Trans. Comput. Soc. Syst. | 3 |
| 2024 | Detecting Spam Movie Review Under Coordinated Attack With Multi-View Explicit and Implicit Relations Semantics FusionabstractSpam reviews have long polluted review systems, undermining their industries. Detecting spam movie reviews faces some brand-new challenges compared to traditional spam detection. These include coordinated spamming attacks during premieres or at advance screenings. However, most of existing studies only use inherent relations among reviews, movies, and users, they do not fully exploit explicit and implicit relations between reviews in coordinated spamming attacks. To address these novel challenges, we propose a spam movie review detection method based on mining explicit and implicit relation semantics and fusing multi-view semantics. To the best of our knowledge, we are the first to enhance spam movie review detection by exploiting both explicit and implicit relations between reviews in coordinated spamming attacks. First, we build an explicit relation movie-review graph with movie synopses and high-quality external reviews. We extract movie factual knowledge embeddings using a Heterogeneous Graph Transformer (HGT) network. Next, we input the factual knowledge embeddings with corresponding review embeddings into a contrastive network to get review credibility features. Additionally, we build an implicit relation graph between reviews using metadata and semantic similarities. We extract relation-enhanced review semantics via another HGT network. Finally, we fuse the three review semantic features through an attention layer before making classification. Experiments show our method achieves higher performance and robustness over state-of-the-art methods. Yicheng Cai, Haizhou Wang 0001, Wenxian Wang, Lei Zhang 0103, Xingshu Chen |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2023 | BGEK: External Knowledge-Enhanced Graph Convolutional Networks for Rumor Detection in Online Social Networks
Xiaoda Wang, Chenxiang Luo, Tengda Guo, Zhangrui Liu, Jiongyan Zhang, Haizhou Wang 0001 |
ICANN (4) | 6 |
| 2023 | DKCS: A Dual Knowledge-Enhanced Abstractive Cross-Lingual Summarization Method Based on Graph Attention Networks
Shuyu Jiang, Dengbiao Tu, Xingshu Chen, Rui Tang 0020, Wenxian Wang, Haizhou Wang 0001 |
ICONIP (13) | 6 |
| 2023 | TPTGAN: Two-Path Transformer-Based Generative Adversarial Network Using Joint Magnitude Masking and Complex Spectral Mapping for Speech Enhancement
Zhaoyi Liu 0002, Zhuohang Jiang, Wendian Luo, Zhuoyao Fan, Haoda Di, Yufan Long, Haizhou Wang 0001 |
ICONIP (9) | 7 |
| 2023 | Fighting Attacks on Large Character Set CAPTCHAs Using Transferable Adversarial ExamplesabstractOver a long period, large character set CAPTCHAs are widely used to defend against automated attack programs on the Internet. However, with the development of deep learning techniques, some attacks for large character set CAPTCHAs have been proposed, proving that they are no longer secure. To defend against black-box attacks on these CAPTCHAs, we propose a novel defense method based on transferable adversarial example techniques. On the one hand, we defend against character recognition attacks by adding adversarial perturbations to the characters of CAPTCHAs combining three strategies: Gradient-based Attacks, Input Transformations and Attention Mechanism. On the other hand, we defend against character detection attacks by leveraging an ensemble method to generate adversarial perturbations on the background of CAPTCHAs. To the best of our knowledge, this is the first study to improve the security of large character set CAPTCHAs against black-box attacks based on transferable adversarial example techniques. Using the eight most popular Chinese CAPTCHA schemes as examples, we conduct comprehensive experiments. Results show that our method improves the security of large character set CAPTCHAs by making the average success rate of black-box attacks significantly drop from 53.33% to 3.49%. Overall, our method can be helpful to the design of more secure large character set CAPTCHAs. Yucheng Fu, Guoheng Sun, Juntian Huang, Haizhou Wang 0001 |
IJCNN | 5 |
| 2023 | Implicit Offensive Speech Detection Based on Multi-feature Fusion
Tengda Guo, Lianxin Lin, Chengping Zheng, Zhijian Tu, Haizhou Wang 0001 |
KSEM (2) | 6 |
| 2023 | VRC-GraphNet: A Graph Neural Network-Based Reasoning Framework for Attacking Visual Reasoning Captchas
Botao Xu, Haizhou Wang 0001 |
SecureComm (1) | 2 |
| 2023 | Depression detection on online social network with multivariate time series feature of user depressive symptoms
Yicheng Cai, Haizhou Wang 0001, Huali Ye, Yanwen Jin, Wei Gao 0055 |
Expert Syst. Appl. | 2 |
| 2023 | Identifying Cantonese rumors with discriminative feature integration in online social networks
Haizhou Wang 0001, Liang Ke, Zhipeng Lu 0001, Hanjian Su, Xingshu Chen |
Expert Syst. Appl. | 2 |
| 2023 | A novel hybrid feature fusion model for detecting phishing scam on Ethereum using deep neural network
Tingke Wen, Yuanxing Xiao, Anqi Wang 0008, Haizhou Wang 0001 |
Expert Syst. Appl. | 4 |
| 2023 | Unveiling Qzone: A measurement study of a large-scale online social network
Haizhou Wang 0001, Yixuan Fang, Shuyu Jiang, Xingshu Chen, Xiaohui Peng 0007, Wenxian Wang |
Inf. Sci. | 1 |
| 2023 | Interlayer Link Prediction in Multiplex Social Networks Based on Multiple Types of Consistency Between Embedding VectorsabstractOnline users are typically active on multiple social media networks (SMNs), which constitute a multiplex social network. With improvements in cybersecurity awareness, users increasingly choose different usernames and provide different profiles on different SMNs. Thus, it is becoming increasingly challenging to determine whether given accounts on different SMNs belong to the same user; this can be expressed as an interlayer link prediction problem in a multiplex network. To address the challenge of predicting interlayer links, feature or structure information is leveraged. Existing methods that use network embedding techniques to address this problem focus on learning a mapping function to unify all nodes into a common latent representation space for prediction; positional relationships between unmatched nodes and their common matched neighbors (CMNs) are not utilized. Furthermore, the layers are often modeled as unweighted graphs, ignoring the strengths of the relationships between nodes. To address these limitations, we propose a framework based on multiple types of consistency between embedding vectors (MulCEVs). In MulCEV, the traditional embedding-based method is applied to obtain the degree of consistency between the vectors representing the unmatched nodes, and a proposed distance consistency index based on the positions of nodes in each latent space provides additional clues for prediction. By associating these two types of consistency, the effective information in the latent spaces is fully utilized. In addition, MulCEV models the layers as weighted graphs to obtain representation. In this way, the higher the strength of the relationship between nodes, the more similar their embedding vectors in the latent representation space will be. The results of our experiments on several real-world and synthetic datasets demonstrate that the proposed MulCEV framework markedly outperforms current embedding-based methods, especially when the number of training iterations is small. Rui Tang 0020, Zhenxiong Miao, Shuyu Jiang, Xingshu Chen, Haizhou Wang 0001, Wei Wang 0070 |
IEEE Trans. Cybern. | 5 |
| 2022 | Fake Restaurant Review Detection Using Deep Neural Networks with Hybrid Feature Fusion Method
Yifei Jian, Xingshu Chen, Haizhou Wang 0001 |
DASFAA (3) | 3 |
| 2022 | GCMK: Detecting Spam Movie Review Based on Graph Convolutional Network Embedding Movie Background Knowledge
Hanyue Li, Yu-Lin He, Haizhou Wang 0001 |
ICANN (2) | 6 |
| 2022 | GAIM: Graph-aware Feature Interactional Model for Spam Movie Review DetectionabstractNowadays, more and more people decide whether to watch a certain movie by reading online movie reviews. Driven by large commercial interest, a growing number of spammers try to manipulate the online word-of-mouth of movies by publishing spam reviews. This has severely destroyed the credibility of movie review platforms and affected the healthy development of the lm industry. However, there is little research on the detection of spam movie reviews. Meanwhile, there are still great challenges for spam movie review detection, such as the lack of publicly available datasets, insufficient features, and low-performance detection models. In this paper, we firstly construct a dataset for spam movie reviews with our proposed labeling strategy and release it publicly. Secondly, we design 28 features to detect spam movie reviews, 13 of which are completely new features. Thirdly, in order to mine the characteristics of collusive attack behavior deeply, we propose a Graph-aware Feature Interactional Model (GAIM), which combines TextCNN, MLP (Multilayer Perceptron), and GAT (Graph Attention Network). After performance evaluation, the experimental results show that GAIM is more effective than the state-of-the-art baselines with an F1-score of 91.88% for spam movie review. Lei Zhang 0103, Xueqiang Song, Yuwei Fang, Dong Li 0051, Haizhou Wang 0001 |
ICPR | 6 |
| 2022 | A Novel Chinese Sarcasm Detection Model Based on Retrospective Reader
Lei Zhang 0103, Xueqiang Song, Yuwei Fang, Dong Li 0051, Haizhou Wang 0001 |
MMM (2) | 6 |
| 2022 | SICM: A Supervised-Based Identification and Classification Model for Chinese Jargons Using Feature Adapter Enhanced BERT
Haochen Su, Yingao Wu, Haizhou Wang 0001 |
PRICAI (2) | 4 |
| 2022 | An Unsupervised Detection Framework for Chinese Jargons in the DarknetabstractWith the continuous development of the darknet technology, the scale of darknet and have increased rapidly in recent years, leading to rampant crime in these anonymous trading markets. Monitoring these markets can effectively combat the criminal forces that hide behind them. One of the difficulties in understanding the darknet is that criminals usually use jargons to disguise transactions and thus avoid surveillance. These jargons usually distort the original meaning of innocent-looking words in obscure ways, posing significant challenges for crime tracking. Current research on Chinese jargon detection mainly adopts the method of keyword filtering, however, such methods have little effect on the complex and ever-changing structure of darknet jargons. We propose a Chinese jargon detection framework based on unsupervised learning. The main idea is to compare similarity with high-dimensional word embedding features from different corpus to find jargons. Firstly, we collect data from six Chinese Tor websites to build a dark corpus dataset. Afterwards, we build a word-based pre-training model called DC-BERT, which can generate high-quality contextual word embeddings. Finally, we construct a cross-corpus jargon detection framework based on similarity analysis, which can effectively detect Chinese jargons in the darknet. The experimental results show that the proposed framework is both innovative and practical, reaching a detection accuracy of 91.5%. Liang Ke, Haizhou Wang 0001 |
WSDM | 3 |
| 2022 | Interlayer link prediction based on multiple network structural attributes
Rui Tang 0020, Xingshu Chen, Chuancheng Wei, Qindong Li, Wenxian Wang, Haizhou Wang 0001, Wei Wang 0070 |
Comput. Networks | 6 |
| 2022 | Detecting fake news on Chinese social media based on hybrid feature fusion method
Haizhou Wang 0001, YuHu Han |
Expert Syst. Appl. | 1 |
| 2022 | Recognizing irrelevant faces in short-form videos based on feature fusion and active learningabstractIn recent years, short-form videos spread rapidly around the world and became a popular way of entertainment for people to share their daily lives. However, many videos record behaviors of other people without their awareness and are uploaded onto the short-form video platforms. Such behavior severely invades personal privacy and can even bring risks of personal information leakage . At present, few studies focus on detecting privacy violations in short-form videos. Meanwhile, due to the difficulty in transferring existing models to the scenario of short-form videos and the lack of reliable datasets, it is very challenging to recognize irrelevant faces in short-form videos. To deal with this problem, we constructed and published an irrelevant faces dataset (IF-Dataset) with 43,965 irrelevant face images and 89,924 relevant face images based on the videos collected from Douyin (the Chinese version of TikTok). In addition, we constructed a framework that implemented our proposed deep learning model M ulti-features M ulti-head F usion Net work (MMFNet) to recognize irrelevant faces from short-form videos. The experimental results show that the F1 score of the MMFNet can reach 87.03%. We also proposed a novel loss function as well as an active learning system to improve the generalization ability of models, which can reach the Relative Error Reduction (RER) up to 29.58%. Our work provides both theoretical and practical support for face protection in short-form videos. Mingcheng Zhu 0001, Rongchuan Zhang, Haizhou Wang 0001 |
Neurocomputing | 3 |
| 2022 | Identification of Chinese dark jargons in Telegram underground markets using context-oriented and linguistic features
Yiwei Hou, Haizhou Wang 0001 |
Inf. Process. Manag. | 3 |
| 2022 | Online social network individual depression detection using a multitask heterogenous modality fusion approach
Chenghao Li 0007, Haizhou Wang 0001 |
Inf. Sci. | 5 |
| 2021 | A novel framework for image-based malware detection with a deep neural network
Yifei Jian, Hongbo Kuang, Chenglong Ren, Zicheng Ma, Haizhou Wang 0001 |
Comput. Secur. | 5 |
| 2021 | End-to-end attack on text-based CAPTCHAs based on cycle-consistent generative adversarial network
Xingshu Chen, Haizhou Wang 0001, Peiming Wang, Wenxian Wang |
Neurocomputing | 3 |
| 2021 | A novel framework for detecting social bots with deep neural networks and active learning
Yuhao Wu 0006, Yuzhou Fang, Shuaikang Shang, Haizhou Wang 0001 |
Knowl. Based Syst. | 6 |
| 2021 | User Identification Based on Integrating Multiple User Information across Online Social NetworksabstractUser identification can help us build more comprehensive user information. It has been attracting much attention from academia. Most of the existing works are profile-based user identification and relationship-based user identification. Due to user privacy settings and social network restrictions on user data crawl, user data may be missing or incomplete in real social networks. User data include profiles, user-generated contents (UGCs), and relationships. The features extracted in previous research may be sparse. In order to reduce the impact of the above problems on user identification, we propose a multiple user information user identification framework (MUIUI). Firstly, we develop multiprocess crawlers to obtain the user data from two popular social networks, Twitter and Facebook. Secondly, we use named entity recognition and entity linking to obtain and integrate locations and organizations from profiles and UGCs. We also extract URLs from profiles and UGCs. We apply the locations jointly with the relationships and develop several algorithms to measure the similarity of the display name, all locations, all organizations, location in profile, all URLs, following organizations, and user ID, respectively. Afterward, we propose a fusion classifier machine learning-based user identification method. The results show that the F1 score of MUIUI reaches 86.46% on the dataset. It proves that MUIUI can reduce the impact of user data that are missing or incomplete. Wenjing Zeng, Rui Tang 0020, Haizhou Wang 0001, Xingshu Chen, Wenxian Wang |
Secur. Commun. Networks | 3 |
| 2020 | A Multimodal Feature Fusion-Based Method for Individual Depression Detection on Sina WeiboabstractExisting studies have shown that various types of information on the online social network (OSN) can help predict the early stage of depression. However, studies using machine learning methods to accomplish depression detection tasks still do not have high classification performance, suggesting that there is much potential for improvement in their feature engineering. In this paper, we first construct a dataset on Sina Weibo (a leading OSN with the largest number of active users in the Chinese community), namely the Weibo User Depression Detection Dataset (WU3D). It includes more than 10,000 depressed users and 20,000 normal users, both of which are manually labeled and rechecked by specialists. Then, we extract text-based word features using the popular pretrained model XLNet and summarize nine statistical features related to user text, social behavior, and pictures. Moreover, we construct a deep neural network classification model, i.e. Multimodal Feature Fusion Network (MFFN), to fuse the above-extracted features from different information sources and further accomplish the classification task. The experimental results show that our approach achieves an F1-Score of 0.9685 on the test dataset, which has a good performance improvement compared to the existing works. In addition, we verify that our multimodal detecting approach is more robust than multimodel ensemble ones. Our work could also provide new research methods for depression detection on other OSN platforms. Chenghao Li 0007, Haizhou Wang 0001 |
IPCCC | 5 |
| 2020 | A Novel Approach for Cantonese Rumor Detection based on Deep Neural NetworkabstractTwitter is a popular social networking platform. While people enjoy the news and anecdotes on Twitter, there are also lots of rumors, which have a negative impact on users and can compromise social order. Among these rumors, many of them are written in Cantonese. At present, the research of English rumor detection is relatively comprehensive, but Cantonese rumors are rarely studied, which brings great challenges to the detection of Cantonese rumors on Twitter. Firstly, there is no available benchmark dataset of Cantonese rumors. Secondly, it is difficult to completely extract the features of rumors. Thirdly, the classical detection approaches are not effective in detecting Cantonese rumors. In this paper, we collected and annotated Cantonese rumors on Twitter and obtained a relatively complete Cantonese rumor dataset. Next, 27 statistical features, involving four categories (user, content, propagation, and comment-based), are extracted to distinguish rumors and non-rumors in Cantonese. Seven of these features are newly proposed in this paper. Then, a novel deep learning model called BLA (namely BERT-based Bi-LSTM network with Attention mechanism) is built for Cantonese rumors detection on Twitter. BLA takes advantage of both statistical and semantic features to effectively detect Cantonese rumors. The experimental results show that the BLA model outperforms other detection models in Cantonese rumor detection. Liang Ke, Zhipeng Lu 0001, Hanjian Su, Haizhou Wang 0001 |
SMC | 5 |
| 2020 | Detecting Social Spammers in Sina Weibo Using Extreme Deep Factorization Machine
Yuhao Wu 0006, Yuzhou Fang, Shuaikang Shang, Haizhou Wang 0001 |
WISE (1) | 6 |
| 2020 | Interlayer link prediction in multiplex social networks: An iterative degree penalty algorithm
Rui Tang 0020, Shuyu Jiang, Xingshu Chen, Haizhou Wang 0001, Wenxian Wang, Wei Wang 0070 |
Knowl. Based Syst. | 4 |
| 2018 | Content pollution propagation in the overlay network of peer-to-peer live streaming systems: modelling and analysisabstractIn the past few years, peer‐to‐peer (P2P) live streaming systems have gained great commercial success and have become a popular way to deliver multimedia content over the Internet, which received more and more attentions from both industry and academia globally. However, the dramatic rise in popularity makes these systems more likely to be vulnerable targets. In this study, mesh‐pull infrastructure architecture and pollution attack principle for P2P live streaming systems were presented firstly, and then the various user behaviours under the pollution attack were analysed. Subsequently, the authors proposed an analytical modelling framework of content pollution attack for P2P live streaming systems. Different from the existing content pollution propagation models, it considers the impact of user behaviours in the attack. Furthermore, to ensure the availability and accuracy of the model, the real‐world experimental attack data for a popular commercial system was used to verify it. The results showed that the model is a feasible and efficient tool to analyse and predict content pollution propagation in real‐world P2P live streaming systems. The authors' work can provide an in‐depth understanding of the content pollution propagation in P2P live streaming systems, and evaluation of restraining illegal content distribution for copyright holders and government. Haizhou Wang 0001, Xingshu Chen, Wenxian Wang, Mei Ya Chan |
IET Commun. | 1 |