VLDB 2026 Research / reviewers in the wild / expert
Wei Wang 0042
dblp:35/7092-42
· DBLP profile ↗
50ranked-venue papers
6as first author
35since 2021 · last 2026
0000-0002-0707-8076ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 1 first-author · 23 since 2021Computer networks · 8 · 1 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 8 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 4 since 2021Systems, architecture and hardware · 3 · 2 since 2021Security and privacy · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Trustworthy Classification for Complex Social Surveys: A Memory-Enhanced Hierarchical Framework with Calibrated UncertaintyabstractAutomated classification of complex social survey questionnaires is crucial for large-scale social science research but faces significant reliability challenges due to intricate hierarchical label structures, severe class imbalance, semantic ambiguity, and incomplete data coverage. Conventional classification methods often struggle with these combined complexities, yielding results that lack trustworthiness. We introduce HOCM, a framework designed for trustworthy classification in complex, real-world taxonomies. It features two synergistic components: (1) memory-enhanced contrastive learning, tailored to learn robust representations from noisy, imbalanced data by leveraging quality-aware category memory banks; and (2) hierarchical uncertainty calibration, which enforces taxonomic consistency while providing reliable confidence estimates and identifying inputs falling outside well-represented known categories. Our evaluation on a large-scale, real-world social survey dataset—a challenging exemplar of our target problem class—demonstrates that HOCM maintains strong accuracy on known classes while effectively identifying uncertain cases, significantly boosting accuracy on confident predictions. Furthermore, it adeptly detects low-resource/unknown categories. HOCM provides a more reliable automated classification tool, enabling efficient expert review and enhancing the trustworthiness of analysis in domains with complex, hierarchical data. Zeqiang Wang, Rebecca Oldroyd, Jiageng Wu, Jie Yang 0039, Wei Wang 0042, Nishanth Sastry, Jon Johnson, Suparna De |
AAAI | 6 |
| 2026 | Aletheia: A Two-Stage Graph-Based Framework for Hallucination Detection in Abstractive Summarization
Tianshi Cai, Guanxu Li, Changyu Zeng, Nijia Han, Ce Huang, Qi Chen 0026, Shuihua Wang, Haiyang Zhang 0004, Wei Wang 0042 |
ICIC (23) | 10 |
| 2026 | MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and Reasoning
Tong Chen 0005, Changyu Zeng, Hongbin Na, Nijia Han, Fuyu Xing, Qi Chen 0026, Qiufeng Wang 0001, Anh Nguyen 0003, Shuihua Wang, Ling Chen 0006, Jionglong Su, Haiyang Zhang 0004, Wei Wang 0042 |
LREC | 15 |
| 2026 | CoRe: Contrast and reconstruction combination self-supervised point cloud representation learning
Changyu Zeng, Jimin Xiao, Anh Nguyen 0003, Xuming Hu, Wei Wang 0042, Yutao Yue |
Expert Syst. Appl. | 5 |
| 2026 | Beyond shortcuts: Mitigating spurious correlations in radiological diagnosis with causal intervention
Xinyi Zeng, Jia Wang 0009, Yi Dong 0002, Wei Wang 0042, Yanji Jiang, Haiyang Zhang 0004 |
Knowl. Based Syst. | 5 |
| 2026 | You look from old classes: Towards accurate few shot class-incremental learning
Yijie Hu, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001 |
Pattern Recognit. | 3 |
| 2026 | Clue and Context Fusion for Sarcasm Detection with Large Multimodal ModelsabstractDetecting sarcasm in social media is fundamentally different from general VLM benchmarks: it is a pragmatic contradiction problem in which the literal signal in one modality is intentionally misaligned with the intended meaning, while dominant pre-training (e.g., CLIP-style contrastive agreement) biases models toward modality alignment rather than incongruity detection. We present SCARF, a contradiction-aware framework that equips large multimodal models with explicit sarcasm cues and context-sensitive retrieval. SCARF constructs coarse scene cues and fine localized evidence via tag-constrained QA, then distills them with visual tokens into a [FUSION] control vector for the LLM; a label-contrastive retriever supplies type- and context-matched exemplars, and a local multi-view encoder surfaces micro-cues. With the same backbone and training data, SCARF attains 87.92% Acc/86.67% F1 on MMSD2.0 and 77.14% Acc/76.44% F1 zero-shot on XDMSD, outperforming a comparably fine-tuned LLaVA-1.5. Ablations show sarcasm clue fusion is the main driver of gains, and tag-constrained QA improves rationale grounding and reduces hallucinations. Yushan Pan, Ding Wang 0006, Wei Wang 0042, Xiaowei Huang 0001, Zhijie Xu |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2025 | Template-Driven LLM-Paraphrased Framework for Tabular Math Word Problem GenerationabstractSolving tabular math word problems (TMWPs) has become a critical role in evaluating the mathematical reasoning ability of large language models (LLMs), where large-scale TMWP samples are commonly required for fine-tuning. Since the collection of high-quality TMWP datasets is costly and time-consuming, recent research has concentrated on automatic TMWP generation. However, current generated samples usually suffer from issues of either correctness or diversity. In this paper, we propose a Template-driven LLM-paraphrased (TeLL) framework for generating high-quality TMWP samples with diverse backgrounds and accurate tables, questions, answers, and solutions. To this end, we first extract templates from existing real samples to generate initial problems, ensuring correctness. Then, we adopt an LLM to extend templates and paraphrase problems, obtaining diverse TMWP samples. Furthermore, we find the reasoning annotation is important for solving TMWPs. Therefore, we propose to enrich each solution with illustrative reasoning steps. Through the proposed framework, we construct a high-quality dataset TabMWP-TeLL by adhering to the question types in the TabMWP dataset, and we conduct extensive experiments on a variety of LLMs to demonstrate the effectiveness of TabMWP-TeLL in improving TMWP-solving performance. Xiaoqiang Kang, Xiao-Bo Jin, Wei Wang 0042, Kaizhu Huang, Qiufeng Wang 0001 |
AAAI | 4 |
| 2025 | Towards Better Robustness Against Natural Corruptions in Document Tampering LocalizationabstractMarvelous advances have been exhibited in recent document tampering localization (DTL) systems. However, confronted with corrupted tampered document images, their vulnerability is fatal in real-world scenarios. While robustness against adversarial attack has been extensively studied by adversarial training (AT), the robustness on natural corruptions remains under-explored for DTL. In this paper, to overcome forensic dependency, we propose the adversarial forensic regularization (AFR) based on min-max optimization to improve robustness. Specifically, we adopt mutual information (MI) to represent forensic dependency between two random variable over tampered and authentic pixels spaces, where the MI can be approximated by Jensen-Shannon-Divergence (JSD) with empirical sampling. To further enable a trade-off between predictive representations in clean tampered document pixels and robust ones in corrupted pixels, an additional regularization term is formulated with divergence between clean and perturbed pixels distribution (DDR). Following min-max optimization framework, our method can also work well against adversarial attacks. To evaluate our proposed method, we collect a dataset (i.e., TSorie-CRP) for evaluating robustness against natural corruptions in real scenarios. Extensive experiments demonstrate the effectiveness of our method against natural corruptions. Without any surprise, our method also achieves good performance against adversarial attack on DTL benchmark datasets. Huiru Shao, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001 |
AAAI | 3 |
| 2025 | Detecting Conversational Mental Manipulation with Intent-Aware PromptingabstractMental manipulation severely undermines mental wellness by covertly and negatively distorting decision-making. While there is an increasing interest in mental health care within the natural language processing community, progress in tackling manipulation remains limited due to the complexity of detecting subtle, covert tactics in conversations. In this paper, we propose Intent-Aware Prompting (IAP), a novel approach for detecting mental manipulations using large language models (LLMs), providing a deeper understanding of manipulative tactics by capturing the underlying intents of participants. Experimental results on the MentalManip dataset demonstrate superior effectiveness of IAP against other advanced prompting strategies. Notably, our approach substantially reduces false negatives, helping detect more instances of mental manipulation with minimal misjudgment of positive cases. The code of this paper is available at https://github.com/Anton-Jiayuan-MA/Manip-IAP. Jiayuan Ma, Hongbin Na, Yining Hua, Wei Wang 0042, Ling Chen 0006 |
COLING | 6 |
| 2025 | Skin Disease Classification with LVLMs: An Empirical StudyabstractSkin diseases pose significant challenges to accurate and efficient diagnosis, often due to their diverse and complex representations. This study investigates the capabilities and limitations of Large Vision-Language Models (LVLMs) in addressing these challenges through skin disease classification tasks. We evaluated LVLMs in zero-shot, few-shot, and finetuning scenarios, exploring their performance, bias, and potential for improvement. Results show that LVLMs lack perceptual granularity in skin disease, though positive signals are also observed. Our findings underscore the necessity for domain- specific optimisation and highlight opportunities for advancing LVLMs in medical diagnostics through innovative strategies and collaborative efforts. Xinyi Zeng, Haiyang Zhang 0004, Wei Wang 0042 |
CSCWD | 5 |
| 2025 | FTCFormer: Fuzzy Token Clustering Transformer for Image ClassificationabstractTransformer-based deep neural networks have achieved remarkable success across various computer vision tasks, largely attributed to their long-range self-attention mechanism and scalability. However, most transformer architectures embed images into uniform, grid-based vision tokens, neglecting the underlying semantic meanings of image regions, resulting in suboptimal feature representations. To address this issue, we propose Fuzzy Token Clustering Transformer (FTCFormer), which incorporates a novel clustering-based downsampling module to dynamically generate vision tokens based on the semantic meanings instead of spatial positions. It allocates fewer tokens to less informative regions and more tokens to represent semantically important regions, regardless of their spatial adjacency or shape irregularity. To further enhance feature extraction and representation, we propose a Density Peak Clustering-Fuzzy K-Nearest Neighbor (DPC-FKNN) mechanism for clustering center determination, a Spatial Connectivity Score (SCS) for token assignment, and a channel-wise merging (Cmerge) strategy for token merging. Extensive experiments on 32 datasets across diverse domains validate the effectiveness of FTCFormer on image classification, showing consistent improvements over the TCFormer baseline, achieving gains of improving 1.43% on five fine-grained datasets, 1.09% on six natural image datasets, 0.97% on three medical datasets and 0.55% on four remote sensing datasets. The code is available at: https://github.com/BaoBao0926/FTCFormer/tree/main. Muyi Bao, Changyu Zeng, Zhengni Yang, Jun Qi 0001, Wei Wang 0042 |
ECAI | 8 |
| 2025 | MedFact: A Large-scale Chinese Dataset for Evidence-based Medical Fact-checking of LLM ResponsesabstractTong Chen, Zimu Wang, Yiyi Miao, Haoran Luo, Sun Yuanfei, Wei Wang, Zhengyong Jiang, Procheta Sen, Jionglong Su. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Tong Chen 0005, Yiyi Miao, Yuanfei Sun, Wei Wang 0042, Zhengyong Jiang, Procheta Sen, Jionglong Su |
EMNLP | 6 |
| 2025 | Can GRPO Boost Complex Multimodal Table Understanding?abstractXiaoqiang Kang, Shengen Wu, Zimu Wang, Yilin Liu, Xiaobo Jin, Kaizhu Huang, Wei Wang, Yutao Yue, Xiaowei Huang, Qiufeng Wang. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Xiaoqiang Kang, Shengen Wu, Xiao-Bo Jin, Kaizhu Huang, Wei Wang 0042, Yutao Yue |
EMNLP | 7 |
| 2025 | EEG-TBSANet: Temporal-Spectral Fusion Network for Robust Epilepsy Diagnosis from EEGabstractThis paper proposes a novel deep learning architecture, EEG-TBSANet, for automatic detection of epileptic seizures from electroencephalogram (EEG) signals. The model integrates temporal convolutional networks (TCN), bidirectional long short-term memory (BiLSTM) networks, and self-attention (SA) mechanisms in one single framework which allows the mechanism to locally and globally extract short and long temporal features while adaptively focusing on relevant components of the signals to detect seizures. EEG-TBSANet was fully tested on three publicly available benchmark datasets: Guinea-Bissau, Bonn, and CHB-MIT, which have different characteristics of signal types and acquisition scenarios. The model achieved classification accuracies of 98.02%, 98.57%, and 99.40% for the corresponding datasets under competitive conditions with several significant baseline and ablation models, including TCN-SA, LSTM-GRU, CNN-BiLSTM, TCN-BiLSTM, MDFLN. Additionally, an extensive ablation study confirmed the critical role of each architectural component in enhancing detection performance. The results show the model’s strong generalization ability across datasets with varied clinical and technical conditions. These results demonstrate that EEG-TBSANet offers a robust and generalizable solution for EEG-based epileptic seizure detection. Shenming Ji, Wei Wang 0042, Jun Qi 0001 |
INDIN | 2 |
| 2025 | Benchmarking and Improving LVLMs on Event Extraction from Multimedia DocumentsabstractThe proliferation of multimedia content necessitates the development of effective Multimedia Event Extraction (M²E²) systems. Though Large Vision-Language Models (LVLMs) have shown strong cross-modal capabilities, their utility in the M²E² task remains underexplored. In this paper, we present the first systematic evaluation of representative LVLMs, including DeepSeek-VL2 and the Qwen-VL series, on the M²E² dataset. Our evaluations cover text-only, image-only, and cross-media subtasks, assessed under both few-shot prompting and fine-tuning settings. Our key findings highlight the following valuable insights: (1) Few-shot LVLMs perform notably better on visual tasks but struggle significantly with textual tasks; (2) Fine-tuning LVLMs with LoRA substantially enhances model performance; and (3) LVLMs exhibit strong synergy when combining modalities, achieving superior performance in cross-modal settings. We further provide a detailed error analysis to reveal persistent challenges in areas such as semantic precision, localization, and cross-modal grounding, which remain critical obstacles for advancing M²E² capabilities. Fuyu Xing, Wei Wang 0042, Haiyang Zhang 0004 |
INLG | 3 |
| 2025 | FedGraphX: Split Federated Graph Learning for Cross-City AIoT Traffic Forecasting with Heterogeneous Sensor NetworksabstractThe rapid proliferation of Artificial Intelligence of Things (AIoT) devices in smart cities, such as roadside sensors and traffic cameras, enables real-time urban traffic monitoring through distributed sensor networks. These systems generate spatio-temporal data critical for Intelligent Transportation Systems (ITS), particularly traffic forecasting, and prediction of congestion, accidents, and travel times. However, existing forecasting methods struggle with cross-city collaboration due to heterogeneous sensor topologies and privacy constraints, limiting their effectiveness. Federated Learning partially mitigates privacy issues but fails to effectively address the topological heterogeneity of sensor networks, impeding robust cross-city collaboration and generalization. To address these limitations, we propose FedGraphX, a Split Federated Graph Learning framework that integrates localized processing with global collaboration. Each city operates as an independent AIoT node, using GRU-based encoders and GraphSAGE models to extract spatio-temporal features from local data. A central Graph Transformer aggregates features across cities, linking sensors via temporal and functional similarities, while a cross-layer attention mechanism aligns local spatial patterns with global dynamics. Experiments on four real-world datasets demonstrate superiority of FedGraphX over centralized and federated baselines, particularly in scenarios with sparse data or topological mismatches. Hanyue Xu, Li-Minn Ang, Kah Phooi Seng, Wei Wang 0042, Jieli Chen |
LCN | 4 |
| 2025 | DvD: Unleashing a Generative Paradigm for Document Dewarping via Coordinates-based Diffusion ModelabstractDocument dewarping aims to rectify deformations in photographic document images, thus improving text readability, which has attracted much attention and made great progress, but it is still challenging to preserve document structures. Given recent advances in diffusion models, it is natural for us to consider their potential applicability to document dewarping. However, it is far from straightforward to adopt diffusion models in document dewarping due to their unfaithful control on highly complex document images (e.g., 2000 × 3000 resolution). In this paper, we propose DvD, the first generative model to tackle document Dewarping via a Diffusion framework. To be specific, DvD introduces a coordinate-level denoising instead of typical pixel-level denoising, generating a mapping for deformation rectification. In addition, we further propose a time-variant condition refinement mechanism to enhance the preservation of document structures. In experiments, we find that current document dewarping benchmarks can not evaluate dewarping models comprehensively. To this end, we present AnyPhotoDoc6300, a rigorously designed large-scale document dewarping benchmark comprising 6,300 real image pairs across three distinct domains, enabling fine-grained evaluation of dewarping models. Comprehensive experiments demonstrate that our proposed DvD can achieve state-of-the-art performance with acceptable computational efficiency on multiple metrics across various benchmarks, including DocUNet, DIR300, and AnyPhotoDoc6300. The new benchmark and code will be publicly available at https://github.com/hanquansanren/DvD. Huangcheng Lu, Maizhen Ning, Xiaowei Huang 0001, Wei Wang 0042, Kaizhu Huang, Qiufeng Wang 0001 |
SIGGRAPH Asia | 5 |
| 2024 | MathAttack: Attacking Large Language Models towards Math Solving AbilityabstractWith the boom of Large Language Models (LLMs), the research of solving Math Word Problem (MWP) has recently made great progress. However, there are few studies to examine the robustness of LLMs in math solving ability. Instead of attacking prompts in the use of LLMs, we propose a MathAttack model to attack MWP samples which are closer to the essence of robustness in solving math problems. Compared to traditional text adversarial attack, it is essential to preserve the mathematical logic of original MWPs during the attacking. To this end, we propose logical entity recognition to identify logical entries which are then frozen. Subsequently, the remaining text are attacked by adopting a word-level attacker. Furthermore, we propose a new dataset RobustMath to evaluate the robustness of LLMs in math solving ability. Extensive experiments on our RobustMath and two another math benchmark datasets GSM8K and MultiAirth show that MathAttack could effectively attack the math solving ability of LLMs. In the experiments, we observe that (1) Our adversarial samples from higher-accuracy LLMs are also effective for attacking LLMs with lower accuracy (e.g., transfer from larger to smaller-size LLMs, or from few-shot to zero-shot prompts); (2) Complex MWPs (such as more solving steps, longer text, more numbers) are more vulnerable to attack; (3) We can improve the robustness of LLMs by using our adversarial samples in few-shot prompts. Finally, we hope our practice and observation can serve as an important attempt towards enhancing the robustness of LLMs in math solving ability. The code and dataset is available at: https://github.com/zhouzihao501/MathAttack. Qiufeng Wang 0001, Mingyu Jin, Jianan Ye, Wei Liu 0131, Wei Wang 0042, Xiaowei Huang 0001, Kaizhu Huang |
AAAI | 7 |
| 2024 | Generating Valid and Natural Adversarial Examples with Large Language ModelsabstractDeep learning-based natural language processing (NLP) models, particularly pre-trained language models (PLMs), have been revealed to be vulnerable to adversarial attacks. However, the adversarial examples generated by many mainstream word-level adversarial attack models are neither valid nor natural, leading to the loss of semantic maintenance, grammaticality, and human imperceptibility. Based on the exceptional capacity of language understanding and generation of large language models (LLMs), we propose LLM-Attack, which aims at generating both valid and natural adversarial examples with LLMs. The method consists of two stages: word importance ranking (which searches for the most vulnerable words) and word synonym replacement (which substitutes them with their synonyms obtained from LLMs). Experimental results on the Movie Review (MR), IMDB, and Yelp Review Polarity datasets against the baseline adversarial attack models illustrate the effectiveness of LLM-Attack, and it outperforms the baselines in human and GPT-4 evaluation by a significant margin. The model can generate adversarial examples that are typically valid and natural, with the preservation of semantic meaning, grammaticality, and human imperceptibility. Wei Wang 0042, Qi Chen 0026, Qiufeng Wang 0001, Anh Nguyen 0003 |
CSCWD | 2 |
| 2024 | Delving into Adversarial Robustness on Document Tampering Localization
Huiru Shao, Zhuang Qian, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001 |
ECCV (65) | 4 |
| 2024 | GANs-based Signal Quality Assessment for Heart Rate Estimation with BallistocardiographabstractThe ballistocardiograph (BCG) is a non-contact technology that monitors the heart and provides detailed cardiovascular parameters. Despite its broad applicability for long-term home monitoring due to Covid-19, BCG signals face challenges from positional changes, body movements, and system noise, which impact detection algorithms. In this paper, we propose a method for detecting inter-beat intervals (IBI) based on signal fusion technology. We utilize a Dynamic Bayesian Network (DBN) to integrate five heartbeat localization features extracted from BCG signals. Additionally, Generative Adversarial Networks (GANs) are used to assess signal quality and select correlated channels, improving heart rate monitoring accuracy. Experimental results demonstrate an average coverage of 95.21% and a mean squared error of 0.05. These results outperform those of methods without channel selection and single-channel BCG, indicating the potential for improving IBI estimation in multichannel BCG signal sensor systems. Ruilin Cai, Jun Qi 0001, Wei Wang 0042, Haiyang Zhang 0004 |
ISPA | 5 |
| 2024 | Domain-specific Guided Summarization for Mental Health Posts
Lu Qian, Haiyang Zhang 0004, Wei Wang 0042, Anh Nguyen 0003 |
PACLIC | 5 |
| 2024 | Self-supervised learning for point cloud data: A surveyabstract3D point clouds are a crucial type of data collected by LiDAR sensors and widely used in transportation applications due to its concise descriptions and accurate localization. Deep neural networks (DNNs) have achieved remarkable success in processing large amount of disordered and sparse 3D point clouds, especially in various computer vision tasks, such as pedestrian detection and vehicle recognition. Among all the learning paradigms, Self-Supervised Learning (SSL), an unsupervised training paradigm that mines effective information from the data itself, is considered as an essential solution to solve the time-consuming and labor-intensive data labelling problems via smart pre-training task design. This paper provides a comprehensive survey of recent advances on SSL for point clouds. We first present an innovative taxonomy, categorizing the existing SSL methods into four broad categories based on the pretexts’ characteristics. Under each category, we then further categorize the methods into more fine-grained groups and summarize the strength and limitations of the representative methods. We also compare the performance of the notable SSL methods in literature on multiple downstream tasks on benchmark datasets both quantitatively and qualitatively. Finally, we propose a number of future research directions based on the identified limitations of existing SSL research on point clouds. Changyu Zeng, Wei Wang 0042, Anh Nguyen 0003, Jimin Xiao, Yutao Yue |
Expert Syst. Appl. | 2 |
| 2024 | Zero-shot text classification with knowledge resources under label-fully-unseen setting
Wei Wang 0042, Qi Chen 0026, Kaizhu Huang, Anh Nguyen 0003, Suparna De |
Neurocomputing | 2 |
| 2024 | Scene Text Recognition via Dual-path Network with Shape-driven Attention AlignmentabstractScene text recognition (STR), one typical sequence-to-sequence problem, has drawn much attention recently in multimedia applications. To guarantee good performance, it is essential for STR to obtain aligned character-wise features from the whole-image feature maps. While most present works adopt fully data-driven attention-based alignment, such practice ignores specific character geometric information. In this article, built upon a group of learnable geometric points, we propose a novel shape-driven attention alignment method that is able to obtain character-wise features. Concretely, we first design a corner detector to generate a shape map to guide the attention alignments explicitly, where a series of points can be learned to represent character-wise features flexibly. We then propose a dual-path network with a mutual learning and cooperating strategy that successfully combines CNN with a ViT-based model, leading to further accuracy improvement. We conduct extensive experiments to evaluate the proposed method on various scene text benchmarks, including six popular regular and irregular datasets, two more challenging datasets (i.e., WordArt and OST), and three Chinese datasets. Experimental results indicate that our method can achieve superior performance with a comparable model size against many state-of-the-art models. Yijie Hu, Bin Dong 0003, Kaizhu Huang, Lei Ding 0012, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2023 | Progressive Supervision for Tampering Localization in Document Images
Huiru Shao, Kaizhu Huang, Wei Wang 0042, Xiaowei Huang 0001, Qiufeng Wang 0001 |
ICONIP (15) | 3 |
| 2023 | Knowledge-embedded Prompt Learning for Zero-shot Social Media Text ClassificationabstractSocial media plays an irreplaceable role in shaping the way information is created shared and consumed. While it provides access to a vast amount of data, extracting and analyzing useful insights from complex and dynamic social media data can be challenging. Deep learning models have shown promise in social media analysis tasks, but such models require a massive amount of labelled data which is usually unavailable in real-world settings. Additionally, these models lack common-sense knowledge which can limit their ability to generate comprehensive results. To address these challenges, we propose a knowledge-embedded prompt learning model for zero-shot social media text classification. Our experimental results on four social media datasets demonstrate that our proposed approach outperforms other well-known baselines. Qi Chen 0026, Wei Wang 0042, Fangyu Wu 0001 |
SMARTCOMP | 3 |
| 2023 | LSNCP: Lightweight and Secure Numeric Comparison Protocol for Wireless Body Area NetworksabstractWireless body area networks (WBANs) have been deployed in numerous applications, where the most common communication technology is Bluetooth. Bluetooth uses the numeric comparison protocol (NCP) to negotiate session keys based on the elliptic curve cryptography (ECC) and Out-of-Band (OoB) channels. However, the scalar multiplication of ECC is a heavy computing operation for devices in WBANs. To address this issue, we propose the lightweight and secure NCP (LSNCP) which requires less scalar multiplication than the NCP in Bluetooth. New logic expressions and rules are proposed to verify the security of LSNCP in GNY logic. The proof shows that LSNCP is secure. We conduct a provable security analysis by integrating the commitment scheme and short hash function. The result shows that LSNCP is secure in the modified Bellare–Rogaway model. Finally, we conduct theoretical analysis and experiments to evaluate the performance of LSNCP. The results confirm that LSNCP has less computation cost than NCP and other benchmark protocols. LSNCP has many potential application scenarios, such as healthcare, Metaverse, and blockchain. Haotian Yin, Xin Huang 0005, Xiaoxin Sun, Jianshuang Li, Sheng Chai, Rana Abubakar, Wei Wang 0042 |
IEEE Internet Things J. | 10 |
| 2022 | Generalised Zero-shot Learning for Entailment-based Text Classification with External KnowledgeabstractText classification techniques have been substantially important to many smart computing applications, e.g. topic extraction and event detection. However, classification is always challenging when only insufficient amount of labelled data for model training is available. To mitigate this issue, zero-shot learning (ZSL) has been introduced for models to recognise new classes that have not been observed during the training stage. We propose an entailment-based zero-shot text classification model, named as S-BERT-CAM, to better capture the relationship between the premise and hypothesis in the BERT embedding space. Two widely used textual datasets are utilised to conduct the experiments. We fine-tune our model using 50% of the labels for each dataset and evaluate it on the label space containing all labels (including both seen and unseen labels). The experimental results demonstrate that our model is more robust to the generalised ZSL and significantly improves the overall performance against baselines. Wei Wang 0042, Qi Chen 0026, Kaizhu Huang, Anh Nguyen 0003, Suparna De |
SMARTCOMP | 2 |
| 2022 | Zero-Shot Text Classification via Knowledge Graph Embedding for Social Media DataabstractThe idea of “citizen sensing” and “human as sensors” is crucial for social Internet of Things, an integral part of cyber–physical–social systems (CPSSs). Social media data, which can be easily collected from the social world, has become a valuable resource for research in many different disciplines, e.g., crisis/disaster assessment, social event detection, or the recent COVID-19 analysis. Useful information, or knowledge derived from social data, could better serve the public if it could be processed and analyzed in more efficient and reliable ways. Advances in deep neural networks have significantly improved the performance of many social media analysis tasks. However, deep learning models typically require a large amount of labeled data for model training, while most CPSS data is not labeled, making it impractical to build effective learning models using traditional approaches. In addition, the current state-of-the-art, pretrained natural language processing (NLP) models do not make use of existing knowledge graphs, thus often leading to unsatisfactory performance in real-world applications. To address the issues, we propose a new zero-shot learning method which makes effective use of existing knowledge graphs for the classification of very large amounts of social text data. Experiments were performed on a large, real-world tweet data set related to COVID-19, the evaluation results show that the proposed method significantly outperforms six baseline models implemented with state-of-the-art deep learning models for NLP. Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Frans Coenen |
IEEE Internet Things J. | 2 |
| 2022 | Ulixes: Facial Recognition Privacy with Adversarial Machine LearningabstractAbstract Facial recognition tools are becoming exceptionally accurate in identifying people from images. However, this comes at the cost of privacy for users of online services with photo management (e.g. social media platforms). Particularly troubling is the ability to leverage unsupervised learning to recognize faces even when the user has not labeled their images. In this paper we propose Ulixes, a strategy to generate visually non-invasive facial noise masks that yield adversarial examples, preventing the formation of identifiable user clusters in the embedding space of facial encoders. This is applicable even when a user is unmasked and labeled images are available online. We demonstrate the effectiveness of Ulixes by showing that various classification and clustering methods cannot reliably label the adversarial examples we generate. We also study the effects of Ulixes in various black-box settings and compare it to the current state of the art in adversarial machine learning. Finally, we challenge the effectiveness of Ulixes against adversarially trained models and show that it is robust to countermeasures. Thomas Cilloni, Wei Wang 0042, Charles Walter, Charles Fleming |
Proc. Priv. Enhancing Technol. | 2 |
| 2021 | Multi-task BERT for Aspect-based Sentiment AnalysisabstractSocial media data are increasingly used for smart computing applications, e.g., social event detection and sentiment analysis. Sentiment analysis, an important natural language processing task, has been applied in many real-world applications such as recommender systems and intelligence business systems. To process such social media data, natural language processing techniques such as BERT can be applied to extract essential language representations and produce state-of-the-art results. In this paper, we utilize the pre-trained BERT model as the backbone network and propose the BERT-SAN model to perform aspect-based sentiment analysis. The result demonstrates that our proposed model has a significant improvement against other baselines. Qi Chen 0026, Wei Wang 0042 |
SMARTCOMP | 3 |
| 2021 | Multi-modal generative adversarial networks for traffic event detection in smart cities
Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Suparna De, Frans Coenen |
Expert Syst. Appl. | 2 |
| 2021 | Automated Social Text Annotation With Joint Multilabel Attention NetworksabstractAutomated social text annotation is the task of suggesting a set of tags for shared documents on social media platforms. The automated annotation process can reduce users' cognitive overhead in tagging and improve tag management for better search, browsing, and recommendation of documents. It can be formulated as a multilabel classification problem. We propose a novel deep learning-based method for this problem and design an attention-based neural network with semantic-based regularization, which can mimic users' reading and annotation behavior to formulate better document representation, leveraging the semantic relations among labels. The network separately models the title and the content of each document and injects an explicit, title-guided attention mechanism into each sentence. To exploit the correlation among labels, we propose two semantic-based loss regularizers, i.e., similarity and subsumption, which enforce the output of the network to conform to label semantics. The model with the semantic-based loss regularizers is referred to as the joint multilabel attention network (JMAN). We conducted a comprehensive evaluation study and compared JMAN to the state-of-the-art baseline models, using four large, real-world social media data sets. In terms of F1, JMAN significantly outperformed bidirectional gated recurrent unit (Bi-GRU) relatively by around 12.8%-78.6% and the hierarchical attention network (HAN) by around 3.9%-23.8%. The JMAN model demonstrates advantages in convergence and training speed. Further improvement of performance was observed against latent Dirichlet allocation (LDA) and support vector machine (SVM). When applying the semantic-based loss regularizers, the performance of HAN and Bi-GRU in terms of F1was also boosted. It is also found that dynamic update of the label semantic matrices (JMANd) has the potential to further improve the performance of JMAN but at the cost of substantial memory and warrants further study. Hang Dong 0002, Wei Wang 0042, Kaizhu Huang, Frans Coenen |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2020 | Poster: An Improvement on Distance based Positioning on Network EdgesabstractDistance based positioning methods have been widely used in today’s wireless networks for positioning network users. In this paper, we present a study on distance based positioning at network edges. We show that existing methods may not be able to find the optimal position at network edges due to the presence of measurement noise and the use of biased estimation. To handle this problem, we propose an improvement on the estimation method. Simulation results show that the proposed improvement can reduce position error by 30% in 20% of a network area. Dawei Liu 0001, Ali H. Al-Bayatti, Wei Wang 0042 |
SEC | 3 |
| 2020 | Multi-modal Adversarial Training for Crisis-related Data Classification on Social MediaabstractSocial media platforms such as Twitter are increasingly used to collect data of all kinds. During natural disasters, users may post text and image data on social media platforms to report information about infrastructure damage, injured people, cautions and warnings. Effective processing and analysing tweets in real time can help city organisations gain situational awareness of the affected citizens and take timely operations. With the advances in deep learning techniques, recent studies have significantly improved the performance in classifying crisis-related tweets. However, deep learning models are vulnerable to adversarial examples, which may be imperceptible to the human, but can lead to model's misclassification. To process multi-modal data as well as improve the robustness of deep learning models, we propose a multi-modal adversarial training method for crisis-related tweets classification in this paper. The evaluation results clearly demonstrate the advantages of the proposed model in improving the robustness of tweet classification. Qi Chen 0026, Wei Wang 0042, Kaizhu Huang, Suparna De, Frans Coenen |
SMARTCOMP | 2 |
| 2020 | Knowledge base enrichment by relation learning from social tagging data
Hang Dong 0002, Wei Wang 0042, Frans Coenen, Kaizhu Huang |
Inf. Sci. | 2 |
| 2019 | Unbalancing Pairing-Free Identity-Based Authenticated Key Exchange Protocols for Disaster ScenariosabstractIn disaster scenarios, such as an area after a terrorist attack, security is a significant problem since communications involve information for the rescue officers, such as polices, militaries, emergency medical technicians, and the survivors. Such information is critically important for the rescue organizations; and protecting the privacy of the survivors is required. Normally, authenticated key exchange (AKE) is an underlying approach for security. However, available AKE protocols are either inconvenient or infeasible in disaster areas due to the very nature of disasters. To address the security problem in disaster scenarios, we propose two pairing-free identity-based AKE (ID-AKE) protocols that have unbalanced computational requirements on the two parties. Compared with existing AKE protocols, the proposed protocols have a number of advantages in disaster scenarios: 1) they are more convenient than symmetric cryptography-based AKE protocols since they do not require any preshared secret between the parties; 2) they are more feasible than asymmetric cryptography-based AKE protocols since they do not require any online server; and 3) they are more friendly to battery-powered and computationally limited devices than pairing-based and pairing-free ID-AKE protocols since they do not involve any bilinear pairing (a time-consuming operation), and have lower computational requirement on the limited party. Security of the proposed protocols are analyzed in detail; and prototypes of them are implemented to evaluate the performance. We also illustrate the application of the protocols through a vivid use case in a terrorist attack scenario. Jie Zhang 0030, Xin Huang 0005, Wei Wang 0042, Yong Yue 0001 |
IEEE Internet Things J. | 3 |
| 2018 | Learning Relations from Social Tagging Data
Hang Dong 0002, Wei Wang 0042, Frans Coenen |
PRICAI (1) | 2 |
| 2017 | A Denial of Service Attack Method for IoT System in Photovoltaic Energy System
Lulu Liang, Kai Zheng 0018, Qiankun Sheng, Wei Wang 0042, Xin Huang 0005 |
NSS | 4 |
| 2017 | Distributed sensor data computing in smart city applicationsabstractWith technologies developed in the Internet of Things, embedded devices can be built into every fabric of urban environments and connected to each other; and data continuously produced by these devices can be processed, integrated at different levels, and made available in standard formats through open services. The data, obviously f a form of `big data', is now seen as the most valuable asset in developing intelligent applications. As the sizes of the IoT data continue to grow, it becomes inefficient to transfer all the raw data to a centralised, cloud-based data centre and to perform efficient analytics even with the state-of-the-art big data processing technologies. To address the problem, this article demonstrates the idea of "distributed intelligence" for sensor data computing, which disperses intelligent computation to the much smaller while autonomous units, e.g., sensor network gateways, smart phones or edge clouds in order to reduce data sizes and to provide high quality data for data centres. As these autonomous units are usually in close proximity to data consumers, they also provide potential for reduced latency and improved quality of services. We present our research on designing methods and apparatus for distributed computing on sensor data, e.g., acquisition, discovery, and estimation, and provide a case study on urban air pollution monitoring and visualisation. Wei Wang 0042, Suparna De, Yuchao Zhou, Xin Huang 0005, Klaus Moessner |
WoWMoM | 1 |
| 2015 | An experimental study on geospatial indexing for sensor service discovery
Wei Wang 0042, Suparna De, Gilbert Cassar, Klaus Moessner |
Expert Syst. Appl. | 1 |
| 2015 | A ranking method for sensor services based on estimation of service access cost
Wei Wang 0042, Suparna De, Klaus Moessner, Zhili Sun |
Inf. Sci. | 1 |
| 2013 | Composition of services in pervasive environments: A Divide and Conquer approachabstractIn pervasive environments, availability and reliability of a service cannot always be guaranteed. In such environments, automatic and dynamic mechanisms are required to compose services or compensate for a service that becomes unavailable during the runtime. Most of the existing works on services composition do not provide sufficient support for automatic service provisioning in pervasive environments. We propose a Divide and Conquer algorithm that can be used at the service runtime to repeatedly divide a service composition request into several simpler sub-requests. The algorithm repeats until for each sub-request we find at least one atomic service that meets the requirements of that sub-request. The identified atomic services can then be used to create a composite service. We discuss the technical details of our approach and show evaluation results based on a set of composite service requests. The results show that our proposed method performs effectively in decomposing a composite service requests to a number of sub-requests and finding and matching service components that can fulfill the service composition request. Gilbert Cassar, Payam M. Barnaghi, Wei Wang 0042, Suparna De, Klaus Moessner |
ISCC | 3 |
| 2013 | Open services for IoT cloud applications in the future internetabstractInternet of Things (IoT) is an emerging area that not only requires development of infrastructure but also deployment of new services capable of supporting multiple, scalable (cloud-based) and interoperable (multi-domain) applications. In the race of designing the IoT as part of the Future Internet architecture, academic and ICT's (Information and Communication Technology) industry communities have realized that a common IoT problem to be tackled is the interoperability of the information. In this paper we review recent trends and challenges on interoperability, and discuss how semantic technologies, open service frameworks and information models can support data interoperability in the design of the Future Internet, taking the IoT and Cloud Computing as reference examples of application domains. Martin Serrano, Hoan Quoc Nguyen-Mau, Manfred Hauswirth, Wei Wang 0042, Payam M. Barnaghi, Philippe Cousin |
WOWMOM | 4 |
| 2012 | A Comprehensive Ontology for Knowledge Representation in the Internet of ThingsabstractSemantic modeling for the Internet of Things has become fundamental to resolve the problem of interoperability given the distributed and heterogeneous nature of the "Things". Most of the current research has primarily focused on devices and resources modeling while paid less attention on access and utilisation of the information generated by the things. The idea that things are able to expose standard service interfaces coincides with the service oriented computing and more importantly, represents a scalable means for business services and applications that need context awareness and intelligence to access and consume the physical world information. We present the design of a comprehensive description ontology for knowledge representation in the domain of Internet of Things and discuss how it can be used to support tasks such as service discovery, testing and dynamic composition. Wei Wang 0042, Suparna De, Ralf Tönjes, Eike Steffen Reetz, Klaus Moessner |
TrustCom | 1 |
| 2012 | Semantics for the Internet of Things: Early Progress and Back to the FutureabstractThe Internet of Things (IoT) has recently received considerable interest from both academia and industry that are working on technologies to develop the future Internet. It is a joint and complex discipline that requires synergetic efforts from several communities such as telecommunication industry, device manufacturers, semantic Web, and informatics and engineering. Much of the IoT initiative is supported by the capabilities of manufacturing low-cost and energy-efficient hardware for devices with communication capacities, the maturity of wireless sensor network technologies, and the interests in integrating the physical and cyber worlds. However, the heterogeneity of the “Things” makes interoperability among them a challenging problem, which prevents generic solutions from being adopted on a global scale. Furthermore, the volume, velocity and volatility of the IoT data impose significant challenges to existing information systems. Semantic technologies based on machine-interpretable representation formalism have shown promise for describing objects, sharing and integrating information, and inferring new knowledge together with other intelligent processing techniques. However, the dynamic and resource-constrained nature of the IoT requires special design considerations to be taken into account to effectively apply the semantic technologies on the real world data. In this article the authors review some of the recent developments on applying the semantic technologies to IoT. Payam M. Barnaghi, Wei Wang 0042, Cory A. Henson, Kerry L. Taylor |
Int. J. Semantic Web Inf. Syst. | 2 |
| 2011 | Rational Research model for ranking semantic entities
Wei Wang 0042, Payam M. Barnaghi, Andrzej Bargiela |
Inf. Sci. | 1 |
| 2010 | Probabilistic Topic Models for Learning Terminological OntologiesabstractProbabilistic topic models were originally developed and utilized for document modeling and topic extraction in Information Retrieval. In this paper, we describe a new approach for automatic learning of terminological ontologies from text corpus based on such models. In our approach, topic models are used as efficient dimension reduction techniques, which are able to capture semantic relationships between word-topic and topic-document interpreted in terms of probability distributions. We propose two algorithms for learning terminological ontologies using the principle of topic relationship and exploiting information theory with the probabilistic topic models learned. Experiments with different model parameters were conducted and learned ontology statements were evaluated by the domain experts. We have also compared the results of our method with two existing concept hierarchy learning methods on the same data set. The study shows that our method outperforms other methods in terms of recall and precision measures. The precision level of the learned ontology is sufficient for it to be deployed for the purpose of browsing, navigation, and information search and retrieval in digital libraries. Wei Wang 0042, Payam M. Barnaghi, Andrzej Bargiela |
IEEE Trans. Knowl. Data Eng. | 1 |