EDBT 2026 Demo / reviewers in the wild / expert
Hua Wang 0002
dblp:33/3535-2
· DBLP profile ↗
72ranked-venue papers in the field
7as first author
25since 2021 · last 2026
0000-0002-8465-0996ORCID · conflict
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 36 (4 first)Data Mining & Knowledge Discovery · 14Database Systems & Data Management · 13 (2 first)Other / Interdisciplinary · 6Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Effective Fairest Community Search Over Heterogeneous Information Networks
Taige Zhao, Jingxian Cheng, Hua Wang 0002 |
ICDE | 7 |
| 2026 | Towards Evolutionary Differential Privacy in Cross-Platform Spatial CrowdsourcingabstractThe development of mobile web services has brought significant attention to spatial crowdsourcing. The uneven distribution of tasks and workers has led to recent research on Cross-Platform Spatial Crowdsourcing (CPSC), aiming for a multi-win situation for platforms, workers, and task requesters. Previous studies on CPSC problems focused on task assignment and worker selection performance, overlooking the importance of privacy preservation. This article addresses the existing challenges of privacy preservation and service quality by formulating a Privacy-Preserving Cross-Platform Spatial Crowdsourcing (PP-CPSC) problem and proves it to be NP-hard. We propose an Evolutionary Differential Privacy (Evo-DP) approach to optimize PP-CPSC. Evo-DP’s evolutionary framework enables efficient and flexible optimization of privacy budget allocation. Within Evo-DP, each solution to the privacy budget allocation is represented as an individual in the population. To approximate the optimal solution, three evolutionary operations—mutation, crossover, and scaling—are employed for population updates, along with a selection process. A hybrid population model is introduced to balance exploration and exploitation abilities. Experimental results demonstrate Evo-DP’s superiority over previous strategies in terms of solution quality, convergence speed, and scalability. Yong-Feng Ge, Hua Wang 0002, Elisa Bertino, Jinli Cao, Yanchun Zhang, Zhonglong Zheng |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2026 | Transformer-Enhanced Adaptive Graph Convolutional Network for Traffic Flow PredictionabstractTraffic flow prediction is vital in urban traffic management, planning, and development. With the continuous advancement of urbanization, there is an increasing demand for traffic flow prediction models to achieve higher accuracy and long-range forecasting capabilities. Against this backdrop, traditional methods that rely on local feature extraction and static spatial graph construction often fall short of expectations. This highlights the urgent need for advanced approaches to dynamically model spatio-temporal features while capturing global dependencies, effectively meeting the demands of complex traffic flow prediction tasks. To achieve this, we propose the Transformer-Enhanced Adaptive Graph Convolutional Network (T-AGCN), a novel model designed to capture global temporal relationships and dynamically extract rich spatial information. T-AGCN incorporates an Adaptive Graph Learner module to model dynamic relationships among traffic nodes and a Transformer-Based Spatio-Temporal graph convolutional module to capture long-range temporal dependencies in historical traffic data effectively. These innovations enable T-AGCN to jointly learn dynamic spatial interactions and complex temporal patterns, offering a comprehensive representation of traffic network dynamics. We evaluate T-AGCN on three real-world datasets, PeMSD7(M), PeMS08, and METR-LA. The experimental results demonstrate that T-AGCN, inspired by the baseline model Spatial-Temporal Graph Convolutional Network (STGCN), significantly enhances its design. Moreover, T-AGCN consistently outperforms state-of-the-art models, including the Transformer-Based Interactive Temporal and Adaptive Network (TITAN) and the Spatial-Temporal Decoupled Masked Autoencoder (STD-MAE). The implementation is available on GitHub at https://github.com/time1722/T-AGCN . Enfu Huang, Zhanshan Zhao, Jiao Yin 0003, Jinli Cao, Hua Wang 0002 |
ACM Trans. Intell. Syst. Technol. | 5 |
| 2026 | ITCoHD-MRec: An Independent Topological Preference-Aware and Cooperative Hypergraph Diffusion-Based Multimodal Recommender ModelabstractMultimodal recommendation provides richer and more accurate personalized recommendations by jointly modeling user’s historical behaviors and different modality of items, such as text, image, audio, and video in online platforms. Most existing work of multimodal recommendation focuses on leveraging modal features and modal correlation graph structures to learn user preferences. Due to insufficient exploration of user collaborative preferences and the noise during high-order multimodal data connections, valuable information may be lost, leading to deviations in understanding user preferences. Therefore, an I ndependent T opological Preference-Aware and Co operative H ypergraph D iffusion-based M ultimodal Rec ommender Model (ITCoHD-MRec) is necessary for online platforms. This article aims to develop an ITCoHD-MRec that incorporates topological perception as well as generative diffusion models in multimodal hypergraph recommendation to make the model more adaptive and robust in complex environments. Firstly, leveraging a Graph Convolutional Network (GCN), the model independently captures user preference representations for both collaborative relevance and modal relevance from the user-item interaction graph, which contains ID embeddings and modal features. This enables the extraction of deeper associations between users and items. Secondly, leveraging topological pruning techniques, the model learns differentiated features in different modal blocks to prevent node representations from becoming homogenized. This helps further identify user preferred connectivity patterns and removes redundant noisy connections. Finally, by employing the diffusion model, information regarding the higher-order interaction patterns between attributes and items within the hypergraph structure is propagated. This effectively captures the potential global dependencies between attributes and items, thereby providing deeper associations enriched with more substantial semantic information for subsequent recommendation tasks. The model autonomously learns different features and higher-order connectivity of nodes, which enables the model to obtain a wider and more accurate perception of user preferences in complex interaction environments. Experimental comparisons with 15 models on four real datasets—Baby, Sports, Clothing, and Electronics show that the model improves the recall by 0.85%–3.57% and the normalized discounted cumulative gain by 2.31%–3.43%, which validates the effectiveness of ITCoHD-MRec. Xiulan Hao, Hua Wang 0002, Zhonglong Zheng, Yunliang Jiang, Yanchun Zhang |
ACM Trans. Inf. Syst. | 3 |
| 2024 | CADIF-OSN: Detecting Cloned Accounts with Missing Profile Attributes on Online Social NetworksabstractThe growth of online social networks (OSNs) has become increasingly significant. Potential cloned accounts on these platforms raise serious concerns due to the risks they pose to user privacy and security. Previous works in the detection of cloned accounts on OSNs do not yield satisfactory results and lack consideration of the impact of missing attributes on the detection process. We propose cloned account detection with imputation framework for online social networks (CADIF-OSN) to accurately find potential cloned accounts on OSNs. This framework enables the accurate identification of potential cloned accounts on OSNs by leveraging their public profile information, even in cases where some of the information may not be accessible. The framework comprises four key components: 1) Fuzzy string matching with Levenshtein Distance that quickly generates suspicious account pairs by matching all the accounts' usernames and screennames; 2) An embedded method Doc2Vec that transforms all existing profile information of accounts into estimable vectors; 3) A HyperImpute model that imputes the missing information; and 4) A deep-forest model that is trained to detect cloned accounts. We evaluated our framework using a Twitter dataset consisting of 3,826 pairs of cloned accounts and 70,000 normal accounts. The evaluation results demonstrate that our framework significantly surpasses existing approaches in terms of Precision and F1-score. Dewei Ning, Yong-Feng Ge, Hua Wang 0002, Changjun Zhou |
CIKM | 3 |
| 2024 | Dynamic-Parameter Genetic Algorithm for Multi-objective Privacy-Preserving Trajectory Data Publishing
Samsad Jahan, Yong-Feng Ge, Hua Wang 0002, Md. Enamul Kabir |
WISE (5) | 3 |
| 2024 | EBUD: Evolving Disaster Burst Detection over Social Streams
Xiyu Qiao, Xiangmin Zhou, Changjun Zhou, Hua Wang 0002, Yanchun Zhang |
WISE (2) | 4 |
| 2024 | Optimising Insider Threat Prediction: Exploring BiLSTM Networks and Sequential FeaturesabstractAbstract Insider threats pose a critical risk to organisations, impacting their data, processes, resources, and overall security. Such significant risks arise from individuals with authorised access and familiarity with internal systems, emphasising the potential for insider threats to compromise the integrity of organisations. Previous research has addressed the challenge by pinpointing malicious actions that have already occurred but provided limited assistance in preventing those risks. In this research, we introduce a novel approach based on bidirectional long short-term memory (BiLSTM) networks that effectively captures and analyses the patterns of individual actions and their sequential dependencies. The focus is on predicting whether an individual would be a malicious insider in a future day based on their daily behavioural records over the previous several days. We analyse the performance of the four supervised learning algorithms on manual features, sequential features, and the ground truth of the day with different combinations. In addition, we investigate the performance of different RNN models, such as RNN, LSTM, and BiLSTM, in incorporating these features. Moreover, we explore the performance of different predictive lengths on the ground truth of the day and different embedded lengths for the sequential features. All the experiments are conducted on the CERT r4.2 dataset. Experiment results show that BiLSTM has the highest performance in combining these features. Phavithra Manoharan, Jiao Yin 0003, Hua Wang 0002, Yanchun Zhang |
Data Sci. Eng. | 4 |
| 2024 | A Comprehensive Survey of Animal Identification: Exploring Data Sources, AI Advances, Classification Obstacles and the Role of TaxonomyabstractWith the rapid development of entity recognition technology, animal recognition has gradually become essential in modern society, supporting labour‐intensive agriculture and animal husbandry tasks. Severe problems such as maintaining biodiversity can also benefit from animal identification technology. However, certain invasive recognition systems have resulted in permanent harm to animals, while noninvasive identification methods also exhibit certain drawbacks. This paper conducts a systematic literature review (SLR), presenting a comprehensive overview of various animal recognition technologies and their applications. Specifically, it examines methodologies such as deep learning, image processing and acoustic analysis used for different animal characteristics and identification purposes. The contribution of machine learning to animal feature extraction is highlighted, emphasising its significance for animal taxonomy and wild species monitoring. Additionally, this review addresses the challenges and limitations of current technologies, including data scarcity, model accuracy and computational requirements, and suggests opportunities for future research to overcome these obstacles. Khandakar Ahmed, Nalin K. Sharda, Hua Wang 0002 |
Int. J. Intell. Syst. | 4 |
| 2024 | Distributed Cooperative Coevolution of Data Publishing Privacy and TransparencyabstractData transparency is beneficial to data participants’ awareness, users’ fairness, and research work’s reproducibility. However, when addressing transparency requirements, we cannot ignore data privacy. This article defines the multi-objective data publishing (MODP) problem, optimizing data privacy and transparency at the same time. Accordingly, we propose a distributed cooperative coevolutionary genetic algorithm (DCCGA) to optimize the MODP problem. In the population of DCCGA, each individual represents an anonymization solution to MODP. Three modules in DCCGA, i.e., grouping module, cooperative coevolutionary module, and evolving module, are proposed for distributed sub-population update and evaluation, improving DCCGA’s optimization performance and parallel efficiency. Moreover, a matrix-based crossover operator and a matrix-based mutation operator are designed to exchange and adjust anonymization information in the individuals efficiently. Experimental results demonstrate that the proposed DCCGA outperforms the competitors with respect to solution accuracy, convergence speed, and scalability. Besides, we verify the effectiveness of all the proposed components in DCCGA. Yong-Feng Ge, Elisa Bertino, Hua Wang 0002, Jinli Cao, Yanchun Zhang |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | A Compact Vulnerability Knowledge Graph for Risk AssessmentabstractSoftware vulnerabilities, also known as flaws, bugs or weaknesses, are common in modern information systems, putting critical data of organizations and individuals at cyber risk. Due to the scarcity of resources, initial risk assessment is becoming a necessary step to prioritize vulnerabilities and make better decisions on remediation, mitigation, and patching. Datasets containing historical vulnerability information are crucial digital assets to enable AI-based risk assessments. However, existing datasets focus on collecting information on individual vulnerabilities while simply storing them in relational databases, disregarding their structural connections. This article constructs a compact vulnerability knowledge graph, VulKG, containing over 276 K nodes and 1 M relationships to represent the connections between vulnerabilities, exploits, affected products, vendors, referred domain names, and more. We provide a detailed analysis of VulKG modeling and construction, demonstrating VulKG-based query and reasoning, and providing a use case of applying VulKG to a vulnerability risk assessment task, i.e., co-exploitation behavior discovery. Experimental results demonstrate the value of graph connections in vulnerability risk assessment tasks. VulKG offers exciting opportunities for more novel and significant research in areas related to vulnerability risk assessment. The data and codes of this article are available at https://github.com/happyResearcher/VulKG.git . Jiao Yin 0003, Hua Wang 0002, Jinli Cao, Yuan Miao 0001, Yanchun Zhang |
ACM Trans. Knowl. Discov. Data | 3 |
| 2024 | XKT: Toward Explainable Knowledge Tracing Model With Cognitive Learning Theories for Questions of Multiple Knowledge ConceptsabstractDeep learning (DL) based knowledge tracing (KT) models have challenges for uninterpretable prediction and parameter representation in educational applications, though they achieved remarkable outcomes in predicting the exercise performance of students. This paper proposes a novel knowledge tracing model of high precision and interpretability (namedXKT) for questions with multiple knowledge concepts based on cognitive learning theories and multidimensional item response theory (MIRT). TheXKTconsists of three differentiable network components: multi-feature embedding, cognition processing network, andMIRT-based neural predictor, which aim to provide an explainable prediction of student exercise performance. Specifically, inXKT, multi-feature embedding learns the rich semantic representation (e.g., knowledge distribution information) to enhance knowledge tracing using a cognition processing network. The cognition processing network performs selective perception, ability memory processing, and long-term knowledge memory processing to ensure the explainable factor representation for theMIRT-based neural predictor. Lastly, theMIRT-based neural predictor employs psychometric parameters to interpret student exercise predictions better. Extensive experiments on four real-world datasets show thatXKToutperforms existingKTmethods in predicting future learner responses. Moreover, ablation studies further show thatXKToffers good interpretability of student performance predictions with multiple knowledge concepts, indicating excellent potential in real-world educational applications. Changqin Huang, Qionghao Huang, Xiaodi Huang 0001, Hua Wang 0002, Ming Li 0065, Kwei-Jay Lin |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Anomaly Detection in Quasi-Periodic Time Series based on Automatic Data Segmentation and Attentional LSTM-CNN (Extended Abstract)abstractQuasi-periodic time series (QTS) exists widely in the real world, and it is important to detect the anomalies of QTS. In this paper, we propose an automatic QTS anomaly detection framework (AQADF) consisting of a two-level clustering-based QTS segmentation algorithm (TCQSA) and a hybrid attentional LSTM-CNN model (HALCM). TCQSA first automatically splits the QTS into quasi-periods which are then classified by HALCM into normal periods or anomalies. Notably, TCQSA integrates a hierarchical clustering and the k-means technique, making itself highly universal and noise-resistant. HALCM hybridizes LSTM and CNN to simultaneously extract the overall variation trends and local features of QTS for modeling its fluctuation pattern. Furthermore, we embed a trend attention gate (TAG) into the LSTM, a feature attention mechanism (FAM) and a location attention mechanism (LAM) into the CNN to finely tune the extracted variation trends and local features according to their true importance to yield a better representation of the fluctuation pattern of the QTS. On four public datasets, HALCM exceeds four state-of-the-art baselines and obtains at least 97.3% accuracy, TCQSA exceeds two cutting-edge QTS segmentation algorithms and can be applied to different types of QTSs. Fan Liu 0007, Xingshe Zhou 0001, Jinli Cao, Zhu Wang 0001, Tianben Wang, Hua Wang 0002, Yanchun Zhang |
ICDE | 6 |
| 2023 | Empowering Vulnerability Prioritization: A Heterogeneous Graph-Driven Framework for Exploitability Prediction
Jiao Yin 0003, Guihong Chen, Hua Wang 0002, Jinli Cao, Yuan Miao 0001 |
WISE | 4 |
| 2023 | TLEF: Two-Layer Evolutionary Framework for t-Closeness Anonymization
Mingshan You, Yong-Feng Ge, Kate N. Wang 0001, Hua Wang 0002, Jinli Cao, Georgios Kambourakis |
WISE | 4 |
| 2022 | An Information-Driven Genetic Algorithm for Privacy-Preserving Data Publishing
Yong-Feng Ge, Hua Wang 0002, Jinli Cao, Yanchun Zhang |
WISE | 2 |
| 2022 | DSGA: A Distributed Segment-Based Genetic Algorithm for Multi-Objective Outsourced Database Partitioning
Yong-Feng Ge, Zhi-hui Zhan, Jinli Cao, Hua Wang 0002, Yanchun Zhang, Kuei-Kuei Lai, Jun Zhang 0003 |
Inf. Sci. | 4 |
| 2022 | Anomaly Detection in Quasi-Periodic Time Series Based on Automatic Data Segmentation and Attentional LSTM-CNNabstractQuasi-periodic time series (QTS) exists widely in the real world, and it is important to detect the anomalies of QTS. In this paper, we propose anautomaticQTSanomalydetectionframework (AQADF) consisting of a two-level clustering-based QTS segmentation algorithm (TCQSA) and a hybrid attentional LSTM-CNN model (HALCM). TCQSA first automatically splits the QTS into quasi-periods which are then classified by HALCM into normal periods or anomalies. Notably, TCQSA integrates a hierarchical clustering and the k-means technique, making itself highly universal and noise-resistant. HALCM hybridizes LSTM and CNN to simultaneously extract the overall variation trends and local features of QTS for modeling its fluctuation pattern. Furthermore, we embed a trend attention gate (TAG) into the LSTM, a feature attention mechanism (FAM) and a location attention mechanism (LAM) into the CNN to finely tune the extracted variation trends and local features according to their true importance to achieve a better representation of the fluctuation pattern of the QTS. On four public datasets, HALCM exceeds four state-of-the-art baselines and obtains at least 97.3 percent accuracy, TCQSA outperforms two cutting-edge QTS segmentation algorithms and can be applied to different types of QTSs. Additionally, the effectiveness of the attention mechanisms is quantitatively and qualitatively demonstrated. Fan Liu 0007, Xingshe Zhou 0001, Jinli Cao, Zhu Wang 0001, Tianben Wang, Hua Wang 0002, Yanchun Zhang |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2022 | MDDE: multitasking distributed differential evolution for privacy-preserving database fragmentation
Yong-Feng Ge, Maria E. Orlowska, Jinli Cao, Hua Wang 0002, Yanchun Zhang |
VLDB J. | 4 |
| 2021 | Developing a Deep Learning Based Approach for Anomalies Detection from EEG Data
Ashik Mostafa Alvi, Siuly Siuly, Hua Wang 0002 |
WISE (1) | 3 |
| 2021 | Data Mining Based Artificial Intelligent Technique for Identifying Abnormalities from Brain Signal Data
Md. Nurul Ahad Tawhid, Siuly Siuly, Kate N. Wang 0001, Hua Wang 0002 |
WISE (1) | 4 |
| 2021 | A Minority Class Boosted Framework for Adaptive Access Control Decision-Making
Mingshan You, Jiao Yin 0003, Hua Wang 0002, Jinli Cao, Yuan Miao 0001 |
WISE (1) | 3 |
| 2021 | Set-Based Adaptive Distributed Differential Evolution for Anonymity-Driven Database FragmentationabstractAbstract By breaking sensitive associations between attributes, database fragmentation can protect the privacy of outsourced data storage. Database fragmentation algorithms need prior knowledge of sensitive associations in the tackled database and set it as the optimization objective. Thus, the effectiveness of these algorithms is limited by prior knowledge. Inspired by the anonymity degree measurement in anonymity techniques such as k-anonymity, an anonymity-driven database fragmentation problem is defined in this paper. For this problem, a set-based adaptive distributed differential evolution (S-ADDE) algorithm is proposed. S-ADDE adopts an island model to maintain population diversity. Two set-based operators, i.e., set-based mutation and set-based crossover, are designed in which the continuous domain in the traditional differential evolution is transferred to the discrete domain in the anonymity-driven database fragmentation problem. Moreover, in the set-based mutation operator, each individual’s mutation strategy is adaptively selected according to the performance. The experimental results demonstrate that the proposed S-ADDE is significantly better than the compared approaches. The effectiveness of the proposed operators is verified. Yong-Feng Ge, Jinli Cao, Hua Wang 0002, Yanchun Zhang |
Data Sci. Eng. | 3 |
| 2021 | Image Preprocessing in Classification and Identification of Diabetic Eye DiseasesabstractDiabetic eye disease (DED) is a cluster of eye problem that affects diabetic patients. Identifying DED is a crucial activity in retinal fundus images because early diagnosis and treatment can eventually minimize the risk of visual impairment. The retinal fundus image plays a significant role in early DED classification and identification. An accurate diagnostic model's development using a retinal fundus image depends highly on image quality and quantity. This paper presents a methodical study on the significance of image processing for DED classification. The proposed automated classification framework for DED was achieved in several steps: image quality enhancement, image segmentation (region of interest), image augmentation (geometric transformation), and classification. The optimal results were obtained using traditional image processing methods with a new build convolution neural network (CNN) architecture. The new built CNN combined with the traditional image processing approach presented the best performance with accuracy for DED classification problems. The results of the experiments conducted showed adequate accuracy, specificity, and sensitivity. Rubina Sarki, Khandakar Ahmed, Hua Wang 0002, Yanchun Zhang, Jiangang Ma, Kate N. Wang 0001 |
Data Sci. Eng. | 3 |
| 2021 | CyberPulse++: A machine learning-based security framework for detecting link flooding attacks in software defined networksabstractA new class of link flooding attacks (LFA) can cut off internet connections of target links by employing legitimate flows to congest these without being detected. LFA is especially powerful in disrupting traffic in software-defined networks if the control channel is targeted. Most of the existing solutions work by conducting a deep packet-level inspection of the physical network links. Therefore these techniques incur a significant performance overhead, are reactive, and result in damage to the network before a delayed defense is mounted. Machine learning (ML) of captured network statistics is emerging as a promising, lightweight, and proactive solution to defend against LFA. In this paper, we propose a ML-based security framework, CyberPulse++, that utilizes a pretrained ML repository to test captured network statistics in real-time to detect abnormal path performance on network links. It effectively tackles several challenges faced by network security solutions such as the practicality of large-scale network-level monitoring and collection of network status information. The framework can use a wide variety of algorithms for training the ML repository and allows the analyst a birds-eye view by generating interactive graphs to investigate an attack in its ramp-up stage. An extensive evaluation demonstrates that the framework offers limited bandwidth and computational overhead in proactively detecting and defending against LFA in real-time. Raihan Ur Rasool, Khandakar Ahmed, Zahid Anwar, Hua Wang 0002, Usman Ashraf, Wajid Rafique |
Int. J. Intell. Syst. | 4 |
| 2020 | An Advanced Two-Step DNN-Based Framework for Arrhythmia Detection
Jinyuan He, Jia Rong, Le Sun 0003, Hua Wang 0002, Yanchun Zhang |
PAKDD (2) | 4 |
| 2020 | Distributed Differential Evolution for Anonymity-Driven Vertical Fragmentation in Outsourced Data Storage
Yong-Feng Ge, Jinli Cao, Hua Wang 0002, Yanchun Zhang |
WISE (2) | 3 |
| 2020 | Adaptive Online Learning for Vulnerability Exploitation Time Prediction
Jiao Yin 0003, MingJian Tang 0001, Jinli Cao, Hua Wang 0002, Mingshan You, Yongzheng Lin |
WISE (2) | 4 |
| 2020 | An Interactive Network for End-to-End Review Helpfulness ModelingabstractAbstract Review helpfulness prediction aims to prioritize online reviews by quality. Existing methods largely combine review texts and star ratings for helpfulness prediction. However, star ratings are used in a way that has either little representation capacity or limited interaction with review texts. As a result, rating information has yet to be fully exploited during the combination. This paper aims to overcome the two drawbacks. A deep interactive architecture is proposed to learn the text–rating interaction (TRI) for helpfulness modeling. TRI enlarges the representation capacity of star ratings while enhancing the influence of rating information on review texts. TRI is evaluated on six real-world domains of the Amazon 5-Core dataset. Extensive experiments demonstrate that TRI can better predict review helpfulness and beat the state of the art. Ablation studies and qualitative analysis are provided to further understand model behaviors and the learned parameters. Jiahua Du, Liping Zheng, Jiantao He, Jia Rong, Hua Wang 0002, Yanchun Zhang |
Data Sci. Eng. | 5 |
| 2019 | Discriminative Regularized Deep Generative Models for Semi-Supervised LearningabstractDeep generative models (DGMs) have shown strong performance in semi-supervised learning (SSL), which incorporate discrete class information into the learning process. Yet existing methods generally overfit to the given labeled data, for only considering the conditional probability of labels. In this paper, we propose a novel discriminative regularized deep generative method for SSL, which fully exploits the discriminative and geometric information of data to address the aforementioned issue. Our method introduces the cluster and manifold assumption that maximizes the classification margin between clusters and simultaneously smooths the predictions of the data which is close in the sub-manifold of each cluster, to regularize the learning of the classifier in DGMs. To derive the regularization based on introduced assumptions, we adopt the generated data of DGMs along with labelled and unlabelled data, to model the data manifold and yield clusters based on the Gumbel-softmax distribution. Experimental results on both text and image datasets demonstrate the effectiveness and flexibility of our method, and prove that two introduced assumptions are complementary in guiding the classification boundary, thus improving the discriminative ability of the classifier. Qianqian Xie, Jimin Huang, Min Peng 0002, Yihan Zhang 0005, Kaifei Peng, Hua Wang 0002 |
ICDM | 6 |
| 2019 | Arrhythmias Classification by Integrating Stacked Bidirectional LSTM and Two-Dimensional CNN
Fan Liu 0007, Xingshe Zhou 0001, Jinli Cao, Zhu Wang 0001, Hua Wang 0002, Yanchun Zhang |
PAKDD (2) | 5 |
| 2019 | Helpfulness Prediction for Online Reviews with Explicit Content-Rating Interaction
Jiahua Du, Jia Rong, Hua Wang 0002, Yanchun Zhang |
WISE | 3 |
| 2019 | Pattern Filtering Attention for Distant Supervised Relation Extraction via Online Clustering
Min Peng 0002, Qingwen Liao, Weilong Hu, Gang Tian, Hua Wang 0002, Yanchun Zhang |
WISE | 5 |
| 2019 | Incorporating word embeddings into topic modeling of short text
Wang Gao 0002, Min Peng 0002, Hua Wang 0002, Yanchun Zhang, Qianqian Xie, Gang Tian |
Knowl. Inf. Syst. | 3 |
| 2019 | Bayesian Sparse Topical CodingabstractSparse topic models (STMs) are widely used for learning a semantically rich latent sparse representation of short texts in large scale, mainly by imposing sparse priors or appropriate regularizers on topic models. However, it is difficult for these STMs to model the sparse structure and pattern of the corpora accurately, since their sparse priors always fail to achieve real sparseness, and their regularizers bypass the prior information of the relevance between sparse coefficients. In this paper, we propose a novel Bayesian hierarchical topic models called Bayesian Sparse Topical Coding with Poisson Distribution (BSTC-P) on the basis of Sparse Topical Coding with Sparse Groups (STCSG). Different from traditional STMs, it focuses on imposing hierarchical sparse prior to leverage the prior information of relevance between sparse coefficients. Furthermore, we propose a sparsity-enhanced BSTC, Bayesian Sparse Topical Coding with Normal Distribution (BSTC-N), via mathematic approximation. We adopt superior hierarchical sparse inducing prior, with the purpose of achieving the sparsest optimal solution. Experimental results on datasets of Newsgroups and Twitter show that both BSTC-P and BSTC-N have better performance on finding clear latent semantic representations. Therefore, they yield better performance than existing works on document classification tasks. Min Peng 0002, Qianqian Xie, Hua Wang 0002, Yanchun Zhang, Gang Tian |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2018 | D-ECG: A Dynamic Framework for Cardiac Arrhythmia Detection from IoT-Based ECGs
Jinyuan He, Jia Rong, Le Sun 0003, Hua Wang 0002, Yanchun Zhang, Jiangang Ma |
WISE (2) | 4 |
| 2018 | Topic-Net Conversation Model
Min Peng 0002, Dian Chen 0004, Qianqian Xie, Yanchun Zhang, Hua Wang 0002, Gang Hu 0003, Wang Gao 0002, Yihan Zhang 0005 |
WISE (1) | 5 |
| 2018 | A Novel Incremental Dictionary Learning Method for Low Bit Rate Speech Streaming
Luyao Teng, Yingxiang Huo, Huan Song, Shaohua Teng, Hua Wang 0002, Yanchun Zhang |
WISE (2) | 5 |
| 2018 | Geographical Proximity Boosted Recommendation Algorithms for Real Estate
Yonghong Yu, Can Wang 0004, Li Zhang 0013, Rong Gao 0001, Hua Wang 0002 |
WISE (2) | 5 |
| 2018 | Gradient Correlation: Are Ensemble Classifiers More Robust Against Evasion Attacks in Practical Settings?
Fuyong Zhang, Yi Wang 0017, Hua Wang 0002 |
WISE (1) | 3 |
| 2018 | Mining Event-Oriented Topics in Microblog Stream with Unsupervised Multi-View Hierarchical EmbeddingabstractThis article presents an unsupervised multi-view hierarchical embedding (UMHE) framework to sufficiently reveal the intrinsic topical knowledge in social events. Event-oriented topics are highly related to such events as it can provide explicit descriptions of what have happened in social community. In many real-world cases, however, it is difficult to include all attributes of microblogs, more often, textual aspects only are available. Traditional topic modelling methods have failed to generate event-oriented topics with the textual aspects, since the inherent relations between topics are often overlooked in these methods. Meanwhile, the metrics in original word vocabulary space might not effectively capture semantic distances. Our UMHE framework overcomes the severe information deficiency and poor feature representation. The UMHE first develops a multi-view Bayesian rose tree to preliminarily generate prior knowledge for latent topics and their relations. With such prior knowledge, we design an unsupervised translation-based hierarchical embedding method to make a better representation of these latent topics. By applying self-adaptive spectral clustering on the embedding space and the original space concomitantly, we eventually extract event-oriented topics in word distributions to express social events. Our framework is purely data-driven and unsupervised, without any external knowledge. Experimental results on TREC Tweets2011 dataset and Sina Weibo dataset demonstrate that the UMHE framework can construct hierarchical structure with high fitness, but also yield topic embeddings with salient semantics; therefore, it can derive event-oriented topics with meaningful descriptions. Min Peng 0002, Hua Wang 0002, Xuhui Li 0001, Yanchun Zhang, Xiuzhen Zhang 0001, Gang Tian |
ACM Trans. Knowl. Discov. Data | 3 |
| 2017 | A Topic Model Based on Poisson DecompositionabstractDetermining appropriate statistical distributions for modeling text corpora is important for accurate estimation of numerical characteristics. Based on the validity of the test on a claim that the data conforms to Poisson distribution we propose Poisson decomposition model (PDM), a statistical model for modeling count data of text corpora, which can straightly capture each document's multidimensional numerical characteristics on topics. In PDM, each topic is represented as a parameter vector with multidimensional Poisson distribution, which can be easily normalized to multinomial term probabilities and each document is represented as measurements on topics and thereby reduced to a measurement vector on topics. We use gradient descent methods and sampling algorithm for parameter estimation. We carry out extensive experiments on the topics produced by our models. The results demonstrate our approach can extract more coherent topics and is competitive in document clustering by using the PDM-based features, compared to PLSI and LDA. Haixin Jiang, Rui Zhou 0001, Limeng Zhang, Hua Wang 0002, Yanchun Zhang |
CIKM | 4 |
| 2017 | A Study on Securing Software Defined Networks
Raihan Ur Rasool, Hua Wang 0002, Wajid Rafique, Jianming Yong, Jinli Cao |
WISE (2) | 2 |
| 2017 | Cryptographic Access Control in Electronic Health Record Systems: A Security Implication
Pasupathy Vimalachandran, Hua Wang 0002, Yanchun Zhang, Guangping Zhuo, Hongbo Kuang |
WISE (2) | 2 |
| 2017 | Parallelization of Massive Textstream Compression Based on Compressed SensingabstractCompressing textstreams generated by social networks can both reduce storage consumption and improve efficiency such as fast searching. However, the compression process is a challenge due to the large scale of textstreams. In this article, we propose a textstream compression framework based on compressed sensing theory and design a series of matching parallel procedures. The new approach uses a linear projection technique in the textstream compression process, achieving fast compression speed and low compression ratio. Two processes are executed by designing elaborated parallel procedures for efficient compressing and decompressing of large-scale textstreams. The decompression process is implemented for approximate solutions of underdetermined linear systems. Experimental results show that the new method can efficiently achieve the compression and decompression tasks on a large amount of text generated by social networks. Min Peng 0002, Wang Gao 0002, Hua Wang 0002, Yanchun Zhang, Qianqian Xie, Gang Hu 0003, Gang Tian |
ACM Trans. Inf. Syst. | 3 |
| 2016 | Improving Distant Supervision of Relation Extraction with Unsupervised Methods
Min Peng 0002, Jimin Huang, Zhaoyu Sun, Shizhong Wang, Hua Wang 0002, Guangping Zhuo, Gang Tian |
WISE (1) | 5 |
| 2016 | Mining Actionable Knowledge Using Reordering Based Diversified Actionable Decision Trees
Sudha Subramani, Hua Wang 0002, Sathiyabhama Balasubramaniam, Rui Zhou 0001, Jiangang Ma, Yanchun Zhang, Frank Whittaker, Yueai Zhao, Sarathkumar Rangarajan |
WISE (1) | 2 |
| 2015 | Central Topic Model for Event-oriented Topics Mining in Microblog StreamabstractTo date, data generates and arrives in the form of stream to propagate discussions of public events in microblog services. Discovering event-oriented topics from the stream will lead to a better understanding of the change of public concern. However, as the massive scale of the data stream, traditional static topic models, such as LDA, are no longer fit for topic detection and tracking tasks. In this paper, we propose a central topic model (CenTM), where a Multi-view Clustering algorithm with Two-phase Random Walk (MC-TRW) is devised to aggregate the LDA's latent topics into central topics. Furthermore, we leverage the aggregation of central topics alternately with MC-TRW and sequential topic inference to improve the scalability in the stream fashion, so as to derive the dynamic central topic model (DCenTM). Specifically, our model is able to uncover the intrinsic characteristics of the central topics and predict the trend of their intensity along a life cycle. Experimental results demonstrate that the proposed central topic model is event-oriented and of high generalization, it therefore can dispose the topic trend prediction effectively and precisely in massive data stream. Min Peng 0002, Xuhui Li 0001, Hua Wang 0002, Yanchun Zhang |
CIKM | 5 |
| 2015 | Multi-Window Based Ensemble Learning for Classification of Imbalanced Streaming Data
Ye Wang 0015, Hua Wang 0002, Bin Zhou 0004, Yanchun Zhang |
WISE (2) | 3 |
| 2015 | Special issue on Security, Privacy and Trust in network-based Big Data
Hua Wang 0002, Xiaohong Jiang 0001, Georgios Kambourakis |
Inf. Sci. | 1 |
| 2014 | Indexing Linked Data in a Wireless Broadcast System with 3D Hilbert Space-Filling CurvesabstractSemantic technologies aim to facilitate machine-to-machine communication and are attracting more and more interest from both academia and industry, especially in the emerging Internet of Things (IoT). In this paper, we consider large-scale information sharing scenarios among mobile objects in IoT by leveraging semantic techniques. We propose to broadcast Linked Data on-air using RDF format to allow simultaneous access to the information and to achieve better scalability. We introduce a novel air indexing method to reduce the information access latency and energy consumption. To build air indexes, we firstly map RDF triples in the Linked Data into points in a 3D space and build B+-trees based on 3D Hilbert curve mappings for all of the 3D points. We then convert these trees into linear sequences so that they can be broadcast over a wireless channel. A novel search algorithm is also designed to efficiently evaluate queries against the air indexes. Experiments show that our indexing method outperforms the air indexing method based on traditional 3D R-trees. Yongrui Qin, Quan Z. Sheng, Nick Falkner, Wei Zhang 0098, Hua Wang 0002 |
CIKM | 5 |
| 2013 | Minimising K-Dominating Set in Arbitrary Network Graphs
Guangyuan Wang, Hua Wang 0002, Xiaohui Tao 0001, Ji Zhang 0001 |
ADMA (2) | 2 |
| 2013 | Effectively Delivering XML Information in Periodic Broadcast Environments
Yongrui Qin, Quan Z. Sheng, Muntazir Mehdi, Hua Wang 0002, Dong Xie 0002 |
DEXA (1) | 4 |
| 2013 | SODIT: An innovative system for outlier detection using multiple localized thresholding and interactive feedbackabstractOutlier detection is an important long-standing research problem in data mining and has enjoyed applications in a wide range of applications in business, engineering, biology and security, etc. However, the traditional outlier detection methods inevitably need to use different parameters for detection such as those used to specify the distance or density cutoff for distinguish outliers from normal data points. Using the trial and error approach, the traditional outlier detection methods are rather tedious in parameter tuning. In this demo proposal, we introduce an innovative outlier detection system, called SODIT, that uses localized thresholding to assist the value specification of the thresholds that reflect closely the local data distribution. In addition, easy-to-use user feedback are employed to further facilitate the determination of optimal parameter values. SODIT is able to make outlier detection much easier to operate and produce more accurate, intuitive and informative results than before. Ji Zhang 0001, Hua Wang 0002, Xiaohui Tao 0001, Lili Sun |
ICDE | 2 |
| 2012 | Data Privacy against Composition Attack
Muzammil M. Baig, Jiuyong Li, Jixue Liu, Xiaofeng Ding 0001, Hua Wang 0002 |
DASFAA (1) | 5 |
| 2012 | Unsupervised Multi-label Text Classification Using a World Knowledge Ontology
Xiaohui Tao 0001, Yuefeng Li 0001, Raymond Y. K. Lau, Hua Wang 0002 |
PAKDD (1) | 4 |
| 2012 | Multi-level delegations with trust management in access control systems
Xiaoxun Sun, Hua Wang 0002, Yanchun Zhang |
J. Intell. Inf. Syst. | 3 |
| 2011 | Cloning for privacy protection in multiple independent data publicationsabstractData anonymization has become a major technique in privacy preserving data publishing. Many methods have been proposed to anonymize one dataset and a series of datasets of a data owner. However, no method has been proposed for the anonymization of data of multiple independent data publications. A data owner publishes a dataset, which contains overlapping population with other datasets published by other independent data owners. In this paper we analyze the privacy risk in the such scenario and vulnerability of partitioned based anonymization methods. We show that no partitioned based anonymization methods can protect privacy in arbitrary data distributions, and identify a case that the privacy can be protected in the scenario. We propose a new generalization principle ε-cloning to protect privacy for multiple independent data publications. We also develop an effective algorithm to achieve the ε-cloning. We experimentally show that the proposed algorithm anonymizes data to satisfy the privacy requirement and preserves good data utility. Muzammil M. Baig, Jiuyong Li, Jixue Liu, Hua Wang 0002 |
CIKM | 4 |
| 2011 | Publishing anonymous survey rating data
Xiaoxun Sun, Hua Wang 0002, Jiuyong Li, Jian Pei 0001 |
Data Min. Knowl. Discov. | 2 |
| 2010 | DISTRO: A System for Detecting Global Outliers from Distributed Data Streams with Privacy Protection
Ji Zhang 0001, Stijn Dekeyser, Hua Wang 0002, Yanfeng Shu |
DASFAA (2) | 3 |
| 2010 | A Pairwise-Systematic Microaggregation for Statistical Disclosure ControlabstractMicrodata protection in statistical databases has recently become a major societal concern and has been intensively studied in recent years. Statistical Disclosure Control (SDC) is often applied to statistical databases before they are released for public use. Micro aggregation for SDC is a family of methods to protect micro data from individual identification. SDC seeks to protect micro data in such a way that can be published and mined without providing any private information that can be linked to specific individuals. Micro aggregation works by partitioning the micro data into groups of at least k records and then replacing the records in each group with the centroid of the group. An optimal micro aggregation method must minimize the information loss resulting from this replacement process. The challenge is how to minimize the information loss during the micro aggregation process. This paper presents a pair wise systematic (P-S) micro aggregation method to minimize the information loss. The proposed technique simultaneously forms two distant groups at a time with the corresponding similar records together in a systematic way and then anonymized with the centroid of each group individually. The structure of P-S problem is defined and investigated and an algorithm of the proposed problem is developed. The performance of the P-S algorithm is compared against the most recent micro aggregation methods. Experimental results show that P-S algorithm incurs less than half information loss than the latest micro aggregation methods for all of the test situations. Md. Enamul Kabir, Hua Wang 0002, Yanchun Zhang |
ICDM | 2 |
| 2010 | Satisfying Privacy Requirements: One Step before Anonymization
Xiaoxun Sun, Hua Wang 0002, Jiuyong Li |
PAKDD (1) | 2 |
| 2009 | Injecting purpose and trust into data anonymisationabstractMost existing works of data anonymisation target at the optimization of the anonymisation metrics to balance the data utility and privacy, whereas they ignore the effects of a requester's trust level and application purposes during the data anonymisation. Our aim of this paper is to propose a much finer level anonymisation scheme with regard to the data requester's trust value and specific application purpose. We prioritize the attributes for anonymisation based on how important and critical they are related to the specified application purposes and propose a trust evaluation strategy to quantify the data requester's reliability, and further build the projection between the trust value and the degree of data anonymiztion, which intends to determine to what extent the data should be anonymized. The decomposition algorithm is developed to find the desired anonymous solution, which guarantees the uniqueness and correctness. Xiaoxun Sun, Hua Wang 0002, Jiuyong Li |
CIKM | 2 |
| 2009 | Optimal Privacy-Aware Path in Hippocratic Databases
Xiaoxun Sun, Hua Wang 0002, Yanchun Zhang |
DASFAA | 3 |
| 2009 | Effective Collaboration with Information Sharing in Virtual UniversitiesabstractA global education system, as a key area in future IT, has fostered developers to provide various learning systems with low cost. While a variety of e-learning advantages has been recognized for a long time and many advances in e-learning systems have been implemented, the needs for effective information sharing in a secure manner have to date been largely ignored, especially for virtual university collaborative environments. Information sharing of virtual universities usually occurs in broad, highly dynamic network-based environments, and formally accessing the resources in a secure manner poses a difficult and vital challenge. This paper aims to build a new rule-based framework to identify and address issues of sharing in virtual university environments through role-based access control (RBAC) management. The framework includes a role-based group delegation granting model, group delegation revocation model, authorization granting, and authorization revocation. We analyze various revocations and the impact of revocations on role hierarchies. The implementation with XML-based tools demonstrates the feasibility of the framework and authorization methods. Finally, the current proposal is compared with other related work. Hua Wang 0002, Yanchun Zhang, Jinli Cao |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2008 | On the Complexity of Restricted k-anonymity Problem
Xiaoxun Sun, Hua Wang 0002, Jiuyong Li |
APWeb | 2 |
| 2006 | Role-Based Delegation with Negative Authorization
Hua Wang 0002, Jinli Cao |
APWeb | 1 |
| 2005 | Towards Secure XML Document with Usage Control
Jinli Cao, Lili Sun, Hua Wang 0002 |
APWeb | 3 |
| 2005 | A Flexible Payment Scheme and Its Role-Based Access ControlabstractThis work proposes a practical payment protocol with scalable anonymity for Internet purchases, and analyzes its role-based access control (RBAC). The protocol uses electronic cash for payment transactions. It is an offline payment scheme that can prevent a consumer from spending a coin more than once. Consumers can improve anonymity if they are worried about disclosure of their identities to banks. An agent provides high anonymity through the issue of a certification. The agent certifies reencrypted data after verifying the validity of the content from consumers, but with no private information of the consumers required. With this new method, each consumer can get the required anonymity level, depending on the available time, computation, and cost. We use RBAC to manage the new payment scheme and improve its integrity. With RBAC, each user may be assigned one or more roles, and each role can be assigned one or more privileges that are permitted to users in that role. To reduce conflicts of different roles and decrease complexities of administration, duty separation constraints, role hierarchies, and scenarios of end-users are analyzed. Hua Wang 0002, Jinli Cao, Yanchun Zhang |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2004 | Specifying Role-Based Access Constraints with Object Constraint Language
Hua Wang 0002, Yanchun Zhang, Jinli Cao |
APWeb | 1 |
| 2002 | Formal Authorization Allocation Approaches for Role-Based Access Control Based on Relational Algebra OperationsabstractWe develop formal authorization allocation algorithms for role-based access control (RBAC). The formal approaches are based on relational structure, and relational algebra and operations. The process of user-role assignments is an important issue in RBAC because it may modify the authorization level or imply high-level confidential information to be derived while users change positions and request different roles. There are two types of problems which may arise in user-role assignment. One is related to the authorization granting process. When a role is granted to a user this role may conflict with other roles of the user or together with this role; the user may have or derive a high level of authority. Another is related to authorization revocation. When a role is revoked from a user, the user may still have the role from other roles. To solve these problems, this paper presents an authorization granting algorithm, and weak revocation and strong revocation algorithms that are based on relational algebra. The algorithms can be used to check conflicts and therefore to help allocate roles without compromising the security in RBAC. We describe how to use the new algorithms with an anonymity scalable payment scheme. Finally, comparisons with other related work are discussed. Hua Wang 0002, Jinli Cao, Yanchun Zhang |
WISE | 1 |
| 2001 | A Consumer Scalable Anonymity Payment Scheme with Role-Based Access ControlabstractThis paper proposes a secure, scalable anonymity and practical payment protocol for Internet purchases, and uses role based access control (RBAC) to manage the new payment scheme. The protocol uses electronic cash for payment transactions. In this new protocol, from the viewpoint of banks, consumers can improve anonymity if they are worried about disclosure of their identities. An agent provides a higher anonymous certificate and improves the security of the consumers. The agent will certify re-encrypted data after verifying the validity of the content from consumers, but with no private information of the consumers required. With this new method, each consumer can get the required anonymity level, depending on the available time, computation and cost. We also analyse how to prevent a consumer from spending a coin more than once. Furthermore, we use RBAC to manage the new payment scheme. Each user may be assigned one or more roles, and each role can be assigned one or more privileges that are permitted to users in that role. Security administration with RBAC consists of determining operations that must be executed by persons in particular jobs, and assigning employees to proper roles. RBAC can improve system security and reduce conflicts of different roles. The complexities with RBAC can be decreased by mutually exclusive roles and role hierarchies. Hua Wang 0002, Jinli Cao, Yanchun Zhang |
WISE (1) | 1 |