Yan Zhu 0007

dblp:143/9985 · DBLP profile ↗
← Back
25ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0002-5804-8926ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Security and privacy · 2Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Fair Link Prediction in Social Network with Intersectional Sensitive Attributes
Yifeng Peng, Yan Zhu 0007, Yiqiang Peng
DEXA (1)3
2025 Identifying Multimodal Sarcasm Based on Incongruous Knowledge Capturing and Contrastive Learning
Yan Zhu 0007, Yiqiang Peng
DEXA (1)1
2025 Text Enhancement-Based Multimodal Fusion for Video Sentiment Analysis
Junhao Luo, Yan Zhu 0007, Yiqiang Peng
PAKDD (7)2
2024 GADEC: Discovering Abnormal Citation Groups Based on Enhanced Local Community Expansion and DQN
Yan Zhu 0007, Xinrui Lin, Yiqiang Peng
DEXA (2)1
2023 TNSEIR: A SEIR pattern-based embedding approach for temporal network
Lei Wang 0289, Yan Zhu 0007, Qiang Peng
Appl. Intell.2
2022 Battering Review Spam Through Ensemble Learning in Imbalanced Datasets
abstract
Abstract Nowadays, people’s buying or availing services decisions are subject to online available reviews/opinions. The authenticity of these reviews/opinions is dubious, as there exist many fake reviews posted to attain monetary benefits by promoting their own or demoting the competitor’s products or services known as review spam. Although the number of spam is relatively less than that of normal reviews in real-life, this class imbalance is a critical concern in review spam detection. The performance degrades when the classifier skew towards the majority class. Moreover, efficient feature selection is essentially needed for this issue. The purpose of this study is to develop a framework based on different effective feature selection along with data balancing techniques. Validation results show that our proposed framework commendably copes up with the review spam issue and a higher precision on the real-life dataset. Further, we tested the sensitivity of our proposed framework using both parametric and non-parametric tests and found it significant.
Faisal Khurshid 0001, Yan Zhu 0007, Jie Hu 0007, Muqeet Ahmad, Mushtaq Ahmad
Comput. J.2
2021 Using deep belief network to demote web spam
Xu Zhuang, Yan Zhu 0007, Qiang Peng, Faisal Khurshid 0001
Future Gener. Comput. Syst.2
2021 A semi-supervised deep learning image caption model based on Pseudo Label and N-gram
Chunping Li, Youfang Han, Yan Zhu 0007
Int. J. Approx. Reason.4
2021 Differentially private data publishing for arbitrarily partitioned data
Rong Wang 0006, Benjamin C. M. Fung, Yan Zhu 0007, Qiang Peng
Inf. Sci.3
2020 Short Text Topic Modeling with Topic Distribution Quantization and Negative Sampling Decoder
abstract
Topic models have been prevailing for many years on discovering latent semantics while modeling long documents.However, for short texts they generally suffer from data sparsity because of extremely limited word cooccurrences; thus tend to yield repetitive or trivial topics with low quality.In this paper, to address this issue, we propose a novel neural topic model in the framework of autoencoding with a new topic distribution quantization approach generating peakier distributions that are more appropriate for modeling short texts.Besides the encoding, to tackle this issue in terms of decoding, we further propose a novel negative sampling decoder learning from negative samples to avoid yielding repetitive topics.We observe that our model can highly improve short text topic modeling performance.Through extensive experiments on real-world datasets, we demonstrate our model can outperform both strong traditional and neural baselines under extreme data sparsity scenes, producing high-quality topics.
Xiaobao Wu, Chunping Li, Yan Zhu 0007, Yishu Miao
EMNLP (1)3
2020 Learning Multilingual Topics with Neural Variational Inference
Xiaobao Wu, Chunping Li, Yan Zhu 0007, Yishu Miao
NLPCC (1)3
2020 Privacy-preserving high-dimensional data publishing for classification
Rong Wang 0006, Yan Zhu 0007, Chin-Chen Chang 0001, Qiang Peng
Comput. Secur.2
2020 Heterogeneous data release for cluster analysis with differential privacy
Rong Wang 0006, Benjamin C. M. Fung, Yan Zhu 0007
Knowl. Based Syst.3
2019 An Attribute-Based Fine-Grained Access Control Mechanism for HBase
Liangqiang Huang, Yan Zhu 0007, Xin Wang 0064, Faisal Khurshid 0001
DEXA (1)2
2019 Triplet-CSSVM: Integrating Triplet-Sampling CNN and Cost-Sensitive Classification for Imbalanced Image Detection
Jiefan Tan, Yan Zhu 0007
DEXA (2)2
2019 Querying Knowledge Graphs with Natural Languages
Xin Wang 0064, Lan Yang 0004, Yan Zhu 0007, Huayi Zhan
DEXA (2)3
2018 An Authentication Method Based on the Turtle Shell Algorithm for Privacy-Preserving Data Mining
abstract
Outsourcing data mining tasks is beneficial for data owners who either lack expertise in data mining or sufficient computing resources. However, directly releasing the original data would leak private information. Research on Privacy-Preserving Data Mining (PPDM) is dedicated to addressing this issue, the aim of this research is to reduce the risk of privacy violations and preserve the knowledge in the original data. However, most existing methods in the literature ignore the case in which service providers want to verify the integrity and authenticity of their clients’ data to avoid data tampering before performing data mining tasks. In this paper, a new method is proposed to extend the turtle shell algorithm of data hiding to protect the privacy of the original data and to acquire authentication functions simultaneously. The act of data perturbation is performed by replacing data values with their closest neighbors according to a reference matrix. Further, a message authentication code is hidden in the perturbed data to verify the integrity and authenticity of the perturbed data. The experimental results showed that the proposed method achieved the purpose of data perturbation and outperformed similar methods in satisfying the PPDM requirement.
Rong Wang 0006, Yan Zhu 0007, Tung-Shou Chen, Chin-Chen Chang 0001
Comput. J.2
2018 Privacy-Preserving Algorithms for Multiple Sensitive Attributes Satisfying t-Closeness
Rong Wang 0006, Yan Zhu 0007, Tung-Shou Chen, Chin-Chen Chang 0001
J. Comput. Sci. Technol.2
2017 Cleaning Out Web Spam by Entropy-Based Cascade Outlier Detection
Sha Wei, Yan Zhu 0007
DEXA (2)2
2017 Feature bundling in decision tree algorithm
abstract
In empirical data modelling, a model of system is built up from a set of cases that the system has observed. Eventually, the performance of the inducted model is dominated by the quality and quantity of observations. Feature transformation methods are widely used to improve quality of knowledge ext racted from observations to build up more accurate and robust model. In the paper, a new feature transformation method named dynamical feature bundling for decision tree algorithm is proposed. Dynamical feature bundling groups a set of features in the tree induction phase and it enables decision tree algorithms to 1) make use of features in one bundle together to make collective judgments in splitting phase; 2) learn more reliable and stable knowledge from feature bundles created based on domain knowledge of experts; 3) embed feature transformation step into tree induction phase, and therefore the extra pre-process step which are necessary for static feature transformation methods is inessential. Our experiments show 2%-9% improvements of AUC value on a very imbalanced dataset. Slight improvements are also obtained on a more balanced data set.
Xu Zhuang, Yan Zhu 0007, Chin-Chen Chang 0001, Qiang Peng
Intell. Data Anal.2
2017 A unified score propagation model for web spam demotion algorithm
Xu Zhuang, Yan Zhu 0007, Chin-Chen Chang 0001, Qiang Peng, Faisal Khurshid 0001
Inf. Retr. J.2
2017 Detection of malicious web pages based on hybrid analysis
Rong Wang 0006, Yan Zhu 0007, Jiefan Tan
J. Inf. Secur. Appl.2
2014 Ascertaining Spam Web Pages Based on Ant Colony Optimization Algorithm
Shouhong Tang, Yan Zhu 0007
DEXA (2)2
2014 Formalizing and validating the web quality model for web source quality evaluation
Yan Zhu 0007
Expert Syst. Appl.2
2002 Evaluating and Selecting Web Sources as External Information Resources of a Data Warehouse
abstract
A company's local data is often insufficient for analyzing market trends and making reasonable business plans. Decision making must also be based on information from suppliers, partners and competitors. Systematically integrating suitable external data from the Web into a data warehouse is a meaningful solution and will benefit the enterprise. However, the autonomy and dynamics of the Web make the task of selecting relevant and qualified external data from the Web challenging. We develop a set of criteria for evaluating and selecting Web resources as external data sources of a data warehouse and discuss how to screen Web data sources using multi-criteria decision making (MCDM) methods. The final decision with respect to selecting Web sources is sensitive to critical factors, i.e., the criterion weight and performance score of alternatives in terms of each criterion. We analyzed the sensitivity of the final rank of alternatives in terms of critical factors in order to gain an insight into the stability of our final decision. The comparison of several MCDM approaches for Web source screening is also presented.
Yan Zhu 0007, Alejandro P. Buchmann
WISE1