VLDB 2026 Research / reviewers in the wild / expert
Zhaoyun Ding
dblp:43/8006
· DBLP profile ↗
30ranked-venue papers
7as first author
19since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 11 · 1 first-author · 8 since 2021Databases, data management, data science and information retrieval · 10 · 4 first-author · 5 since 2021Security and privacy · 5 · 1 first-author · 5 since 2021Computer networks · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CyberNER-LLM: Cyber Threat Intelligence Named Entity Recognition With Large Language Model
Xinzheng Liu, Wangqun Lin, Zhaoyun Ding |
ICICS (2) | 3 |
| 2025 | The Local Minimum Strategy: Accelerating Relocation in Cuckoo FilterabstractEfficient set representation and membership testing are important in high-speed network measurement. Fast insertions, space efficiency, fast query, and low false positive rate are the core requirements of traffic measurement, but existing solutions, such as hash tables and Bloom filters(BFs), cannot satisfy these requirements simultaneously. The state-of-the-art cuckoo filter and its variants(CFs) rely on the random eviction relocation strategy to resolve hash collisions, improving space utilization while reducing false positives and maintaining high query efficiency. However, CFs suffer a critical challenge in practical applications: insertion could trigger multiple evictions when all candidate buckets for an element are saturated, leading to insertion performance degradation, especially when space utilization exceeds 0.8. To solve the above problem, we propose a novel relocation strategy based on a random graph model, called the local minimum strategy. Our core idea is to use the implicit meaning of the number of evictions in each bucket as an indication to minimize the relocation of elements. We theoretically and experimentally prove that the eviction threshold for each bucket is$O(\log m)$, where$m$is the number of buckets. The threshold establishes the bounds for the probability of a successful insertion. The experimental results show that, the local minimum strategy significantly reduces the number of relocations by 67 %, as well as increasing the insertion throughput by more than 10 %. Niuniu Zhang, Lailong Luo, Qianzhen Zhang, Shangsen Li, Zhaoyun Ding, Xiang Zhao 0002, Deke Guo |
IWQoS | 5 |
| 2025 | Boosting Discriminability for Robust Multimodal Entity Linking with Visual Modality MissingabstractMultimodal Entity Linking (MEL) aims to retrieve ambiguous mentions within multimodal contexts to the referent entities in a multimodal knowledge base, typically based on the assumption of modality completeness. However, when deployed in open-world applications, MEL systems may encounter uncertainly missing of visual modalities from user-proposed mentions. In this paper, we propose a novel setting dubbed MEL-MM to simulate the practical challenge, and reveal that the semantic discriminability is a crucial factor to enhance the anti-missingness resilience. To this end, we introduce an innovative yet efficient approach termed Cross-View Introspective Ranking Distillation (CVIRD), which seeks to sufficiently align the linking similarities between teacher and student models trained from modality-complete and incomplete data. To be specific, as the first concept in CVIRD, Missing-Aware Ranking Distillation (MARD) focuses on modeling the discriminability by formulating the similarity rankings between mention and entities in a missing-sensitive and differentiable manner. Moreover, the second concept of Cross-View Distillation with Introspection (CVDI) aims to improve discriminability extraction in MARD through multi-level distillation, considering both cross-view retrieval and self-consistency. Experiments verify the effectiveness and model-agnostic ability of our method, which achieves superior performance in contrast to competitive missingness-resilient strategies. Mingrui Lao, Yanming Guo, Xueyi Zhang 0001, Siqi Cai 0002, Zhaoyun Ding, Haizhou Li 0001 |
SIGIR | 6 |
| 2025 | LLM-Driven APT Analysis and Detection Based on Provenance Graph by Threat PatternabstractWe propose a threat detection framework that enhances provenance-based APT analysis by combining event clustering with LLM-driven threat pattern extraction refinement. Our method is designed to work in conjunction with existing systems like Kairos, which able to detect anomaly time windows. We first group fine-grained events into semantic clusters to preserve behavioral context, then apply structured threat patterns to identify suspicious sequences. To improve precision, we introduce a two-stage refinement: relaxing matching constraints for high recall, followed by LLM-driven filtering to assess semantic plausibility in framework of MITRE based on "Part Chain of Pattern" assumption. Evaluated on CADETS E3 and THEIA E3 using defender-observable ground truth, our approach outperforms Kairos in precision and F1-score. The results show that integrating structured pattern matching with contextual language models can effectively enhance existing detection pipelines, offering a practical path toward more accurate and interpretable threat hunting. Yibin Fu, Zhaoyun Ding, Deqi Cao |
TrustCom | 2 |
| 2025 | CyberSOIE-LLM: Cybersecurity Semi-Open Information Extraction with Large Language ModelsabstractAdvanced Persistent Threats (APTs) are increasing in frequency and sophistication, rendering traditional defenses inadequate against the pronounced asymmetry inherent in contemporary cybersecurity. Cyber Threat Intelligence (CTI) is widely viewed as pivotal for transitioning from reactive to proactive defence; however, the exponential growth of unstructured CTI far exceeds the speed of manual analysis. We introduce CyberSOIE-LLM, the first cybersecurity-oriented Semi-Open Information Extraction framework that employs Large Language Models as domain experts to extract security knowledge from CTI efficiently, automatically and accurately. The framework unites three synergistic modules. Generator blends external knowledge enhancement with Chain-of-Thought (CoT) prompting to produce initial triples. Refiner enforces precision through multi-level rule- and semantics-based validation coupled with self-correction. Unifier externally aligns and internally normalises relation labels, yielding high-quality, low-redundancy cybersecurity triples. Extensive experiments on four Chinese–English CTI benchmarks show that CyberSOIE-LLM significantly surpasses state-of-the-art baselines, providing a scalable and trustworthy knowledge substrate for dynamic, proactive defence. Xinzheng Liu, Zhaoyun Ding |
TrustCom | 2 |
| 2025 | Psycholinguistic knowledge-guided graph network for personality detection of silent users
Houjie Qiu, Xingkong Ma, Bo Liu 0014, Yiqing Cai, Zhaoyun Ding |
Inf. Process. Manag. | 6 |
| 2024 | The Optimal Sampling Strategy for Few-shot Named Entity Recognition Based with Prototypical NetworkabstractFew-shot Named Entity Recognition (NER) is the task of identifying and classifying entities under conditions of low resources. The preceding approach relies on meta-learning, wherein the classification model is trained through N-way K-shot. However, It neglects essential information: the model’s input is organized at the sentence level. To tackle the aforementioned issues, we introduce a sampling strategy — the Loose Nway K˜shot sampling algorithm. This approach combines the number of entities and sentence level which is the smallest unit of NER models’ input. Experimental results demonstrate that our strategy currently stands as the most effective method for Few-shot NER. Furthermore, our investigation reveals that various classes of entities within the same sentence can impede the prototype representation of each class. Consequently, we introduce a training method that involves utilizing sentences with single entity class for pre-training purposes. The experimental outcomes substantiate that this training methodology maximizes the utilization of labeled data, enabling the pre-training model to swiftly adapt to new domains. This approach significantly enhances the performance of NER. Junqi Chen 0006, Zhaoyun Ding, Guang Jin, Guoli Yang |
IJCNN | 2 |
| 2023 | TGR: Neural-symbolic ontological reasoner for domain-specific knowledge graphs
Xixi Zhu, Bin Liu 0026, Zhaoyun Ding |
Appl. Intell. | 4 |
| 2023 | Multiclass imbalanced and concept drift network traffic classification framework based on online active learningabstractThe complex problems of multiclass imbalance, virtual or real concept drift, concept evolution, high-speed traffic streams and limited label cost budgets pose severe challenges in network traffic classification tasks. In this paper, we propose a multiclass imbalanced and concept drift network traffic classification framework based on online active learning (MicFoal), which includes a configurable supervised learner for the initialization of a network traffic classification model, an active learning method with a hybrid label request strategy, a label sliding window group, a sample training weight formula and an adaptive adjustment mechanism for the label cost budget based on a periodic performance evaluation. In addition, a novel uncertain label request strategy based on a variable least confidence threshold vector is designed to address the problems of a variable multiclass imbalance ratio or even the number of classes changing over time. Experiments performed based on eight well-known real-world network traffic datasets demonstrate that MicFoal is more effective and efficient than several state-of-the-art learning algorithms. Weike Liu, Cheng Zhu 0002, Zhaoyun Ding, Hang Zhang 0008, Qingbao Liu |
Eng. Appl. Artif. Intell. | 3 |
| 2022 | OTE: An Optimized Chinese Short Text Matching Algorithm Based on External Knowledge
Zhaoyun Ding |
KSEM (1) | 2 |
| 2022 | Implementing Large-Scale ABox Materialization Using Subgraph Reasoning
Xixi Zhu, Zhaoyun Ding |
KSEM (1) | 3 |
| 2022 | A General Personality Analysis Model Based on Social Posts and Links
Xingkong Ma, Houjie Qiu, Shujia Yao, Jingsong Zhang, Zhaoyun Ding, Bo Liu 0014 |
PRICAI (1) | 6 |
| 2022 | Pre-training Fine-tuning data Enhancement method based on active learningabstractWith the development of Internet technology, the number of Internet users increases rapidly, and the amount of data generated on the Internet is very large every day. At the same time, with the development of storage technology and query technology, it is very easy to collect massive data, but the information value contained in these data is uneven, and most of them are unmarked. However, traditional supervised learning has a great demand for labeled samples. Faced with a large number of unlabeled samples, there is a problem of the lack of effective automatic labeling methods, and manual labeling costs are high. If the strategy of simple random sampling is used for annotation, it may lead to the selection of noisy information and waste of resources, and low-quality training data could also have an influence on the prediction accuracy of the model. Meanwhile, the training effect of traditional deep learning methods is very limited for small sample labeled training sets.This paper takes the text emotion analysis task in natural language processing as the background, selects IMDB film review data as the training set and test set, starts with the design of active learning algorithm based on clustering analysis, combined with the appropriate pre-training fine-tuning model, constructs a data enhancement method based on active learning. In the experiment, it is found that when the labeled training set is reduced by 90%, the prediction accuracy of the pre-training model is reduced by no more than 2%, which verifies the effectiveness of the data enhancement method combining active learning with the pre-training model. Deqi Cao, Zhaoyun Ding |
TrustCom | 2 |
| 2022 | Research on NER Based on Register Migration and Multi-task Learning
Zhaoyun Ding, ShuoShuo Niu |
WASA (3) | 2 |
| 2022 | CNsum: Automatic Summarization for Chinese News Text
Songping Huang, Zhaoyun Ding, Aixin Nian |
WASA (2) | 4 |
| 2021 | Drug-Target Interaction Prediction Based on Gaussian Interaction Profile and Information Entropy
Lina Liu 0011, Shuang Yao, Zhaoyun Ding, Maozu Guo 0001, Donghua Yu, Keli Hu |
ISBRA | 3 |
| 2021 | A Semantic Textual Similarity Calculation Model Based on Pre-training Model
Zhaoyun Ding, Kai Liu 0037, Bin Liu 0026 |
KSEM | 1 |
| 2021 | Detection of Anomaly User Behaviors Based on Deep Neural NetworksabstractIn order to predict the user's anomaly operation behaviors, we perform deep learning modeling on the user's UNIX command line operation sequence. The use of deep neural networks in anomaly detection is to build a model through the training set, enabling the model to predict the user's next action or command based on the given first$n$actions or commands. The network trains the command set commonly used by users. After a period, the network can match the real commands according to the existing user characteristic files in the network, and any mismatched events or commands are regarded as anomalies. The detection of anomaly user behaviors is an imbalanced classification problem. To address imbalanced classification problem, we propose an imbalanced self-paced sampling method to improve the efficiency of anomaly user behavior detection. The results show that the DNNs model can usually find anomaly user behaviors that are not easily detectable by other models in anomaly detection. Zhaoyun Ding, Lina Liu 0011, Donghua Yu, Songping Huang, Hang Zhang 0008, Kai Liu 0037 |
TrustCom | 1 |
| 2021 | A comprehensive active learning method for multiclass imbalanced data streams with concept driftabstractA challenge to many real-world applications is multiclass imbalance with concept drift. In this paper, we propose a comprehensive active learning method for multiclass imbalanced streaming data with concept drift (CALMID). First, we design a comprehensive online active learning framework that includes an ensemble classifier, a drift detector, a label sliding window, sample sliding windows and an initialization training sample sequence. Next, a variable threshold uncertainty strategy based on an asymmetric margin threshold matrix is designed to comprehensively address the problem that a given class can simultaneously be a majority to a given subset of classes while also being a minority to others. Last but not least, we design a novel sample weight formula that comprehensively considers the class imbalance ratio of the sample’s category and the prediction difficulty. On 10 multiclass synthetic streams with different imbalance ratios and concept drifts, and on 5 real-world imbalanced streams with 7 to 55 classes and unknown drifts, the experimental results demonstrate that the proposed CALMID is more effective and efficient than several state-of-the-art learning algorithms. Weike Liu, Hang Zhang 0008, Zhaoyun Ding, Qingbao Liu, Cheng Zhu 0002 |
Knowl. Based Syst. | 3 |
| 2020 | Integrating Multisourced Texts in Online Business Intelligence SystemsabstractOnline business intelligence systems often collect the texts from different sources, such as social media and news websites that can be heterogeneous in practice. These collections bring the difficulties of managing and organizing the comprehensive information hidden in different texts of the system. To more effectively organize the multisourced texts and help online users acquire wider knowledge, we propose a business intelligence system which integrates the multisourced texts from multisources. Regarding in many occasions, multisourced texts share some common contents with respect to the same topics. For example, a tweet and a news report may talk about the same event. Therefore, our goal is to correlate such texts of different sources with respect to the similar topics and get integrated more comprehensive information to facilitate other data mining tasks as well as online applications. To handle the problem, we propose a heterogeneous information network-based text aligning (HINTA) framework in this paper. HINTA applies meta-paths to calculate the text similarities, and constructs correlated pairs between the two types of texts. Next, HINTA first applies anchored pairs as bridges to combine the different types of texts. Finally, three different inference methods are employed to align the multisourced texts. Experimental results on real-world dataset show the effectiveness and efficiency of the framework in addressing the texts alignment problem. Jianping Cao, Senzhang Wang, Benxian Li, Xiao Wang 0002, Zhaoyun Ding, Fei-Yue Wang 0001 |
IEEE Trans. Syst. Man Cybern. Syst. | 5 |
| 2018 | Node Importance Evaluation Based on Background Error ReconstructionabstractNetworks portray interactions about people communication, goods transportation and information exchange within a society, etc. How to effectively quantify the node importance is a fundamental problem in complex networks research community. In this paper, we propose a node importance evaluation method called Error Reconstruction (ER), which is mainly based on the measurement and reconstruction of errors (strong correlation with difference) between nodes in the network. First, we introduced the Network Representation Learning (NRL) approach to obtain the feature matrix about the node information and network structure, and thus we extracted the background nodes (less important nodes) to reconstruct the dense and sparse error in a single-scale network. Then we calculated the reconstruction error in a series of transformed multi-scale networks by analyzing the error with different network scales and integrated the reconstruction error with Bayesian estimation to get the final quantization of all nodes in the network. Experimental results show that the proposed method could effectively discover the vital nodes in the network. Yunzhi Han, Xianqiang Zhu, Xiaofeng Cao 0001, Zhaoyun Ding, Cheng Zhu 0002 |
SMC | 4 |
| 2018 | Predicting the popularity of topics based on user sentiment in microblogging websites
Xiang Wang 0015, Chen Wang 0026, Zhaoyun Ding, Jiumin Huang |
J. Intell. Inf. Syst. | 3 |
| 2016 | Mentioning the Optimal Users in the Appropriate Time on Twitter
Zhaoyun Ding, Xueqing Zou, Su He, Fengcai Qiao, Hui Wang 0030 |
APWeb (2) | 1 |
| 2016 | Finding the Optimal Users to Mention in the Appropriate Time on Twitter
Dayong Shen, Zhaoyun Ding, Fengcai Qiao, Hui Wang 0030 |
KSEM | 2 |
| 2016 | Exploring sentiment parsing of microblogging texts for opinion polling on chinese public figures
Xin Zhang 0018, Pei Li 0001, Sheng Zhang 0022, Zhaoyun Ding, Hui Wang 0030 |
Appl. Intell. | 5 |
| 2015 | Finding Influential Users and Popular Contents on Twitter
Zhaoyun Ding, Hui Wang 0030, Fengcai Qiao, Jianping Cao, Dayong Shen |
WISE (2) | 1 |
| 2013 | An Influence Strength Measurement via Time-Aware Probabilistic Generative Model for Microblogs
Zhaoyun Ding, Yan Jia 0001, Bin Zhou 0004, Yi Han 0006, Chunfeng Yu |
APWeb | 1 |
| 2013 | Measuring the spreadability of users in microblogsabstractMessage forwarding (e.g., retweeting on Twitter.com) is one of the most popular functions in many existing microblogs, and a large number of users participate in the propagation of information, for any given messages. While this large number can generate notable diversity and not all users have the same ability to diffuse the messages, this also makes it challenging to find the true users with higher spreadability, those generally rated as interesting and authoritative to diffuse the messages. In this paper, a novel method called SpreadRank is proposed to measure the spreadability of users in microblogs, considering both the time interval of retweets and the location of users in information cascades. Experiments were conducted on a real dataset from Twitter containing about 0.26 million users and 10 million tweets, and the results showed that our method is consistently better than the PageRank method with the network of retweets and the method of retweetNum which measures the spreadability according to the number of retweets. Moreover, we find that a user with more tweets or followers does not always have stronger spreadability in microblogs. Zhaoyun Ding, Yan Jia 0001, Bin Zhou 0004, Yi Han 0006 |
J. Zhejiang Univ. Sci. C | 1 |
| 2011 | A Word Position-Related LDA ModelabstractLDA (Latent Dirichlet Allocation) proposed by Blei is a generative probabilistic model of a corpus, where documents are represented as random mixtures over latent topics, and each topic is characterized by a distribution over words, but not the attributes of word positions of every document in the corpus. In this paper, a Word Position-Related LDA Model is proposed taking into account the attributes of word positions of every document in the corpus, where each word is characterized by a distribution over word positions. At the same time, the precision of the topic-word's interpretability is improved by integrating the distribution of the word-position and the appropriate word degree, taking into account the different word degree in the different word positions. Finally, a new method, a size-aware word intrusion method is proposed to improve the ability of the topic-word's interpretability. Experimental results on the NIPS corpus show that the Word Position-Related LDA Model can improve the precision of the topic-word's interpretability. And the average improvement of the precision in the topic-word's interpretability is about 9.67%. Also, the size-aware word intrusion method can interpret the topic-word's semantic information more comprehensively and more effectively through comparing the different experimental data. Lidong Zhai, Zhaoyun Ding, Yan Jia 0001, Bin Zhou 0004 |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2010 | A Web Service Discovery Method Based on TagabstractWith the growth of web service, traditional web service discovery mechanisms have become inefficient because of their low precision. Though the current semantic-based service discovery methods enhance the recall rate and precision in a way, most of the semantic-based service discovery methods are based on the new model of the semantic web and the ontology language, the number of the services based on the model of the semantic web and the ontology language is so little that these methods can not been used effectively. In this paper, we bring forward a new service discovery method based on the service tags. In order to discover service based on the services tags, we bring forward the algorithm QEBT and QPBT. Our experimental study shows that the algorithm QPBT is better than QEBT. Furthermore, both QEBT and QPBT are effective on discovering service. Zhaoyun Ding, Deng Lei, An Lun |
CISIS | 1 |