VLDB 2026 Research / reviewers in the wild / expert
Yan Jia 0001
dblp:24/1403-1
· DBLP profile ↗
35ranked-venue papers in the field
0as first author
9since 2021 · last 2026
0009-0000-9471-5936ORCID · conflict
Domains — venue-derived; a paper can count in several
Database Systems & Data Management · 17Information Retrieval & Web Search · 12Data Mining & Knowledge Discovery · 3Big Data, Cloud & Distributed Data Systems · 1Knowledge Engineering, Semantic Web & Information Systems · 1Other / Interdisciplinary · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SecKQL-Agent: A Real-World APT29 Events Benchmark and Framework for Reliable Text-to-KQL in Security Analytics
Haiyan Wang 0009, Shaofang Long, Yan Jia 0001, Zhaoquan Gu |
DASFAA (6) | 6 |
| 2025 | ADMatcher: Self-supervised Subgraph Matching via Adaptive Dense-Aware Graph Contrastive Learning
Yan Jia 0001, Liyi Zeng, Zhaoquan Gu |
DASFAA (3) | 5 |
| 2025 | TaylorS: A Multi-Order Expansion Structure for Urban Spatio-Temporal ForecastingabstractAlthough a variety of models have been proposed for urban spatio-temporal forecasting, most existing forecasting models are developed manually for specific tasks. By investigating the correlation between multi-order derivative and spatio-temporal data, we propose a generic yet simple plug-in structure, namedTaylorS, to improve the performance and generalization of existing forecasting models. The TaylorS converts the non-linear regression problem into a multi-order non-linear approximation problem by plugging a Taylor expansion into the forecasting task. To achieve this, we design a two-step training framework, including a training step and an adjusting step. During training, we train a given forecasting model as a base model to be equipped with prior knowledge. During adjusting, we fine-tune the base model while plugging an adjustment model into the base model. The adjustment model, as a multi-order expansion, takes the multi-order derivative of data to evaluate data uncertainty for further forecasting approximation and adjustment. Extensive experimental results demonstrate that the proposed TaylorS framework can consistently improve the performance of existing state-of-the-art methods and generalize these methods to different forecasting tasks. Jianyang Qin, Yan Jia 0001, Binxing Fang, Qing Liao 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2024 | MUSE-Net: Disentangling Multi-Periodicity for Traffic Flow ForecastingabstractAccurate forecasting of traffic flow plays a crucial role in building smart cities in the new era. Previous work has achieved success in learning inherent spatial and temporal patterns of traffic flow. However, existing works investigated the multiple periodicities (e.g., hourly, daily, and weekly) of traffic via entanglement learning, which has not yet dealt with distribution shift and interaction shift problems in traffic flow. In this paper, we propose a novel disentanglement learning network, called MUSE-Net, to tackle the limitations of entanglement learning by simultaneously factorizing the exclusiveness and interaction of multi-periodic patterns in traffic flow. Grounded in the theory of mutual information, we first learn and dis-entangle exclusive and interactive representations of traffics from multi-periodic patterns. Then, we utilize semantic-pushing and semantic-pulling regularizations to encourage the learned representations to be independent and informative. Moreover, we derive a lower bound estimator to tractably optimize the disentanglement problem with multiple variables and propose a joint training model for traffic forecasting. Extensive experimental results on several real-world traffic datasets demonstrate the effectiveness of the proposed framework. The code is available at: https://github.com/JianyangQin/MUSE-Net. Jianyang Qin, Yan Jia 0001, Yongxin Tong, Heyan Chai 0001, Ye Ding 0002, Xuan Wang 0002, Binxing Fang, Qing Liao 0001 |
ICDE | 2 |
| 2024 | A Multi-modal Prompt Learning Framework for Early Detection of Fake NewsabstractInformation spreads quickly through social media platforms, especially fake news with negative or even malicious intentions. In recent years, psychological studies have found that explicit reminders of fake news would diminish its consequence. Therefore, it is crucial to identify their authenticity at an early stage to avoid serious consequences. However, existing methods for fake news detection either utilize auxiliary information including users’ profiles and related events propagation networks or require sufficient and high-quality training data, which is not suitable for early fake news detection in real. An increasing number of social media news not only involves natural language content but also visual content such as images and videos, which give us a new view of fake news detection at an early stage by multi-modal data. In this paper, we propose a Multi-modal Prompt Learning framework (MPL) based on the multi-modal pre-trained model CLIP for early detection of fake news. A learnable prompt module is developed to adaptively and efficiently generate prompt representations to boost the semantic context. MPL can be implemented in supervised or few-shot settings. Extensive experiments show that the proposed MPL obtains substantial performance and efficiency improvement for the early-stage fake news detection task. The results demonstrate that MPL performs considerably well compared to both the state-ofthe-art supervised multi-modal models and the latest promptbased few-shot multi-modal models. Especially, the high recall of fake news and the high precision of real news that MPL achieved compared to other baselines verify that it will better approach one of the motivations that providing early notification of “maybe real” or “maybe fake” with the release of the news. Weiqi Hu, Ye Wang 0015, Yan Jia 0001, Qing Liao 0001, Bin Zhou 0004 |
ICWSM | 3 |
| 2024 | Co-Engaged Location Group Search in Location-Based Social NetworksabstractSearching for well-connected user communities in a Location-based Social Network (LBSN) has been extensively investigated. However, very few studies focus on finding a group of locations in an LBSN which are significantly engaged with socially cohesive user groups. In this work, we investigate the problem ofCo-engagedLocation groupSearch (CLS) from LBSNs where the selected locations are visited frequently by the members of the socially cohesive user groups, and the locations are reachable within a given distance threshold. To the best of our knowledge, this is the first work to search for socially co-engaged location groups in LBSNs. We devise a score function to measure the co-engagement of the location groups by combining social connectivity of the cohesive user groups and check-in density of the users to the selected locations. To solve theCLSproblem, we propose aFilter-and-Verifyalgorithm that effectively filters out ineligible locations, and their corresponding check-in users. Further, we derive a lower bound on the number of check-ins to prune the insignificant locations and develop a novel greedy forward expansion algorithm (GFA). To accelerate the computation ofCLS, we propose a ranking function and devise an incremental algorithm,GIA, that can filter the unqualified location groups. We establish the effectiveness of our solutions by conducting extensive experiments on three real-world datasets. Nur Al Hasan Haldar, Jianxin Li 0001, Naveed Akhtar, Yan Jia 0001, Ajmal Mian |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Temporal-Relational Matching Network for Few-Shot Temporal Knowledge Graph Completion
Xing Gong, Jianyang Qin, Heyan Chai 0001, Ye Ding 0002, Yan Jia 0001, Qing Liao 0001 |
DASFAA (2) | 5 |
| 2021 | MSSF-GCN: Multi-scale Structural and Semantic Information Fusion Graph Convolutional Network for Controversy Detection
Bin Zhou 0004, Ye Wang 0015, Liqun Gao, Yan Jia 0001 |
WISE (1) | 6 |
| 2021 | Performance Evaluation of Pre-trained Models in Sarcasm Detection Task
Bin Zhou 0004, Ye Wang 0015, Liqun Gao, Yan Jia 0001 |
WISE (2) | 6 |
| 2020 | A Graph Data Privacy-Preserving Method Based on Generative Adversarial Networks
Aiping Li, Qianye Jiang, Bin Zhou 0004, Yan Jia 0001 |
WISE (2) | 5 |
| 2017 | A Refined Method for Detecting Interpretable and Real-Time Bursty Topic in Microblog Stream
Tao Zhang 0164, Bin Zhou 0004, Jiuming Huang, Yan Jia 0001 |
WISE (1) | 4 |
| 2017 | Big Search in CyberspaceabstractWith the rapid development of big data analytics, mobile computing, Internet of Things, cloud computing, and social networking, cyberspace has expanded to a cross-fused and ubiquitous space made up of human beings, things, and information. Internet applications have evolved from Web 1.0 to Web 2.0 and Web 3.0, and web information has seen an explosive growth, which is strongly promoting the advent of a global era of big data. In this ubiquitous cyberspace, traditional search engines can no longer fully satisfy the evolving needs of various types of users. Therefore, search engines must make completely innovative, revolutionary changes for the next generation of search, which is referred to as “big search”. This paper first studies the development needs of big search. Then, big search is defined, and the 5S properties (Sourcing, Sensing, Synthesizing, Solution, and Security) of big search, which are different from those of traditional search engines, are elaborated. Also, the paper provides a system architecture for big search, explores the key technologies that support the 5S properties, and describes potential application fields of big search technology. Finally, the research opportunities of big search are discussed. Binxing Fang, Yan Jia 0001, Xiaoyong Li 0003, Aiping Li, Xindong Wu 0001 |
IEEE Trans. Knowl. Data Eng. | 2 |
| 2015 | Top (k1, k2) Distance-based outliers detection in an uncertain datasetabstractIn this paper, we focus on distance-based outliers detection in an uncertain dataset, which is very useful in large social network. Based on the x-tuple model and the possible world semantics, we propose the concept of tuple outlier score, top k1probability and top (k1, k2) distance-based outlier. We then design an algorithm using dynamic programming technique to calculate tuple outlier scores and detect top (k1, k2) distance-based outliers. The local neighbor region is proposed to detect approximate outliers with high precision efficiently. We also propose two pruning strategies to avoid additional computation overhead and prune data objects that cannot be outliers. After theory analysis, we conduct experiments in two real datasets to verify good performance of our method. Fei Liu 0016, Yan Jia 0001 |
IEEE BigData | 2 |
| 2015 | Detecting Internet Hidden Paid Posters Based on Group and Individual Characteristics
Xiang Wang 0015, Bin Zhou 0004, Yan Jia 0001 |
WISE (2) | 3 |
| 2014 | Do neighbor buddies make a difference in reblog likelihood? An analysis on SINA Weibo dataabstractReblogging, also known as retweeting in Twitter parlance, is a major type of activities in many online social networks. Although there are many studies on reblogging behaviors and potential applications, whether neighbors who are well connected with each other (called “buddies” in our study) may make a difference in reblog likelihood has not been examined systematically. In this paper, we tackle the problem by conducting a systematic statistical study on a large SINA Weibo data set, which is a sample of 135, 859 users, 10, 129, 028 followers, and 2, 296, 290, 930 reblog messages in total. To the best of our knowledge, this data set has more reblog messages than any data sets reported in literature. We examine a series of hypotheses about how essential neighborhood structures may help to boost the likelihood of reblogging, including buddy neighbors versus buddyless neighbors, traffic between buddy neighbors, activeness (i.e., the total number of blog messages a user sends), and the number of buddy triangles a user participates in. Our empirical study discloses several interesting phenomena that are not reported in literature, which may imply interesting and valuable new applications. Lumin Zhang, Jian Pei 0001, Yan Jia 0001, Bin Zhou 0004, Xiang Wang 0015 |
ASONAM | 3 |
| 2013 | An Influence Strength Measurement via Time-Aware Probabilistic Generative Model for Microblogs
Zhaoyun Ding, Yan Jia 0001, Bin Zhou 0004, Yi Han 0006, Chunfeng Yu |
APWeb | 2 |
| 2013 | An Efficient Approach on Answering Top-k Queries with Grid Dominant Graph Index
Aiping Li, Jinghu Xu, Liang Gan, Bin Zhou 0004, Yan Jia 0001 |
APWeb | 5 |
| 2012 | Adaptive Topic Community Tracking in Social Network
Yan Jia 0001, Bin Zhou 0004 |
APWeb | 2 |
| 2012 | Community detection in incomplete information networksabstractWith the recent advances in information networks, the problem of community detection has attracted much attention in the last decade. While network community detection has been ubiquitous, the task of collecting complete network data remains challenging in many real-world applications. Usually the collected network is incomplete with most of the edges missing. Commonly, in such networks, all nodes with attributes are available while only the edges within a few local regions of the network can be observed. In this paper, we study the problem of detecting communities in incomplete information networks with missing edges. We first learn a distance metric to reproduce the link-based distance between nodes from the observed edges in the local information regions. We then use the learned distance metric to estimate the distance between any pair of nodes in the network. A hierarchical clustering approach is proposed to detect communities within the incomplete information networks. Empirical studies on real-world information networks demonstrate that our proposed method can effectively detect community structures within incomplete information networks. Wangqun Lin, Xiangnan Kong, Philip S. Yu, Quanyuan Wu, Yan Jia 0001, Chuan Li 0002 |
WWW | 5 |
| 2012 | Contextual correlation based thread detection in short text message streams
Jiuming Huang, Bin Zhou 0004, Quanyuan Wu, Yan Jia 0001 |
J. Intell. Inf. Syst. | 5 |
| 2011 | Link-based hidden attribute discovery for objects on WebabstractInformation extraction from the Web is of growing importance. Objects on the Web are often associated with many attributes that describe the objects. It is essential to extract these attributes and map them to their corresponding objects. However, much attribute information about an object is hidden in the dynamic user interaction and is not on the Web page that describes the object. Existing information extraction approaches focus on getting information from the object Web page only, which means a lot of attribute information is lost. In this paper, we study the dynamic user interaction on exploratory search Websites and propose a novel link-based approach to discover attributes and map them to objects. We build an exploratory search model for exploratory Web sites, and we propose algorithms for identifying, clustering, and relationship mining of related Web pages based on the model. Using the unsupervised method in our approach, we are able to discover hidden attributes not explicitly shown on object Web pages. We test our approach on two online shopping Websites. We achieve high precision and recall: For entirely crawled Web sites the precision and recall are 98% and 97% respectively. For randomly crawled (sampled) Web sites the precision and recall are 98% and 80% respectively. Jiuming Huang, Haixun Wang, Yan Jia 0001, Ariel Fuxman |
EDBT | 3 |
| 2010 | Join Directly on Heavy-Weight Compressed Data in Column-Oriented Database
Liang Gan, Runheng Li, Yan Jia 0001 |
WAIM | 3 |
| 2009 | Continuous privacy preserving publishing of data streamsabstractRecently, privacy preserving data publishing has received a lot of attention in both research and applications. Most of the previous studies, however, focus on static data sets. In this paper, we study an emerging problem of continuous privacy preserving publishing of data streams which cannot be solved by any straightforward extensions of the existing privacy preserving publishing methods on static data. To tackle the problem, we develop a novel approach which considers both the distribution of the data entries to be published and the statistical distribution of the data stream. An extensive performance study using both real data sets and synthetic data sets verifies the effectiveness and the efficiency of our methods. Bin Zhou 0002, Yi Han 0006, Jian Pei 0001, Bin Jiang 0009, Yufei Tao 0001, Yan Jia 0001 |
EDBT | 6 |
| 2009 | Effective Feature Selection on Data with Uncertain LabelsabstractNowadays, various learning technologies are required on uncertain data. As an important pre-processing step in data mining, feature selection needs to consider this vagueness or uncertainty. In this paper, we propose a novel algorithm to evaluate the correlation between features and uncertain class labels on the basis of Hilbert-Schmidt Independence Criterion. Consequently, the features can be ranked according to this criterion. Experimental results on extensive datasets demonstrate the benefits of our method. Yan Jia 0001, Yi Han 0006, Weihong Han |
ICDE | 2 |
| 2009 | Understanding Importance of Collaborations in Co-authorship Networks: A Supportiveness Analysis ApproachabstractCo-authorship networks, an important type of social networks, have been studied extensively from various angles such as degree distribution analysis, social community extraction and social entity ranking. Most of the previous studies consider the co-authorship relation between two authors as a collaboration. In this paper, we introduce a novel and interesting “supportiveness” measure on co-authorship relation. The fact that two authors co-author one paper can be regarded as one author supports the other's scientific work. We propose several supportiveness measures, and exploit a supportiveness-based author ranking scheme. Several efficient algorithms are developed to compute the top-n most supportive authors. Moreover, we extend the supportiveness analysis to community extraction, and develop feasible solutions to identify the most supportive groups of authors. The empirical study conducted on a large real data set indicates that the supportiveness measures are interesting and meaningful, and our methods are effective and efficient in practice. Yi Han 0006, Bin Zhou 0002, Jian Pei 0001, Yan Jia 0001 |
SDM | 4 |
| 2008 | Leakage-Aware Energy Efficient Scheduling for Fixed-Priority Tasks with Preemption Thresholds
XiaoChuan He, Yan Jia 0001 |
ADMA | 2 |
| 2008 | Network Dynamic Risk Assessment Based on the Threat Stream AnalysisabstractThis paper considers the problem of the dynamic risk assessment for the network based on the threat stream analysis. We analyze the general approach to do the network dynamic risk assessment. A stream based cube model is built to analyze the characteristics of the threat stream. Then combining with the research about the description and analysis of the threat effect, we propose the architecture of the network dynamic risk assessment. Wencong Cheng, Xishan Xu, Yan Jia 0001 |
WAIM | 3 |
| 2008 | Counting Data Stream Based on Improved Counting Bloom FilterabstractBurst detection is an inherent problem for data streams, so it has attracted extensive attention in research community due to its broad applications. One of the basic problems in burst detection is how to count frequencies of all elements in data stream. This paper presents a novel solution based on Improved Counting Bloom Filter, which is also called BCBF+HSet. Comparing with intuitionistic approach such as array and list, our solution significantly reduces space complexity though it introduces few error rates. Further, we discuss space/time complexity and error rate of our solution, and compare it with two classic Counting Bloom Filters, CBF and DCF. Theoretical analysis and simulation results demonstrate the efficiency of the proposed solution. Zhijian Yuan, Jiajia Miao, Yan Jia 0001, Le Wang 0008 |
WAIM | 3 |
| 2007 | Middleware Based Context Management for the Component-Based Pervasive Computing
Jun Wang 0063, Yan Jia 0001, Weihong Han |
ATC | 3 |
| 2006 | Closed Queueing Network Model for Multi-tier Data Stream Processing Center
Huaimin Wang 0001, Yan Jia 0001, Bixin Liu |
APWeb | 3 |
| 2006 | Efficient Non-Blocking Top-k Query Processing in Distributed Networks
Yan Jia 0001, Shuqiang Yang |
DASFAA | 2 |
| 2006 | Supporting Efficient Distributed Top-k Monitoring
Yan Jia 0001, Shuqiang Yang |
WAIM | 2 |
| 2005 | Stratus: A Distributed Web Service Discovery Infrastructure Based on Double-Overlay Network
Jian-Qiang Hu, Changguo Guo, Yan Jia 0001 |
APWeb | 3 |
| 2005 | Parallel Mining of Top-K Frequent Itemsets in Very Large Text Database
Yongheng Wang, Yan Jia 0001, Shuqiang Yang |
WAIM | 2 |
| 2004 | An Efficient Hierarchical Failure Recovery Algorithm Ensuring Semantic Atomicity for Workflow Applications
Yi Ren 0008, Quanyuan Wu, Yan Jia 0001, Jianbo Guan |
WAIM | 3 |