Tieyun Qian

dblp:17/5583 · DBLP profile ↗
in reviewer pool ← Back
69ranked-venue papers in the field
12as first author
33since 2021 · last 2026
0000-0003-4667-5794ORCID · verified

Domains — venue-derived; a paper can count in several

Database Systems & Data Management · 32 (3 first)Information Retrieval & Web Search · 22 (6 first)Data Mining & Knowledge Discovery · 8 (2 first)Knowledge Engineering, Semantic Web & Information Systems · 6 (1 first)Other / Interdisciplinary · 1
YearPublicationVenuePosition
2026 Debiasing LLMs in Knowledge-Intensive Tasks via Information-Gain Guided Front-Door Adjustment
Yongqi Li 0002, Hankun Kang, Mayi Xu, Jintao Wen, Yuanyuan Zhu 0001, Ming Zhong 0002, Jiawei Jiang 0001, Tieyun Qian
DASFAA (3)9
2026 Discrimination Matters: A Simple but Effective Method for Zero-Shot Relation Triplet Extraction
Tieyun Qian, Lixin Zou, Xuming Hu, Wanli Li 0002, Zaiwen Feng
DASFAA (6)3
2026 DM-RAG: Enhancing User Support in Dameng Databases with Retrieval-Augmented Generation
Qiang Huang 0009, Ke Liu 0014, Liang Deng, Sijing Zhang, Chuang Hu, Tieyun Qian, Xiao Yan 0002, Jiawei Jiang 0001
ICDE6
2026 HAL: Accurate, Private, and Efficient Sample Alignment for Multimodal Federated Learning
abstract
Vertical multimodal federated learning (VMFL) enables multiple clients holding data from different modalities to conduct collaboratively model training. Existing methods typically assume that multimodal data samples (i.e., text and image) from the same entity (i.e., person) are paired across the clients (i.e., aligned). However, this assumption rarely holds in practice, as data is often collected independently with no shared identifiers. To address this challenge, we propose hashing-based alignment (HAL), a new VMFL framework that works without pre-aligned samples. HAL consists of two key components. The first component is an efficient and privacy-preserving method to identify similar samples from different modalities as aligned pairs. It adopts locality sensitive hashing (LSH) for the efficient retrieval of similar samples, introduces a shift-orthogonal hashing scheme to tackle the gaps between different modalities, and uses a bloom-style method for secure Hamming distance estimation. We prove that the shift-orthogonal hashing reduces distance estimation errors and secure Hamming distance estimation satisfies differential privacy. The second component is a neighbor-aware fusion strategy, which applies cross-attention to aggregate informative signals from the aligned samples without relying on explicit similarity scores. Experimental results on two real-world datasets show that compared with five state-of-the-art (SOTA) baselines, HAL improves the cross-modal retrieval accuracy by over 63%, while also achieving up to 154× speedup.
Xiaokai Zhou, Xiao Yan 0002, Yuxiang Wang 0013, Quanqing Xu, Chuang Hu, Tieyun Qian, Jiawei Jiang 0001
KDD (1)7
2026 Recursive Short-to-Long Generalization for Multi-hop Reasoning
Mayi Xu, Ke Sun 0010, Jianhao Chen 0003, Qiankun Pi, Guixin Su, Yunfeng Ning, Yongqi Li 0002, Yuanyuan Zhu 0001, Ming Zhong 0002, Jiawei Jiang 0001, Tieyun Qian
SIGIR11
2026 ContiGuard: A Framework for Continual Toxicity Detection Against Evolving Evasive Perturbations
abstract
Toxicity detection mitigates the dissemination of toxic content (e.g., hateful comments, posts, and messages within online social actions) to safeguard a healthy online social environment. However, malicious users persistently develop evasive perturbations to disguise toxic content and evade detectors. Traditional detectors or methods are static over time and are inadequate in addressing these evolving evasion tactics. Thus, continual learning emerges as a logical approach to dynamically update detection ability against evolving perturbations. Nevertheless, disparities across perturbations hinder the detector's continual learning on perturbed text. More importantly, perturbation-induced noises distort semantics to degrade comprehension and also impair critical feature learning to render detection sensitive to perturbations. These amplify the challenge of continual learning against evolving perturbations.
Hankun Kang, Jianhao Chen 0003, Jintao Wen, Mayi Xu, Weiyu Zhang 0001, Wenpeng Lu, Tieyun Qian
WWW8
2026 Tracing Belief-Driven Thoughts with Theory-of-Mind Agents: An Opinion Analysis Framework
Jintao Wen, Yunfeng Ning, Hankun Kang, Tieyun Qian
WWW5
2026 Developing continuous toxicity detection against increasing types of perturbed toxic text
Hankun Kang, Jianhao Chen 0003, Yongqi Li 0002, Mayi Xu, Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian
Inf. Process. Manag.9
2026 Reasoning based on symbolic and parametric knowledge bases: A survey
Mayi Xu, Yunfeng Ning, Yongqi Li 0002, Jianhao Chen 0003, Jintao Wen, Birong Pan, Zepeng Bao, Hankun Kang, Ke Sun 0010, Tieyun Qian
Inf. Process. Manag.13
2026 Local and Global Exploration for Next New POI Recommendation
abstract
The next Point-of-Interest (POI) recommendation is a hotspot for both industry and academia, which helps users better experience the physical world. However, existing methods suffer from a severe bias towards recommending repeat POIs that have been visited by the target user before, and perform inefficiently when recommending new POIs that have not been visited by the target user yet. To overcome this issue, we delve into the next new POI recommendation and uncover the coexistence of local and global exploration patterns in users’ visits to new POIs, showing their willingness to explore not only nearby new POIs but also those distant ones. Subsequently, we develop a novel Local and Global Exploration ( LGE ) framework for the next new POI recommendation. In particular, LGE involves three key modules: (1) a Zone-Aware Local Exploration (ZLE) module, which encourages users to explore POIs in the local area by learning zone-aware POI representations and regularizing POI prediction with zone information; (2) an Intention-Aware Global Exploration (IGE) module, which recommends POIs that meet user intentions without distance constraints by extracting static and dynamic intentions from category information; (3) a fusion module, which contains a Mean Pooling (MP) strategy and a Weighted Pooling (WP) strategy to aggregate the outputs of local and global exploration modules for the final recommendation. Experiments carried out on real-world datasets have shown the effectiveness of LGE in recommending new POIs.
Ke Sun 0010, Liyu Zhou, Mayi Xu, Tieyun Qian
ACM Trans. Knowl. Discov. Data4
2025 Efficient Frequency-Aware k-Core Query on Temporal Graphs
abstract
In temporal graphs, time and topology are considered to be intertwined. As an evidence, it is observed that the vertices in more cohesive subgraphs have more frequent and more numerous interactions between each other in the history. Motivated by that, we study a novel frequency-aware k-core query problem. Different from previous studies that focus on finding k-cores in the projected subgraphs of given time intervals, we look for the subgraphs of k-core in which neighbor vertices have at least a certain number of high-frequency interactions. To address the problem, we propose 1) a minimum slope algorithm for computing the frequency in linear time, 2) a space-efficient index that stores the distinct “core frequency” of vertices for addressing arbitrary queries, 3) a propagation algorithm that collects core frequencies by message passing for index construction, and 4) efficient algorithms for retrieving a specific or all skyline results from the index respectively. The experimental results show that, our algorithms achieve several orders of magnitude improvement on efficiency compared to corresponding baselines, and meanwhile, the size of index is even smaller than that of graph unless the graph has very few timestamps on each edge. More importantly, by both statistics and case study, it is verified that the frequency-aware k-core query indeed find more cohesive subgraphs in the static k-core.
Zhongfan Du, Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian, Mengchi Liu, Jeffrey Xu Yu
ICDE4
2025 TDT: Tensor Based Directed Truss Decomposition
abstract
Truss decomposition is to find the hierarchy of all the k-trusses in a graph for$k\geq 2$. Existing GPU-based algorithms first compute edge support by parallelly counting the number of triangles each edge is contained in, and then iteratively peel off edges with the smallest support and update support of the affected edges in parallel. However, these algorithms perform truss decomposition on undirected graphs, which causes large storage space and numerous triangle existence checks during support update. Moreover, they are developed based on CUDA, which cannot naturally adapt to emerging hardware accelerators and support the end-to-end downstream graph machine learning (ML) tasks. In this paper, we propose a truss decomposition framework based on tensors (TDT), which can leverage the parallelism of heterogeneous hardware backends to speed up the computation and seamlessly integrate with downstream graph ML tasks. We first convert the original input graph into a directed graph and represent it by compacted tensors. Then we perform truss decomposition on the tensorized directed graph by efficient tensor operators. Such a directed-graph storage model not only saves the storage space but also naturally supports efficient support computation/update during the truss decomposition. To further accelerate truss decomposition, we also partition vertex neighbors into blocks to balance the computation workload and optimize key steps such as support computation/update in our framework. Extensive experimental studies show that our Python-based TDT algorithm not only achieves$2.3\times-8.5\times$speedup in most cases compared with the state-of-the-art CUDA-based algorithms, but also can efficiently deal with large graphs with hundreds of millions of nodes and billions of edges while the baseline fails due to large storage cost. Our source code is publicly available at https://github.com/LiGuojing194/TDTdecomposition.
Guojing Li, Yuanyuan Zhu 0001, Ming Zhong 0002, Tieyun Qian, Jeffrey Xu Yu
ICDE5
2025 Hounding Data Diversity: Towards Participant Selection in Vertical Federated Learning
abstract
Due to the rising concerns on privacy protection, how to build machine learning models from distributed databases with privacy guarantees has gained more popularity. Vertical federated learning (VFL) trains machine learning models in a privacy-preserving way when the data features are scattered over distributed databases. We study the participant selection problem (PSP) for VFL, which chooses a given number of participants to conduct training while maximizing model accuracy. Compared to training with all participants, PSP can filter out hitch-riders that contribute marginally to model quality and reduce training time by involving fewer participants. To achieve good model accuracy, we formulate PSP as choosing a set of participants that maximizes the likelihood of the data samples. Then, utilizing the k-nearest neighbors (KNN) classifier as the proxy model, we express the likelihood as a function of the selected participants and prove that the function is sub modular. The submodular property is favorable as it can account for the feature diversity among the participants and allows to greedily select the participant with the maximum gain in each step. However, the selection process requires finding the top-k neighbors of a data sample as the basic operation, which is expensive in VFL setting as it involves encrypted communication. As such, we adapt the Fagin's algorithm, a famous top-k query algorithm, to reduce the amount of encrypted communication. We deploy our solution VFPS-SM across five distributed nodes and conduct experiments with 10 datasets and 3 models to evaluate its performance. The results show that VFPS-SM can reduce the end-to-end running time by up to$35\times$, selection time$365\times$and improve model accuracy by 6.0% compared with state-of-the-art baselines.
Xiaokai Zhou, Xiao Yan 0002, Fangcheng Fu, Hao Huang 0001, Quanqing Xu, Chuanhui Yang, Bo Du 0001, Tieyun Qian, Jiawei Jiang 0001
ICDE9
2025 Generative Meta-Learning for Zero-Shot Relation Triplet Extraction
abstract
Zero-shot Relation Triplet Extraction (ZeroRTE) aims to extract relation triplets from texts containing unseen relation types. This capability benefits various downstream information retrieval (IR) tasks. The primary challenge lies in enabling models to generalize effectively to unseen relation categories. Existing approaches typically leverage the knowledge embedded in pre-trained language models to accomplish the generalization process. However, these methods focus solely on fitting the training data during training, without specifically improving the model's generalization performance, resulting in limited generalization capability. For this reason, we explore the integration of bi-level optimization (BLO) with pre-trained language models for learning generalized knowledge directly from the training data, and propose a generative meta-learning framework which exploits the 'learning-to-learn' ability of meta-learning to boost the generalization capability of generative models.
Wanli Li 0002, Tieyun Qian, Zeyu Zhang 0004, Jiawei Li 0008, Zhuang Chen 0002, Lixin Zou
SIGIR2
2025 TQEx: Tensor-based Query Engine Enhanced by Bridging the Gap
abstract
With the development of AI and the growing demand for computational power, hardware is becoming increasingly specialized and heterogeneous. The emergence of diverse specialized hardware architectures, each with distinct characteristics and programming abstractions, poses significant portability and sustainability challenges for existing data processing systems. Tensor Computation Runtimes (TCRs) abstract away the low-level hardware complexities by providing users with a hardware-independent tensor-based interface, enabling data scientists to effectively leverage the powerful capabilities of new hardware accelerators (collectively referred to as XPU). Built on TCRs, the existing relational query engine TQP demonstrates portability across a wide range of target hardware and sustainability along with the ongoing evolution of TCRs and hardware. However, it neglects the big gap between irregular SQL workloads and uniform tensor operations when mapping SQL operators to tensor programs, which causes significant storage and computation overhead. In this paper, for the first time, we analyze the underlying gap between SQL and tensors, and provide guidelines to bridge it. Following these guidelines, we build a new Tensor-based Query Engine Enhanced (TQEx) by bridging the gap from multiple aspects: develop efficient storage and computation strategies for variable-length data, and design efficient SQL operators such as join and aggregate based on tensors. We also extend TQEx to multi-XPUs for large-scale data processing. Extensive experimental studies show that our query engine, TQEx, achieves a 9.6× speedup (with a peak of 41.9×) over TQP on TPC-H, and it is also 27.9× faster than leading GPU databases such as HeavyDB. On TPC-H at scale factor 100, TQEx outperforms DuckDB by 12.2× and HeavyDB by 22.7× on supported queries.
Yuanyuan Zhu 0001, Hao Zhang 0098, Congli Gao, Ming Zhong 0002, Jiawei Jiang 0001, Tieyun Qian, Jeffrey Xu Yu
Proc. ACM Manag. Data8
2025 TGraph: A Tensor-centric Graph Processing Framework
abstract
Graph is ubiquitous in various real-world applications, and many graph processing systems have been developed. Recently, hardware accelerators have been exploited to speed up graph systems. However, such hardware-specific systems are hard to migrate across different hardware backends. In this paper, we propose the first tensor-based graph processing framework, Tgraph, which can be smoothly deployed and run on any powerful hardware accelerators (uniformly called XPU) that support Tensor Computation Runtimes (TCRs). TCRs, which are deep learning frameworks along with their runtimes and compilers, provide tensor-based interfaces to users to easily utilize specialized hardware accelerators without delving into the complex low-level programming details. However, building an efficient tensor-based graph processing framework is non-trivial. Thus, we make the following efforts: (1) propose a tensor-centric computation model for users to implement graph algorithms with easy-to-use programming interfaces; (2) provide a set of graph operators implemented by tensor to shield the computation model from the detailed tensor operators so that Tgraph can be easily migrated and deployed across different TCRs; (3) design a tensor-based graph compression and computation strategy and an out-of-XPU-memory computation strategy to handle large graphs. We conduct extensive experiments on multiple graph algorithms (BFS, WCC, SSSP, etc.), which validate that Tgraph not only outperforms seven state-of-the-art graph systems, but also can be smoothly deployed and run on multiple DL frameworks (PyTorch and TensorFlow) and hardware backends (Nvidia GPU, AMD GPU, and Apple MPS).
Yuanyuan Zhu 0001, Hao Zhang 0098, Congli Gao, Guojing Li, Ming Zhong 0002, Jiawei Jiang 0001, Tieyun Qian, Chenyi Zhang 0002, Jeffrey Xu Yu
Proc. ACM Manag. Data10
2025 On More Efficiently and Versatilely Querying Historical k-Cores
abstract
The recently proposed historical k -core query introduces a new paradigm of structure analysis for temporal graphs. However, the query processing based on the existing PHC-index, which preserves the distinct "core time" of each vertex, needs to traverse all vertices for each query, even though the results usually contain only a small subset of vertices. Inspired by the traditional k -shell that ensures the optimal k -core query processing, we propose a novel concept called "core time shell", which reveals the hierarchical structure of vertices with respect to their core time. Based on the core time shell, we design a time-space balanced Merged Core Time Shell index (MCTS-index). It is theoretically guaranteed that, the MCTS-index provides the approximately optimal query performance, and has the approximately same space complexity as the PHC-index. Moreover, we leverage the MCTS-index to efficiently address the brand-new "when" historical k -core queries orthogonal to the current "what" historical k -core queries. Our experimental results on ten real-world temporal graphs demonstrate both the superior efficiency of processing "what" queries and the effectiveness of processing versatile "when" queries for the MCTS-index.
Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian, Mengchi Liu, Jeffrey Xu Yu
Proc. VLDB Endow.4
2025 PS-MI: Accurate, Efficient, and Private Data Valuation in Vertical Federated Learning
abstract
Vertical federated learning (VFL) trains models when multiple databases (a.k.a participants) hold different features of the same set of samples. By quantifying each participant's contribution to model training, data valuation can prevent hitch-riders and reward the instrumental parties. However, vertical federated data valuation (VFDV) is challenging because it needs to be accurate and efficient while protecting participant data privacy. In this paper, we propose a method meeting all three requirements by using projection and sampling for mutual information estimation (thus dubbed PS-MI). In particular, we first show that the utility of a participant set (a.k.a a coalition ) can be expressed as the mutual information (MI) between their features and the target labels. MI is favorable because it does not depend on the model to train (i.e., model-agnostic ) and can be estimated via k -nearest neighbor (KNN). To run KNN, instead of using costly homomorphic encryption to protect data privacy, we apply simple random projection to participant features before distance computation. We prove that random projection ensures differential privacy and preserves unbiased distance estimates. Since the contribution of a participant involves many coalitions, we adopt stratified sampling to reduce the number of coalitions while controlling estimation variance. To further improve efficiency, we incorporate optimizations including using locality sensitive hashing (LSH) to prune kNN candidates, batching kNN candidate checking for multiple coalitions, and adaptive early termination for utility evaluation. We compare PS-MI with 5 state-of-the-art VFDV methods. The results show that PS-MI yields higher accuracy and shorter running time than the baselines, and the maximum speedup can be 592×.
Xiaokai Zhou, Xiao Yan 0002, Fangcheng Fu, Ziwen Fu, Tieyun Qian, Yuanyuan Zhu 0001, Qinbo Zhang, Bin Cui 0001, Jiawei Jiang 0001
Proc. VLDB Endow.5
2024 Querying Cohesive Subgraph Regarding Span-Constrained Triangles on Temporal Graphs
abstract
The recent prosperity of temporal graph research redefines many traditional concepts on static graphs, such as triangle, motif,$k$-core, etc. Inspired by that, we propose a novel$(k, \delta)$-truss on temporal graphs, which requires its triangles to exist in short enough time windows ever. The$(k,\delta)$-truss satisfies both static and temporal cohesion, while the original$k$-truss is its special case when$\delta=\infty$. In order to address the$(k, \delta)$-truss query, we propose both index-free and index-based approaches. By leveraging the dual containment relation on$(k, \delta)$-trusses, our indexes can compress all$(k, \delta)$-trusses losslessly into map or tree structures with dramatically less space, so that a specific$(k,\ \delta)$-truss can be retrieved from indexes in the optimal time. To enable our index to scale to large temporal graphs, we develop two index construction algorithms that can reduce redundant computation significantly, based on truss decomposition and truss maintenance respectively. The experimental results demonstrate that index-based approaches process queries in interactive time and outperform the index-free approach by 2~4 orders of magnitude, while indexes achieve compression ratios up to 10-4.
Chuhan Hu, Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian, Ting Yu 0004, Hongyang Chen 0001, Mengchi Liu, Jeffrey Xu Yu
ICDE4
2024 Evolution Forest Index: Towards Optimal Temporal $k$-Core Component Search via Time-Topology Isomorphic Computation
abstract
For a temporal graph like transaction network, finding a densely connected subgraph that contains a vertex like a suspicious account during a period is valuable. Thus, we study the Temporal k -Core Component Search (TCCS) problem, which aims to find a connected component of temporal k -core for any given vertex and time interval. Towards this goal, we propose a novel Evolution Forest Index (EF-Index) that can address TCCS in optimal time. Essentially, EF-Index leverages the evolutionary order on temporal k -cores to both compress the connectivity between vertices in temporal k -cores of all time intervals into a minimum set of compactest Minimum Temporal Spanning Forests (MTSFs) and retrieve MTSF for a given time interval rapidly. Here, a crucial innovation is that, we extend the temporal k -core evolution theory by introducing a pair of time-topology isomorphic relations, on top of which the evolutionary order in topology domain can be simply computed by a "kernel function" in time domain. Moreover, we design an efficient mechanism to update EF-Index incrementally for dynamic edge streams. The experimental results on a variety of real-world temporal graphs demonstrate that, EF-Index outperforms the state-of-the-art approach by 1--3 orders of magnitude on processing TCCS, and its space overhead is reduced by 4--5 orders of magnitude compared with preserving connectivity uncompressedly.
Junyong Yang, Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian, Mengchi Liu, Jeffrey Xu Yu
Proc. VLDB Endow.4
2024 A Unified and Scalable Algorithm Framework of User-Defined Temporal $(k,\mathcal {X})$(k,X)-Core Query
abstract
Querying cohesive subgraphs on temporal graphs (e.g., social network, finance network, etc.) with various conditions has attracted intensive research interests recently. In this paper, we study a novel Temporal$(k,\mathcal {X})$-Core Query (TXCQ) that extends a fundamental Temporal$k$-Core Query (TCQ) proposed in our conference paper by optimizing or constraining an arbitrary metric$\mathcal {X}$of$k$-core, such as size, engagement, interaction frequency, time span, burstiness, periodicity, etc. Our objective is to address specific TXCQ instances with conditions on different$\mathcal {X}$in a unified algorithm framework that guarantees scalability. For that, this journal paper proposes a taxonomy of measurement$\mathcal {X}(\cdot )$and achieve our objective using a two-phase framework while$\mathcal {X}(\cdot )$is time-insensitive or time-monotonic. Specifically, Phase 1 still leverages the query processing algorithm of TCQ to induce all distinct$k$-cores during a given time range, and meanwhile locates the “time zones” in which the cores emerge. Then, Phase 2 conducts fast local search and$\mathcal {X}$evaluation in each time zone with respect to the time insensitivity or monotonicity of$\mathcal {X}(\cdot )$. By revealing two insightful concepts named tightest time interval and loosest time interval that bound time zones, the redundant core induction and unnecessary$\mathcal {X}$evaluation in a zone can be reduced dramatically. Our experimental results demonstrate that TXCQ can be addressed as efficiently as TCQ, which achieves the latest state-of-the-art performance, by using a general algorithm framework that leaves$\mathcal {X}(\cdot )$as a user-defined function.
Ming Zhong 0002, Junyong Yang, Yuanyuan Zhu 0001, Tieyun Qian, Mengchi Liu, Jeffrey Xu Yu
IEEE Trans. Knowl. Data Eng.4
2024 City Matters! A Dual-Target Cross-City Sequential POI Recommendation Model
abstract
Existing sequential Point of Interest (POI) recommendation methods overlook a fact that each city exhibits distinct characteristics and totally ignore the city signature. In this study, we claim that city matters in sequential POI recommendation and fully exploring city signature can highlight the characteristics of each city and facilitate cross-city complementary learning. To this end, we consider the two-city scenario and propose a Dual-Target Cross-City Sequential POI Recommendation model (DCSPR) to achieve the purpose of complementary learning across cities. On one hand, DCSPR respectively captures geographical and cultural characteristics for each city by mining intra-city regions and intra-city functions of POIs. On the other hand, DCSPR builds a transfer channel between cities based on intra-city functions, and adopts a novel transfer strategy to transfer useful cultural characteristics across cities by mining inter-city functions of POIs. Moreover, to utilize these captured characteristics for sequential POI recommendation, DCSPR involves a new region- and function-aware network for each city to learn transition patterns from multiple views. Extensive experiments conducted on two real-world datasets with four cities demonstrate the effectiveness of DCSPR .
Ke Sun 0010, Chenliang Li 0005, Tieyun Qian
ACM Trans. Inf. Syst.3
2023 Attribute Graph Neural Networks for Strict Cold Start Recommendation : Extended Abstract
abstract
Recently, deep learning based methods, especially graph neural network (GNN), have made impressive progress on rating prediction problem in recommender systems. However, the performance of existing methods drops quickly in the cold start scenario. More importantly, such methods are unable to learn the preference embedding of a strict cold start user/item since there is no interaction for this user/item. In this work, we develop a novel framework Attribute Graph Neural Networks (AGNN) by exploiting the attribute graph rather than the commonly used interaction graph. AGNN can produce the preference embedding for a strict cold user/item by learning on the distribution of attributes with an extended variational auto-encoder (eVAE) structure. It also contains a new graph neural network variant (gated-GNN) to effectively aggregate various attributes of different dimensions in a neighborhood. Empirical results demonstrate that AGNN achieves the new state-of-the-art performance.
Tieyun Qian, Yile Liang, Qing Li 0001, Hui Xiong 0001
ICDE1
2023 Cold-Start Multi-hop Reasoning by Hierarchical Guidance and Self-verification
Mayi Xu, Ke Sun 0010, Yongqi Li 0002, Tieyun Qian
ECML/PKDD (2)4
2023 Towards Model Robustness: Generating Contextual Counterfactuals for Entities in Relation Extraction
abstract
The goal of relation extraction (RE) is to extract the semantic relations between/among entities in the text. As a fundamental task in information systems, it is crucial to ensure the robustness of RE models. Despite the high accuracy current deep neural models have achieved in RE tasks, they are easily affected by spurious correlations. One solution to this problem is to train the model with counterfactually augmented data (CAD) such that it can learn the causation rather than the confounding. However, no attempt has been made on generating counterfactuals for RE tasks.
Mi Zhang 0006, Tieyun Qian
WWW2
2023 Scalable Time-Range k-Core Query on Temporal Graphs
abstract
Querying cohesive subgraphs on temporal graphs with various time constraints has attracted intensive research interests recently. In this paper, we study a novel Temporal k -Core Query (TCQ) problem: given a time interval, find all distinct k -cores that exist within any subintervals from a temporal graph, which generalizes the previous historical k -core query. This problem is challenging because the number of subintervals increases quadratically to the span of time interval. For that, we propose a novel Temporal Core Decomposition (TCD) algorithm that decrementally induces temporal k -cores from the previously induced ones and thus reduces "intra-core" redundant computation significantly. Then, we introduce an intuitive concept named Tightest Time Interval (TTI) for temporal k -core, and design an optimization technique with theoretical guarantee that leverages TTI as a key to predict which subintervals will induce duplicated k -cores and prunes the subintervals completely in advance, thereby eliminating "inter-core" redundant computation. The complexity of optimized TCD (OTCD) algorithm no longer depends on the span of query time interval but only the scale of final results, which means OTCD algorithm is scalable. Moreover, we propose a compact in-memory data structure named Temporal Edge List (TEL) to implement OTCD algorithm efficiently in physical level with bounded memory requirement. TEL organizes temporal edges in a "timeline" and can be updated instantly when new edges arrive in dynamical temporal graphs. We compare OTCD algorithm with the incremental historical k -core query on several real-world temporal graphs, and observe that OTCD algorithm outperforms it by three orders of magnitude, even though OTCD algorithm needs none precomputed index.
Junyong Yang, Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian, Mengchi Liu, Jeffrey Xu Yu
Proc. VLDB Endow.4
2023 Intent Disentanglement and Feature Self-Supervision for Novel Recommendation
abstract
One key property in recommender systems is the long-tail distribution in user-item interactions where most items only have few user feedback. Improving the recommendation of tail items can promote novelty and bring positive effects to both users and providers, and thus is a desirable property of recommender systems. Current novel recommendation methods over-emphasize the importance of tail items without differentiating the degree of users’ intent on popularity and often incur a sharp decline of accuracy. Moreover, none of existing studies has ever taken the extreme case of tail items, i.e., cold-start items without any interaction, into consideration. In this work, we first disclose the mechanism that drives a user's interaction towards popular or niche items by disentangling her intent into conformity influence (popularity) and personal interests (preference). We then present a unified end-to-end framework to simultaneously optimize accuracy and novelty targets based on the disentangled intent of popularity and that of preference. We further develop a new paradigm for novel recommendation of cold-start items which exploits the self-supervised learning technique to model the correlation between collaborative features and content features. We conduct extensive experiments on three real-world datasets. The results demonstrate that our proposed model yields significant improvements over the state-of-the-art baselines in terms of the trade-off between accuracy and novelty.
Tieyun Qian, Yile Liang, Qing Li 0001, Ke Sun 0010, Zhiyong Peng 0001
IEEE Trans. Knowl. Data Eng.1
2023 Pre-Training Across Different Cities for Next POI Recommendation
abstract
The Point-of-Interest (POI) transition behaviors could hold absolute sparsity and relative sparsity very differently for different cities. Hence, it is intuitive to transfer knowledge across cities to alleviate those data sparsity and imbalance problems for next POI recommendation. Recently, pre-training over a large-scale dataset has achieved great success in many relevant fields, like computer vision and natural language processing. By devising various self-supervised objectives, pre-training models can produce more robust representations for downstream tasks. However, it is not trivial to directly adopt such existing pre-training techniques for next POI recommendation, due to thelacking of common semantic objects (users or items) across different cities. Thus in this paper, we tackle such a new research problem ofpre-training across different citiesfor next POI recommendation. Specifically, to overcome the key challenge that different cities do not share any common object, we propose a novel pre-training model namedCATUS, by transferring thecategory-leveluniversal transition knowledge over different cities. Firstly, we build two self-supervised objectives inCATUS:next category predictionandnext POI prediction, to obtain the universal transition-knowledge across different cities and POIs. Then, we design acategory-transition oriented sampleron the data level and animplicit and explicit transfer strategyon the encoder level to enhance this transfer process. At the fine-tuning stage, we propose adistance oriented samplerto better align the POI representations into the local context of each city. Extensive experiments on two large datasets consisting of four cities demonstrate the superiority of our proposedCATUSover the state-of-the-art alternatives. The code and datasets are available at https://github.com/NLPWM-WHU/CATUS.
Ke Sun 0010, Tieyun Qian, Chenliang Li 0005, Qing Li 0001, Ming Zhong 0002, Yuanyuan Zhu 0001, Mengchi Liu
ACM Trans. Web2
2022 Enhancing Graph Convolution Network for Novel Recommendation
Tieyun Qian, Yile Liang, Ke Sun 0010, Hang Yun, Mi Zhang 0006
DASFAA (2)2
2022 Attribute Graph Neural Networks for Strict Cold Start Recommendation
abstract
Rating prediction is a classic problem underlying recommender systems. It is traditionally tackled with matrix factorization. Recently, deep learning based methods, especially graph neural networks, have made impressive progress on this problem. Despite their effectiveness, existing methods focus on modeling the user-item interaction graph. The inherent drawback of such methods is that their performance is bound to the density of the interactions, which is however usually of high sparsity. More importantly, for a strict cold start user/item that neither appears in the training data nor has any interactions in the test stage, such methods are unable to learn the preference embedding of the user/item since there is no link to this user/item in the graph. In this work, we develop a novel frameworkAttribute Graph Neural Networks(AGNN) by exploiting the attribute graph rather than the commonly used interaction graph. This leads to the capability of learning embeddings for the strict cold start users/items. Our AGNN can produce the preference embedding for a strict cold user/item by learning on the distribution of attributes with an extended variational auto-encoder (eVAE) structure. Moreover, we propose a new graph neural network variant, i.e., gated-GNN, to effectively aggregate various attributes of different modalities in a neighborhood. Empirical results on three real-world datasets demonstrate that our model yields significant improvements for strict cold start recommendations and outperforms or matches the state-of-the-art performance in the warm start scenario.
Tieyun Qian, Yile Liang, Qing Li 0001, Hui Xiong 0001
IEEE Trans. Knowl. Data Eng.1
2021 Enhancing Domain-Level and User-Level Adaptivity in Diversified Recommendation
abstract
Recommender systems are playing a vital role in online platforms due to the ability of incorporating users' personal tastes. Beyond accuracy, diversity has been recognized as a key factor to broaden users' horizons as well as to promote enterprises' sales. However, the trade-off between accuracy and diversity remains to be a big challenge. More importantly, none of existing methods has explored the domain and user biases toward diversity.
Yile Liang, Tieyun Qian, Qing Li 0001, Hongzhi Yin
SIGIR2
2021 Context-aware seq2seq translation model for sequential recommendation
Ke Sun 0010, Tieyun Qian, Ming Zhong 0002
Inf. Sci.2
2021 Business location planning based on a novel geo-social influence diffusion model
Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian, Jianxin Li 0001
Inf. Sci.4
2020 Adversarial Generation of Target Review for Rating Prediction
Huilin Yu, Tieyun Qian, Yile Liang, Bing Liu 0001
DASFAA (2)2
2020 AGTR: Adversarial Generation of Target Review for Rating Prediction
abstract
Abstract Recent years have witnessed a growing trend of utilizing reviews to improve the performance and interpretability of recommender systems. Almost all existing methods learn the latent representations from the user’s and the item’s historical reviews and then combine these two representations for rating prediction. The fatal limitation in these methods is that they are unable to utilize the most predictive review of the target user for the target item since such a review is not available at test time. In this paper, we propose a novel recommendation model, called AGTR, which cangenerate the unseen target review with adversarial training for rating prediction. To this end, we develop a unified framework to combinethe rating tailored generative adversarial netsfor synthetic review generation andthe neural latent factor moduleusing the generated target review along with historical reviews for rating prediction. Extensive experiments on four real-world datasets demonstrate that our model achieves the state-of-the-art performance in both rating prediction and review generation tasks.
Huilin Yu, Tieyun Qian, Yile Liang, Bing Liu 0001
Data Sci. Eng.2
2020 Relation constrained attributed network embedding
Tieyun Qian
Inf. Sci.2
2020 Generating behavior features for cold-start spam review detection with adversarial learning
Xiaoya Tang, Tieyun Qian, Zhenni You
Inf. Sci.2
2019 What Can History Tell Us?
abstract
Recommendation systems have been widely applied to many E-commerce and online social media platforms. Recently, sequential item recommendation, especially session-based recommendation, has aroused wide research interests. However, existing sequential recommendation approaches either ignore the historical sessions or consider all historical sessions without any distinction that whether the historical sessions are relevant or not to the current session, which motivates us to distinguish the effect of each historical session and identify relevant historical sessions for recommendation. In light of this, we propose a novel deep learning based sequential recommender framework for session-based recommendation, which takes Nonlocal Neural Network and Recurrent Neural Network as the main building blocks. Specifically, we design a two-layer nonlocal architecture to identify historical sessions that are relevant to the current session and learn the long-term user preferences mostly from these relevant sessions. Besides, we also design a gated recurrent unit (GRU) enhanced by the nonlocal structure to learn the short-term user preferences from the current session. Finally, we propose a novel approach to integrate both long-term and short-term user preferences in a unified way to facilitate training the whole recommender model in an end-to-end manner. We conduct extensive experiments on two widely used real-world datasets, and the experimental results show that our model achieves significant improvements over the state-of-the-art methods.
Ke Sun 0010, Tieyun Qian, Hongzhi Yin, Tong Chen 0005, Ling Chen 0006
CIKM2
2019 Towards both Local and Global Query Result Diversification
Ming Zhong 0002, Yuanyuan Zhu 0001, Tieyun Qian, Jianxin Li 0001
DASFAA (2)5
2019 Keyword Search Based Mashup Construction with Guaranteed Diversity
Ming Zhong 0002, Jian Wang 0018, Tieyun Qian
DEXA (2)4
2019 Aspect Aware Learning for Aspect Category Sentiment Analysis
abstract
Aspect category sentiment analysis (ACSA) is an underexploited subtask in aspect level sentiment analysis. It aims to identify the sentiment of predefined aspect categories. The main challenge in ACSA comes from the fact that the aspect category may not occur in the sentence in most of the cases. For example, the review “ they have delicious sandwiches ” positively talks about the aspect category “ food ” in an implicit manner. In this article, we propose a novel aspect aware learning (AAL) framework for ACSA tasks. Our key idea is to exploit the interaction between the aspect category and the contents under the guidance of both sentiment polarity and predefined categories. To this end, we design a two-way memory network for integrating AAL into the framework of sentiment classification. We further present two algorithms to incorporate the potential impacts of aspect categories. One is to capture the correlations between aspect terms and the aspect category like “sandwiches” and “food.” The other is to recognize the aspect category for sentiment representations like “food” for “delicious.” We conduct extensive experiments on four SemEval datasets. The results reveal the essential role of AAL in ACSA by achieving the state-of-the-art performance.
Peisong Zhu, Zhuang Chen 0002, Haojie Zheng, Tieyun Qian
ACM Trans. Knowl. Discov. Data4
2019 Spatiotemporal Representation Learning for Translation-Based POI Recommendation
abstract
The increasing proliferation of location-based social networks brings about a huge volume of user check-in data, which facilitates the recommendation of points of interest (POIs). Time and location are the two most important contextual factors in the user’s decision-making for choosing a POI to visit. In this article, we focus on the spatiotemporal context-aware POI recommendation, which considers the joint effect of time and location for POI recommendation. Inspired by the recent advances in knowledge graph embedding, we propose a spatiotemporal context-aware and translation-based recommender framework (STA) to model the third-order relationship among users, POIs, and spatiotemporal contexts for large-scale POI recommendation. Specifically, we embed both users and POIs into a “transition space” where spatiotemporal contexts (i.e., a < time, location > pair) are modeled as translation vectors operating on users and POIs. We further develop a series of strategies to exploit various correlation information to address the data sparsity and cold-start issues for new spatiotemporal contexts, new users, and new POIs. We conduct extensive experiments on two real-world datasets. The experimental results demonstrate that our STA framework achieves the superior performance in terms of high recommendation accuracy, robustness to data sparsity, and effectiveness in handling the cold-start problem.
Tieyun Qian, Nguyen Quoc Viet Hung, Hongzhi Yin
ACM Trans. Inf. Syst.1
2018 "Bridge": Enhanced Signed Directed Network Embedding
abstract
Signed directed networks with positive or negative links convey rich information such as like or dislike, trust or distrust. Existing work of sign prediction mainly focuses on triangles (triadic nodes) motivated by balance theory to predict positive and negative links. However, real-world signed directed networks can contain a good number of "bridge'' edges which, by definition, are not included in any triangles. Such edges are ignored in previous work, but may play an important role in signed directed network analysis.%Such edges serve as fundamental building blocks and may play an important role in signed network analysis.
Tieyun Qian, Huan Liu 0001, Ke Sun 0010
CIKM2
2018 Sample Location Selection for Efficient Distance-Aware Influence Maximization in Geo-Social Networks
Ming Zhong 0002, Yuanyuan Zhu 0001, Jianxin Li 0001, Tieyun Qian
DASFAA (1)5
2018 BASSI: Balance and Status Combined Signed Network Embedding
Tieyun Qian, Ming Zhong 0002, Xuhui Li 0001
DASFAA (1)2
2017 Exploit Label Embeddings for Enhancing Network Classification
Tieyun Qian, Ming Zhong 0002, Xuhui Li 0001
DEXA (2)2
2016 Fast Rare Category Detection Using Nearest Centroid Neighborhood
Hao Huang 0001, Yunjun Gao, Tieyun Qian, Liang Hong 0001, Zhiyong Peng 0001
APWeb (1)4
2016 Modeling for Noisy Labels of Crowd Workers
Qian Yan 0001, Hao Huang 0001, Yunjun Gao, Chen Ying, Qingyang Hu, Tieyun Qian, Qinming He
APWeb (2)6
2016 Inferring Lurkers' Gender by Their Interest Tags
Peisong Zhu, Tieyun Qian, Zhenni You, Xuhui Li 0001
DEXA (2)2
2016 Fast and accurate identification of implicit enterprise users in social media
abstract
There are a large number of implicit enterprise users in social media, who register as ordinary users but act like enterprise ones. It is a fundamental task to eliminate this type of users from the user set since their existing will severely affect the performance of many applications like network demographic investigation and targeted advertisement. Despite of its importance, this problem is surprisingly unexplored.
Zhenni You, Diqian Wu, Tieyun Qian
WebDB4
2016 Identifying Implicit Enterprise Users from the Imbalanced Social Data
Zhenni You, Tieyun Qian, Baochao Zhang
WISE (2)2
2015 Age Detection for Chinese Users in Weibo
Li Chen 0031, Tieyun Qian, Fei Wang 0082, Zhenni You, Qingxi Peng, Ming Zhong 0002
WAIM2
2015 LDM: A DTD Schema Mapping Language Based on Logic Patterns
Xuhui Li 0001, Yijun Guan, Mengchi Liu, Ming Zhong 0002, Tieyun Qian
WAIM6
2015 Coherent Topic Hierarchy: A Strategy for Topic Evolutionary Analysis on Microblog Feeds
Xuhui Li 0001, Min Peng 0002, Tieyun Qian, Jimin Huang, Jiping Liu, Ri Hong, Pinglan Liu
WAIM5
2014 Co-training on authorship attribution with very fewlabeled examples: methods vs. views
abstract
Authorship attribution (AA) aims to identify the authors of a set of documents. Traditional studies in this area often assume that there are a large set of labeled documents available for training. However, in the real life, it is hard or expensive to collect a large set of labeled data. For example, in the online review domain, most reviewers (authors) only write a few reviews, which are not enough to serve as the training data for accurate classification. In this paper, we present a novel two-view co-training framework to iteratively identify the authors of a few unlabeled data to augment the training set. The key idea is to first represent each document as several distinct views, and then a co-training technique is adopted to exploit the large amount of unlabeled documents. Starting from 10 training texts per author, we systematically evaluate the effectiveness of co-training for authorship attribution with limited labeled data. Two methods and three views are investigated: logistic regression (LR) and support vector machines (SVM) methods, and character, lexical, and syntactic views. The experimental results show that LR is particularly effective for improving co-training in AA, and the lexical view performs the best among three views when combined with a LR classifier. Furthermore, the co-training framework does not make much difference between one classifier from two views and two classifiers from one view. Instead, it is the learning approach and the view that plays a critical role.
Tieyun Qian, Bing Liu 0001, Ming Zhong 0002
SIGIR1
2014 Authorship Attribution with Very Few Labeled Data: A Co-training Approach
Mengdi Fan, Tieyun Qian, Li Chen 0031, Ming Zhong 0002
WAIM2
2013 Detecting Professional Spam Reviewers
Junlong Huang, Tieyun Qian, Ming Zhong 0002, Qingxi Peng
ADMA (2)2
2013 A Local Greedy Search Method for Detecting Community Structure in Weighted Social Networks
Tieyun Qian
ADMA (1)2
2013 Early prediction on imbalanced multivariate time series
abstract
Multivariate time series (MTS) classification is an important topic in time series data mining, and lots of efficient models and techniques have been introduced to cope with it. However, early classification on imbalanced MTS data largely remains an open problem. To deal with this issue, we adopt a multiple under-sampling and dynamical subspace generation method to obtain initial training data, and each training data is used to learn a base learner. Finally, an ensemble classifier is introduced for early classification on imbalanced MTS data. Experimental results show that our proposed methods can achieve effective early prediction on imbalanced MTS data.
Tieyun Qian
CIKM3
2013 MVP Index: Towards Efficient Known-Item Search on Large Graphs
Ming Zhong 0002, Mengchi Liu, Zhifeng Bao, Xuhui Li 0001, Tieyun Qian
DASFAA (1)5
2012 Leveraging Network Structure for Incremental Document Clustering
Tieyun Qian, Jianfeng Si, Qing Li 0001
APWeb1
2011 Finding Relevant Papers Based on Citation Relations
Yicong Liang, Qing Li 0001, Tieyun Qian
WAIM3
2010 Refining Graph Partitioning for Social Network Clustering
Tieyun Qian, Yang Yang 0009, Shuo Wang 0011
WISE1
2009 What's behind topic formation and development: a perspective of community core groups
abstract
Over the past several years, there has been a great interest in topic detection and tracking (TDT). Recently, analyzing general research trend from the huge amount of history documents also arouses considerable attention. However, existing work on TDT mainly focuses on overall trend analysis, and is unable to address questions such as "what determines the evolution of a topic?" and "when and how does a new topic get formed?".
Tieyun Qian, Qing Li 0001, Bing Liu 0001, Hui Xiong 0001, Jaideep Srivastava, Phillip C.-Y. Sheu
CIKM1
2009 Simultaneously Finding Fundamental Articles and New Topics Using a Community Tracking Method
Tieyun Qian, Jaideep Srivastava, Zhiyong Peng 0001, Phillip C.-Y. Sheu
PAKDD1
2008 Efficient strategies for tough aggregate constraint-based sequential pattern mining
Enhong Chen, Huanhuan Cao, Qing Li 0001, Tieyun Qian
Inf. Sci.4
2007 On the strength of hyperclique patterns for text categorization
Tieyun Qian, Hui Xiong 0001, Yuanzhen Wang, Enhong Chen
Inf. Sci.1
2006 Adapting association patterns for text categorization: weaknesses and enhancements
abstract
The use of association patterns for text categorization has attracted great interest and a variety of useful methods have been developed. However, the key characteristics of pattern-based text categorization remain unclear. Indeed, there are still no concrete answers for the following two questions: First, what kind of association patterns are the best candidate for pattern-based text categorization? Second, what is the most desirable way to use patterns for text categorization? In this paper, we focus on answering the above two questions. Specifically, we show that hyperclique patterns are more desirable than frequent patterns for text categorization. Along this line, we develop an algorithm for text categorization using hyperclique patterns. The experimental results show that our method provides better performance than state-of-the-art methods in terms of both computational performance and classification accuracy.
Tieyun Qian, Hui Xiong 0001, Yuanzhen Wang, Enhong Chen
CIKM1
2005 2-PS Based Associative Text Classification
Tieyun Qian, Yuanzhen Wang, Jianlin Feng
DaWaK1