Jie Yang 0009

dblp:12/1198-9 · also Jack Yang 0003, Jie (Jack) Yang · DBLP profile ↗
← Back
13ranked-venue papers in the field
2as first author
12since 2021 · last 2025
0000-0003-1317-8142ORCID · conflict

Domains — venue-derived; a paper can count in several

Data Mining & Knowledge Discovery · 4 (1 first)Big Data, Cloud & Distributed Data Systems · 3Information Retrieval & Web Search · 2Other / Interdisciplinary · 2Database Systems & Data Management · 1Knowledge Engineering, Semantic Web & Information Systems · 1 (1 first)
YearPublicationVenuePosition
2025 Every Lie Has a Grain of Truth: Disentangling Deception from Authentic Content for Fake News Detection
Junping Liu, Zhenhao Hu, Xinrong Hu, Wangli Yang, Wanqing Li 0009, Jie Yang 0009, Yi Guo 0001
IEEE Big Data6
2025 Impact-Aware Retrieval Defense: Mitigating Word Substitution Ranking Attacks for Enhanced Stability
Junping Liu, Xinrong Hu, Wangli Yang, Wanqing Li 0009, Jie Yang 0009, Wenbin Zhang 0002, Yi Guo 0001
IEEE Big Data6
2025 Negative-Free Graph Contrastive Learning for Recommendation
abstract
Graph Contrastive Learning (GCL) emerges as a powerful approach in recommendation systems, leveraging graph structures to learn effective representations. However, existing contrastive sampling strategies often introduce unintended biases, most notably, the misclassification of genuine positive samples as negatives, which undermines representation quality and overall recommendation performance. Accordingly, this paper revisits the conventional contrastive sampling and introduces Negative-Free Sampling for Graph Contrastive Learning (NFS). NFS adopts a two-stage sampling strategy that selectively identifies and utilizes only positive instances during training. By removing reliance on negative samples, it effectively mitigates misclassification bias and improves the semantic alignment between related representations. In addition, a comprehensive theoretical analysis is also provided to establish the robustness of NFS against representation collapse. Experimental results on three benchmarks demonstrate that NFS consistently outperforms or performs state-of-the-art methods, achieving up to a 14.2% relative improvement across evaluated datasets. In addition, a detailed ablation study is also provided to examine how exclusively leveraging positive samples contributes to the efficiency of GCL. The results further demonstrate the plug-and-play nature of the proposed method and its resilience to noisy data.
Junping Liu, Mingchao Yu, Xinrong Hu, Jie Yang 0009, Yi Guo 0001, Wanqing Li 0001, Wenbin Zhang 0002
ICDM4
2024 Extractive Question Answering with Contrastive Puzzles and Reweighted Clues
Jie Yang 0009, Wanqing Li 0001
ICDAR (6)2
2024 ConClue: Conditional Clue Extraction for Multiple Choice Question Answering
Wangli Yang, Jie Yang 0009, Wanqing Li 0009, Yi Guo 0001
ICDAR (6)2
2024 SCAD: Subspace Clustering based Adversarial Detector
abstract
Adversarial examples pose significant challenges for Natural Language Processing (NLP) model robustness, often causing notable performance degradation. While various detection methods have been proposed with the aim of differentiating clean and adversarial inputs, they often require fine-tuning with ample data, which is problematic for low-resource scenarios. To alleviate this issue, a Subspace Clustering based Adversarial Detector (termed SCAD) is proposed in this paper, leveraging a union of subspaces to model the clean data distribution. Specifically, SCAD estimates feature distribution across semantic subspaces, assigning unseen examples to the nearest one for effective discrimination. The construction of semantic subspaces does not require many observations and hence ideal for the low-resource setting.
Xinrong Hu, Wushuan Chen, Jie Yang 0009, Yi Guo 0001, Xun Yao, Bangchao Wang, Junping Liu
WSDM3
2024 COTER: Conditional Optimal Transport meets Table Retrieval
abstract
Ad hoc table retrieval refers to the task of performing semantic matching between given queries and candidate tables. In recent years, the approach to addressing this retrieval task has undergone significant shifts, transitioning from utilizing hand-crafted features to leveraging the power of Pre-trained Language Models (PLMs). However, key challenges arise when candidate tables contain shared items, and/or queries may refer to only a subset of table items rather than the entire one. Existing models often struggle to distinguish the most informative items and fail to accurately identify the relevant items required to match with the query.
Xun Yao, Xinrong Hu, Jie Yang 0009, Yi Guo 0001, Daniel (Dianliang) Zhu
WSDM4
2023 Improving Adversarially Robust Sequential Recommendation through Generalizable Perturbations
abstract
Sequential recommendation is of great importance for a variety of purposes, such as application engineering, resource optimization, and marketing. Yet, existing sequence-based recommendation models are susceptible to adversarial attacks, which aim to perturb input sequences and mislead trained models, resulting in incorrect predictions. Defense methods are accordingly adopted to enhance model robustness. Nevertheless, these methods encounter challenges, such as error propagation (from the model output to generate adversarial samples), the high system complexity, and the difficulty of maintaining the model generalizability. To bridge this gap, this paper introduces a simple yet effective adversarial defense algorithm, termed Perturbation-Driven Sequential Recommendation (PDSR). In the training process, PDSR leverages a simple perturbation-generation module to create adversarial samples, eliminating the need for gradient estimation, thus streamlining the process. Additionally, it also incorporates a robust encoder designed to increase tolerance towards representation variations by ensuring alignment between original and perturbed representations, thereby boosting model generalizability. Comprehensive experiments are conducted based on a combination of five benchmark datasets, two attack methods, and four sequential recommendation models. When compared to four state-of-the-art defense baselines, PDSR demonstrates notable improvements in defense performance.
Xun Yao, Ruyi He, Xinrong Hu, Jie Yang 0009, Yi Guo 0001, Zijian Huang 0012
IEEE Big Data4
2023 MIRS: [MASK] Insertion Based Retrieval Stabilizer for Query Variations
Junping Liu, Mingkang Gong, Xinrong Hu, Jie Yang 0009, Yi Guo 0001
DEXA (1)4
2023 Towards Robust Token Embeddings for Extractive Question Answering
Xun Yao, Junlong Ma, Xinrong Hu, Jie Yang 0009, Yi Guo 0001, Junping Liu
WISE4
2023 CREAM: Named Entity Recognition with Concise query and REgion-Aware Minimization
Xun Yao, Xinrong Hu, Jie Yang 0009, Yi Guo 0001
WISE4
2022 Low-rank and sparse representation based learning for cancer survivability prediction
Jie Yang 0009, Jun Ma 0002, Khin Than Win, Junbin Gao, Zhenyu Yang 0004
Inf. Sci.1
2014 A sparsity-based training algorithm for Least Squares SVM
abstract
We address the training problem of the sparse Least Squares Support Vector Machines (SVM) using compressed sensing. The proposed algorithm regards the support vectors as a dictionary and selects the important ones that minimize the residual output error iteratively. A measurement matrix is also introduced to reduce the computational cost. The main advantage is that the proposed algorithm performs model training and support vector selection simultaneously. The performance of the proposed algorithm is tested with several benchmark classification problems in terms of number of selected support vectors and size of the measurement matrix. Simulation results show that the proposed algorithm performs competitively when compared to existing methods.
Jie Yang 0009, Jun Ma 0002
CIDM1