Jianyang Qin

dblp:291/9397 · DBLP profile ↗
← Back
16ranked-venue papers
6as first author
14since 2021 · last 2026
0000-0003-1444-430XORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 3 first-author · 6 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 4 since 2021
YearPublicationVenuePosition
2026 Improving Heterogeneous Graph Contrastive Learning Robustness via Hierarchical Vulnerability Protection
abstract
Recently, Heterogeneous Graph Contrastive Learning (HGCL) has received significant attention due to its impressive capability to represent heterogeneous graphs without detailed annotations. However, the inherent fragility of heterogeneous graph structures makes HGCL vulnerable to perturbation attacks. Most existing defense works for heterogeneous graphs primarily focus on supervised scenarios, which protect all nodes equally via structural pruning. This defensive mechanism can result in insufficient structure information for HGCL, thus degrading performance in self-supervised scenarios without labels. In this paper, we argue that some nodes are more susceptible to attacks, and the influence of the perturbation attack will accumulate across layers during representation aggregation. To tackle these problems, we propose a novel Heterogeneous Graph Contrastive Learning with Hierarchical Vulnerability Protection (HVP-HGCL), which identifies the most vulnerable nodes to perturbation attack and protects them across different aggregation layers to improve the robustness of HGCL. Specifically, we first design the Vulnerability Detection (VD) based on the HGCL framework to determine which nodes are more sensitive to attack in self-supervised scenarios. Subsequently, we propose a simple but efficient Hierarchical Protection (HP) to safeguard those vulnerable nodes from attack noise during different layers. Combining the above two modules, HVP-HGCL can not only improve the robustness of HGCL but also ensure sufficient structural information for effective contrastive learning. Extensive experiments demonstrate that HVP-HGCL improves robustness against adversarial attacks and achieves competitive performance on downstream tasks.
Jinhao Cui, Jianyang Qin, Lingzhi Wang 0001, Cuiyun Gao 0001, Qing Liao 0001
KDD (1)3
2025 Joint Scheduling of Causal Prompts and Tasks for Multi-Task Learning
abstract
Multi-task prompt learning has emerged as a promising technique for fine-tuning pre-trained Vision-Language Models (VLMs) to various downstream tasks. However, existing methods ignore challenges caused by spurious correlations and dynamic task relationships, which may reduce the model performance. To tackle these challenges, we propose JSCPT, a novel approach for Joint Scheduling of Causal Prompts and Tasks to enhance multi-task prompt learning. Specifically, we first design a Multi-Task Vison-Language Prompt (MTVLP) model, which learns task-shared and task-specific vison-language prompts and selects useful prompt features via causal intervention, alleviating spurious correlations. Then, we propose the task-prompt scheduler that models inter-task affinities and assesses the causal effect of prompt features to optimize the multi-task prompt learning process. Finally, we formulate the scheduler and the multi-task prompt learning process as a bi-level optimization problem to optimize prompts and tasks adaptively. In the lower optimization, MTVLP is updated with the scheduled gradient, while in the upper optimization, the scheduler is updated with the implicit gradient. Extensive experiments show the superiority of our proposed JSCPT approach over several baselines in terms of multi-task prompt learning for pre-trained VLMs.
Jianyang Qin, Jinhao Cui, Qing Liao 0001
CVPR2
2025 Label Prediction Inherited Hashing for Cross-Modal Retrieval: Applying Supervised Hashing to Unsupervised Tasks
abstract
Supervised cross-modal hashing has achieved remarkable progress in retrieving related items across different modalities. However, in practical applications, a significant portion of data remains unlabeled, such as online data on websites, which must be included for effective retrieval. To address this challenge, while maintaining the high accuracy and efficiency of supervised methods, few works have attempted to adapt existing supervised techniques to handle unsupervised tasks through a general modular approach. To this end, we introduce a novel cross-modal hashing method, termed Label Prediction Inherited Hashing (LPIH). Initially, LPIH leverages labeled data to learn high-quality general label functions using supervised methods. Subsequently, it inherits the existing hash codes from existing supervised methods to further refine the pseudo-label information. Finally, LPIH integrates the refined pseudo-label information with the existing hash functions to learn new hash functions specifically tailored for unsupervised tasks. Extensive experimental results on three public datasets demonstrate the superior performance of LPIH compared to state-of-the-art (SOTA) cross-modal hashing methods. Specifically, LPIH achieves an average precision improvement of 5% over SOTA methods, highlighting its effectiveness in bridging the gap between supervised and unsupervised learning in the context of cross-modal retrieval.
Kaihang Jiang, Wai Keung Wong, Jianyang Qin, Xiaozhao Fang, Jie Wen 0001, Bingzhi Chen, Hongbo Gao 0001
ACM Multimedia3
2025 Turning the Tables: Enabling Backward Transfer via Causal-Aware LoRA in Continual Learning
abstract
Current parameter-efficient fine-tuning (PEFT) methods have shown superior performance in continual learning. However, most existing PEFT-based methods focus on mitigating catastrophic forgetting by limiting modifications to the old task model caused by new tasks. This hinders backward knowledge transfer, as when new tasks have a strong positive correlation with old tasks, appropriately training on new tasks can transfer beneficial knowledge to old tasks. Critically, achieving backward knowledge transfer faces two fundamental challenges: (1) some parameters may be ineffective on task performance, which constrains the task solution space and model capacity; (2) since old task data are inaccessible, modeling task correlation via shared data is infeasible. To address these challenges, we propose CaLoRA, a novel \textbf{c}ausal-\textbf{a}ware \textbf{lo}w-\textbf{r}ank \textbf{a}daptation framework that is the first PEFT-based continual learning work with backward knowledge transfer. Specifically, we first propose \textbf{p}ar\textbf{a}meter-level \textbf{c}ounterfactual \textbf{a}ttribution (PaCA) that estimates the causal effect of LoRA parameters via counterfactual reasoning, identifying effective parameters from a causal view. Second, we propose \textbf{c}ross-t\textbf{a}sk \textbf{g}radient \textbf{a}daptation (CaGA) to quantify task correlation by gradient projection and evaluate task affinity based on gradient similarity. By incorporating causal effect, task correlation, and affinity, CaGA adaptively adjusts task gradients, facilitating backward knowledge transfer without relying on data replay. Extensive experiments across multiple benchmarks and continual learning settings show that CaLoRA outperforms state-of-the-art methods. In particular, CaLoRA better mitigates catastrophic forgetting by enabling positive backward knowledge transfer.
Runze Ye, Jianyang Qin, Jinhao Cui, Lingzhi Wang 0001, Qing Liao 0001
NeurIPS3
2025 Bridging Time and Linguistics: LLMs as Time Series Analyzer through Symbolization and Segmentation
abstract
Recent studies reveal that Large Language Models (LLMs) exhibit strong sequential reasoning capabilities, allowing them to replace specialized time-series models and serve as foundation models for complex time-series analysis. To activate the capabilities of LLMs for time-series tasks, numerous studies have attempted to bridge the gap between time series and linguistics by aligning textual representations with time-series patterns. However, it is a non-trivial endeavor to losslessly capture the infinite time-domain variability using natural language, leading to suboptimal alignment performance. Beyond representation, contextual differences, where semantics in time series are conveyed by consecutive points, unlike in text by individual tokens, are often overlooked by existing methods. To address these, we propose S$^2$TS-LLM, a simple yet effective framework to repurpose LLMs for universal time series analysis through the following two main paradigms: (i) a spectral symbolization paradigm transforms time series into frequency-domain representations characterized by a fixed number of components and prominent amplitudes, which enables a limited set of symbols to effectively abstract key frequency features; (ii) a contextual segmentation paradigm partitions the sequence into blocks based on temporal patterns and reassigns positional encodings accordingly, thereby mitigating the structural mismatch between time series and natural language. Together, these paradigms bootstrap the LLMs' perception of temporal patterns and structures, effectively bridging time series and linguistics. Extensive experiments show that S$^2$TS-LLM can serve as a powerful time series analyzer, outperforming state-of-the-art methods across time series tasks.
Jianyang Qin, Jinhao Cui, Lingzhi Wang 0001, Zhao Liu 0006, Qing Liao 0001
NeurIPS1
2025 TaylorS: A Multi-Order Expansion Structure for Urban Spatio-Temporal Forecasting
abstract
Although a variety of models have been proposed for urban spatio-temporal forecasting, most existing forecasting models are developed manually for specific tasks. By investigating the correlation between multi-order derivative and spatio-temporal data, we propose a generic yet simple plug-in structure, namedTaylorS, to improve the performance and generalization of existing forecasting models. The TaylorS converts the non-linear regression problem into a multi-order non-linear approximation problem by plugging a Taylor expansion into the forecasting task. To achieve this, we design a two-step training framework, including a training step and an adjusting step. During training, we train a given forecasting model as a base model to be equipped with prior knowledge. During adjusting, we fine-tune the base model while plugging an adjustment model into the base model. The adjustment model, as a multi-order expansion, takes the multi-order derivative of data to evaluate data uncertainty for further forecasting approximation and adjustment. Extensive experimental results demonstrate that the proposed TaylorS framework can consistently improve the performance of existing state-of-the-art methods and generalize these methods to different forecasting tasks.
Jianyang Qin, Yan Jia 0001, Binxing Fang, Qing Liao 0001
IEEE Trans. Knowl. Data Eng.1
2025 Heterogeneous Pairwise-Semantic Enhancement Hashing for Large-Scale Cross-Modal Retrieval
abstract
Cross-modal hash learning has drawn widespread attention for large-scale multimodal retrieval because of its stability and efficiency in approximate similarity searches. However, most existing cross-modal hashing approaches employ discrete label-guided information to coarsely reflect intra- and intermodality correlations, making them less effective to measuring the semantic similarity of data with multiple modalities. In this paper, we propose a new heterogeneous pairwise-semantic enhancement hashing (HPsEH) for large-scale cross-modal retrieval by distilling higher-level pairwise-semantic similarity from supervision information. First, we adopt a supervised self-expression to learn a data-specific quantified semantic matrix, which uses real values to measure both the similarity and dissimilarity ranks of paired instances, such that the intrinsic semantics of the data can be well captured. Then, we fuse the label-based information and quantified semantic similarity to collaboratively learn the hash codes of multimodal data, such that both the intermodality consistency and modality-specific features can be simultaneously obtained during hash code learning. Moreover, we employ effective iterative optimization to address the discrete binary solution and massive pairwise matrix calculation, making the HPsEH scalable to large-scale datasets. Extensive experimental results on three widely used datasets demonstrate the superiority of our proposed HPsEH method over most state-of-the art approaches.
Wai Keung Wong, Lunke Fei, Jianyang Qin, Shuping Zhao, Jie Wen 0001
IEEE Trans. Multim.3
2025 Random Online Hashing for Cross-Modal Retrieval
abstract
In the past decades, supervised cross-modal hashing methods have attracted considerable attentions due to their high searching efficiency on large-scale multimedia databases. Many of these methods leverage semantic correlations among heterogeneous modalities by constructing a similarity matrix or building a common semantic space with the collective matrix factorization method. However, the similarity matrix may sacrifice the scalability and cannot preserve more semantic information into hash codes in the existing methods. Meanwhile, the matrix factorization methods cannot embed the main modality-specific information into hash codes. To address these issues, we propose a novel supervised cross-modal hashing method called random online hashing (ROH) in this article. ROH proposes a linear bridging strategy to simplify the pair-wise similarities factorization problem into a linear optimization one. Specifically, a bridging matrix is introduced to establish a bidirectional linear relation between hash codes and labels, which preserves more semantic similarities into hash codes and significantly reduces the semantic distances between hash codes of samples with similar labels. Additionally, a novel maximum eigenvalue direction (MED) embedding method is proposed to identify the direction of maximum eigenvalue for the original features and preserve critical information into modality-specific hash codes. Eventually, to handle real-time data dynamically, an online structure is adopted to solve the problem of dealing with new arrival data chunks without considering pairwise constraints. Extensive experimental results on three benchmark datasets demonstrate that the proposed ROH outperforms several state-of-the-art cross-modal hashing methods.
Kaihang Jiang, Wai Keung Wong, Xiaozhao Fang, Jiaxing Li 0009, Jianyang Qin, Shengli Xie 0001
IEEE Trans. Neural Networks Learn. Syst.5
2024 MUSE-Net: Disentangling Multi-Periodicity for Traffic Flow Forecasting
abstract
Accurate forecasting of traffic flow plays a crucial role in building smart cities in the new era. Previous work has achieved success in learning inherent spatial and temporal patterns of traffic flow. However, existing works investigated the multiple periodicities (e.g., hourly, daily, and weekly) of traffic via entanglement learning, which has not yet dealt with distribution shift and interaction shift problems in traffic flow. In this paper, we propose a novel disentanglement learning network, called MUSE-Net, to tackle the limitations of entanglement learning by simultaneously factorizing the exclusiveness and interaction of multi-periodic patterns in traffic flow. Grounded in the theory of mutual information, we first learn and dis-entangle exclusive and interactive representations of traffics from multi-periodic patterns. Then, we utilize semantic-pushing and semantic-pulling regularizations to encourage the learned representations to be independent and informative. Moreover, we derive a lower bound estimator to tractably optimize the disentanglement problem with multiple variables and propose a joint training model for traffic forecasting. Extensive experimental results on several real-world traffic datasets demonstrate the effectiveness of the proposed framework. The code is available at: https://github.com/JianyangQin/MUSE-Net.
Jianyang Qin, Yan Jia 0001, Yongxin Tong, Heyan Chai 0001, Ye Ding 0002, Xuan Wang 0002, Binxing Fang, Qing Liao 0001
ICDE1
2023 Sparse Graph Hashing with Spectral Regression
Jianyang Qin, Lunke Fei, Shuping Zhao, Jie Wen 0001
CGI (4)2
2023 Temporal-Relational Matching Network for Few-Shot Temporal Knowledge Graph Completion
Xing Gong, Jianyang Qin, Heyan Chai 0001, Ye Ding 0002, Yan Jia 0001, Qing Liao 0001
DASFAA (2)2
2023 Adaptive Multi-hop Neighbor Selection for Few-Shot Knowledge Graph Completion
Xing Gong, Jianyang Qin, Ye Ding 0002, Yan Jia 0001, Qing Liao 0001
ICONIP (11)2
2022 Joint Specifics and Consistency Hash Learning for Large-Scale Cross-Modal Retrieval
abstract
With the dramatic increase in the amount of multimedia data, cross-modal similarity retrieval has become one of the most popular yet challenging problems. Hashing offers a promising solution for large-scale cross-modal data searching by embedding the high-dimensional data into the low-dimensional similarity preserving Hamming space. However, most existing cross-modal hashing usually seeks a semantic representation shared by multiple modalities, which cannot fully preserve and fuse the discriminative modal-specific features and heterogeneous similarity for cross-modal similarity searching. In this paper, we propose a joint specifics and consistency hash learning method for cross-modal retrieval. Specifically, we introduce an asymmetric learning framework to fully exploit the label information for discriminative hash code learning, where 1) each individual modality can be better converted into a meaningful subspace with specific information, 2) multiple subspaces are semantically connected to capture consistent information, and 3) the integration complexity of different subspaces is overcome so that the learned collaborative binary codes can merge the specifics with consistency. Then, we introduce an alternatively iterative optimization to tackle the specifics and consistency hashing learning problem, making it scalable for large-scale cross-modal retrieval. Extensive experiments on five widely used benchmark databases clearly demonstrate the effectiveness and efficiency of our proposed method on both one-cross-one and one-cross-two retrieval tasks.
Jianyang Qin, Lunke Fei, Zheng Zhang 0006, Jie Wen 0001, Yong Xu 0001, David Zhang 0001
IEEE Trans. Image Process.1
2021 Scalable Discriminative Discrete Hashing For Large-Scale Cross-Modal Retrieval
abstract
Cross-modal hashing has received increasing research attentions due to its less storage and efficient retrieval. However, most existing cross-modal hashing methods focus only on exploring multi-modal information, while underestimate the significance of local and Euclidean structure information on the hashing learning procedure. In this paper, we propose a supervised discrete-based cross-modal hashing method, named Scalable Discriminative Discrete Hashing (SDDH), for cross-modal retrieval, where 1) the discrete hash codes are directly obtained by multi-modal features and semantic labels so that the quantization errors are dramatically reduced, and 2) the discrete hash codes simultaneously preserve the heterogeneous similarity and manifold information in the original space by employing matrix factoring with orthogonal and balanced constraints. Moreover, an efficient optimization is introduced to tackle the discrete solution, which makes the SDDH scalable to large-scale cross-modal retrieval. Empirical results on three widely-used benchmark databases clearly demonstrate the effectiveness and efficiency of the proposed method in comparison with state-of-the-arts.
Jianyang Qin, Lunke Fei, Jian Zhu 0001, Jie Wen 0001, Chunwei Tian, Shuai Wu 0001
ICASSP1
2020 Jointly Learning Multiple Curvature Descriptor for 3D Palmprint Recognition
abstract
3D palmprint-based biometric recognition has drawn growing research attention due to its several merits over 2D counterpart such as robust structural measurement of a palm surface and high anti-counterfeiting capability. However, most existing 3D palmprint descriptors are hand-crafted that usually extract stationary features from 3D palmprint images. In this paper, we propose a feature learning method to jointly learn compact curvature feature descriptor for 3D palmprint recognition. We first form multiple curvature data vectors to completely sample the intrinsic curvature information of 3D palmprint images. Then, we jointly learn a feature projection function that project curvature data vectors into binary feature codes, which have the maximum inter-class variances and minimum intra-class distance so that they are discriminative. Moreover, we learn the collaborative binary representation of the multiple curvature feature codes by minimizing the information loss between the final representation and the multiple curvature features, so that the proposed method is more compact in feature representation and efficient in matching. Experimental results on the baseline 3D palmprint database demonstrate the superiority of the proposed method in terms of recognition performance in comparison with state-of-the-art 3D palmprint descriptors.
Lunke Fei, Jianyang Qin, Peng Liu 0045, Jie Wen 0001, Chunwei Tian, Bob Zhang 0001, Shuping Zhao
ICPR2
2020 Discrete Semantic Matrix Factorization Hashing for Cross-Modal Retrieval
abstract
Hashing has been widely studied for cross-modal retrieval due to its promising efficiency and effectiveness in massive data analysis. However, most existing supervised hashing has the limitations of inefficiency for very large-scale search and intractable discrete constraint for hash codes learning. In this paper, we propose a new supervised hashing method, namely, Discrete Semantic Matrix Factorization Hashing (DSMFH), for cross-modal retrieval. First, we conduct the matrix factorization via directly utilizing the available label information to obtain a latent representation, so that both the inter-modality and intra-modality similarities are well preserved. Then, we simultaneously learn the discriminative hash codes and corresponding hash functions by deriving the matrix factorization into a discrete optimization. Finally, we adopt an alternatively iterative procedure to efficiently optimize the matrix factorization and discrete learning. Extensive experimental results on three widely used image-tag databases demonstrate the superiority of the DSMFH over state-of-the-art cross-modal hashing methods.
Jianyang Qin, Lunke Fei, Shaohua Teng, Wei Zhang 0005, Dongning Liu, Genping Zhao
ICPR1