Zijian Wang 0010

dblp:03/4540-10 · DBLP profile ↗
← Back
19ranked-venue papers
2as first author
18since 2021 · last 2026
0000-0002-8147-5109ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 1 first-author · 11 since 2021Databases, data management, data science and information retrieval · 10 · 1 first-author · 10 since 2021Theory of computation · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Multi-armed bandits in recommender systems: advances, challenges, and future prospects
Cairong Yan, Jiaxin Nan, Zijian Wang 0010, Yongquan Wan
Knowl. Inf. Syst.3
2025 Compensating Information and Capturing Modal Preferences in Multimodal Recommendation: A Dual-Path Representation Learning Framework
abstract
In the context of information explosion, multimodal recommender systems (MMRS) have demonstrated great potential in capturing users' complex preferences and enhancing recommendation performance by integrating multimodal data such as images and text. However, multimodal data inherently suffers from semantic inconsistency, which can introduce information conflicts or noise. Moreover, users' reliance on different modalities varies dynamically with context and time (multimodal dynamic preferences). These challenges may lead to truth deviation and deep semantic mismatch, ultimately degrading recommendation performance. To address these issues, we propose Dual-Path Multimodal Recommendation (DPRec), a novel model that improves precision and robustness of the recommendation through cross-modal information compensation and dynamic modal preference learning. Specifically, DPRec first employs a cross-modal attention mechanism to dynamically model inter-modal correlations, effectively exploring complementary and shared features for robust user and item representations. Second, it integrates feature projection, modality alignment, and dynamic weighting mechanisms to adaptively adjust modality importance based on user context, ensuring flexibility in handling preference dynamics. Lastly, a modality contrastive loss is utilized to maximize mutual information between modalities, mitigating semantic mismatch by enhancing deep collaborative representations. Extensive experiments on three public datasets show that DPRec consistently outperforms state-of-the-art (SOTA) methods, achieving average improvements of 3.94% in Recall@20 and 3.84% in NDCG@20. Our code is publicly available at: https://anonymous.4open.science/r/DPRec-4D15.
Cairong Yan, Xubin Mao, Zijian Wang 0010, Xicheng Zhao, Linlin Meng
CIKM3
2025 CoCoB: Adaptive Collaborative Combinatorial Bandits for Online Recommendation
Cairong Yan, Jinyi Han, Jin Ju, Yanting Zhang 0001, Zijian Wang 0010, Xuan Shao
DASFAA (5)5
2025 KG-TS: Knowledge Graph-Driven Thompson Sampling for Online Recommendation
Cairong Yan, Hualu Xu, Yanting Zhang 0001, Zijian Wang 0010, Xuan Shao
DASFAA (5)4
2025 Dynamic Bidirectional Attentional Mamba Model for EEG-Based Motor Imagery Classification
Qianzi Shen, Zijian Wang 0010, Yanting Zhang 0001, Cairong Yan
ICIC (27)3
2025 MCCVM: Multi-Scale Cross-axes Conv-VMamba for Medical Image Classification
abstract
The performance of medical image classification relies on the effective capture and balance of local and global features. Although hybrid models combining Convolutional Neural Networks (CNNs) and Transformers have achieved notable success, they face two critical challenges: insufficient cross-axes information modeling, which hampers spatial dependency capture, and difficulty dynamically balancing within multi-scale features essential for complex medical images. In addition, the quadratic complexity of the self-attention mechanism in Transformers limits efficiency on high-resolution imaging data. To tackle these issues, this paper proposes the Multi-scale Cross-axes Convolutional VMamba Model (MCCVM), which combines CNNs for local feature extraction, the Mamba model for efficient global feature processing, and novel mechanisms for cross-axes information modeling and feature fusion. The MCCVM incorporates a Convolutional VMamba Fusion Block (CMF), which replaces Transformers with the Mamba framework to enhance computational efficiency while maintaining global feature extraction. A Cross-axes Attention SSM Block (CASSM) is also introduced within the Mamba structure to better model cross-axes spatial dependencies. Finally, a channel-based deep convolutional gated feature fusion network (CGFFN) is employed to dynamically balance local and global features, ensuring a comprehensive representation of medical images. Extensive experiments on medical image datasets demonstrate the superiority and effectiveness of our MCCVM.
Zijian Wang 0010, Qianzi Shen, Yanting Zhang 0001, Cairong Yan
IJCNN2
2025 Precise spiking neurons for fitting any activation function in ANN-to-SNN Conversion
Qianzi Shen, Xuhang Li, Yanting Zhang 0001, Zijian Wang 0010, Cairong Yan
Appl. Intell.5
2023 Thompson Sampling with Time-Varying Reward for Contextual Bandits
Cairong Yan, Hualu Xu, Haixia Han, Yanting Zhang 0001, Zijian Wang 0010
DASFAA (2)5
2023 Learning Golf Swing Key Events from Gaussian Soft Labels Using Multi-Scale Temporal MLPFormer
abstract
A complete golf swing includes several key events. The standardization of poses in each key event is directly related to the hitting effect. Thus, it is meaningful for the players to analyze their poses, especially at key frames, so as to improve swing performances. With the rapid development of deep learning techniques in computer vision, we are able to detect key frames during a golf swing. In this paper, we propose a framework to recognize key events in golf swing based on pure monocular video data. To achieve this, we have combined attention mechanism in the backbone network to extract concise features and leveraged the transformer structure to fuse multi-scale temporal information to enhance the feature representation. Besides, we also introduce Gaussian kernels into the label generation process, which can effectively solve the problem of ambiguity in detecting key events within their neighbouring similar frames. Notably, our method achieves an average recognition accuracy of 83.4% (+7.3% compared with SwingNet) for eight golf swing events on GoIfDB dataset.
Yanting Zhang 0001, Fuyu Tu, Zijian Wang 0010, Dandan Zhu 0001
IJCNN3
2023 Intralayer-Connected Spiking Neural Network with Hybrid Training Using Backpropagation and Probabilistic Spike-Timing Dependent Plasticity
abstract
Spiking neural networks (SNNs) are highly computationally efficient artificial intelligence methods due to their advantages in having a biologically plausible computational framework. Recent research has shown that SNN trained using backpropagation (SNN‐BP) exhibits excellent performance and has shown great potential in tasks such as image classification and security detection. However, the backpropagation method limits the dynamics and biological plausibility of the neural models in SNN, which will limit the recognition and simulation performance of SNN. In order to make neural models more similar to biological neurons, this study proposes a leaky integrate‐and‐fire (LIF) neuron model with dense intralayer connections, as well as efficient forward and backward processes in BP training. The new model will make the interaction between neurons within the layer more frequent, enhancing the intrinsic information exchange capability of SNN. An effective probabilistic spike‐timing dependent plasticity (STDP) method is also proposed to reduce the overweighted connections between neurons, as well as a hybrid training method using BP and probabilistic STDP. The training method combines the advantages of BP and STDP to improve the performance of SNN models. An intralayer‐connected SNN with hybrid training (ISNN‐HY) is proposed with the combination of these improvements. The proposed model was evaluated on three static image datasets and one neuromorphic dataset. The results showed that the performance of ISNN‐HY is superior to that of other SNN‐BP models. The proposed method also makes it possible to accurately simulate biological neural systems.
Xuhang Li, Yaqin Zhu, Jiayong Li, Zijian Wang 0010
Int. J. Intell. Syst.7
2023 MIN: multi-dimensional interest network for click-through rate prediction
Cairong Yan, Xiaoke Li, Yanting Zhang 0001, Zijian Wang 0010, Yongquan Wan
Knowl. Inf. Syst.4
2022 Automatic Moving Pose Grading for Golf Swing in Sports
abstract
Swing is a critical important part in golf and it involves the whole body movement when players hit the ball. A good swing requires proper posture, which needs a lot of practice to achieve the standardized full-body coordination. Considering that the amateur players often lack necessary supervision during self-practice, we introduce an automatic moving pose grading method based on monocular swing videos. Specifically, given a swing video, the 3D human poses are firstly extracted. Then, dynamic time warping (DTW) is introduced to perform moving pose alignment between a query and a reference video. Finally, a distance-based grading strategy is proposed based on the temporally aligned swing videos. The system can effectively solve the increasing need in golf teaching, where proper posture can be greatly aided by adopting the feedback when golf players practice swing by themselves.
Yanting Zhang 0001, Qing'Ao Wang, Fuyu Tu, Zijian Wang 0010
ICIP4
2022 Detection-by-tracking of traffic signs in videos
Yanting Zhang 0001, Zijian Wang 0010, Ruoning Song, Cairong Yan, Yonggang Qi
Appl. Intell.2
2022 Recurrent spiking neural network with dynamic presynaptic currents based on backpropagation
abstract
In recent years, spiking neural networks (SNNs), which originated from the theoretical basis of neuroscience, have attracted neuromorphic computing and brain-like computing due to their advantages, such as neural dynamics and coding mechanism, which are similar to biological neurons. SNNs have become one of the mainstream frameworks in the field of brain-like computing. However, most of the Leaky Integrate-and-Fire (LIF) neuron models currently used by SNNs based on direct training of backpropagation (BP) do not consider the changes in the recurrent connections and the dynamic strength of neuron connections over time. This study presented the LIF neuron model with recurrent connections and a method for dynamically changing the presynaptic currents. Recurrent LIF neurons have an additional cyclic connection compared with classic LIF neurons. Their postsynaptic current stimulates a change in membrane potential at the next time point. Their dynamics were more similar to the activities of biological neurons. We also proposed an efficient and flexible BP training method for recurrent LIF neurons. On the basis of the above methods, we proposed the recurrent SNN with dynamic presynaptic currents based on backpropagation (RDS-BP). We test the proposed RDS-BP on three image data sets (MNIST, Fashion-MNIST and CIFAR-10) and two text data sets (IMDB and TREC). The results showed that the performance of RDS-BP not only exceeded the naive SNN models based on BP but also exceeded the SNN methods proposed in previous studies in recent years, which had excellent performance in previous experiments. Our work provides a new LIF neuron model with a recurrent connection and dynamic presynaptic current and a BP training arrangement for the proposed neuron, which could merit developments with neuromorphic and brain-like computing.
Zijian Wang 0010, Yanting Zhang 0001, Haibo Shi, Lei Cao 0002, Cairong Yan
Int. J. Intell. Syst.1
2021 Learning Fashion Similarity Based on Hierarchical Attribute Embedding
abstract
Embedding items directly into a common feature space, and then measuring the similarity by calculating the feature distance in this space, has become the main method for similarity learning in current fashion retrieval tasks. The method is simple and efficient, but it ignores the correlation among fashion attributes and the impact of these correlations on the feature space, thereby reducing the accuracy of retrieval. Since the number of fashion attributes is large and the semantic granularity is also different, how to capture the relationship between fashion attributes and perform refined embedding to accurately represent fashion items is a challenge. In this paper, by constructing an attribute tree, we propose a hierarchical attribute embedding method for representing fashion items to enhance the relationship between attributes and use masking technology to disentangle different attributes. Based on these modules, we propose a hierarchical attribute-aware embedding network (HAEN) which takes images and attributes as input, learns multiple attribute-specific embedding spaces, and measures fine-grained similarity in the corresponding spaces. The extensive experimental result on two fashion-related public datasets FashionAI and DARN shows the superiority (+5.11% and +3.09% in MAP, respectively) of our proposed HAEN compared with state-of-the-art methods.
Cairong Yan, Anan Ding, Yanting Zhang 0001, Zijian Wang 0010
DSAA4
2021 Two-Phase Multi-armed Bandit for Online Recommendation
abstract
Personalized online recommendations strive to adapt their services to individual users by making use of both item and user information. Despite recent progress, the issue of balancing exploitation-exploration (EE) [1] remains challenging. In this paper, we model the personalized online recommendation of e-commence as a two-phase multi-armed bandit problem. This is the first time that “big arm” and “small arm” are introduced into multi-armed bandit (MAB), and a two-stage strategy is adopted to provide target users with the most suitable recommendation list. In the first phase, MAB is used to obtain an item subset that users may be interested in from a large number of items. We use item categories as arms instead of individual items in existing related models to control the arm scale and reduce computational complexity. In the second phase, we directly use the items generated in the first phase as arms of MAB and obtain rewards through fine-grained implicit feedback from users. Empirical studies on three real-world datasets show that our proposed method TPBandit performs better than state-of-the-art bandit-based recommendation methods in several evaluation metrics such as Precision, Recall, and Hit Ratio. Moreover, the two-phase method improves the recommendation performance by nearly 50% compared to the one-phase method in the best case.
Cairong Yan, Haixia Han, Zijian Wang 0010, Yanting Zhang 0001
DSAA3
2021 A Multi-Task Learning Approach for Recommendation based on Knowledge Graph
abstract
Sparsity and cold start problem are two classic problems of collaborative filtering. To alleviate these issues, researchers usually add side information to the recommendation models to boost the performance. In this paper, we propose a multi-task learning approach for recommendation based on knowledge graph (KGeRec), which takes recommendation as the main task and the knowledge graph as an auxiliary task to provide side information for recommendation. To fully capture the correlation information between these two tasks, a feature interaction layer (FlU) based on cross networks is designed to share features between them. Besides, a side information embedding layer (SIE) is also designed in the recommendation task to exploit more feature information. We apply KGeRec to three public datasets about movie, book, and music. Experimental results show that the proposed KGeRec outperforms the state-of-the-art approaches (+2.2% in AUC, +2.6% in Accuracy, +2.5% and in F1-score, compared to the maximum value in Type I models; +1.3% in AUC, +0.8% in Accuracy, and +2% in F1-score, compared to the maximum value in Type II models) and it performs well in sparse datasets. We also validate the effectiveness of knowledge graphs in improving recommendation performance.
Cairong Yan, Yanting Zhang 0001, Zijian Wang 0010, Pengwei Wang 0001
IJCNN4
2021 Modeling Long- and Short-Term User Behaviors for Sequential Recommendation with Deep Neural Networks
abstract
In e-commerce platforms, a user's next behavior will be affected by his long-term constant interests and short-term temporal needs. Such information is usually hidden in the users' historical online behavior data, so how to capture long-term and short-term patterns becomes the key to design better recommendation models or algorithms. Current mainstream methods such as Markov chain, convolutional neural network, and recurrent neural network cannot well express the mixed dynamic characteristics. In this paper, we propose an attention-based deep neural network (ADNNet) to solve the problem. In ADNNet, a convolutional neural network is used to extract the short-term patterns in the behavior sequences, and a gated recurrent unit is used to mine the long-term patterns in the behavior sequences. The attention mechanism is adopted to help the network automatically learn the best fusion coefficient of these two patterns. Our experimental result on four real public datasets (+0.69% in Hit Ratio and +3.49% in MRR) shows the superiority of our proposed ADNNet compared with other state-of-the-art methods.
Cairong Yan, Yanting Zhang 0001, Zijian Wang 0010, Pengwei Wang 0001
IJCNN4
2018 Short time Fourier transformation and deep neural networks for motor imagery brain computer interface recognition
abstract
Summary Motor imagery (MI) is an important control paradigm in the field of brain‐computer interface (BCI), which enables the recognition of personal intention. So far, numerous methods have been designed to classify EEG signal features for MI task. However, deep neural networks have been seldom applied to analyze EEG signals. In this study, two novel kinds of deep learning schemes based on convolutional neural networks (CNN) and Long Short‐Term Memory (LSTM) were proposed for MI‐classification. The frequency domain representations of EEG signals were obtained using short time Fourier transform (STFT) to train models. Classification results were compared between conventional algorithm, CNN, and LSTM models. Compared with two other methods, CNN algorithms had shown better performance. These conclusions verified that CNN method was promising for MI‐based BCIs.
Zijian Wang 0010, Lei Cao 0002, Zuo Zhang 0001, Xiaoliang Gong, Yaoru Sun, Haoran Wang 0010
Concurr. Comput. Pract. Exp.1