VLDB 2026 Research / reviewers in the wild / expert
Jing Li 0055
dblp:181/2820-55
· DBLP profile ↗
55ranked-venue papers
0as first author
38since 2021 · last 2026
0000-0002-8181-0886ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 34 · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 12 · 10 since 2021Databases, data management, data science and information retrieval · 7 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Multi-representation space recommendation with graph contrastive learning
Xiaoda Li, Yue He 0005, Jing Li 0055, Kai Zhu 0009, Jiaheng Yu |
Expert Syst. Appl. | 3 |
| 2026 | JCLRec: Joint diffusion model and dual contrastive learning for sequential recommendation
Kai Zhu 0009, Jing Li 0055, Yue He 0005, Mingfeng Wang, Jiaheng Yu, Jun Wan 0005 |
Knowl. Based Syst. | 2 |
| 2026 | Efficient Oriented Object Detection via Wavelet-Based Energy Label Reassignment and Dual Prediction StrategyabstractArbitrary-oriented object detection remains a pivotal research focus due to its practical significance and inherent challenges. Existing methods often extend frameworks and sampling strategies designed for horizontal object detectors, which struggle to handle the arbitrary orientations, high aspect ratios, and diverse scales of oriented objects. To overcome these limitations, we propose a novel and efficient method for arbitrary-oriented object detection. This approach dynamically assigns prediction layers by object pixel area, then leverages wavelet transform-based energy weighting for bottom-up sample reassignment, optimizing feature representation for oriented targets. In addition, a robust framework integrates heatmap keypoint prediction on feature maps of a quarter-sized image, along with sparse predictions on other scales. By querying small-object regions within deep feature maps, a progressive top-down feature fusion strategy further enhances the perception of fine-grained details. Extensive evaluations on four benchmark datasets demonstrate the method's substantial improvements in detection performance, establishing its potential for broader applications in oriented object detection. Beihang Song, Jing Li 0055, Jia Wu 0001, Xuefei Li 0001, Jun Wan 0005 |
IEEE Trans. Multim. | 2 |
| 2025 | Joint Content semantic relation Learning with Mamba for Multimodal RecommendationabstractMultimodal recommendation systems have received widespread attention in recent years. The existing methods mainly treat multimodal features as a supplement to collaborative ID features to build alignment representation. However, these methods only use the user-item collaborative relation and item-item latent relation to implicitly represent multimodal features without explicitly using multimodal content, which ignores the intrinsic semantic relations within different modal content and makes the model not pay insufficient attention to user preferences, causing not great recommendation performance. Considering this problem, we argue the complete fine-grained representation of multimodal content, which includes collaboration relation, latent relation, and semantic relation. To address this, this paper proposes the Joint Content Semantic relation learning with Mamba for Multimodal Recommendation (JCSMRec). First, we construct Semantic relation representation with Mamba (SR2M) Module to explore the inter-modal correlation and provide an excellent aligned explicit representation for multimodal content. Then, we design a User preference-aware Module, according to user preference to guide learning the representation granularity of different modalities in the latent relations. Then, we conduct a joint representation to align and fusion multimodal features, which use the cyclic KL loss to align different semantic spaces and alleviate two joint losses to let user preference guide the learning process of the SR2M Module. This design effectively avoids over-reliance on ID features and enables a more comprehensive learning of textual and visual semantic information suitable for recommendation. Our method has achieved good results on multiple datasets and can be effectively inserted into other recommendation methods to improve results. Yue He 0005, Jing Li 0055, Kai Zhu 0009, Guohao Li 0009 |
IJCNN | 3 |
| 2025 | Adaptive User Dynamic Interest Guidance for Generative Sequential RecommendationabstractRecently, diffusion model-based methods have utilized user interest features as guidance conditions to achieve stable generation results in sequential recommendation tasks. However, these models struggle to capture users' dynamic interests, as the interests of different users are often inconsistent. Moreover, the fixed number of interests predefined by existing models cannot adapt to the diverse preferences of users, making it difficult to further improve recommendation performance. To address these issues, we propose a novel generative sequential recommendation framework named ADIGRec (Adaptive User Dynamic Interest Guidance for Generative Sequential Recommendation), which adaptively focuses on users' dynamic interest features. Specifically, our framework combines users' dynamic features and inherent interest features encoded from historical sequences as new guidance conditions. Furthermore, we introduce a module that injects dynamic interest features into the noise item embeddings, enabling explicit interaction with the guidance conditions during the generation phase. This approach essentially fits the noise in the target space rather than the user preference space, leading to improved recommendation diversity. Additionally, we propose a novel regularization method to mitigate the impact of user interest routing collapse on the generation results. Extensive experiments on three publicly available datasets demonstrate that our method achieves superior performance compared to established baseline methods. Kai Zhu 0009, Jing Li 0055, Jia Wu 0001, Yue He 0005, Guohao Li 0009 |
SIGIR | 2 |
| 2025 | HEART: Historically Information Embedding and Subspace Re-Weighting Transformer-Based TrackingabstractTransformers-based trackers offer significant potential for integrating semantic interdependence between template and search features in tracking tasks. Transformers possess inherent capabilities for processing long sequences and extracting correlations within them. Several researchers have explored the feasibility of incorporating Transformers to model continuously changing search areas in tracking tasks. However, their approach has substantially increased the computational cost of an already resource-intensive Transformer. Additionally, existing Transformers-based trackers rely solely on mechanically employing multi-head attention to obtain representations in different subspaces, without any inherent bias. To address these challenges, we propose HEART (Historical Information Embedding And Subspace Re-weighting Tracker). Our method embeds historical information into the queries in a lightweight and Markovian manner to extract discriminative attention maps for robust tracking. Furthermore, we develop a multi-head attention distribution mechanism to retrieve the most promising subspace weights for tracking tasks. HEART has demonstrated its effectiveness on five datasets, including OTB-100, LaSOT, UAV123, TrackingNet, and GOT-10k. Tianpeng Liu, Jing Li 0055, Amin Beheshti, Jia Wu 0001, Beihang Song, Lezhi Lian |
IEEE Trans. Big Data | 2 |
| 2025 | Facial Expression Recognition With Heatmap Neighbor Contrastive LearningabstractMany supervised learning-based facial expression recognition (FER) methods achieve good performance with the assistance of expression labels and a complex framework. However, there are inconsistent annotations in different expression datasets, making the above methods disadvantageous for new expression datasets or datasets with limited training data. The objective of this paper is to learn self-supervised facial expression features that enable the FER model not to rely on the annotation consistency of the different datasets. Most current self-supervised learning algorithms based on contrastive learning learn the representation by forcing different augmented views of the same image close in the embedding space, but they cannot cover all variances within a semantic class. We propose a heatmap neighbor contrastive learning (HNCL) method for FER. It treats the images corresponding to the heatmap nearest neighbors of expressions as other positives, providing more semantic variations than pre-defined augmented transformations. Therefore, our HNCL can learn better expression features covering more intra-class variances, improving the performance of the FER model based on self-supervised learning. After fine-tuning, HNCL with a simple framework achieves top-three performance on the in-the-lab datasets and even matches the performance of state-of-the-art supervised learning methods on the in-the-wild datasets. Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Yibing Zhan, Dapeng Tao, Jun Wan 0005 |
IEEE Trans. Multim. | 2 |
| 2024 | Bilateral Unsymmetrical Graph Contrastive Learning for RecommendationabstractRecent methods utilize graph contrastive Learning within graph-structured user-item interaction data for collaborative filtering and have demonstrated their efficacy in recommendation tasks. However, they ignore that the difference relation density of nodes between the user- and item-side causes the adaptability of graphs on bilateral nodes to be different after multi-hop graph interaction calculation, which limits existing models to achieve ideal results. To solve this issue, we propose a novel framework for recommendation tasks called Bilateral Unsymmetrical Graph Contrastive Learning (BusGCL) that consider the bilateral unsymmetry on user-item node relation density for sliced user and item graph reasoning better with bilateral slicing contrastive training. Especially, taking into account the aggregation ability of hypergraph-based graph convolutional network (GCN) in digging implicit similarities is more suitable for user nodes, embeddings generated from three different modules: hypergraph-based GCN, GCN and perturbed GCN, are sliced into two subviews by the user- and item-side respectively, and selectively combined into subview pairs bilaterally based on the characteristics of inter-node relation structure. Furthermore, to align the distribution of user and item embeddings after aggregation, a dispersing loss is leveraged to adjust the mutual distance between all embeddings for maintaining learning ability. Comprehensive experiments on two public datasets have proved the superiority of BusGCL in comparison to various recommendation methods. Other models can simply utilize our bilateral slicing contrastive learning to enhance recommending performance without incurring extra expenses. Jiaheng Yu, Jing Li 0055, Yue He 0005, Kai Zhu 0009 |
IJCNN | 2 |
| 2024 | Contrastive graph learning long and short-term interests for POI recommendationabstractModeling users’ short-term dynamic and long-term static interests to enhance Point-of-Interests (POI) recommendation performance has shown lots of advantages. Since users’ check-in records can be viewed as a graph network, methods based on Graph Neural Networks (GNNs) have recently shown promising applicability for POI recommendation. However, existing GNN-based works have the following shortcomings: (1) ignoring the impact of complex higher-order relationships between user-POI dynamics over time; and (2) ignoring the difference in POI importance that cannot effectively capture the imbalances of geographical influence among POIs. To address these challenges, we propose a novel Self-supervised Long-and Short-term model (SLS-REC) for POI recommendation. Specifically, we first design a spatio-temporal Hawkes attention hypergraph neural network to capture the spatial dependence and temporal evolution in users’ short-term dynamic interests. Then we introduce a dynamic propagation mechanism of GNNs to learn the geographic influences underlying geographic imbalances among POIs. In addition, the contrastive learning framework over a fine-grained node dropout strategy is applied to maximize the mutual information of long and short-term interest representations. Finally, we adaptively unify the recommendation and self-supervised task with an attention-based mechanism to optimize the proposed SLS-REC model for POI recommendation. Experiments on real-world datasets show that the proposed model significantly outperforms state-of-the-art methods. Jia-Run Fu, Rong Gao 0001, Yonghong Yu, Jia Wu 0001, Jing Li 0055, Donghua Liu, Zhiwei Ye |
Expert Syst. Appl. | 5 |
| 2024 | Flexibly utilizing syntactic knowledge in aspect-based sentiment analysis
Xiaosai Huang, Jing Li 0055, Jia Wu 0001, Donghua Liu, Kai Zhu 0009 |
Inf. Process. Manag. | 2 |
| 2024 | PTMB: An online satellite task scheduling framework based on pre-trained Markov decision process for multi-task scenario
Guohao Li 0009, Xuefei Li 0001, Jing Li 0055, Xin Shen 0001 |
Knowl. Based Syst. | 3 |
| 2024 | Single-stage oriented object detection via Corona Heatmap and Multi-stage Angle Prediction
Beihang Song, Jing Li 0055, Jia Wu 0001, Shan Xue 0001, Jun Wan 0005 |
Knowl. Based Syst. | 2 |
| 2024 | Quality-aware face alignment using high-resolution spatial dependencies
Jinyan Ma, Xuefei Li 0001, Jing Li 0055, Jun Wan 0005, Tong Liu 0039, Guohao Li 0009 |
Multim. Tools Appl. | 3 |
| 2024 | Confusable facial expression recognition with geometry-aware conditional network
Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Jun Wan 0005 |
Pattern Recognit. | 2 |
| 2024 | Direction Prediction Redefinition: Transfer Angle to Scale in Oriented Object DetectionabstractOriented object detection has garnered significant attention. However, rotational symmetry and discontinuity at boundaries can confuse networks, leading to discontinuous loss and regression inconsistency. In this paper, we propose an efficient multi-directional object detection framework named Direction Prediction Redefinition (DPR). We describe the angle variation of rotated bounding boxes ($B_{r}$) as changes in the dimensions of horizontal bounding boxes ($B_{h}$). Specifically, we generate two sets of horizontal bounding boxes by predicting the center points of the corresponding boundaries within the rotated bounding box, thereby avoiding boundary issues caused by angle prediction. To further achieve robust rotated boundary representation, we propose the Joint Scale Representation method and the State Feature Encoding module, which are used to eliminate outliers in rotated boundaries and guide the correct selection of horizontal bounding box vertices, respectively. Moreover, we further abstract DPR as Multiple Trigonometric functions based DPR (DPR-MT). This method maps a single angle into four sets of trigonometric functions and considers them as the four sides of the horizontal bounding box. This approach predicts angles in the form of horizontal bounding boxes without complex operations, making it plug-and-play. Experimental results and visual analysis on challenging datasets further verify the effectiveness and competitiveness of our proposed method. Beihang Song, Jing Li 0055, Jia Wu 0001, Jun Wan 0005 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Tracking With Saliency Region TransformerabstractTransformers show a great impact on visual tracking thanks to their powerful representation learning capabilities. As the capacity of the model grows, the speed of the tracker tends to decrease gradually. Our work focuses on dealing with massively redundant information in tracking sequences with the Saliency Region Tracker (SRTrack). SRTrack is a heuristic two-stage tracker consisting of a lightweight tracking stage and a saliency stage. The former can handle simple tracking sequences while the latter is designed to perform delicate tracking on challenging frames with more discriminative features. However, the two-stage design leads to feature extrapolation, creating inconsistencies between training and inference features. In order to mitigate this problem, we develop an attention scaling factor that guarantees model robustness while yielding a slight performance gain. Our SRTrack achieves a state-of-the-art 0.699 AUC running at 61 FPS on LaSOT. Several experiments on large benchmarks demonstrate the high efficiency and accuracy of SRTrack. Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Lefei Zhang, Jun Wan 0005, Lezhi Lian |
IEEE Trans. Image Process. | 2 |
| 2023 | Cross-Domain Facial Expression Recognition via Disentangling Identity RepresentationabstractMost existing cross-domain facial expression recognition (FER) works require target domain data to assist the model in analyzing distribution shifts to overcome negative effects. However, it is often hard to obtain expression images of the target domain in practical applications. Moreover, existing methods suffer from the interference of identity information, thus limiting the discriminative ability of the expression features. We exploit the idea of domain generalization (DG) and propose a representation disentanglement model to address the above problems. Specifically, we learn three independent potential subspaces corresponding to the domain, expression, and identity information from facial images. Meanwhile, the extracted expression and identity features are recovered as Fourier phase information reconstructed images, thereby ensuring that the high-level semantics of images remain unchanged after disentangling the domain information. Our proposed method can disentangle expression features from expression-irrelevant ones (i.e., identity and domain features). Therefore, the learned expression features exhibit sufficient domain invariance and discriminative ability. We conduct experiments with different settings on multiple benchmark datasets, and the results show that our method achieves superior performance compared with state-of-the-art methods. Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Lefei Zhang, Jun Wan 0005 |
IJCAI | 2 |
| 2023 | Two-way Cross-domain Recommendation with Central Social InfluenceabstractAs an effective method to solve the cold start problem of recommender systems, cross-domain recommendation has received more and more attention and research. Currently, most cross-domain recommendation models rely on the unidirectional knowledge transfer between the same users in the source and target domains. The rich information in the source domain is transferred to the target domain with sparse information to achieve better recommendation effect. However, the performance of cross-domain recommendation heavily depends on the number of overlapping active users, and these models cannot fully utilize the useful knowledge behind active users in a single domain. This limitation makes it difficult for the model to achieve ideal results in real-world scenarios. To solve the above problems and optimize cross-domain recommendation, we propose a Two-way Cross-domain Recommendation with Central Social Influence(CST-CDR) and the concept of super-user. Through the idea of clustering, user circles are formed centering on active super-users, so the potential common preferences of single-field active users, dual-field active users, and super-users are fully explored, alleviating the dependence of the model on overlapping and active users. At the same time, the cross-domain recommendation is extended to realize two-way information migration, so that users who are active in only one domain can have a preliminary preference judgment. Finally, the effectiveness of the proposed method is demonstrated on two real datasets. Jing Li 0055, Mingfeng Wang, Kai Zhu 0009 |
IJCNN | 2 |
| 2023 | Worldwide COVID-19 Topic Knowledge Graph Analysis From Social MediaabstractCurrent research on online public opinion regarding the coronavirus disease in 2019 (COVID-19) leverages keyword extraction, sentiment analysis, and topic modeling to analyze online public opinion. The multi-granularity features of online public opinions and semantic relations between the features, how-ever, remain less explored. Reliance on only topics or keywords for measuring public opinion is insufficient as topics are too broad and keywords too narrow. Analysis at an intermediate level, which most studies overlook, is crucial in gaining a clearer insight into public opinion. Additionally, exploring the semantic relationships between components of online public opinion can shed light on the logical connections between them and help understand how they interact in the dissemination of online public opinion, leading to a better understanding of its evolution mechanism. We conducted a public opinion analysis on Worldwide COVID-19 outbreaks via Topic Knowledge Graph. Specifically, we first use the Combined Topic Model to extract public opinion topics. Then multi-dimensions attributes of the topic such as [subject, predicate, object] triples, topic popularity, and topic emotion intensity are extracted. Subsequently, the semantic relations of different public opinion topics are calculated from the two levels of predicates co-occurrence and subject-object sharing. Finally, we applied the constructed framework to the public opinion information related to COVID-19 and analyzed the characteristics in the evolution of online public opinion. Hao Fan 0003, Jing Li 0055, Jia Wu 0001 |
IJCNN | 4 |
| 2023 | Byzantine-Robust Federated Learning via Server-Side Mixtue of Experts
Zheyuan Shen, Keke Yang, Jing Li 0055 |
PRICAI (2) | 5 |
| 2023 | Visual tracking with dumbbell selection network
Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Yafu Xiao, Yan Hong 0005 |
Neurocomputing | 2 |
| 2023 | Self-supervised Dual Hypergraph learning with Intent Disentanglement for session-based recommendationabstractExisting works on session-based recommendation have shown the advantage in enhancing the prediction ability of recommendation with various deep learning techniques. However, the following challenges need to be addressed: (1) the hierarchy of item transition patterns is overlooked; (2) existing works fail to distinguish various factors of item transition within a single session for disentangling user intents. To cope with the above challenges, we propose a novel session based recommendation model called S elf-supervised D ual H ypergraph learning with I ntent D isentanglement model ( SDHID ). Specifically, we first propose a disentangled capsule hypergraph convolutional channel for ne-grained intent learning to capture the intra-session pattern. Accordingly, we introduce the hypergraph and capsule networks in disentangling to learn the item embedding for different factors, and then the representation of the intra-session pattern is obtained by aggregating item embedding with attention weights. Moreover, we build a novel dual–primal hypergraph convolutional channel by mapping the hypergraph to a dual–primal graph for learning the item transition pattern of inter-session. In addition, the above two channels are combined into a self-supervised contrastive learning framework by maximizing mutual information between the learned session representations. We unify the recommendation and the self-supervised tasks under a primary and auxiliary learning framework. The combined optimization of two tasks leads to a hierarchical joint learning item transition for intra- and inter-session. Extensive experiments on real datasets show that the proposed model outperforms several state-of-the-art models. Rong Gao 0001, Yuhe Tao, Yonghong Yu, Jia Wu 0001, Xiongkai Shao, Jing Li 0055, Zhiwei Ye |
Knowl. Based Syst. | 6 |
| 2023 | Multi-Aspect enhanced Graph Neural Networks for recommendation
Chenyan Zhang, Shan Xue 0001, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Donghua Liu |
Neural Networks | 3 |
| 2023 | Transfer Learning With Document-Level Data Augmentation for Aspect-Level Sentiment ClassificationabstractAspect-level sentiment classification (ASC) seeks to reveal the emotional tendency of a designated aspect of a text. Some researchers have recently tried to exploit large amounts of document-level sentiment classification (DSC) data available to help improve the performance of ASC models through transfer learning. However, these studies often ignore the difference in sentiment distribution between document-level and aspect-level data without preprocessing the document-level knowledge. Our study provides a transfer learning with document-level data augmentation (TL-DDA) framework to transfer more accurate document-level knowledge to the ASC model by means ofdocument-level data augmentationandattention fusion. First, we usedocument data selectionandtext concatenationto produce document-level data with various sentiment distributions. The augmented document data is then utilized for pre-training a well-designed DSC model. Finally, afterattention adjustment, wefuse the word attentionobtained from this DSC model into the ASC model. Results of experiments utilizing two publicly available datasets suggest that TL-DDA is reliable. Xiaosai Huang, Jing Li 0055, Jia Wu 0001, Donghua Liu |
IEEE Trans. Big Data | 2 |
| 2023 | SRDF: Single-Stage Rotate Object Detector via Dense Prediction and False Positive SuppressionabstractOriented object detection has made astonishing progress. However, existing methods neglect to address the issue of false positives caused by the background or nearby clutter objects. Meanwhile, class imbalance and boundary overflow issues caused by the predicting rotation angles may affect the accuracy of rotated bounding box predictions. To address the above issues, we propose a Single-stage Rotate object detector via Dense prediction and False positive suppression (SRDF). Specifically, we design an Instance-level False Positive Suppression Module (IFPSM), IFPSM acquires the weight information of target and non-target regions by supervised learning of spatial feature encoding, and applies these weight values to the deep feature map, thereby attenuating the response signals of non-target regions within the deep feature map. Compared to commonly used attention mechanisms, this approach more accurately suppresses false positive regions. Then, we introduce a hybrid classification and regression method to represent the object orientation, the proposed mothed divide the angle into two segments for prediction, reducing the number of categories and narrowing the range of regression. This alleviates the issue of class imbalance caused by treating one degree as a single category in classification prediction, as well as the problem of boundary overflow caused by directly regressing the angle. In addition, we transform the traditional post-processing steps based on matching and searching to a two-dimensional probability distribution mathematical model, which accurately and quickly extracts the bounding boxes from dense prediction results. Extensive experiments on Remote Sensing, Synthetic Aperture Radar, and Scene Text benchmarks demonstrate the superiority of the proposed SRDF method over state-of-the-art rotated object detection methods. Our codes are available at https://github.com/TomZandJerryZ/SRDF. Beihang Song, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Jun Wan 0005, Tianpeng Liu |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Facial Expression Recognition on the High Aggregation SubgraphsabstractWith the development of deep learning technology, the performance of facial expression recognition (FER) has been significantly improved. The current main challenge comes from the confusion of facial expressions caused by the highly nonlinear changes of facial expressions. However, the existing FER methods based on Convolutional Neural Networks (CNN) often ignore the underlying relationship between expressions which is crucial to meliorate the performance of recognition for confusable expressions. And the methods based on Graph Convolutional Networks (GCN) can capture the relationship between vertices, but the aggregation degree of subgraphs generated by these methods is low. They are easy to include unconfident neighbors, which increases the learning difficulty of the network. To solve the above problems, this paper proposes a method to recognize facial expressions on the high aggregation subgraphs (HASs) by combing the advantages of CNN extracting features and GCN modeling complex graph patterns. Specifically, we formulate FER as a vertex prediction problem. Considering the importance of high-order neighbors and higher efficiency, we utilize vertex confidence to find high-order neighbors. Then we construct the HASs based on the top embedding features of these high-order neighbors. And we utilize the GCN to perform reasoning and infer the class of vertices for HASs without a large number of overlapping subgraphs. Our method captures the underlying relationship between expressions on the HASs and improves the accuracy and efficiency of FER. Experimental results on both the in-the-lab datasets and the in-the-wild datasets show that our method achieves higher recognition accuracy than several state-of-the-art methods. This highlights the benefit of the underlying relationship between expressions for FER. Tong Liu 0039, Jing Li 0055, Jia Wu 0001, Bo Du 0001, Yi Liu 0038 |
IEEE Trans. Image Process. | 2 |
| 2023 | Tracking With Mutual Attention NetworkabstractVisual tracking is a visual task that tracks a specific target by only giving its first frame location and size. To punish the low-quality but high-scoring tracking results, researchers resorted to foreground reinforcement learning to suppress the scores of positive samples near edges. However, for training with negative samples, all backgrounds are equally labeled as false. In this way, the interdependence and difference between the foreground and the background are not considered. We interpret the underlying reason for drifts as the imbalance between the embedding of background and foreground information. Specifically, some catastrophic tracking results and common tracking errors should not be treated equally but should strengthen the implicit connection between the foreground and background. In this paper, we propose a Mutual Attention (MA) module to strengthen the interdependence between positive and negative samples. It can aggregate the rich contextual interdependence between the target template and the search area, thereby providing an implicit way to update the target template accordingly. As for the difference, we design a background training enhancement (BTE) mechanism to distinguish negative samples with varying degrees of error, that is, to down-weight outrageous and absurd tracking results to improve the robustness of the tracker. The results on a large number of benchmarks indicate the validity of our results, such as OTB-100, VOT-2018, VOT-2019, and LaSOT. Tianpeng Liu, Jing Li 0055, Jia Wu 0001, Beihang Song |
IEEE Trans. Multim. | 2 |
| 2022 | Global Context-Aware Graph Neural Networks for Session-based RecommendationabstractSession-based recommendation, which uses the interactive information of anonymous users in a period to predict user preferences, has attracted extensive attention. Existing methods mainly exploit the strengths of graph neural networks (GNNs) in capturing structured data features to model complex item transitions for recommendations. However, there still remain some challenges: 1) many methods neglect to exploit cross-session interactions and fail to integrate the latent intra- and inter-session information effectively. 2) some methods barely explore fine-grained contextual factors underlying in the sessions for session-based recommendation, thus failing to capture more comprehensive user preferences. To solve these issues, we propose a novel method called Global Context-Aware Graph Neural Networks (GCA-GNN) which captures the local-session and cross-session preferences respectively, and captures fine-grained global contextual factors to complement user preferences. Specifically, GCA-GNN models user preferences from two different views: the cross-session view and the local-session view. The former learns collaborative user preferences from a cross-session graph, while the latter is designed to learn users' personal preferences from local-session graphs. Furthermore, a context-aware Capsule Graph Neural Network is employed to extract fine-grained contextual factors, serving as complementary information. And we introduce an auxiliary self-supervised learning task to enhance user preferences. Experiments on benchmark datasets demonstrate the strength of our model over the state-of-the-art methods. Mingfeng Wang, Jing Li 0055, Donghua Liu, Chenyan Zhang, Xiaosai Huang |
IJCNN | 2 |
| 2022 | Interest Evolution-driven Gated Neighborhood aggregation representation for dynamic recommendation in e-commerce
Donghua Liu, Jing Li 0055, Jia Wu 0010, Bo Du 0001, Xuefei Li 0001 |
Inf. Process. Manag. | 2 |
| 2022 | GARAT: Generative Adversarial Learning for Robust and Accurate Tracking
Jing Li 0055, Shan Xue 0001, Jia Wu 0001, Huanmei Guan, Zhiquan Ding |
Neural Networks | 2 |
| 2022 | Robust face alignment by dual-attentional spatial-aware capsule networks
Jinyan Ma, Jing Li 0055, Bo Du 0001, Jia Wu 0001, Jun Wan 0005, Yafu Xiao |
Pattern Recognit. | 2 |
| 2022 | Adaptive Hierarchical Attention-Enhanced Gated Network Integrating Reviews for Item RecommendationabstractMany studies focusing on integrating reviews with ratings to improve recommendation performance have been quite successful. However, these works still face several shortcomings: (1) The importance of dynamically integrating review and interaction data features is typically ignored, yet treating these fusion features equally may lead to an incomplete understanding of user preferences. (2) Some forms of soft attention methods are adopted to model the local semantic information of words. As features thus captured may contain irrelevant information, the generated attention map is neither discriminatory nor detailed. In this paper, we propose a novelAdaptiveHierarchicalAttention-enhancedGated network integrating reviews for item recommendation, named AHAG. AHAG is a unified framework to capture the hidden intentions of users by adaptively incorporating reviews. Specifically, we design a gated network to dynamically fuse the extracted features and select the features that are most relevant to user preferences. To capture distinguishing fine-grained features, we introduce a hierarchical attention mechanism to learn important semantic information features and the dynamic interaction of these features. Besides, the high-order non-linear interaction of neural factorization machines is utilized to derive the rating prediction. Experiments on seven real-world datasets show that the proposed AHAG significantly outperforms state-of-the-art methods. Furthermore, the attention mechanism can highlight the relevant information in reviews to increase the interpretability of the recommendation task. Source codes are available inhttps://github.com/luojia527/AHAG. Donghua Liu, Jia Wu 0001, Jing Li 0055, Bo Du 0001, Xuefei Li 0001 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2022 | Drift-Proof Tracking With Deep Reinforcement LearningabstractObject tracking is an essential and challenging sub-domain in the field of computer vision owing to its wide range of applications and complexities of real-life situations. It has been studied extensively over the last decade, leading to the proposal of several tracking frameworks and approaches. Recently, the introduction of reinforcement learning and the ‘Actor-Critic’ framework has effectively improved the tracking speed of deep learning trackers. However, most existing deep reinforcement learning trackers experience a slight performance degradation mainly owing to the drift issues. Drifts pose a threat to the tracking performance, which may lead to losing the tracked target. Herein, we propose a drift-proof tracker with deep reinforcement learning that aims to improve the tracking performance by counteracting drifts while maintaining its real-time advantage. We utilize a reward function with the Distance-IoU (DIoU) metric to guide the reinforcement learning to alleviate the drifts caused by the trained model. Furthermore, double negative samples (hard negative and drift samples) are constructed in tracking for network initialization, which is followed by calculating the loss by a small error-friendly loss function. Therefore, our tracker can better discriminate between the positive and negative samples and correct the predicted bounding boxes when the drift occurs. Meanwhile, a generative adversarial network is adopted for positive sample augmentation. Extensive experimental results on multiple popular benchmarks show that our algorithm effectively reduces the occurrences of drift and boosts the tracking performance, compared to those of other state-of-the-art trackers. Zhongze Chen, Jing Li 0055, Jia Wu 0001, Yafu Xiao |
IEEE Trans. Multim. | 2 |
| 2022 | Robust Facial Landmark Detection by Multiorder Multiconstraint Deep NetworksabstractRecently, heatmap regression has been widely explored in facial landmark detection and obtained remarkable performance. However, most of the existing heatmap regression-based facial landmark detection methods neglect to explore the high-order feature correlations, which is very important to learn more representative features and enhance shape constraints. Moreover, no explicit global shape constraints have been added to the final predicted landmarks, which leads to a reduction in accuracy. To address these issues, in this article, we propose a multiorder multiconstraint deep network (MMDN) for more powerful feature correlations and shape constraints' learning. Especially, an implicit multiorder correlating geometry-aware (IMCG) model is proposed to introduce the multiorder spatial correlations and multiorder channel correlations for more discriminative representations. Furthermore, an explicit probability-based boundary-adaptive regression (EPBR) method is developed to enhance the global shape constraints and further search the semantically consistent landmarks in the predicted boundary for robust facial landmark detection. It is interesting to show that the proposed MMDN can generate more accurate boundary-adaptive landmark heatmaps and effectively enhance shape constraints to the predicted landmarks for faces with large pose variations and heavy occlusions. Experimental results on challenging benchmark data sets demonstrate the superiority of our MMDN over state-of-the-art facial landmark detection methods. Jun Wan 0005, Zhihui Lai 0001, Jing Li 0055, Jie Zhou 0009, Can Gao |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2021 | Siamese Guided Anchoring Network for Visual TrackingabstractRecently, the Siamese Region Proposal Network (SiamRPN) has been widely explored in tracking and achieved remarkable performance. However, the existing SiamRPN-based method uses a predefined and highly dependent on prior knowledge anchor, which limits the tracking accuracy. Besides, when the target changes drastically, the anchor box obtained by the SiamRPN-based method also has some negative samples, which leads to a decrease inaccuracy. To address these issues, this paper proposes a siamese guided anchoring network for visual tracking, which can obtain more representative anchors by estimating the position and shape of the target, reducing the adverse effects of negative samples. At the same time, a feature adaption module is proposed to adapt to the target scale change for learning more discriminative and useful features and achieving more accurate visual tracking. Extensive experiments on challenging OTB100 and VOT2018 datasets demonstrate the competitive performance of the proposed algorithm in comparison with the state-of-the-art trackers. Jing Li 0055, Yafu Xiao, Jun Wan 0005 |
IJCNN | 2 |
| 2021 | Multi-task adversarial autoencoder network for face alignment in the wild
Xiaoqian Yue, Jing Li 0055, Jia Wu 0001, Jun Wan 0005, Jinyan Ma |
Neurocomputing | 2 |
| 2021 | A hybrid neural network approach to combine textual information and rating information for item recommendation
Donghua Liu, Jing Li 0055, Bo Du 0001, Rong Gao 0001, Yujia Wu |
Knowl. Inf. Syst. | 2 |
| 2021 | Learning adaptive updating siamese network for visual tracking
Jing Li 0055, Bo Du 0001, Zhiquan Ding, Tianqi Qin |
Multim. Tools Appl. | 2 |
| 2020 | Support Correlation Filters Tracking using Mask MatrixabstractSupport correlation filter tracking method uses cyclic sampling to transform the calculation into frequency domain, which solves the problems of sampling and large computation of support vector machine. However, the current method can not exploit the information of backgrounds because all samples are generated by cyclic sampling around the target in the tracking process. To solve this problem, this paper proposes a background awareness support correlation filter tracking method using mask matrix. In the tracking process, the mask matrix is used to extract the patchs densely from background as negative samples, so the background information is used effectively. Experiments on OTB100 database show that compared with Scale Kerneling Supported Correlation Filtering (SKSCF), the proposed algorithm achieves a gain of 4.2% in mean OP and 6.2% AUC score respectively. Zhenyang Su, Jing Li 0055, Zhiquan Ding, Tianqi Qin, Yafu Xiao |
IJCNN | 2 |
| 2020 | Text Classification using Triplet Capsule NetworksabstractMost existing methods only consider the local features of the samples, and their experimental results show better performance than traditional Non-deep learning methods. However, in these methods, the global features of the sample space are usually ignored, and these ignored global features will affect the classification accuracy. To solve this problem, a novel triple capsule network framework is proposed to text classification. The training in the first stage, to obtain a basic capsule network for obtaining local features. Then, three capsule networks sharing parameters are combined spatially, and the triplet loss function is used in the second stage of training. By comparative learning, the capsule network can learn global features that can represent the spatial distance between different categories. Through comparison experiments on six datasets and ten general benchmark algorithms, the results show that our results is the first in the four datasets. Yujia Wu, Jing Li 0055, Zhiquan Ding |
IJCNN | 2 |
| 2020 | Real-time visual tracking using complementary kernel support correlation filters
Zhenyang Su, Jing Li 0055, Bo Du 0001, Yafu Xiao |
Frontiers Comput. Sci. | 2 |
| 2020 | Siamese capsule networks with global and local features for text classification
Yujia Wu, Jing Li 0055, Jia Wu 0001 |
Neurocomputing | 2 |
| 2020 | Learning spatial-temporally regularized complementary kernelized correlation filters for visual tracking
Zhenyang Su, Jing Li 0055, Chengfang Song, Yafu Xiao, Jun Wan 0005 |
Multim. Tools Appl. | 2 |
| 2020 | A target response adaptive correlation filter tracker with spatial attention
Jing Li 0055, Bo Du 0001, Yafu Xiao |
Multim. Tools Appl. | 2 |
| 2020 | Robust face alignment by cascaded regression and de-occlusion
Jun Wan 0005, Jing Li 0055, Zhihui Lai 0001, Bo Du 0001, Lefei Zhang |
Neural Networks | 2 |
| 2020 | MeMu: Metric correlation Siamese network and multi-class negative sampling for visual tracking
Yafu Xiao, Jing Li 0055, Bo Du 0001, Jia Wu 0001, Wenfan Zhang |
Pattern Recognit. | 2 |
| 2019 | DRCGR: Deep Reinforcement Learning Framework Incorporating CNN and GAN-Based for Interactive RecommendationabstractRecently, the application of deep reinforcement learning into the field of session-based interactive recommendation has attracted great attention from researchers. However, despite that some interactive recommendation models based on deep reinforcement learning have been proposed, they still suffer to the following limitations: (1) these works ignore the skip behaviors of sequential patterns in users' clicking behavior; (2) these works fail to incorporate positive feedback and negative feedback into the proposed deep reinforcement recommender system when the positive feedback is sparse. Therefore, to solve the problems mentioned above, a novel Deep Q-Network based recommendation framework incorporating CNN and GAN-based models is proposed to acquire robust performance, named DRCGR. Specifically, in DRCGR, a CNN model is used to capture the sequential features for positive feedback. Then, an adversarial training is adopted to learn optimal negative feedback representations Then, positive/negative representations are fed into DQN simultaneously, which are conducive to generating better action-value function The experimental results based on real-world e-commerce data demonstrate our framework's superiority over some state-of-the-art recommendation models. Rong Gao 0001, Haifeng Xia, Jing Li 0055, Donghua Liu, Gang Chun |
ICDM | 3 |
| 2019 | Correlation Filter Tracking Method via Metric Learning and Adaptive Multi-stage Appearance
Yan Hong 0005, Jing Li 0055, Yafu Xiao, Wenfan Zhang, Chengfang Song, Shan Xue 0001 |
IJCNN | 2 |
| 2019 | DAML: Dual Attention Mutual Learning between Ratings and Reviews for Item RecommendationabstractDespite the great success of many matrix factorization based collaborative filtering approaches, there is still much space for improvement in recommender system field. One main obstacle is the cold-start and data sparseness problem, requiring better solutions. Recent studies have attempted to integrate review information into rating prediction. However, there are two main problems: (1) most of existing works utilize a static and independent method to extract the latent feature representation of user and item reviews ignoring the correlation between the latent features, which may fail to capture the preference of users comprehensively. (2) there is no effective framework that unifies ratings and reviews. Therefore, we propose a novel d ual a ttention m utual l earning between ratings and reviews for item recommendation, named DAML. Specifically, we utilize local and mutual attention of the convolutional neural network to jointly learn the features of reviews to enhance the interpretability of the proposed DAML model. Then the rating features and review features are integrated into a unified neural network model, and the higher-order nonlinear interaction of features are realized by the neural factorization machines to complete the final rating prediction. Experiments on the five real-world datasets show that DAML achieves significantly better rating prediction accuracy compared to the state-of-the-art methods. Furthermore, the attention mechanism can highlight the relevant information in reviews to increase the interpretability of rating prediction. Donghua Liu, Jing Li 0055, Bo Du 0001, Rong Gao 0001 |
KDD | 2 |
| 2019 | Face alignment by Component Adaptive Mechanism
Jun Wan 0005, Jing Li 0055, Yujia Wu, Yafu Xiao, Xuefei Li 0001 |
Neurocomputing | 2 |
| 2019 | Robust correlation filter tracking with multi-scale spatial view
Yafu Xiao, Jing Li 0055, Bo Du 0001, Jia Wu 0001, Xuefei Li 0001 |
Neurocomputing | 2 |
| 2018 | Correlation Filter Tracking with Multiscale Spatial ViewabstractVisual tracking has already become one of the most important research focuses in the field of computer vision. Due to such interference as serious occlusion or severe illumination change and so on, the appearance model of the target tends to vary heavily, posing great challenges on tracking. In this paper, a correlation filter tracking with multiscale spatial view (CFMSV) is proposed in which a group of multiscale spatial filters of different view areas is established. We adopt the filters to perform collaborative location by introducing the method of pre-location and exploiting the multiscale spatial view around the target. A large number of tracking experiments have been made, which confirmed that the CFMSV tracking method proposed in our work is superior to the state-of-the-art methods in tracking performance. Yafu Xiao, Jing Li 0055, Wenfan Zhang |
IJCNN | 2 |
| 2018 | STSCR: Exploring spatial-temporal sequential influence and social information for location recommendation
Rong Gao 0001, Jing Li 0055, Xuefei Li 0001, Chengfang Song, Donghua Liu |
Neurocomputing | 2 |
| 2018 | A personalized point-of-interest recommendation model via fusion of geo-social information
Rong Gao 0001, Jing Li 0055, Xuefei Li 0001, Chengfang Song |
Neurocomputing | 2 |
| 2016 | Efficient compressive sensing tracking via mixed classifier decision
Jing Li 0055, Bo Du 0001, Zhenyang Su |
Sci. China Inf. Sci. | 2 |