EDBT 2026 Demo / reviewers in the wild / expert
Jie Cao 0014
dblp:39/6191-14
· DBLP profile ↗
38ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0003-0481-5170ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 6 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 2 first-author · 10 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A sparse large-scale multi-objective optimization algorithm based on growing neural gas network and variable sparsity analysis
Jie Cao 0014, Zuohan Chen, Jianlin Zhang 0002 |
Inf. Sci. | 1 |
| 2026 | Attention-Enhanced Cross-Modality Alignment for Adapting Vision-Language ModelsabstractAbstract Prompt learning is an effective approach for adapting pre-trained vision-language models (VLMs) to a variety of downstream tasks. However, prompts designed manually or generated by large language models may not effectively capture key discriminative visual features. In addition, pre-trained VLMs may not align images and text well at a fine-grained level. To address these two issues, we propose an attention-enhanced cross-modality alignment network, which includes an adaptive channel attention (ACA) module and a cross-modal measurement (CMM) module. The ACA module adapts the existing efficient channel attention to highlight discriminative visual and textual features. The CMM module leverages four pairs of image-text similarities across both frozen and learnable branches, improving the alignment of fine-grained discriminative visual and textual features. Experiments show that the proposed method outperforms state-of-the-art methods on two representative tasks: base-to-novel generalization and cross-dataset evaluation. Our code is available at https://github.com/xueshaoying/XSY_AECA.git . Shaoying Xue, Jie Cao 0014, Zhanyu Ma |
Mach. Learn. | 4 |
| 2026 | Guiding Hyperspectral Tracking: A Cross-Modal Spatial Prompt and Dynamic Semantic Memory FrameworkabstractHyperspectral video tracking benefits from material-aware spectral cues, but still suffers from spatial ambiguity, noisy/redundant bands, and temporal drift under appearance changes. To address these issues, we propose CST-Track, a DaSSP-Net-based framework that jointly performs intra-frame spatial calibration and inter-frame semantic updating. Specifically, the Cross-Modal Prompt Attention Module (CPAM) uses same-frame false-color cues as spatial prompts and applies soft spectral gating to enhance reliable hyperspectral responses. The Temporal Spectral-Semantic Updater (TSSU) further maintains a confidence-gated dynamic prompt through domain-aware projection, reducing unreliable memory updates. Experiments on HOT2022 and HOT2024 show that CST-Track achieves state-of-the-art performance. Compared with the DaSSP-Net baseline, CST-Track improves AUC/DP@20P from 0.682/0.917 to 0.705/0.958 on HOT2022. On HOT2024-VIS/NIR/RedNIR, it further yields AUC gains of 4.9/3.5/7.8 and DP@20P gains of 3.9/2.2/7.4 percentage points over the baseline, respectively, and surpasses the second-best competing SOTA tracker by 2.7 percentage points in AUC while running at about 22 FPS on HOT2024 subsets. Yuan Shanqin, Jie Cao 0014, Haopeng Liang, Jinhua Wang 0005 |
IEEE Signal Process. Lett. | 2 |
| 2026 | SCNNnet: Self-Attention Enhancement and Cross-Correlation Nearest Neighbor Network for Few-Shot Fine-Grained Social Visual Recognition
Yalin Wang 0012, Jie Cao 0014 |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2026 | Fine-Tuning via Linked Domains: A Closed-Form Dual Alignment Mechanism for Transferring Vision-Language ModelsabstractAdapters and prompt learning have become two de facto strategies to fine-tune pre-trained vision-language models, mitigating the high computational cost of fine-tuning an entire model for downstream tasks. They can align the prediction from the fine-tuned model with that from the pre-trained model. However, the existing methods of these strategies primarily focus on aligning within a single modality, and the exploration of bidirectional interactions between modalities remains limited. To address this issue, we propose a closed-form dual alignment mechanism (DAM) thatnot only ensures the consistency in predictions within a single modality but also achieves the alignment of features across different modalities. In DAM, all alignments are achieved by closed-form solutions to ridge regression, without inducing a massive number of learnable parameters. Experimental results demonstrate that DAM outperforms the state-of-the-art methods on 11 benchmarks over various evaluation metrics. Our codes are available at https://github.com/Peiy-Lu/DAM. Peiyu Lu, Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | PMFNet: collaborative prompt enhancement and dynamic fusion for robust RGB-thermal tracking
Jie Cao 0014, Haopeng Liang, Xiaoyang Shi, Yuan Shanqin |
Vis. Comput. | 1 |
| 2025 | A neural network guided dual-space search evolutionary algorithm for large scale multi-objective optimization
Jie Cao 0014, Zuohan Chen, Jianlin Zhang 0002 |
Eng. Appl. Artif. Intell. | 1 |
| 2025 | A Deep Q-Network-driven multi-objective evolutionary algorithm for distributed heterogeneous hybrid flow shop scheduling with worker fatigue
Jianlin Zhang 0002, Longbin Ma, Jie Cao 0014, Zuohan Chen, Tianpeng Xu |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Dual-space distribution metric-based evolutionary algorithm for multimodal multi-objective optimization
Jie Cao 0014, Zuohan Chen, Jianlin Zhang 0002 |
Expert Syst. Appl. | 1 |
| 2025 | A multi-task optimization algorithm via reinforcement learning for multimodal multi-objective optimization
Jie Cao 0014, Yuze Yang, Jianlin Zhang 0002, Zuohan Chen, Zongli Liu |
Expert Syst. Appl. | 1 |
| 2025 | SRML: Structure-relation mutual learning network for few-shot image classificationabstractFew-shot image classification aims at tackling a challenging but practical classification setting, where only few labelled images are available for training. Metric-based methods are main-stream solutions for few-shot image classification, but many of them extract features that are either irrelevant to target objects in the query images or insufficient to describe the local shape or structural patterns within images, which can lead to mis-identification of the target objects, especially when the images are of multiple objects. To resolve this issue, we propose the structure-relation mutual learning (SRML) network, which first learns both the intra-image structural features and the inter-image relational features in a parallel fashion via two parallel branches, the structural feature extractor (SFE) and the relational feature extractor (RFE), and then harnesses mutual learning to enable knowledge exchange between them. In such a manner, the structural features learnt from the SFE branch not only contain the structural patterns within the images, but also focus more on the target objects, guided by the relational knowledge from the RFE branch. In return, the RFE branch can exploit the more-focused structural knowledge to better match the target objects in the support and query images. We conduct extensive experiments on four few-shot classification benchmark datasets to showcase the superior classification of the proposed SRML network, achieving a 3.17% improvement in classification accuracy over the leading competitor, RENet Kang et al. (2021). The code of this work can be found in https://github.com/Rilliant7/SRML . Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
Pattern Recognit. | 5 |
| 2025 | Rise by Lifting Others: Interacting Features to Uplift Few-Shot Fine-Grained ClassificationabstractFew-shot fine-grained classification entails notorious subtle inter-class variation. Recent works address this challenge by developing attention mechanisms, such as the task discrepancy maximization (TDM) that can highlight discriminative channels. This paper, however, aims to reveal that, besides designing sophisticated attention modules, a well-designed input scheme, which simply blends two types of features and their interactions capturing different properties of the target object, can also greatly promote the quality of the learnt weights. To illustrate, we design a bi-feature interactive TDM (BiFI-TDM) module to serve as a strong foundation for TDM to discover the most discriminative channels with ease. Specifically, we design a novel mixing strategy to produce four sets of channel weights with different focuses, reflecting the properties of the corresponding input features and their interactions, as well as a proper feature re-weighting scheme. Extensive experiments on four benchmark fine-grained image datasets showcase superior performance of BiFI-TDM in metric-based few-shot methods. Our codes are available athttps://github.com/Peiy-Lu/BiFI-TDM. Peiyu Lu, Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Selectively Augmented Attention Network for Few-Shot Image ClassificationabstractFew-shot image classification is a challenging task that aims to learn from a limited number of labelled training images a classification model that can be generalised to unseen classes. Two strategies are usually taken to improve the classification performances of few-shot image classifiers: either applying data augmentation to enlarge the sample size of the training set and reduce overfitting, or involving attention mechanisms to highlight discriminative spatial regions or channels. However, naively applying them to few-shot classifiers directly and separately may lead to undesirable results; for example, some augmented images may focus majorly on the background rather than the object, which brings additional noises to the training process. In this paper, we propose a unified framework, the selectively augmented attention (SAA) network, that carefully integrates the best of the two approaches in an end-to-end fashion via a selective best match module to select the most representative images from the augmented training set. The selected images tend to concentrate on the objects with less irrelevant background, which can assist the subsequent calculation of attentions by alleviating the interference from background. Moreover, we design a joint attention module to jointly learn both the spatial and channel-wise attentions. Experimental results on four benchmark datasets showcase the superior classification performance of the proposed SAA network compared with the state-of-the-arts. Rui Zhu 0006, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2024 | A Hyperheuristic and Reinforcement Learning Guided Meta-heuristic Algorithm RecommendationabstractAutomatic selection of the most appropriate algorithms for complex optimization problems has emerged as a cutting-edge trend in artificial intelligence. This approach circumvents the interpretability challenges posed through trial and error. A hyperheuristic and reinforcement learning-guided meta-heuristic algorithm recommendation (HHRL-MAR) is proposed to facilitate the adaptive selection of a diverse array of meta-heuristic algorithms tailored to the unique characteristics of various problems in this paper. To this end, four meta-heuristics with distinct advantages are integrated to form the action space within the reinforcement learning, serving as the low-level heuristic for hyperheuristic. The incorporated reward mechanism based on the real-time state of the population enhances both the flexibility and accuracy of the algorithm. Three selection strategies in light of simulated annealing and ε–greedy are avoid premature convergence associated with designed to a singular selection approach. The experimental results show the efficacy of HHRL-MAR for large-scale complex continuous optimization in terms of accuracy, stability, and convergence speed. Ningning Zhu, Fuqing Zhao, Jie Cao 0014 |
CSCWD | 3 |
| 2024 | Unsupervised Signal Anomaly Transformer method: Achieving bearing life anomaly detection without the need for failure samples
Ping Yu 0004, Mengmeng Ping, Jialin Ma, Jie Cao 0014 |
Eng. Appl. Artif. Intell. | 4 |
| 2024 | A constrained multi-objective evolutionary algorithm with Pareto estimation via neural network
Zongli Liu, Jie Cao 0014, Jianlin Zhang 0002, Zuohan Chen |
Expert Syst. Appl. | 3 |
| 2024 | Bi-Directional Ensemble Feature Reconstruction Network for Few-Shot Fine-Grained ClassificationabstractThe main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting - a quick pilot study reveals that they in fact push for the opposite (i.e., lower inter-class variations and higher intra-class variations). To alleviate this problem, prior works predominately use a support set to reconstruct the query image and then utilize metric learning to determine its category. Upon careful inspection, we further reveal that such unidirectional reconstruction methods only help to increase inter-class variations and are not effective in tackling intra-class variations. In this paper, we introduce a bi-reconstruction mechanism that can simultaneously accommodate for inter-class and intra-class variations. In addition to using the support set to reconstruct the query set for increasing inter-class variations, we further use the query set to reconstruct the support set for reducing intra-class variations. This design effectively helps the model to explore more subtle and discriminative features which is key for the fine-grained problem in hand. Furthermore, we also construct a self-reconstruction module to work alongside the bi-directional module to make the features even more discriminative. We introduce the snapshot ensemble method in the episodic learning strategy - a simple trick to further improve model performance without increasing training costs. Experimental results on three widely used fine-grained image classification datasets, as well as general and cross-domain few-shot image datasets, consistently show considerable improvements compared with other methods. Jijie Wu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jie Cao 0014, Jun Guo 0002, Yi-Zhe Song |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2023 | Bi-directional Feature Reconstruction Network for Fine-Grained Few-Shot Image ClassificationabstractThe main challenge for fine-grained few-shot image classification is to learn feature representations with higher inter-class and lower intra-class variations, with a mere few labelled samples. Conventional few-shot learning methods however cannot be naively adopted for this fine-grained setting -- a quick pilot study reveals that they in fact push for the opposite (i.e., lower inter-class variations and higher intra-class variations). To alleviate this problem, prior works predominately use a support set to reconstruct the query image and then utilize metric learning to determine its category. Upon careful inspection, we further reveal that such unidirectional reconstruction methods only help to increase inter-class variations and are not effective in tackling intra-class variations. In this paper, we for the first time introduce a bi-reconstruction mechanism that can simultaneously accommodate for inter-class and intra-class variations. In addition to using the support set to reconstruct the query set for increasing inter-class variations, we further use the query set to reconstruct the support set for reducing intra-class variations. This design effectively helps the model to explore more subtle and discriminative features which is key for the fine-grained problem in hand. Furthermore, we also construct a self-reconstruction module to work alongside the bi-directional module to make the features even more discriminative. Experimental results on three widely used fine-grained image classification datasets consistently show considerable improvements compared with other methods. Codes are available at: https://github.com/PRIS-CV/Bi-FRN. Jijie Wu, Dongliang Chang, Aneeshan Sain, Zhanyu Ma, Jie Cao 0014, Jun Guo 0002, Yi-Zhe Song |
AAAI | 6 |
| 2023 | A Pareto front estimation-based constrained multi-objective evolutionary algorithm
Jie Cao 0014, Zesen Yan, Zuohan Chen, Jianlin Zhang 0002 |
Appl. Intell. | 1 |
| 2023 | An estimation of distribution algorithm with multiple intensification strategies for two-stage hybrid flow-shop scheduling problem with sequence-dependent setup time
Huan Liu 0001, Fuqing Zhao, Ling Wang 0001, Jie Cao 0014, Jianxin Tang, Jonrinaldi |
Appl. Intell. | 4 |
| 2023 | A knowledge-driven co-evolutionary algorithm assisted by cross-regional interactive learning
Ningning Zhu, Fuqing Zhao, Jie Cao 0014, Jonrinaldi |
Eng. Appl. Artif. Intell. | 3 |
| 2023 | A dual-stage large-scale multi-objective evolutionary algorithm with dynamic learning strategy
Jie Cao 0014, Kaiyue Guo, Jianlin Zhang 0002, Zuohan Chen |
Expert Syst. Appl. | 1 |
| 2023 | ReNAP: Relation network with adaptiveprototypical learning for few-shot classification
Yalan Li, Yixiao Zheng, Rui Zhu 0006, Zhanyu Ma, Jing-Hao Xue, Jie Cao 0014 |
Neurocomputing | 7 |
| 2022 | An Angle-based Many-Objective evolutionary algorithm with Shift-based density estimation and sum of objectives
Jianlin Zhang 0002, Jie Cao 0014, Fuqing Zhao, Zuohan Chen |
Expert Syst. Appl. | 2 |
| 2022 | A multipopulation cooperative coevolutionary whale optimization algorithm with a two-stage orthogonal learning mechanism
Fuqing Zhao, Haizhu Bao, Ling Wang 0001, Jie Cao 0014, Jianxin Tang, Jonrinaldi |
Knowl. Based Syst. | 4 |
| 2022 | A novel hyperchaotic image encryption algorithm with simultaneous shuffling and diffusion
Xiangquan Gui, Shouliang Li, Jie Cao 0014 |
Multim. Tools Appl. | 5 |
| 2021 | A hierarchical knowledge guided backtracking search algorithm with self-learning strategy
Fuqing Zhao, Ling Wang 0001, Jie Cao 0014, Jianxin Tang |
Eng. Appl. Artif. Intell. | 4 |
| 2021 | A two-stage evolutionary strategy based MOEA/D to multi-objective problems
Jie Cao 0014, Jianlin Zhang 0002, Fuqing Zhao, Zuohan Chen |
Expert Syst. Appl. | 1 |
| 2021 | A hierarchical guidance strategy assisted fruit fly optimization algorithm with cooperative learning mechanism
Fuqing Zhao, Ruiqing Ding, Ling Wang 0001, Jie Cao 0014, Jianxin Tang |
Expert Syst. Appl. | 4 |
| 2021 | Candidate box fusion based approach to adjust position of the candidate box for object detectionabstractAbstract The method of object detection has been applied to all aspects in our lives. Although object detection methods based on deep learning have been widely used in various fields, there are still some overlooked problems in the candidate box selection stage. The detection results of traditional candidate box selection methods can only select a relatively optimal maximum candidate box. If the maximum candidate box is still not accurate enough, this type of methods will not be able to do adjust it. To solve this problem, an object detection method based on the multiple candidate box fusion is proposed. The method can not only retain the maximum candidate box and delete the non‐maximum candidate box, but also adjust the position of the maximum candidate box again. Thereby a more accurate maximum candidate box can be obtained. In order to verify the generalization ability of the method, the candidate box fusion method is combined with the two object detection frameworks: faster R‐CNN model and YOLOv3 model. The results of these experiments prove that the proposed method can achieve higher detection accuracy and complete the object detection task more effectively. Jie Cao 0014, Zuohan Chen |
IET Image Process. | 1 |
| 2021 | Deep InterBoost networks for small-sample image classification
Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002 |
Neurocomputing | 6 |
| 2021 | QoS-oriented joint optimization of concurrent scheduling and power control in millimeter wave mesh backhaul network
Zhongyu Ma, Jie Cao 0014, Qun Guo 0001, Xiangwei Li, Hongfeng Ma |
J. Netw. Comput. Appl. | 2 |
| 2021 | ReMarNet: Conjoint Relation and Margin Learning for Small-Sample Image ClassificationabstractDespite achieving state-of-the-art performance, deep learning methods generally require a large amount of labeled data during training and may suffer from overfitting when the sample size is small. To ensure good generalizability of deep networks under small sample sizes, learning discriminative features is crucial. To this end, several loss functions have been proposed to encourage large intra-class compactness and inter-class separability. In this paper, we propose to enhance the discriminative power of features from a new perspective by introducing a novel neural network termed Relation-and-Margin learning Network (ReMarNet). Our method assembles two networks of different backbones so as to learn the features that can perform excellently in both of the aforementioned two classification mechanisms. Specifically, a relation network is used to learn the features that can support classification based on the similarity between a sample and a class prototype; at the meantime, a fully connected network with the cross entropy loss is used for classification via the decision boundary. Experiments on four image datasets demonstrate that our approach is effective in learning discriminative features from a small set of labeled samples and achieves competitive performance against state-of-the-art methods. Code is available at https://github.com/liyunyu08/ReMarNet. Liyun Yu, Zhanyu Ma, Jing-Hao Xue, Jie Cao 0014, Jun Guo 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | BSNet: Bi-Similarity Network for Few-shot Fine-grained Image ClassificationabstractFew-shot learning for fine-grained image classification has gained recent attention in computer vision. Among the approaches for few-shot learning, due to the simplicity and effectiveness, metric-based methods are favorably state-of-the-art on many tasks. Most of the metric-based methods assume a single similarity measure and thus obtain a single feature space. However, if samples can simultaneously be well classified via two distinct similarity measures, the samples within a class can distribute more compactly in a smaller feature space, producing more discriminative feature maps. Motivated by this, we propose a so-called Bi-Similarity Network (BSNet) that consists of a single embedding module and a bi-similarity module of two similarity measures. After the support images and the query images pass through the convolution-based embedding module, the bi-similarity module learns feature maps according to two similarity measures of diverse characteristics. In this way, the model is enabled to learn more discriminative and less similarity-biased features from few shots of fine-grained images, such that the model generalization ability can be significantly improved. Through extensive experiments by slightly modifying established metric/similarity based networks, we show that the proposed approach produces a substantial improvement on several fine-grained image benchmark datasets. Codes are available at: https://github.com/PRIS-CV/BSNet. Jijie Wu, Zhanyu Ma, Jie Cao 0014, Jing-Hao Xue |
IEEE Trans. Image Process. | 5 |
| 2020 | A novel parallel accelerated CRPF algorithm
Jinhua Wang 0005, Jie Cao 0014, Ping Yu 0004, Kaijie Huang |
Appl. Intell. | 2 |
| 2020 | OSLNet: Deep Small-Sample Classification With an Orthogonal Softmax LayerabstractA deep neural network of multiple nonlinear layers forms a large function space, which can easily lead to overfitting when it encounters small-sample data. To mitigate overfitting in small-sample classification, learning more discriminative features from small-sample data is becoming a new trend. To this end, this paper aims to find a subspace of neural networks that can facilitate a large decision margin. Specifically, we propose the Orthogonal Softmax Layer (OSL), which makes the weight vectors in the classification layer remain orthogonal during both the training and test processes. The Rademacher complexity of a network using the OSL is only 1/K, where K is the number of classes, of that of a network using the fully connected classification layer, leading to a tighter generalization error bound. Experimental results demonstrate that the proposed OSL has better performance than the methods used for comparison on four small-sample benchmark datasets, as well as its applicability to large-sample datasets. Codes are available at: https://github.com/dongliangchang/OSLNet. Dongliang Chang, Zhanyu Ma, Zheng-Hua Tan, Jing-Hao Xue, Jie Cao 0014, Jingyi Yu 0001, Jun Guo 0002 |
IEEE Trans. Image Process. | 6 |
| 2018 | A hybrid short-term traffic flow forecasting model based on time series multifractal characteristics
Jie Cao 0014, Minan Tang, Yirong Guo |
Appl. Intell. | 3 |
| 2018 | A multivariate short-term traffic flow forecasting method based on wavelet analysis and seasonal time series
Jie Cao 0014, Minan Tang, Yirong Guo |
Appl. Intell. | 3 |