EDBT 2026 Demo / reviewers in the wild / expert
Jialiang Tang
dblp:279/0611
· DBLP profile ↗
39ranked-venue papers
12as first author
37since 2021 · last 2025
0000-0001-7347-4790ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 26 · 7 first-author · 24 since 2021Systems, architecture and hardware · 9 · 3 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Hybrid Data-Free Knowledge DistillationabstractData-free knowledge distillation aims to learn a compact student network from a pre-trained large teacher network without using the original training data of the teacher network. Existing collection-based and generation-based methods train student networks by collecting massive real examples and generating synthetic examples, respectively. However, they inevitably become weak in practical scenarios due to the difficulties in gathering or emulating sufficient real-world data. To solve this problem, we propose a novel method called Hybrid Data-Free Distillation (HiDFD), which leverages only a small amount of collected data as well as generates sufficient examples for training student networks. Our HiDFD comprises two primary modules, i.e., the teacher-guided generation and student distillation. The teacher-guided generation module guides a Generative Adversarial Network (GAN) by the teacher network to produce high-quality synthetic examples from very few real-world collected examples. Specifically, we design a feature integration mechanism to prevent the GAN from overfitting and facilitate the reliable representation learning from the teacher network. Meanwhile, we drive a category frequency smoothing technique via the teacher network to balance the generative training of each category. In the student distillation module, we explore a data inflation strategy to properly utilize a blend of real and synthetic data to train the student network via a classifier-sharing-based feature alignment technique. Intensive experiments across multiple benchmarks demonstrate that our HiDFD can achieve state-of-the-art performance using 120 times less collected data than existing methods. Jialiang Tang, Shuo Chen 0003, Chen Gong 0002 |
AAAI | 1 |
| 2025 | Learn from Balance: Rectifying Knowledge Transfer for Long-Tailed ScenariosabstractKnowledge Distillation (KD) transfers knowledge from a large pre-trained teacher network to a compact and efficient student network, making it suitable for deployment on resource-limited media terminals. However, traditional KD methods require balanced data to ensure robust training, which is often unavailable in practical applications. In such scenarios, a few head categories occupy a substantial proportion of examples. This imbalance biases the trained teacher network towards the head categories, resulting in severe performance degradation on the less represented tail categories for both the teacher and student networks. In this paper, we propose a novel framework called Knowledge Rectification Distillation (KRDistill) to address the imbalanced knowledge inherited in the teacher network through the incorporation of the balanced category priors. Furthermore, we rectify the biased predictions produced by the teacher network, particularly focusing on the tail categories. Consequently, the teacher network can provide balanced and accurate knowledge to train a reliable student network. Intensive experiments conducted on various long-tailed datasets demonstrate that our KRDistill can effectively train reliable student networks in realistic scenarios of data imbalance. Xinlei Huang, Jialiang Tang, Xubin Zheng, Jinjia Zhou, Wenxin Yu 0001, Ning Jiang 0002 |
ICASSP | 2 |
| 2025 | Empowering Large Language Models for Time Series Forecasting with Patterns and SemanticsabstractTime Series Forecasting (TSF) is critical in many real-world domains like financial planning and health mon-itoring. Recent studies have revealed that Large Language Models (LLMs), with their powerful in-contextual modeling capabilities, hold significant potential for TSF. However, existing LLM-based methods usually perform suboptimally because they neglect the inherent characteristics of time series data. Unlike the textual data used in LLM pre-training, the time series data is semantically sparse and comprises distinctive tempo-ral patterns. To address this problem, we propose LLM-PS to empower the LLM for TSF by learning the fundamental Patterns and meaningful Semantics from time series data. Our LLM-PS incorporates a new multi-scale convolutional neural network adept at capturing both short-term fluctuations and long-term trends within the time series. Meanwhile, we introduce a time-to-text module for extracting valuable semantics across continuous time intervals rather than isolated time points. By integrating these patterns and semantics, LLM - PS effectively models temporal dependencies, enabling a deep comprehension of time series and delivering accurate forecasts. Intensive exper-imental results demonstrate that LLM-PS achieves state-of-the-art performance in both short- and long-term forecasting tasks, as well as in few- and zero-shot settings. Code is available at https://github.com/tangjialiang97ILLMPS. Jialiang Tang, Shuo Chen 0003, Chen Gong 0002, Jing Zhang 0037, Dacheng Tao |
ICDM | 1 |
| 2024 | Open-World Semi-Supervised Learning under Compound Distribution Shifts
Shijia Xu, Lin Zhao 0003, Jialiang Tang, Chen Gong 0002 |
BMVC | 3 |
| 2024 | Direct Distillation Between Different Domains
Jialiang Tang, Shuo Chen 0003, Gang Niu 0001, Hongyuan Zhu 0002, Joey Tianyi Zhou, Chen Gong 0002, Masashi Sugiyama |
ECCV (80) | 1 |
| 2024 | ClearKD: Clear Knowledge Distillation for Medical Image ClassificationabstractIn recent years, computer-aided diagnosis (CAD) systems employing convolutional neural networks (CNNs) have achieved remarkable performance in medical image classification tasks. Despite this, deploying CNN-based CAD systems on medical equipment presents challenges due to their enormous computational and storage resource requirements. In this case, knowledge distillation reduces the cost of deploying CNNs by guiding a lightweight student network to learn from a robust teacher network. However, medical images have higher inter-class similarity than natural images, which makes it difficult for the teacher network to provide clear and accurate classification knowledge to the student network, resulting in the performance degradation of the student network. To address this problem, we divide the teacher predictions into clear predictions, ambiguous predictions, and misclassified predictions to analyze the interference caused by the similarity of medical images on knowledge distillation and propose a novel knowledge distillation frame-work, termed ClearKD. By enhancing ambiguous predictions and misclassified predictions with clear predictions as a reference, our ClearKD consistently provides high-quality teacher classification knowledge to the student network, increasing the ability of the student network to distinguish medical images. The experimental results on the skin lesions classification datasets (ISIC2019) and the brain tumor dataset demonstrate that our ClearKD outperforms existing state-of-the-art knowledge distillation methods in medical image classification tasks. Xinlei Huang, Ning Jiang 0002, Jialiang Tang |
IJCNN | 3 |
| 2024 | Decoupled Multi-teacher Knowledge Distillation based on EntropyabstractMulti-teacher knowledge distillation (MKD) aims to leverage the valuable and diverse knowledge presented by multiple teacher networks to improve the performance of the student network. Existing approaches typically rely on simple methods such as averaging the prediction logits or using sub-optimal weighting strategies to combine knowledge from multiple teachers. However, employing these techniques cannot fully reflect the importance of teachers and may even mislead student’s learning. To address these issues, we propose a novel Decoupled Multi-teacher Knowledge Distillation based on Entropy (DE-MKD). DE-MKD decomposes the vanilla KD loss and assigns weights to each teacher to reflect its importance based on the entropy of their predictions. Furthermore, we extend the proposed approach to distill the intermediate features from teachers to further improve the performance of the student network. Extensive experiments conducted on the publicly available CIFAR-100 image classification dataset demonstrate the effectiveness and flexibility of our proposed approach. Xin Cheng 0004, Jialiang Tang, Wenxin Yu 0001, Ning Jiang 0002, Jinjia Zhou |
ISCAS | 2 |
| 2024 | Amalgamating Knowledge for Comprehensive Classification with Uncertainty SuppressionabstractKnowledge distillation(KD) aims to obtain a lightweight student network with the target dataset's pre-trained network(s). In practical applications, the student network distilled on one dataset may fail to make fine-grained classifications of multiple categories(such as birds and dogs). To this end and to make better use of various datasets' pre-trained models, knowledge amalgamation (KA) strives to integrate the knowledge of multiple expert models trained on different datasets to attain a student network with multi-expert knowledge. Proposed KA methods for image classification ignore the problem that teacher networks may encounter with untrained class samples and provide misleading guidance to the student network. To address this problem, we propose a knowledge amalgamation framework based on uncertainty suppression. A series of experiments demonstrate the effectiveness of our framework; some of the experiments yield an accuracy improvement of 2% compared to the proposed methods. Lebin Li, Ning Jiang 0002, Jialiang Tang, Xinlei Huang |
ISCAS | 3 |
| 2024 | Adaptive Informative Semantic Knowledge Transfer for Knowledge DistillationabstractKnowledge distillation aims to improve the generalization capacity of the student model by transferring knowledge from the teacher model. Existing feature-based methods explore knowledge transfer through hand-crafted feature mappings between teacher-student pairs. However, in different layers, the knowledge volume varies, and the knowledge exhibits semantic gaps. This leads to the possibility that hand-crafted layer associations may not enable the student model to effectively learn knowledge from the teacher model. We address this problem from two angles. On one hand, to ensure maximum knowledge transfer, we propose adaptive feature mapping based on the effective receptive field, which can quantify the knowledge volume of different layers and thus establish the optimal knowledge transfer paths between teacher-student pairs. On the other hand, to enhance the student model's ability to learn knowledge with semantic gaps from the teacher model, we propose adaptive feature fusion that fuses multiple intermediate layers of the teacher model as additional supervision. Experimental results demonstrate that the proposed method can significantly improve the performance of the student model. Ruijian Xu, Ning Jiang 0002, Jialiang Tang, Xinlei Huang |
ISCAS | 3 |
| 2024 | Virtual Student Distribution Knowledge Distillation for Long-Tailed Recognition
Xinlei Huang, Jialiang Tang, Ning Jiang 0002 |
PRCV (4) | 3 |
| 2024 | Learning Student Network Under Universal Label NoiseabstractData-free knowledge distillation aims to learn a small student network from a large pre-trained teacher network without the aid of original training data. Recent works propose to gather alternative data from the Internet for training student network. In a more realistic scenario, the data on the Internet contains two types of label noise, namely: 1) closed-set label noise, where some examples belong to the known categories but are mislabeled; and 2) open-set label noise, where the true labels of some mislabeled examples are outside the known categories. However, the latter is largely ignored by existing works, leading to limited student network performance. Therefore, this paper proposes a novel data-free knowledge distillation paradigm by utilizing a webly-collected dataset under universal label noise, which means both closed-set and open-set label noise should be tackled. Specifically, we first split the collected noisy dataset into clean set, closed noisy set, and open noisy set based on the prediction uncertainty of various data types. For the closed-set noisy examples, their labels are refined by teacher network. Meanwhile, a noise-robust hybrid contrastive learning is performed on the clean set and refined closed noisy set to encourage student network to learn the categorical and instance knowledge inherited by teacher network. For the open-set noisy examples unexplored by previous work, we regard them as unlabeled and conduct self-supervised learning on them to enrich the supervision signal for student network. Intensive experimental results on image classification tasks demonstrate that our approach can achieve superior performance to state-of-the-art data-free knowledge distillation methods. Jialiang Tang, Ning Jiang 0002, Hongyuan Zhu 0002, Joey Tianyi Zhou, Chen Gong 0002 |
IEEE Trans. Image Process. | 1 |
| 2023 | Distribution Shift Matters for Knowledge Distillation with Webly Collected ImagesabstractKnowledge distillation aims to learn a lightweight student network from a pre-trained teacher network. In practice, existing knowledge distillation methods are usually infeasible when the original training data is unavailable due to some privacy issues and data management considerations. Therefore, data-free knowledge distillation approaches proposed to collect training instances from the Internet. However, most of them have ignored the common distribution shift between the instances from original training data and webly collected data, affecting the reliability of the trained student network. To solve this problem, we propose a novel method dubbed "Knowledge Distillation between Different Distributions" (KD3), which consists of three components. Specifically, we first dynamically select useful training instances from the webly collected data according to the combined predictions of teacher network and student network. Subsequently, we align both the weighted features and classifier parameters of the two networks for knowledge memorization. Meanwhile, we also build a new contrastive learning block called MixDistribution to generate perturbed data with a new distribution for instance alignment, so that the student network can further learn a distribution-invariant representation. Intensive experiments on various benchmark datasets demonstrate that our proposed KD3can outperform the state-of-the-art data-free knowledge distillation approaches. Jialiang Tang, Shuo Chen 0003, Gang Niu 0001, Masashi Sugiyama, Chen Gong 0002 |
ICCV | 1 |
| 2023 | Dynamic Feature Distillation
Xinlei Huang, Ning Jiang 0002, Jialiang Tang |
ICONIP (13) | 3 |
| 2023 | Feature Reconstruction Distillation with Self-attention
Ning Jiang 0002, Jialiang Tang, Xinlei Huang |
ICONIP (12) | 3 |
| 2023 | Dy-KD: Dynamic Knowledge Distillation for Reduced Easy Examples
Ning Jiang 0002, Jialiang Tang, Xinlei Huang |
ICONIP (12) | 3 |
| 2023 | Joint Regularization Knowledge Distillation
Haifeng Qing, Ning Jiang 0002, Jialiang Tang, Xinlei Huang, Wengqing Wu |
ICONIP (12) | 3 |
| 2023 | Correlation Guided Multi-teacher Knowledge Distillation
Luyao Shi, Ning Jiang 0002, Jialiang Tang, Xinlei Huang |
ICONIP (4) | 3 |
| 2023 | Knowledge Distillation via Information Matching
Ning Jiang 0002, Jialiang Tang, Xinlei Huang |
ICONIP (4) | 3 |
| 2023 | Positive-Unlabeled Learning for Knowledge Distillation
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001 |
Neural Process. Lett. | 2 |
| 2022 | Stimulates Potential for Knowledge Distillation
Haifeng Qing, Jialiang Tang, Xinlei Huang, Ning Jiang 0002 |
ICANN (4) | 2 |
| 2022 | Optimizing Knowledge Distillation via Shallow Texture Knowledge Transfer
Xinlei Huang, Jialiang Tang, Haifeng Qing, Ning Jiang 0002 |
ICONIP (4) | 2 |
| 2022 | Cross-Layer Fusion for Feature Distillation
Ning Jiang 0002, Jialiang Tang, Xinlei Huang, Haifeng Qing |
ICONIP (4) | 3 |
| 2022 | Improve 3D Feature Extraction and Fusion for Stage Diagnosis of Alzheimer's DiseaseabstractAlzheimer’s disease (AD) is typical dementia, which is progressive and irreversible. Usually, the clinical diagnosis of patients is at a later stage, so early diagnosis can control the patient’s condition in time. The doctor is usually diagnosing the patient’s condition by 3D brain magnetic resonance imaging (MRI). However, because the 3D MRI structures of adjacent stages are almost similar, the multi-class diagnosis of AD becomes difficult. Therefore, there is a need to enhance the ability to extract more discriminative features from 3DMRI, promoting more accurate diagnosis. In addition, not only the entire MRI changes but also local areas in the MRI. Therefore, it is necessary to pay attention to changes in the entire image and local areas and to fuse features of different scales. In this paper, we propose an innovative convolutional network architecture for feature extraction and feature fusion. It consists of three modules: 1) a network based on ResNet-10, 2) 3D asymmetric convolution block (ACB), 3) multi-scale channel attentional feature fusion (MS-CAFF) module. The proposed model has been tested on the ADNI dataset and achieved an accuracy of 88.33%, which is nearly 2% higher than the latest research. Mingjin Liu, Wenxin Yu 0001, Jialiang Tang, Ning Jiang 0002, Kang Xu 0002 |
ISCAS | 3 |
| 2021 | Gradient Local Binary Pattern For Convolutional Neural NetworksabstractConvolutional neural networks(CNNs) have achieved a performance significantly superior to traditional machine learning methods. However, in the traditional machine learning methods, the feature extraction algorithms are compelling and beneficial for CNNs. This paper introduces the classic feature extraction algorithm gradient local binary pattern(GLBP) to the CNNs. More specially, the GLBP extractor weights will be fixed into the $3\times 3$ sized kernels to construct the GLBP layer to replace the first layer of CNNs. In the GLBP layer, the features extracted by the GLBP kernels will concate or add to the feature process by the convolutional kernels. Through extensive experiments, we demonstrated that the GLBP layer could efficiently improve CNNs performance. When training on the ImageNet dataset, the ResNet18 with GLBP layer obtained 1.19% Top-1 accuracy improvement and 0.87% Top-5 accuracy improvement, respectively. Jialiang Tang, Ning Jiang 0002, Wenxin Yu 0001 |
ICIP | 1 |
| 2021 | Using a Two-Stage GAN to Learn Image Degradation for Image Super-Resolution
Jiarui Cheng, Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001 |
ICONIP (5) | 3 |
| 2021 | Improving Shallow Neural Networks via Local and Global Normalization
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001 |
ICONIP (1) | 2 |
| 2021 | Attention-Based 3D ResNet for Detection of Alzheimer's Disease Process
Mingjin Liu, Jialiang Tang, Wenxin Yu 0001, Ning Jiang 0002 |
ICONIP (1) | 2 |
| 2021 | Data-Free Knowledge Distillation with Positive-Unlabeled Learning
Jialiang Tang, Xin Cheng 0004, Ning Jiang 0002, Wenxin Yu 0001 |
ICONIP (2) | 1 |
| 2021 | Consistent Knowledge Distillation Based on Siamese Networks
Jialiang Tang, Xin Cheng 0004, Ning Jiang 0002, Wenxin Yu 0001 |
ICONIP (5) | 1 |
| 2021 | Triplet Mapping for Continuously Knowledge Distillation
Jialiang Tang, Ning Jiang 0002, Wenxin Yu 0001 |
ICONIP (1) | 2 |
| 2021 | Triplet Knowledge Distillation Networks for Model CompressionabstractKnowledge distillation is a widely used neural network model compression technique. In general, the knowledge distillation transfer the knowledge from a large pre-trained teacher network with superior performance to a small student network enables the student network to achieve better performance. This paper proposes a triplet knowledge distillation framework (abbreviated as TKD), which introduces a smaller assistant network into the knowledge distillation structure. The performance of the assistant network is lower than that of the student network. During the training of the TKD, by minimizing the Mean Squared Error(MSE) loss function, the output of the student network will closer to the output of the teacher network and further from that of the assistant network. Therefore, the student network can learn more expressive knowledge from the teacher network while throwing away mistaken knowledge in the assistant network. Finally, the student network achieves a surprising performance even superior to the teacher network. We have demonstrated the effectiveness of TKD by extensive experiments on benchmark datasets(CIFAR-10, CIFAR-100, SVHN, STL-10). When using VGGNet as an experimental model, the student network VGGNet13 achieving 94.29%, 75.30%, 95.53%, and 87.61% accuracy on the CIFAR-10, CIFAR-100, SVHN, and STL-10 datasets, improved by 1.24%, 2.81%, 0.40%, and 2.32%, respectively. Jialiang Tang, Ning Jiang 0002, Wenxin Yu 0001, Wenqin Wu |
IJCNN | 1 |
| 2021 | The Detection of Attentive Mental State Using a Mixed Neural Network ModelabstractThe application of deep learning (DL) in various brain computer interface (BCI) systems has achieved great success, but the results on the attention classification task are still not satisfactory. In this paper, an end-to-end mixed neural network model was proposed to classify the attention and non- attention mental states from multi-channel electroencephalography (EEG) data. During the experiment, a cross-subject strategy was performed on the attention detection task. Evaluated on a different electrodes combination of a publicly available dataset, the proposed model outperforms these baseline methods while maintaining relatively low computational complexity. The improved performance is meaningful for the attentive mental state classification task and is useful for the process of attention enhancement. Huan Cai, Jialiang Tang, Yihan Wu 0005, Min Xia 0005, Gang He 0001, Yangsong Zhang 0001 |
ISCAS | 2 |
| 2021 | Gradient Local Binary Pattern Layer to Initialize the Convolutional Neural NetworksabstractDeep neural network technology is a milestone achievement in the field of computer vision. It obtained the performance that the shallow network cannot achieve through the multi-layer network structure and the learning method of reverse adjustment parameters. However, the feature extraction algorithm of the shallow network is very effective and also is more beneficial for deep neural networks. In this paper, we combine the shallow network algorithm to proposes the gradient local binary pattern layer(GLBP layer) to replace the first layer of Convolutional Neural Networks(CNNs). The GLBP layer plays a role in initializing the CNNs and can improve network performance without increasing the number and complexity of network layers. In the experiment, using the extracted layer modified by the GLBP feature algorithm to replace other classic deep neural networks, 2.65% and 2.9% performance improvements were obtained in the WideResNet16-2 and ResNet-101 respectively when training on CIFAR-100 dataset. Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001, Jinjia Zhou, Liuwei Mai |
ISCAS | 2 |
| 2021 | Data-Free Network Pruning for Model CompressionabstractConvolutional neural networks(CNNs) are often over-parameterized and cannot apply to existing resource-limited artificial intelligence(AI) devices. Some methods are proposed to model compress the CNNs, but these methods are data-driven and often unable when lacking data. To solve this problem, in this paper, we propose a data-free model compression and acceleration method based on generative adversarial networks and network pruning(named DFNP), which can train a compact neural network only needs a pre-trained neural network. The DFNP consists of the source network, generator, and target network. First, the generator will generate the pseudo data under the supervise of the source network. Then the target network will get by pruning the source network and use these generated data for training. And the source network will transfer knowledge to the target network to promote the target network to achieve a similar performance of the source network. When the VGGNet- 19 is select as the source network, the target network trained by DFNP contains only 25% parameters and 65% calculations of the source network. Still, it retains 99.4% accuracy on the CIFAR-10 dataset without any real data. Jialiang Tang, Mingjin Liu, Ning Jiang 0002, Huan Cai, Wenxin Yu 0001, Jinjia Zhou |
ISCAS | 1 |
| 2021 | Spatial and Channel Dimensions Attention Feature Transfer for Better Convolutional Neural NetworksabstractKnowledge distillation is an extensively researched model compression technology, which uses a large teacher network to transmit information to a small student network. The critical point of the knowledge distillation method to improve the performance of the student network is to find an effective method to extract the information from the feature. The attention mechanism is a widely used feature processing method to process features effectively and obtain more expressive information. In this paper, we propose to use the dual attention mechanism in knowledge distillation to improve the performance of student networks, which extracts information from the spatial and channel dimensions of the feature. The channel dimension attention is search 'what' channel is more meaningful, and the spatial dimension attention is determine 'where' part of the feature is more expressive in a feature map. We have conducted extensive experiments on different datasets, shown that by implementing a dual attention mechanism to extract more expressive information for knowledge transfer, the student network can achieve performance beyond the teacher network. Jialiang Tang, Mingjin Liu, Ning Jiang 0002, Wenxin Yu 0001, Changzheng Yang |
ISCAS | 1 |
| 2021 | Knowledge Distillation Based on Positive-Unlabeled Classification and Attention MechanismabstractWith the rapid development of deep learning, convolutional neural networks(CNNs) have achieved great success. But these high-capability CNNs often with a huge burden of computation and memory, which hinders these CNNs from applying to practical application. To solve this problem, in this paper, we proposed a method to train a compact model with high-capacity. The student network with fewer parameters and calculations will learning from the knowledge of the teacher network with more parameters and calculations. To promote the ability of the student network, the more expressive knowledge is extracted from the middle-layer feature of neural networks by attention mechanism, and the knowledge transforms more effective from the teacher network to the student network by the positive- unlabeled(PU) classifier. We validate our method in extensive experiments, showing that it can train the student network to achieve significant performance superior to the teacher network. Jialiang Tang, Mingjin Liu, Ning Jiang 0002, Wenxin Yu 0001, Changzheng Yang, Jinjia Zhou |
ISCAS | 1 |
| 2021 | Local Feature Normalization
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001, Jinjia Zhou |
KSEM | 2 |
| 2020 | Search-and-Train: Two-Stage Model Compression and Acceleration
Ning Jiang 0002, Jialiang Tang, Wenxin Yu 0001, Jinjia Zhou |
ICONIP (5) | 2 |
| 2020 | Customizable GAN: Customizable Image Synthesis Based on Adversarial Learning
Wenxin Yu 0001, Jinjia Zhou, Xuewen Zhang, Jialiang Tang, Siyuan Li 0004, Ning Jiang 0002, Gang He 0001, Gang He 0002 |
ICONIP (4) | 5 |