Zhifeng Wang 0001

dblp:55/497-1 · also Zhi-Feng Wang 0001 · DBLP profile ↗
← Back
22ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0001-6960-509XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 first-author · 3 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 GCKT: Context-Aware Gating of Heterogeneous Learning Features With Transformer for Cognitive Knowledge Tracing in Intelligent Tutoring Systems
abstract
With the rapid growth of online education, Knowledge Tracing (KT) has become central to adaptive learning systems. Yet existing models struggle to integrate the multidimensional and heterogeneous signals generated during learning—such as exercise attributes, response behaviors, temporal factors, and hierarchical knowledge structure. Many methods rely on naive feature concatenation or fixed weighting, limiting their ability to capture synergistic interactions among features. We propose Gated full‐features Transformer Cognitive Knowledge Tracing (GCKT), a Transformer‐based model with a gated fusion mechanism that dynamically integrates multiple inputs. The model first embeds exercise, response correctness, response time, and hierarchical knowledge features (topics and concepts). Topic and concept embeddings are linearly projected into a unified knowledge representation. The exercise, time, correctness, and unified knowledge embeddings are then concatenated and passed through a learnable gating network (linear layer with sigmoid) to produce context‐aware importance weights. These weights are applied element‐wise to adaptively scale each feature before projection into a fused representation for the sequence encoder, enabling the Transformer to more accurately model the evolution of students’ cognitive states. Extensive experiments on public datasets, including MOOCRadar and Math, show that GCKT consistently outperforms strong baselines—such as DKT, AKT, and SAINT+—on key metrics (AUC and F1), delivering robust gains across settings. The results demonstrate that dynamic, fine‐grained feature fusion substantially improves KT performance and that GCKT offers a general, effective approach for modeling complex learning scenarios.
Zhifeng Wang 0001
Int. J. Intell. Syst.1
2026 Leveraging Contrastive Vision-Language Pretraining for Zero-Shot Open Learning Behavior Detection in Smart Education
abstract
With the rapid advancement of industrial informatics, educational behavior detection has become crucial for enabling intelligent, personalized teaching in smart education. Traditional methods, reliant on predefined behavior categories, struggle to meet the demands of open-vocabulary learning behavior detection in real-world scenarios. To address this, we introduce education open behavior detection (Edu-OBD), an open-vocabulary learning behavior detection model designed for smart education using a pretraining approach based on region-text pairs. The model employs a contrastive language-image pretraining encoder and contrastive loss to enable effective open-vocabulary behavior detection. To enhance the integration of image and text features, we propose a novel visual-language integrator comprising two modules: text-guided sparse attention (TGSA) and multiscale text enhancer (MSTE). TGSA strengthens category-relevant semantic information, while MSTE refines multiscale feature embeddings. In addition, the position-aware context fusion module (PCFM) uses relative position encoding to improve feature extraction. For evaluation, we developed the zero-shot behavior evaluation dataset (ZBED), tailored for zero-shot evaluation on novel behaviors. Experimental results show that Edu-OBD achieves 33.6% average precision (AP) at 81.2 frames per second (FPS) on ZBED and 28.1% AP for novel category detection, demonstrating its effectiveness as a flexible and efficient solution for educational behavior detection in complex environments, paving the way for future research in smart education.
Zhifeng Wang 0001, Mingzhang Zuo
IEEE Trans. Comput. Soc. Syst.1
2025 DS-BTIAN: A Novel Deep-Shallow Bidirectional Transformer Interactive Attention Network for Multimodal Emotion Recognition
abstract
In this work, we propose a novel Deep-Shallow Bidirectional Transformer Interactive Attention Network (DS-BTIAN) designed for robust multimodal emotion recognition. DS-BTIAN leverages pre-trained Wav2Vec2.0 and BERT models for efficient feature extraction without the need for finetuning, enhancing representation by integrating outputs from both shallow and deep layers. A novel bidirectional cross-modal interactive attention mechanism based on multi-head attention is introduced, adaptively modeling the relationship between speech and text, thus promoting effective multimodal fusion. Additionally, we incorporate an adaptive speech segmentation and emotion aggregation strategy to better capture emotional dynamics. The network is optimized using a joint training strategy combining center contrastive loss and angular Softmax loss, enhancing discriminative feature learning. We conducted extensive speaker-independent ten-fold cross-validation experiments on the IEMOCAP dataset. The results demonstrate that DS-BTIAN significantly outperforms state-of-the-art methods, achieving a remarkable average weighted accuracy of 76.37% and an unweighted accuracy of 77.21% on the test set, representing a substantial improvement over baseline models. Ablation studies further confirm the effectiveness of each component, underscoring DS-BTIAN’s outstanding capability in capturing and integrating multimodal emotional information. By utilizing pre-trained models for feature extraction without fine-tuning, our approach not only delivers superior performance but also offers a resource-efficient solution, particularly valuable in scenarios with limited computational resources or annotated data.
Zengzhao Chen, Chuanxu Zhao, Zhifeng Wang 0001, Qiuyu Zheng
ICASSP3
2025 DWMGrad: an innovative neural network optimization approach using dynamic window data for adaptive updating of momentum and learning rate
Zhifeng Wang 0001
Appl. Intell.1
2025 MTLSER: Multi-task learning enhanced speech emotion recognition with pre-trained acoustic model
Zengzhao Chen, Zhifeng Wang 0001, Chuanxu Zhao, Mengting Lin, Qiuyu Zheng
Expert Syst. Appl.3
2025 SLBDetection-Net: Towards closed-set and open-set student learning behavior detection in smart classroom of K-12 education
Zhifeng Wang 0001, Shi Dong 0004
Expert Syst. Appl.1
2025 SMDRL: Self-supervised mobile device representation learning framework for recording source identification from unlabeled data
Zhifeng Wang 0001
Expert Syst. Appl.3
2025 SAU-Net: Saliency-Based Adaptive Unfolding Network for Interpretable High-Quality Image Compressed Sensing in Internet of Things
abstract
Deep unfolding networks have emerged as a prominent approach for IoT image Compressed Sensing (CS) due to their interpretability and exceptional performance, leveraging iterative optimization algorithms as design guides. However, the uniform sampling strategy fails to adapt to the diverse information density distributions of image blocks, resulting in inadequate information extraction. Additionally, the reconstruction network neglects to integrate the characteristics of the sampling network, hindering the establishment of an adaptive CS framework based on end-to-end deep neural networks. In light of these limitations, this paper presents a novel image CS network called the Saliency-based Adaptive Unfolding Network (SAU-Net), which addresses non-uniform sampling and adaptive image reconstruction. During the sampling phase, the saliency-based non-uniform sampling module is devised to extract input image information comprehensively by dynamically assigning sampling rates based on block saliency. Subsequently, the reconstruction phase incorporates a saliency enhancement block that augments feature representation and image reconstruction capabilities by amplifying the saliency-based deep reconstruction features. To facilitate efficient end-to-end sampling and reconstruction, we introduce an adaptive multi-channel soft threshold block that tailors the soft thresholding function to multi-channel features, thereby enhancing the function’s generalization capability. Extensive experiments on three benchmark datasets, Set11, CBSD68 and Urban100, validate the superior performance of the proposed SAU-Net compared to state-of-the-art CS methods. Our code is publicly available at https://github.com/CCNUZFW/SAU-Net.
Shiyan Xia, Zhifeng Wang 0001
IEEE Internet Things J.4
2025 Leveraging label semantics and meta-label refinement for multi-label question classification
Shi Dong 0004, Xiaobei Niu, Rui Zhong 0005, Zhifeng Wang 0001, Mingzhang Zuo
Knowl. Based Syst.4
2025 MMKT: Multimodal Knowledge Tracing in Personalized E-Learning Systems for Supporting Lifelong Learning
abstract
To effectively support lifelong learning, personalized e-learning systems aim to provide customized educational services tailored to individual learners’ needs. A crucial aspect for enabling these intelligent services is knowledge tracing, which involves dynamically modeling learners’ knowledge acquisition based on their past learning records. Traditional deep knowledge tracing models primarily focus on learners’ responses and the knowledge components covered in exercises, often neglecting other valuable multimodal learning features. In response to this limitation, our contribution is the introduction of the multimodal knowledge tracing (MMKT) model, which leverages four distinct learning modalities: textual, image, cognitive parametric, and knowledge association modalities. MMKT initially employs a combination of pretrained models and fine-tuning to obtain representations for text and image modalities. Using the EM algorithm, it acquires the cognitive parameter modality for exercises and learners. Subsequently, a multimodal fusion is applied to obtain a comprehensive exercise representation. In the next step, the model explores knowledge impact weights through the knowledge association modality. This information is combined with the cognitive parameter modality of the learner and the hidden states tracked by the gated recurrent unit, resulting in a multimodal representation of the learner’s knowledge state. Experiments conducted on publicly available datasets demonstrate the significant improvement of the proposed MMKT model. The proposed model outperforms the widely used baselines, deep knowledge tracing and dynamic key–value memory networks, achieving relative AUC gains of 3.57% and 6.37%, respectively, in predicting student learning performance. This study provides valuable insights that assist learners in understanding their knowledge mastery and contribute to the realization of lifelong learning from a technical standpoint.
Zhifeng Wang 0001, Zixin Lu, Shi Dong 0004, Mingzhang Zuo
IEEE Trans. Comput. Soc. Syst.1
2025 SCB-DETR: Multiscale Deformable Transformers for Occlusion-Resilient Student Learning Behavior Detection in Smart Classroom
abstract
The integration of artificial intelligence (AI) into the modern educational system is rapidly evolving, particularly in monitoring student behavior in classrooms—a task traditionally dependent on manual observation. This conventional method is notably inefficient, prompting a shift toward more advanced solutions such as computer vision. However, existing target detection models face significant challenges such as occlusion, blurring, and scale disparity, which are exacerbated by the dynamic and complex nature of classroom settings. Furthermore, these models must adeptly handle multiple target detection. To overcome these obstacles, we introduce the student classroom behavior detection with multiscale deformable transformers (SCB-DETR), an innovative approach that utilizes large convolutional kernels for upstream feature extraction, and multiscale feature fusion. This technique significantly improves the detection capabilities for multiscale and occluded targets, offering a robust solution for analyzing student behavior. SCB-DETR establishes an end-to-end system that simplifies the detection process and consistently outperforms other deep learning methods. Employing our custom student classroom behavior (SCBehavior) dataset, SCB-DETR achieves a mean Average Precision (mAP) of 0.626, which is a 1.5% improvement over the baseline model’s mAP and a 6% increase in AP50. These results demonstrate SCB-DETR’s superior performance in handling the uneven distribution of student behaviors and ensuring precise detection in dynamic classroom environments. The source code of this study is publicly available athttps://github.com/CCNUZFW/SCB-DETR.
Zhifeng Wang 0001
IEEE Trans. Comput. Soc. Syst.1
2024 MEConformer: Highly representative embedding extractor for speaker verification via incorporating selective convolution into deep speaker encoder
Qiuyu Zheng, Zengzhao Chen, Zhifeng Wang 0001, Mengting Lin
Expert Syst. Appl.3
2024 Spatio-temporal representation learning enhanced source cell-phone recognition from speech recordings
Shixiong Feng, Zhifeng Wang 0001, Xiangkui Wan, Yunfan Chen, Nan Zhao 0006
J. Inf. Secur. Appl.3
2024 ENFformer: Long-short term representation of electric network frequency for digital audio tampering detection
Zhifeng Wang 0001
Knowl. Based Syst.3
2024 Digital audio tampering detection based on spatio-temporal representation learning of electrical network frequency
Shuai Kong, Zhifeng Wang 0001, Xiangkui Wan, Yunfan Chen
Multim. Tools Appl.3
2024 GSISTA-Net: generalized structure ISTA networks for image compressed sensing based on optimized unrolling algorithm
Zhifeng Wang 0001, Shiyan Xia, Xiangkui Wan
Multim. Tools Appl.3
2024 Deletion and insertion tampering detection for speech authentication based on fluctuating super vector of electrical network frequency
Shuai Kong, Zhifeng Wang 0001, Shixiong Feng, Nan Zhao 0006, Juan Wang 0019
Speech Commun.3
2023 Knowledge Graph-Enhanced Intelligent Tutoring System Based on Exercise Representativeness and Informativeness
abstract
In the realm of online tutoring intelligent systems, e‐learners are exposed to a substantial volume of learning content. The extraction and organization of exercises and skills hold significant importance in establishing clear learning objectives and providing appropriate exercise recommendations. Presently, knowledge graph‐based recommendation algorithms have garnered considerable attention among researchers. However, these algorithms solely consider knowledge graphs with single relationships and do not effectively model exercise‐rich features, such as exercise representativeness and informativeness. Consequently, this paper proposes a framework, namely, the Knowledge Graph Importance‐Exercise Representativeness and Informativeness Framework, to address these two issues. The framework consists of four intricate components and a novel cognitive diagnosis model called the Neural Attentive Cognitive Diagnosis model to recommend the proper exercises. These components encompass the informativeness component, exercise representation component, knowledge importance component, and exercise representativeness component. The informativeness component evaluates the informational value of each exercise and identifies the candidate exercise set (EC) that exhibits the highest exercise informativeness. Moreover, the exercise representation component utilizes a graph neural network to process student records. The output of the graph neural network serves as the input for exercise‐level attention and skill‐level attention, ultimately generating exercise embeddings and skill embeddings. Furthermore, the skill embeddings are employed as input for the knowledge importance component. This component transforms a one‐dimensional knowledge graph into a multidimensional one through four class relations and calculates skill importance weights based on novelty and popularity. Subsequently, the exercise representativeness component incorporates exercise weight knowledge coverage to select exercises from the candidate exercise set for the tested exercise set. Lastly, the cognitive diagnosis model leverages exercise representation and skill importance weights to predict student performance on the test set and estimate their knowledge state. To evaluate the effectiveness of our selection strategy, extensive experiments were conducted on two types of publicly available educational datasets. The experimental results demonstrate that our framework can recommend appropriate exercises to students, leading to improved student performance.
Linqing Li, Zhifeng Wang 0001
Int. J. Intell. Syst.2
2023 A Unified Interpretable Intelligent Learning Diagnosis Framework for Learning Performance Prediction in Intelligent Tutoring Systems
abstract
Intelligent learning diagnosis is a critical engine of intelligent tutoring systems, which aims to estimate learners’ current knowledge mastery status and predict their future learning performance. The significant challenge with traditional learning diagnosis methods is the inability to balance diagnostic accuracy and interpretability. Although the existing psychometric‐based learning diagnosis methods provide some domain interpretation through cognitive parameters, they have insufficient modeling capability with a shallow structure for large‐scale learning data. While the deep learning‐based learning diagnosis methods have improved the accuracy of learning performance prediction, their inherent black‐box properties lead to a lack of interpretability, making their results untrustworthy for educational applications. To settle the abovementioned problem, the proposed unified interpretable intelligent learning diagnosis framework, which benefits from the powerful representation learning ability of deep learning and the interpretability of psychometrics, achieves a better performance of learning prediction and provides interpretability from three aspects: cognitive parameters, learner‐resource response network, and weights of self‐attention mechanism. Within the proposed framework, this paper presents a two‐channel learning diagnosis mechanism LDM‐ID as well as a three‐channel learning diagnosis mechanism LDM‐HMI. Experiments on two real‐world datasets and a simulation dataset show that our method has higher accuracy in predicting learners’ performances compared with the state‐of‐the‐art models and can provide valuable educational interpretability for applications such as precise learning resource recommendation and personalized learning tutoring in intelligent tutoring systems.
Zhifeng Wang 0001, Wenxing Yan, Shi Dong 0004
Int. J. Intell. Syst.1
2023 Spatio-temporal representation learning enhanced speech emotion recognition with multi-head attention mechanisms
Zengzhao Chen, Mengting Lin, Zhifeng Wang 0001, Qiuyu Zheng
Knowl. Based Syst.3
2021 Registration and occlusion handling based on the FAST ICP-ORB method for augmented reality systems
Xuefan Wang, Zhifeng Wang 0001, Huang Yao
Multim. Tools Appl.4
2018 Occlusion handling using moving volume and ray casting techniques for augmented reality systems
Xuefan Wang, Huang Yao, Jia Chen 0026, Zhifeng Wang 0001, Liu Yi
Multim. Tools Appl.5