Zhizhong Ma

dblp:222/3461 · DBLP profile ↗
← Back
8ranked-venue papers
1as first author
8since 2021 · last 2026
0000-0002-0282-4375ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
YearPublicationVenuePosition
2026 SLATSCOG: A secure authentication framework via federated data generation and temporally-enhanced split learning
Tiantian Zhu 0001, Zhengqiu Weng, Zhizhong Ma, Suyu Zhang
Knowl. Based Syst.4
2026 ProGrasp: Storage-efficient provenance graph compression for APT forensics via structure prediction and attribute aggregation
Tiantian Zhu 0001, Yiqian Yang, Zhengqiu Weng, Haofei Sun, Zhizhong Ma, Guolang Chen
Knowl. Based Syst.5
2025 Laplacian eigenmaps based manifold regularized CNN for visual recognition
Ming Zong, Zhizhong Ma, Fangyi Zhu, Yujun Ma, Ruili Wang 0001
Inf. Sci.2
2024 Advancing Music Emotion Recognition: A Transformer Encoder-Based Approach
abstract
Music Emotion Recognition (MER) involves identifying the emotional content conveyed by music. This field is becoming increasingly significant due to its broad range of applications, including music recommendation systems, mood-based playlists, and therapeutic tools. This paper presents a novel MER model designed for song-level analysis, leveraging the Transformer Encoder architecture. The model incorporates various embedding techniques to capture both local and global contexts within musical data, thereby improving the extraction of crucial features for emotion recognition. Additionally, a Self-Attention Pooling Layer is used to effectively integrate and interpret complex musical features. Experiments using the DEAM dataset reveal that this model excels in emotion identification, surpassing existing approaches and offering promising directions for future research in the field of MER.
Yangyuan Chen, Zhizhong Ma, Mingjing Wang, Mingzhe Liu 0001
MMAsia2
2024 PIAENet: Pyramid integration and attention enhanced network for object detection
Xiangyan Tang, Wenhang Xu, Keqiu Li, Mengxue Han, Zhizhong Ma, Ruili Wang 0001
Inf. Sci.5
2022 Determining the best Acoustic Features for Smoker Identification
abstract
Speech-based automatic smoker identification (also known as smoker/non-smoker classification) aims to identify speakers’ smoking status from their speech. In the COVID-19 pandemic, speech-based automatic smoker identification approaches have received more attention in smoking cessation research due to low cost and contactless sample collection. This study focuses on determining the best acoustic features for smoker identification. In this paper, we investigate the performance of four acoustic feature sets/representations extracted using three feature extraction/learning approaches: (i) hand-crafted feature sets including the extended Geneva Minimalistic Acoustic Parameter Set and the Computational Paralinguistics Challenge Set, (ii) the Bag-of-Audio-Words representations, (iii) the neural representations extracted from raw waveform signals by SincNet. Experimental results show that: (i) SincNet feature representations are the most effective for smoker identification and outperform the MFCC baseline features by 16% in absolute accuracy; (ii) the performance of hand-crafted feature sets and the Bag-of-Audio-Words representations rely on the scale of the dimensions of feature vectors.
Zhizhong Ma, Yuanhang Qiu, Feng Hou, Ruili Wang 0001, Joanna Ting Wai Chu, Chris Bullen
ICASSP1
2022 CyclicAugment: Speech Data Random Augmentation with Cosine Annealing Scheduler for Auotmatic Speech Recognition
Zhihan Wang, Feng Hou, Yuanhang Qiu, Zhizhong Ma, Satwinder Singh, Ruili Wang 0001
INTERSPEECH4
2021 Self-Supervised Learning Based Phone-Fortified Speech Enhancement
Yuanhang Qiu, Ruili Wang 0001, Satwinder Singh, Zhizhong Ma, Feng Hou
Interspeech4