Huaicheng Zhang

dblp:295/2194 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
9since 2021 · last 2025
0000-0002-6874-5530ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Computer networks · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 LeVo: High-Quality Song Generation with Multi-Preference Alignment
abstract
Recent advances in large language models (LLMs) and audio language models have significantly improved music generation, particularly in lyrics-to-song generation. However, existing approaches still struggle with the complex composition of songs and the scarcity of high-quality data, leading to limitations in audio quality, musicality, instruction following, and vocal-instrument harmony. To address these challenges, we introduce LeVo, a language model based framework consisting of LeLM and Music Codec. LeLM is capable of parallel modeling of two types of tokens: mixed tokens, which represent the combined audio of vocals and accompaniment to achieve better vocal-instrument harmony, and dual-track tokens, which separately encode vocals and accompaniment for high-quality song generation. It employs two decoder-only transformers and a modular extension training strategy to prevent interference between different token types. To further enhance musicality and instruction following ability, we introduce a multi-preference alignment method based on Direct Preference Optimization (DPO). This method handles diverse human preferences through a semi-automatic data construction process and post-training. Experimental results demonstrate that LeVo significantly outperforms existing open-source methods in both objective and subjective metrics, while performing competitively with industry systems. Ablation studies further justify the effectiveness of our designs. Audio examples and source code are available at https://levo-demo.github.io and https://github.com/tencent-ailab/songgeneration.
Shun Lei, Yaoxun Xu, Huaicheng Zhang, Wei Tan 0011, Hangting Chen, Yixuan Zhang 0005, Haina Zhu, Shuai Wang 0016, Zhiyong Wu 0001, Dong Yu 0001
NeurIPS4
2025 JAMC: A jigsaw-based autoencoder with masked contrastive learning for cardiovascular disease diagnosis
Yue Ge, Huaicheng Zhang, Jiguang Shi, Deyu Luo, Sheng Chang 0003, Jin He 0002, Qijun Huang, Hao Wang 0046
Knowl. Based Syst.2
2024 CELL: Supplementing Context for Multi-modal Few-shot Cardiologist
abstract
Language-based multi-modal learning has showcased remarkable efficacy across various domains. However, certain areas such as electrocardiogram (ECG) analysis face challenges due to incomplete and scarce textual data. Analyzing incomplete corpora poses difficulties for pre-trained language models, while data scarcity makes it difficult tre alignment and classification respectively, classification loss combined with consistent contexts are added in pre-training to alleviate gaps between these targets. Meanwhile, consistent contexts narrow gaps between prompts, allowing models to focus on genuinely informative features. Consequently, ECG feature spaces align more closely with semantic spaces, ensuring robust classification performance and enhancing the quality of multi-modal representations. Through extensive experiments involving zero-shot and few-shot learning, CELL demonstrates superior performances in near-distribution, out-of-distribution, and clinical domains. Its robust and generalized performance positions CELL as a promiso harmonize multimodal pretraining and downstream goals. To address these problems, this study introduces a novel ECG-Language multi-modal learning method named Context ECG-Language Learning (CELL). To supplement context and complete corpora, dynamic contexts composed of learnable vectors are incorporated into language model embedding, with other parameters fixed. Since pre-training and downstream targets are featuing approach for multi-modal learning in fields with scarce and incomplete corpora.
Huaicheng Zhang, Jiguang Shi, Yue Ge, Sheng Chang 0003, Hao Wang 0046, Qijun Huang
BIBM1
2024 DDDG: A dual bi-directional knowledge distillation method with generative self-supervised pre-training and its hardware implementation on SoC for ECG
Huaicheng Zhang, Wenhan Liu, Qianxi Guo, Jiguang Shi, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang
Expert Syst. Appl.1
2024 SIGxCL: A Signal-Image-Graph Cross-Modal Contrastive Learning Framework for CVD Diagnosis Based on Internet of Medical Things
abstract
Recently, contrastive learning (CL) has garnered wide interest because it enables unsupervised pretraining to alleviate conventional deep learning methods’ strong reliance on artificial labels. While CL-based methods have been applied to cardiovascular disease (CVD) diagnosis with noninvasive electrocardiogram (ECG), most of these methods are limited within the 1-D signal modality and primarily focus on temporal features like amplitude and time sequence. The morphological features derived from clinically significant image-like ECGs are ignored. Furthermore, the relationships among different leads are neglected as well, describing the activities and interaction of various heart regions that are essential in CVD diagnosis and lesion localization. To address these limitations, this work proposes a novel cross-modal CL framework named signal–image–graph cross-modal contrastive learning (SIGxCL), which represents and jointly analyzes ECGs in signal, image, and graph modalities. Crucial for CL, modality-specific transformations are introduced for ECGs in the three modalities. SIGxCL enables signal–image–graph correspondence by maximizing the agreement of the accordant cross-modal ECGs in the invariant space. Consequently, SIGxCL could capture and leverage temporal, morphological, and spatially physiological features simultaneously. Compared to random initial and conventional supervised methods, SIGxCL achieves remarkable enhancements. Considering the best performances of the existing CL-based methods, SIGxCL outperforms them by up to 4.72%, 9.41%, and 4.31% across three data sets. SIGxCL is designed to be compatible with Internet of Medical Things (IoMT) and can be deployed on resource-limited portable devices. The deployment includes pretraining, online/offline tuning, and real-time inference modules. In conclusion, SIGxCL demonstrates superior performance and provides a promising approach for real-time IoMT-based diagnosis.
Huaicheng Zhang, Wenhan Liu, Zhoutong Li, Jiguang Shi, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang
IEEE Internet Things J.1
2024 An adaptive threshold-based semi-supervised learning method for cardiovascular disease detection
Jiguang Shi, Zhoutong Li, Wenhan Liu, Huaicheng Zhang, Deyu Luo, Yue Ge, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang
Inf. Sci.4
2024 Honest-GE: 2-step heuristic optimization and node-level embedding empower spatial-temporal graph model for ECG
Huaicheng Zhang, Wenhan Liu, Deyu Luo, Jiguang Shi, Qianxi Guo, Yue Ge, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang
Inf. Sci.1
2024 ST-ReGE: A Novel Spatial-Temporal Residual Graph Convolutional Network for CVD
abstract
Recently, deep learning (DL) has enabled rapid advancements in electrocardiogram (ECG)-based automatic cardiovascular disease (CVD) diagnosis. Multi-lead ECG signals have lead systems based on the potential differences between electrodes placed on the limbs and the chest. When applying DL models, ECG signals are usually treated as synchronized signals arranged in Euclidean space, which is the abstraction and generalization of real space. However, conventional DL models typically merely focus on temporal features when analyzing Euclidean data. These approaches ignore the spatial relationships of different leads, which are physiologically significant and useful for CVD diagnosis because different leads represent activities of specific heart regions. These relationships derived from spatial distributions of electrodes can be conveniently created in non-Euclidean data, making multi-lead ECGs better conform to their nature. Considering graph convolutional network (GCN) adept at analyzing non-Euclidean data, a novel spatial-temporal residual GCN for CVD diagnosis is proposed in this work. ECG signals are firstly divided into single-channel patches and transferred into nodes, which will be connected by spatial-temporal connections. The proposed model employs residual GCN blocks and feed-forward networks to alleviate over-smoothing and over-fitting. Moreover, residual connections and patch dividing enable the capture of global and detailed spatial-temporal features. Experimental results reveal that the proposed model achieves at least a 5.85% and 6.80% increase inF1over other state-of-the-art algorithms with similar parameters and computations in both PTB-XL and Chapman databases. It indicates that the proposed model provides a promising avenue for intelligent diagnosis with limited computing resources.
Huaicheng Zhang, Wenhan Liu, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang
IEEE J. Biomed. Health Informatics1
2023 Dense lead contrast for self-supervised representation learning of multilead electrocardiograms
Wenhan Liu, Zhoutong Li, Huaicheng Zhang, Sheng Chang 0003, Hao Wang 0046, Jin He 0002, Qijun Huang
Inf. Sci.3