Jiachen Zheng

dblp:305/0307 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2025
0009-0003-3374-2991ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Speech recognition and synthesis · 59% Generative modeling · 27% Representation and self-supervised learning · 14%

Topics — the 7 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
masked generative modeling
1.722025
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training · NeurIPS 2025
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer · ICLR 2025
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked token prediction
0.912025
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer · ICLR 2025
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
non-autoregressive text-to-speech
0.912025
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer · ICLR 2025
Natural language and speech › Speech recognition and synthesis
speech synthesis
0.912025
Metis: A Foundation Speech Generation Model with Masked Generative Pre-training · NeurIPS 2025
Natural language and speech › Speech recognition and synthesis
text-to-speech synthesis
0.912025
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer · ICLR 2025
Natural language and speech › Speech recognition and synthesis › text-to-speech synthesis
zero-shot text-to-speech
0.912025
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer · ICLR 2025
Natural language and speech › Speech recognition and synthesis › speech representation learning
self-supervised speech representation
0.312025
MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer · ICLR 2025

Methods — techniques the papers use, named apart from their topics

self-supervised learning · 1.7masked generative pretraining · 0.9masked generative codec transformer · 0.9discrete token prediction · 0.9discrete speech representation · 0.9
YearPublicationVenuePosition
2025 MaskGCT: Zero-Shot Text-to-Speech with Masked Generative Codec Transformer
abstract
The recent large-scale text-to-speech (TTS) systems are usually grouped as autoregressive and non-autoregressive systems. The autoregressive systems implicitly model duration but exhibit certain deficiencies in robustness and lack of duration controllability. Non-autoregressive systems require explicit alignment information between text and speech during training and predict durations for linguistic units (e.g. phone), which may compromise their naturalness. In this paper, we introduce $\textbf{Mask}$ed $\textbf{G}$enerative $\textbf{C}$odec $\textbf{T}$ransformer (MaskGCT), a fully non-autoregressive TTS model that eliminates the need for explicit alignment information between text and speech supervision, as well as phone-level duration prediction. MaskGCT is a two-stage model: in the first stage, the model uses text to predict semantic tokens extracted from a speech self-supervised learning (SSL) model, and in the second stage, the model predicts acoustic tokens conditioned on these semantic tokens. MaskGCT follows the mask-and-predict learning paradigm. During training, MaskGCT learns to predict masked semantic or acoustic tokens based on given conditions and prompts. During inference, the model generates tokens of a specified length in a parallel manner. Experiments with 100K hours of in-the-wild speech demonstrate that MaskGCT outperforms the current state-of-the-art zero-shot TTS systems in terms of quality, similarity, and intelligibility. Audio samples are available at https://maskgct.github.io/. We release our code and model checkpoints at https://github.com/open-mmlab/Amphion/blob/main/models/tts/maskgct.
Yuancheng Wang, Haoyue Zhan, Liwei Liu 0008, Ruihong Zeng, Jiachen Zheng, Xueyao Zhang, Shunsi Zhang, Zhizheng Wu 0001
ICLR6
2025 Metis: A Foundation Speech Generation Model with Masked Generative Pre-training
abstract
We introduce ***Metis***, a foundation model for unified speech generation. Unlike previous task-specific or multi-task models, Metis follows a pre-training and fine-tuning paradigm. It is pre-trained on large-scale unlabeled speech data using masked generative modeling and then fine-tuned to adapt to diverse speech generation tasks. Specifically, (1) Metis utilizes two discrete speech representations: SSL tokens derived from speech self-supervised learning (SSL) features, and acoustic tokens directly quantized from waveforms. (2) Metis performs masked generative pre-training on SSL tokens, utilizing 300K hours of diverse speech data, without any additional condition. (3) Through fine-tuning with task-specific conditions, Metis achieves efficient adaptation to various speech generation tasks while supporting multimodal input, even when using limited data and trainable parameters. Experiments demonstrate that Metis can serve as a foundation model for unified speech generation: Metis outperforms state-of-the-art task-specific or multi-task systems across five speech generation tasks, including zero-shot text-to-speech, voice conversion, target speaker extraction, speech enhancement, and lip-to-speech, even with fewer than 20M trainable parameters or 300 times less training data. Audio samples are are available at https://metis-demo.github.io/. We release the code and model checkpoints at https://github.com/open-mmlab/Amphion.
Yuancheng Wang, Jiachen Zheng, Junan Zhang, Xueyao Zhang, Huan Liao, Zhizheng Wu 0001
NeurIPS2
2023 Federated Clique Percolation for Privacy-preserving Overlapping Community Detection
abstract
Community structure is a typical characteristic of complex networks. Finding communities in complex networks has many important applications, such as the advertisement and recommendation based on social networks and the discovery of new protein molecules in biological networks, which make it a hot topic in the field of complex network analysis. With the increasing concerns about the leakage of personal privacy, discovering communities spread across the local networks owned by multiple participants accurately while preserving each participant’s privacy has become an emerging challenge in distributed community detection. In this article, we propose a general federated graph learning model for privacy-preserving distributed graph learning and develop two federated clique percolation algorithms (CPAs) based on it to discover overlapping communities distributed across multiple participants’ local networks without disclosing any participant’s network privacy. Homomorphic encryption and hash operation are used in combination to protect the privacy of the vertices and edges of each local network. Furthermore, vertex attributes are involved in the calculation of clique similarity and clique percolation when dealing with attributed networks. The experimental results on real-world and artificial datasets demonstrate that the proposed algorithms achieve identical results to those of their stand-alone counterparts and more than 200% higher accuracy than the simple distributed CPAs without federating learning.
Kun Guo 0003, Wenzhong Guo, Enjie Ye, Yutong Fang, Jiachen Zheng, Ximeng Liu, Kai Chen 0005
ACM Trans. Intell. Syst. Technol.5
2022 An ensemble framework for interpretable malicious code detection
abstract
Malicious code is an ever-growing security threats to computer systems and networks, while malware detection provides effective defense against malicious codes. In this paper, a brief overview is presented on currently prevalent methods to detect malicious codes, including signature-based methods, behavioral-based detection and machine learning (ML) based ones. More specifically, the potentially effective malicious features are summarized and the novel methods using ML are deeply discussed. Furthermore, an ensemble interpretable framework is explored for automatic and efficient malicious code detection. Based on the knowledge graph of malware, the novel framework inclines to achieve robust malware detection even confronted with unseen malicious codes. Finally, both advantages and disadvantages are discussed and experimental results are outlined to verify the effectiveness of the novel methods.
Jieren Cheng, Jiachen Zheng, Xiaomei Yu
Int. J. Intell. Syst.2