Yicheng Zhong

dblp:177/8441 · DBLP profile ↗
← Back
11ranked-venue papers
3as first author
9since 2021 · last 2026
0000-0001-7661-5384ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 5 since 2021Systems, architecture and hardware · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 HPSU: A Benchmark for Human-Level Perception in Real-World Spoken Speech Understanding
abstract
Recent advances in Speech Large Language Models (Speech LLMs) have led to great progress in speech understanding tasks such as Automatic Speech Recognition (ASR) and Speech Emotion Recognition (SER). However, whether these models can achieve human-level auditory perception, particularly in terms of their ability to comprehend latent intentions and implicit emotions in real-world spoken language, remains underexplored. To this end, we introduce the Human-level Perception in Spoken Speech Understanding (HPSU), a new benchmark for fully evaluating the human-level perceptual and understanding capabilities of Speech LLMs. HPSU comprises over 20,000 expert-validated spoken language understanding samples in English and Chinese. It establishes a comprehensive evaluation framework by encompassing a spectrum of tasks, ranging from basic speaker attribute recognition to complex inference of latent intentions and implicit emotions. To address the issues of data scarcity and high cost of manual annotation in real-world scenarios, we developed a semi-automatic annotation process. This process fuses audio, textual, and visual information to enable precise speech understanding and labeling, thus enhancing both annotation efficiency and quality. We systematically evaluate various open-source and proprietary Speech LLMs. The results demonstrate that even top-performing models still fall considerably short of human capabilities in understanding genuine spoken interactions. Consequently, HPSU will be useful for guiding the development of Speech LLMs toward human-level perception and cognition.
Peiji Yang, Yicheng Zhong, Jianxing Yu, Zhisheng Wang 0001, Zihao Gou, Wenqing Chen, Jian Yin 0001
AAAI3
2026 TellWhisper: Tell Whisper Who Speaks When
abstract
Multi-speaker automatic speech recognition (MASR) aims to predict "who spoke when and what" from multi-speaker speech, a key technology for multi-party dialogue understanding.However, most existing approaches decouple temporal modeling and speaker modeling when addressing "when" and "who": some inject speaker cues before encoding (e.g., speaker masking), which can cause irreversible information loss; others fuse identity by mixing speaker posteriors after encoding, which may entangle acoustic content with speaker identity.This separation is brittle under rapid turntaking and overlapping speech, often leading to degraded performance.To address these limitations, we propose TellWhisper, a unified framework that jointly models speaker identity and temporal within the speech encoder.Specifically, we design TS-RoPE, a time-speaker rotary positional encoding: time coordinates are derived from frame indices, while speaker coordinates are derived from speaker activity and pause cues.By applying region-specific rotation angles, the model explicitly captures per-speaker continuity, speaker-turn transitions, and state dynamics, enabling the attention mechanism to simultaneously attend to "when" and "who".Moreover, to estimate framelevel speaker activity, we develop Hyper-SD, which casts speaker classification in hyperbolic space to enhance inter-class separation and refine speaker-activity estimates.Extensive experiments demonstrate the effectiveness of the proposed approach.The project webpage is available at https://walker-hyf.github. io/TellWhisper.
Peiji Yang, Yicheng Zhong
ACL (1)4
2026 SimFuzz: Similarity-guided Block-level Mutation for RISC-V Processor Fuzzing
abstract
The Instruction Set Architecture (ISA) defines processor operations and serves as the interface between hardware and software. As an open ISA, RISC-V lowers the barriers to processor design and encourages widespread adoption, but also exposes processors to security risks such as functional bugs. Processor fuzzing is a powerful technique for automatically detecting these bugs. However, existing fuzzing methods suffer from two main limitations. First, their emphasis on redundant test case generation causes them to overlook cross-processor corner cases. Second, they rely too heavily on coverage guidance. Current coverage metrics are biased and inefficient, and become ineffective once coverage growth plateaus.To overcome these limitations, we propose SimFuzz, a fuzzing framework that constructs a high-quality seed corpus from historical bug-triggering inputs and employs similarity-guided, block-level mutation to efficiently explore the processor input space. By introducing instruction similarity, SimFuzz expands the input space around seeds while preserving control-flow structure, enabling deeper exploration without relying on coverage feedback. We evaluate SimFuzz on three widely used open-source RISC-V processors: Rocket, BOOM, and XiangShan, and discover 17 bugs in total, including 14 previously unknown issues, 7 of which have been assigned CVE identifiers. These bugs affect the decode and memory units, cause instruction and data errors, and can lead to kernel instability or system crashes. Experimental results show that SimFuzz achieves up to 73.22% multiplexer coverage on the high-quality seed corpus. Our findings highlight critical security bugs in mainstream RISC-V processors and offer actionable insights for improving functional verification.
Hao Lyu 0002, JingZheng Wu, Xiang Ling 0001, Yicheng Zhong, Tianyue Luo
DATE4
2024 ExpCLIP: Bridging Text and Facial Expressions via Semantic Alignment
abstract
The objective of stylized speech-driven facial animation is to create animations that encapsulate specific emotional expressions. Existing methods often depend on pre-established emotional labels or facial expression templates, which may limit the necessary flexibility for accurately conveying user intent. In this research, we introduce a technique that enables the control of arbitrary styles by leveraging natural language as emotion prompts. This technique presents benefits in terms of both flexibility and user-friendliness. To realize this objective, we initially construct a Text-Expression Alignment Dataset (TEAD), wherein each facial expression is paired with several prompt-like descriptions. We propose an innovative automatic annotation method, supported by CahtGPT, to expedite the dataset construction, thereby eliminating the substantial expense of manual annotation. Following this, we utilize TEAD to train a CLIP-based model, termed ExpCLIP, which encodes text and facial expressions into semantically aligned style embeddings. The embeddings are subsequently integrated into the facial animation generator to yield expressive and controllable facial animations. Given the limited diversity of facial emotions in existing speech-driven facial animation training data, we further introduce an effective Expression Prompt Augmentation (EPA) mechanism to enable the animation generator to support unprecedented richness in style control. Comprehensive experiments illustrate that our method accomplishes expressive facial animation generation and offers enhanced flexibility in effectively conveying the desired style.
Yicheng Zhong, Huawei Wei, Peiji Yang, Zhisheng Wang 0001
AAAI1
2024 Spontaneous Style Text-to-Speech Synthesis with Controllable Spontaneous Behaviors Based on Language Models
Peiji Yang, Yicheng Zhong, Yixuan Zhou 0002, Zhisheng Wang 0001, Zhiyong Wu 0001, Xixin Wu, Helen M. Meng
INTERSPEECH3
2023 Semi-supervised Speech-driven 3D Facial Animation via Cross-modal Encoding
abstract
Existing Speech-driven 3D facial animation methods typically follow the supervised paradigm, involving regression from speech to 3D facial animation. This paradigm faces two major challenges: the high cost of supervision acquisition, and the ambiguity in mapping between speech and lip movements. To address these challenges, this study proposes a novel cross-modal semi-supervised framework, comprising a Speech-to-Image Transcoder and a Face-to-Geometry Regressor. The former jointly learns a common representation space from speech and image domains, enabling the transformation of speech into semantically-consistent facial images. The latter is responsible for reconstructing 3D facial meshes from the transformed images. Both modules require minimal effort to acquire the necessary training data, thereby obviating the dependence on costly supervised data. Furthermore, the joint learning scheme enables the fusion of intricate visual features into speech encoding, thereby facilitating the transformation of subtle speech variations into nuanced lip movements, ultimately enhancing the fidelity of 3D face reconstructions. Consequently, the ambiguity of the direct mapping of speech-to-animation is significantly reduced, leading to coherent and high-fidelity generation of lip motion. Extensive experiments demonstrate that our approach produces competitive results compared to supervised methods.
Peiji Yang, Huawei Wei, Yicheng Zhong, Zhisheng Wang 0001
ICCV3
2023 Bi-Graph Reasoning for Masticatory Muscle Segmentation From Cone-Beam Computed Tomography
abstract
Automated segmentation of masticatory muscles is a challenging task considering ambiguous soft tissue attachments and image artifacts of low-radiation cone-beam computed tomography (CBCT) images. In this paper, we propose a bi-graph reasoning model (BGR) for the simultaneous detection and segmentation of multi-category masticatory muscles from CBCTs. The BGR exploits the local and long-range interdependencies of regions of interest and category-specific prior knowledge of masticatory muscles by reasoning on the category graph and the region graph. The category graph of the learnable muscle prior knowledge handles high-level dependencies of muscle categories, enhancing the feature representation with noise-agnostic category knowledge. The region graph models both local and global dependencies of the candidate muscle regions of interest. The proposed BGR accommodates the high-level dependencies and enhances the region features in the presence of entangled soft tissue and image artifacts. We evaluated the proposed approach by segmenting masticatory muscles on clinically acquired CBCTs. Extensive experimental results show that the BGR effectively segments masticatory muscles with state-of-the-art accuracy.
Yicheng Zhong, Yuru Pei, Kaichen Nie, Yungeng Zhang, Tianmin Xu, Hongbin Zha
IEEE Trans. Medical Imaging1
2022 CreaGAN: An Automatic Creative Generation Framework for Display Advertising
abstract
Creatives are an effective form of delivering product information on the E-commerce platform. Designing an exquisite creative is a time-consuming but crucial task for sellers. In order to accelerate the process, we propose an automatic creative generation framework, named CreaGAN, to make the design procedure easier, faster, and more accurate. Given a well-designed creative for one product, our method can generalize to other product materials by utilizing existing design elements (e.g., background material). The framework consists of two major parts: aesthetics-aware placement and creative inpainting model. The placement model aims to generate plausible locations for new products by considering aesthetic principles. And the inpainting model focus on filling the mismatched regions through contextual information. We conduct experiments on both the public dataset and the real-world creative dataset. Quantitative and qualitative results demonstrate that our method outperforms current state-of-the-art methods and obtains more reasonable and aesthetic visualization results.
Shiyao Wang 0001, Qi Liu 0003, Yicheng Zhong, Zhilong Zhou, Tiezheng Ge, Defu Lian, Yuning Jiang 0001
ACM Multimedia3
2021 Robust 3D face reconstruction from single noisy depth image through semantic consistency
abstract
Abstract This paper addresses the 3D face reconstruction and semantic annotation from a single‐view noisy depth image. A deep neural network‐based coarse‐to‐fine framework is presented to take advantage of 3D morphable model (3DMM) regression and per‐vertex geometry refinement. The low‐dimensional subspace coefficients of the 3DMM initialize the global facial geometry, being prone to be over‐smooth because of the low‐pass characteristics of the shape subspace. The proposed geometry refinement subnetwork predicts per‐vertex displacements to enrich local details, which is learned from unlabelled noisy depth images based on the registration‐like loss. In order to guarantee the semantic correspondence between the resultant 3D face and the depth image, a semantic consistency constraint is introduced to adapt an annotation model learned from the synthetic data to real noisy depth images. The resultant depth annotations are required to be consistent with the label propagation from the coarse and refined parametric 3D faces. The proposed coarse‐to‐fine reconstruction scheme and the semantic consistency constraint are evaluated on the depth‐based 3D face reconstruction and semantic annotation. The series of experiments demonstrate that the proposed approach achieves the performance improvements over compared methods regarding 3D face reconstruction and depth image annotation.
Peixin Li, Yuru Pei, Yicheng Zhong, Yuke Guo, Hongbin Zha
IET Comput. Vis.3
2020 An Unsupervised Approach for 3D Face Reconstruction from a Single Depth Image
Peixin Li, Yuru Pei, Yicheng Zhong, Yuke Guo, Gengyu Ma, Wenhai Wu, Hongbin Zha
CGI3
2020 Face Denoising and 3D Reconstruction from A Single Depth Image
abstract
The reconstruction of 3D face shapes and expressions from a single depth image obtained by a consumer depth camera is a challenging issue considering device-specific noise, the data missing, and the lack of textual constraints. In order to relieve the computationally-intensive nonlinear optimization of traditional template-fitting-based methods, we aim to build an end-to-end regression framework between a depth image and a 3D face encoded by the identity, the expression, and the pose parameters. Concerning the lack of paired depth images and 3D faces, we utilize the unsupervised CycleGAN-based network to adapt the regression model learned from the synthetic data to the real-captured noisy depth images. Instead of separate depth image denoising and 3D face inference, we present a task-specific coupled loss for end-to-end 3D face estimation. We propose a three-tier constraint for the shape consistency in the joint embedding, the depth image, and the surface space to avoid shape distortions in the unsupervised domain adaptation network. We report promising qualitative results for the task of the face denoising and 3D face reconstruction from a single depth image.
Yicheng Zhong, Yuru Pei, Peixin Li, Yuke Guo, Gengyu Ma, Wenhai Wu, Hongbin Zha
FG1