Ziyu Yao 0001

dblp:178/8600-1 · DBLP profile ↗
← Back
8ranked-venue papers
4as first author
8since 2021 · last 2025
0000-0003-1310-0169ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 3 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021
YearPublicationVenuePosition
2025 CountLLM: Towards Generalizable Repetitive Action Counting via Large Language Model
abstract
Repetitive action counting, which aims to count periodic movements in a video, is valuable for video analysis applications such as fitness monitoring. However, existing methods largely rely on regression networks with limited representational capacity, which hampers their ability to accurately capture variable periodic patterns. Additionally, their supervised learning on narrow, limited training sets leads to overfitting and restricts their ability to generalize across diverse scenarios. To address these challenges, we propose CountLLM, the first large language model (LLM)-based framework that takes video data and periodic text prompts as inputs and outputs the desired counting value. CountLLM leverages the rich clues from explicit textual instructions and the powerful representational capabilities of pre-trained LLMs for repetitive action counting. To effectively guide CountLLM, we develop a periodicity-based structured template for instructions that describes the properties of periodicity and implements a standardized answer format to ensure consistency. Additionally, we propose a progressive multimodal training paradigm to enhance the periodicity-awareness of the LLM. Empirical evaluations on widely recognized benchmarks demonstrate CountLLM’s superior performance and generalization, particularly in handling novel and out-of-domain actions that deviate significantly from the training data, offering a promising avenue for repetitive action counting.
Ziyu Yao 0001, Xuxin Cheng, Zhiqi Huang 0001
CVPR1
2024 Soul-Mix: Enhancing Multimodal Machine Translation with Manifold Mixup
abstract
Multimodal machine translation (MMT) aims to improve the performance of machine translation with the help of visual information, which has received widespread attention recently.It has been verified that visual information brings greater performance gains when the textual information is limited.However, most previous works ignore to take advantage of the complete textual inputs and the limited textual inputs at the same time, which limits the overall performance.To solve this issue, we propose a mixup method termed Soul-Mix to enhance MMT by using visual information more effectively.We mix the predicted translations of complete textual input and the limited textual inputs.Experimental results on the Multi30K dataset of three translation directions show that our Soul-Mix significantly outperforms existing approaches and achieves new state-of-the-art performance with fewer parameters than some previous models.Besides, the strength of Soul-Mix is more obvious on more challenging MSCOCO dataset which includes more out-of-domain instances with lots of ambiguous verbs.
Xuxin Cheng, Ziyu Yao 0001, Yifei Xin, Hongxiang Li 0004, Yaowei Li 0001, Yuexian Zou
ACL (1)2
2024 Recovering Global Data Distribution Locally in Federated Learning
Ziyu Yao 0001
BMVC1
2024 PoseRAC: Enhancing Repetitive Action Counting with Salient Poses
Ziyu Yao 0001, Yuexian Zou
ICONIP (5)1
2024 FD2Talk: Towards Generalized Talking Head Generation with Facial Decoupled Diffusion Model
abstract
Talking head generation is a significant research topic that still faces numerous challenges. Previous works often adopt generative adversarial networks or regression models, which are plagued by generation quality and average facial shape problem. Although diffusion models show impressive generative ability, their exploration in talking head generation remains unsatisfactory. This is because they either solely use the diffusion model to obtain an intermediate representation and then employ another pre-trained renderer, or they overlook the feature decoupling of complex facial details, such as expressions, head poses and appearance textures. Therefore, we propose a Facial Decoupled Diffusion model for Talking head generation called FD2Talk, which fully leverages the advantages of diffusion models and decouples the complex facial details through multi-stages. Specifically, we separate facial details into motion and appearance. In the initial phase, we design the Diffusion Transformer to accurately predict motion coefficients from raw audio. These motions are highly decoupled from appearance, making them easier for the network to learn compared to high-dimensional RGB images. Subsequently, in the second phase, we encode the reference image to capture appearance textures. The predicted facial and head motions and encoded appearance then serve as the conditions for the Diffusion UNet, guiding the frame generation. Benefiting from decoupling facial details and fully leveraging diffusion models, extensive experiments substantiate that our approach excels in enhancing image quality and generating more accurate and diverse results compared to previous state-of-the-art methods.
Ziyu Yao 0001, Xuxin Cheng, Zhiqi Huang 0001
ACM Multimedia1
2023 C²A-SLU: Cross and Contrastive Attention for Improving ASR Robustness in Spoken Language Understanding
Xuxin Cheng, Ziyu Yao 0001, Zhihong Zhu 0001, Yaowei Li 0001, Hongxiang Li 0004, Yuexian Zou
INTERSPEECH2
2023 FC-MTLF: A Fine- and Coarse-grained Multi-Task Learning Framework for Cross-Lingual Spoken Language Understanding
Xuxin Cheng, Wanshi Xu, Ziyu Yao 0001, Zhihong Zhu 0001, Yaowei Li 0001, Hongxiang Li 0004, Yuexian Zou
INTERSPEECH3
2023 GhostT5: Generate More Features with Cheap Operations to Improve Textless Spoken Question Answering
Xuxin Cheng, Zhihong Zhu 0001, Ziyu Yao 0001, Hongxiang Li 0004, Yaowei Li 0001, Yuexian Zou
INTERSPEECH3