Jiayin Zhu

dblp:217/3683 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Generative modeling · 50% Vision and language · 12% Knowledge representation and reasoning · 12%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Computing education · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.012026
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation · AAAI 2026
Knowledge, reasoning and agents › Knowledge representation and reasoning › spatial reasoning
geometry problem solving
1.012026
GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines · ACL (1) 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines · ACL (1) 2026
Machine learning › Generative modeling › diffusion model
score distillation sampling
1.012026
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation · AAAI 2026
Machine learning › Generative modeling › 3d generative model
text-conditioned 3d generation
1.012026
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation · AAAI 2026
Machine learning › Generative modeling › diffusion model › 3d shape generation
text-to-3d generation
1.012026
AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation · AAAI 2026
Computer vision › Video understanding and tracking › video analytics › behavior analysis › human behavior analysis
driver behavior recognition
0.912025
MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception · CVPR 2025
Machine learning › Transfer learning and domain adaptation › negative transfer
negative transfer mitigation
0.312025
MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception · CVPR 2025

Methods — techniques the papers use, named apart from their topics

auxiliary line reasoning · 2.0score distillation sampling · 1.0fine-tuning · 1.0conditional latent space · 1.0multi-axis region attention · 0.9multi-attention mechanism · 0.9dual-branch multimodal embedding · 0.9
YearPublicationVenuePosition
2026 AnchorDS: Anchoring Dynamic Sources for Semantically Consistent Text-to-3D Generation
abstract
Optimization‐based text‑to‑3D methods distill guidance from 2D generative models via Score Distillation Sampling (SDS), but implicitly treat this guidance as static. This work shows that ignoring source dynamics yields inconsistent trajectories that suppress or merge semantic cues, leading to "semantic over-smoothing" artifacts. As such, we reformulate text‑to‑3D optimization as mapping a *dynamically evolving source* distribution to a fixed target distribution. We cast the problem into a dual‑conditioned latent space, conditioned on both the text prompt and the intermediately rendered image. Given this joint setup, we observe that the image condition naturally anchors the current source distribution. Building on this insight, we introduce AnchorDS, an improved score distillation mechanism that provides state‑anchored guidance with image conditions and stabilizes generation. We further penalize erroneous source estimates and design a lightweight filter strategy and fine‑tuning strategy that refines the anchor with negligible overhead. AnchorDS produces finer-grained detail, more natural colours, and stronger semantic consistency, particularly for complex prompts, while maintaining efficiency. Extensive experiments show that our method surpasses previous methods in both quality and efficiency.
Jiayin Zhu, Linlin Yang 0001, Yicong Li 0004, Angela Yao
AAAI1
2026 GeoLaux: A Benchmark for Evaluating MLLMs' Geometry Performance on Long-Step Problems Requiring Auxiliary Lines
abstract
Yumeng Fu, Jiayin Zhu, Lingling Zhang, Wenjun Wu, Bo Zhao, Shaoxuan Ma, Yushun Zhang, Jun Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yumeng Fu, Jiayin Zhu, Lingling Zhang 0001, Shaoxuan Ma, Yushun Zhang, Jun Liu 0002
ACL (1)2
2026 InstructHumans: Editing Animated 3D Human Textures With Instructions
abstract
We present InstructHumans, a novel framework for instruction-driven animatable 3D human texture editing. Existing text-based 3D editing methods often directly apply Score Distillation Sampling (SDS). SDS, designed for generation tasks, cannot account for the defining requirement of editing – maintaining consistency with the source avatar. This work shows that naively using SDS harms editing, as it may destroy consistency. We propose a modified SDS for Editing (SDS-E) that selectively incorporates subterms of SDS across diffusion timesteps. We further enhance SDS-E with spatial smoothness regularization and gradient-based viewpoint sampling for edits with sharp and high-fidelity detailing. Incorporating SDS-E into a 3D human texture editing framework allows us to outperform existing 3D editing methods. Our avatars faithfully reflect the textual edits while remaining consistent with the original avatars. Project page: <uri xmlns:mml="http://www.w3.org/1998/Math/MathML" xmlns:xlink="http://www.w3.org/1999/xlink">https://jyzhu.top/instruct-humans/</uri>.
Jiayin Zhu, Linlin Yang 0001, Angela Yao
IEEE Trans. Multim.1
2025 MMTL-UniAD: A Unified Framework for Multimodal and Multi-Task Learning in Assistive Driving Perception
abstract
Advanced driver assistance systems require a comprehensive understanding of the driver’s mental/physical state and traffic context but existing works often neglect the potential benefits of joint learning between these tasks. This paper proposes MMTL-UniAD, a unified multi-modal multitask learning framework that simultaneously recognizes driver behavior (e.g., looking around, talking), driver emotion (e.g., anxiety, happiness), vehicle behavior (e.g., parking, turning), and traffic context (e.g., traffic jam, traffic smooth). A key challenge is avoiding negative transfer between tasks, which can impair learning performance. To address this, we introduce two key components into the framework: one is the multi-axis region attention network to extract global context-sensitive features, and the other is the dual-branch multimodal embedding to learn multi-modal embeddings from both task-shared and task-specific features. The former uses a multi-attention mechanism to extract task-relevant features, mitigating negative transfer caused by task-unrelated features. The latter employs a dual-branch structure to adaptively adjust task-shared and task-specific parameters, enhancing cross-task knowledge transfer while reducing task conflicts. We assess MMTL-UniAD on the AIDE dataset, using a series of ablation studies, and show that it outperforms state-of-the-art methods across all four tasks. The code is available on https://github.com/Wenzhuo-Liu/MMTL-UniAD.
Wenzhuo Liu, Wenshuo Wang 0001, Yicheng Qiao, Qiannan Guo, Jiayin Zhu, Zilong Chen, Huiming Yang, Zhiwei Li 0011, Tiao Tan, Huaping Liu 0001
CVPR5
2025 UMD-Net: A Unified Multi-Task Assistive Driving Network Based on Multimodal Fusion
abstract
In recent years, researchers have focused on identifying tasks related to driver state, traffic environment, and others to enhance the safety of autonomous driving assistance systems. However, current research on these tasks is conducted independently, neglecting the interconnections between the driver, traffic environment, and vehicle. In this paper, we propose a Unified Multi-task Assistive Driving Network Based on Multimodal Fusion (UMD-Net), the first unified model capable of recognizing four tasks simultaneously by utilizing multimodal data: driver behavior recognition, driver emotion recognition, traffic context recognition, and vehicle behavior recognition. In order to better enhance the synergistic effects between multiple tasks, we designed the position-sensitive multi-directional attention feature extraction subnetwork and recursive dynamic feature fusion module. The former captures the key features of multi-view images by different directions of attention mechanism to improve the generalization of the model across multiple tasks. The latter dynamically adjusts the fusion weight according to the multimodal features to enhance the representation ability of important features in multi-task learning. Our model was evaluated on the public dataset AIDE, achieving the best performance across all four tasks and a high accuracy of 95.31% in the traffic context recognition task, demonstrating the superiority of our approach. The code is available on https://github.com/Wenzhuo-Liu/UMD-Net.
Wenzhuo Liu, Yicheng Qiao, Zhiwei Li 0011, Wenshuo Wang 0001, Wei Zhang 0012, Jiayin Zhu, Yanhuan Jiang, Li Wang 0092, Hong Wang 0014, Huaping Liu 0001, Kunfeng Wang
IEEE Trans. Intell. Transp. Syst.6
2024 FMDNet: Feature-Attention-Embedding-Based Multimodal-Fusion Driving-Behavior-Classification Network
abstract
Driving behavior classification is a critical component of social transportation systems and advanced driver assistance systems, and it has gained increasing attention in recent years. Accurate classification algorithms for driving behavior play a significant role in enhancing traffic safety, energy conservation, and related fields. In this article, we propose a novel driving behavior classification network named feature-attention-embedding-based multimodal-fusion driving-behavior-classification network (FMDNet). FMDNet incorporates eight types of data, including acceleration along the x-axis, y-axis, z-axis, roll angle, pitch angle, yaw angle, roadside image, and vehicle speed, to classify driving behavior. To effectively fuse features extracted from different modalities, taking into account their varying importance, we introduce the feature attention embedding-based fusion module (FAEF) as our fusion strategy. This fusion strategy enhances the network's capability to capture meaningful features by incorporating two feature attention embedding units that delve deeper into the interplay between different modes. Furthermore, we provide further validation of the effectiveness of our approach through extensive ablation experiments to investigate and analyze the impact of various modal data on the classification of driving behavior. Our proposed FMDNet achieves state-of-the-art performance on the public UAH-DriveSet dataset, demonstrating its effectiveness with an impressive F1-score of 99.0%. Additionally, the robustness of our model is confirmed on distracted dataset, achieving a remarkable F1-score of 99.7%. The model's outstanding performance on both the UAH-DriveSet dataset and the distracted-dataset highlights its capabilities and potential for real-world applications.https://github.com/Wenzhuo-Liu/FMDNet
Wenzhuo Liu, Jianli Lu, Junbin Liao, Yicheng Qiao, Guoying Zhang, Jiayin Zhu, Bozhang Xu, Zhiwei Li 0011
IEEE Trans. Comput. Soc. Syst.6