Yiyi Zhang 0002

dblp:140/8590-2 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
13since 2021 · last 2026
0000-0001-9180-8512ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 2 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Security and privacy · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 High-Quality Full-Head 3D Avatar Generation from Any Single Portrait Image
abstract
In this work, we introduce a novel high-fidelity full-head 3D avatar generation method from a single image, regardless of perspective, style, expression, or accessories. Prior works often fail to preserve consistent head geometry and facial details, primarily due to their limited capacity in modeling fine-grained facial textures and maintaining identity information. To address these challenges, we construct a new high-quality dataset containing 227 sequences of digital human portraits captured from 96 different perspectives, totalling 21,792 frames, featuring high-quality facial texture details. To further improve performance, we propose a novel multi-view diffusion named ID-TS diffusion model, which integrate identity and expression information into the two-stage multi-view diffusion process. The low-resolution stage ensures structural consistency of heads across multiple views, while the high-resolution stage preserves facial detail fidelity and coherence. Finally, we propose an enhanced feed-forward Gaussian avatar reconstruction method that optimizes the network on multi-view images of each single subject, significantly improving 3D facial texture details. Extensive experiments show that our method demonstrates robust performance across challenging scenarios, while showcasing broad applicability across numerous downstream tasks.
Yujie Gao 0001, Chencheng Wang, Xianbing Sun, Jiahui Zhan, Wentao Wang 0009, Yiyi Zhang 0002, Haohua Zhao 0001, Liqing Zhang 0001, Jianfu Zhang 0003
AAAI6
2026 The Simulator's Blueprint: Automata Learning from Cybersecurity Logs
abstract
Abstract We show how to use passive automata learning to infer models of attacker-defender interactions in cybersecurity. By treating system event logs as words in a formal language, we can apply algorithms such as RPNI to infer compact deterministic finite automata from observed traces. We evaluate this approach through a case study on the Cyber Operations Research Gymnasium (CybORG), a widely-used simulation framework for training defensive agents using machine learning. We analyze the structural properties of the inferred automata and assess their empirical fidelity with respect to the semantics of CybORG. Our results show that accurate formal models can be learned from a relatively small number of traces, suggesting a promising path toward more automated and data-driven approaches to cybersecurity.
Tudor Braicu, Benjamin Ylvisaker, Nicolas A. Espinosa Dice, Yiding Chen, Yiyi Zhang 0002, Nate Foster, Hossein Hojjat
CAV (1)5
2025 Exploring the Efficacy of Multi-Agent Reinforcement Learning for Autonomous Cyber Defence: A CAGE Challenge 4 Perspective
abstract
As cyber threats become increasingly automated and sophisticated, novel solutions must be introduced to improve defence of enterprise networks. Deep Reinforcement Learning (DRL) has demonstrated potential in mitigating these advanced threats. Single DRL Agents have proven utility toward execution of autonomous cyber defence. Despite the success of employing single DRL Agents, this approach presents significant limitations, especially regarding scalability within large enterprise networks. An attractive alternative to the single agent approach is the use of Multi-Agent Reinforcement Learning (MARL). However, developing MARL agents is costly with few options for examining MARL cyber defence techniques against adversarial agents. This paper presents a MARL network security environment, the fourth iteration of the Cyber Autonomy Gym for Experimentation (CAGE) challenges. This challenge was specifically designed to test the efficacy of MARL algorithms in an enterprise network. Our work aims to evaluate the potential of MARL as a robust and scalable solution for autonomous network defence.
Mitchell Kiely, Metin Ahiskali, Etienne Borde, Benjamin Bowman, David Bowman, Dirk Van Bruggen, KC Cowan, Prithviraj Dasgupta, Erich Devendorf, Ben Edwards, Alex Fitts, Sunny Fugate, Ryan Gabrys, Wayne Gould, H. Howie Huang, Jules Jacobs, Ryan Kerr, Isaiah J. King, Li Li 0009, Luis Martinez, Christopher Moir, Craig Murphy, Olivia Naish, Claire Owens, Miranda Purchase, Ahmad Ridley, Adrian Taylor, Sara Farmer, William John Valentine, Yiyi Zhang 0002
AAAI30
2025 Warp Gait Across Ages: Cross-age Gait Video Translation with Part-aware Flow Warping
abstract
Cross-age gait video translation aims to translate a gait video of a subject captured at a certain age into another age while preserving the individual identity and realism. This has a wide range of applications, including age-invariant gait recognition. In this paper, we propose a method for cross-age gait video translation using a spatially smooth geometric warping field. More specifically, we employ a variant of the spatial transformer network (STN), called flow warping, which achieves accurate coarse-to-fine warping field inference. In addition, we extend the flow warping framework by introducing gait stance-dependent warping fields to better represent the temporal variations in the warping fields. We also develop an approach called Part-STN to infer part-dependent warping fields and then merge them into a unified warping field to retain consistency within each body part. We then train Part-STN with loss functions considering three aspects: aging effect (using an age group classification loss), identity preservation (using a cycle-consistency loss and triplet loss on the representation learned from a pretrained gait recognition model), and realism (using an adversarial loss and smoothness loss on estimated flow warping). Our framework outperforms state-of-the-art cross-age gait video translation methods quantitatively and qualitatively on the largest gait database OULP-Age, with respect to both age group classification, identity recognition and realism.
Yiyi Zhang 0002, Hanchong Yan, Liqing Zhang 0001, Yasushi Yagi
IJCB1
2025 Perceiving and Acting in First-Person: A Dataset and Benchmark for Egocentric Human-Object-Human Interactions
Liang Xu 0012, Chengqun Yang, Zili Lin, Fei Xu 0008, Congsheng Xu, Yiyi Zhang 0002, Jie Qin 0004, Xingdong Sheng, Yunhui Liu 0006, Xin Jin 0014, Yichao Yan, Wenjun Zeng 0001, Xiaokang Yang 0001
ICCV7
2025 Pedestrian Motion Reconstruction: A Large-scale Benchmark via Mixed Reality Rendering with Multiple Perspectives and Modalities
abstract
Reconstructing pedestrian motion from dynamic sensors, with a focus on pedestrian intention, is crucial for advancing autonomous driving safety. However, this task is challenging due to data limitations arising from technical complexities, safety, and cost concerns. We introduce the Pedestrian Motion Reconstruction (PMR) dataset, which focuses on pedestrian intention to reconstruct behavior using multiple perspectives and modalities. PMR is developed from a mixed reality platform that combines real-world realism with the extensive, accurate labels of simulations, thereby reducing costs and risks. It captures the intricate dynamics of pedestrian interactions with objects and vehicles, using different modalities for a comprehensive understanding of human-vehicle interaction. Analyses show that PMR can naturally exhibit pedestrian intent and simulate extreme cases. PMR features a vast collection of data from 54 subjects interacting across 12 urban settings with 7 objects, encompassing 12,138 sequences with diverse weather conditions and vehicle speeds. This data provides a rich foundation for modeling pedestrian intent through multi-view and multi-modal insights. We also conduct comprehensive benchmark assessments across different modalities to thoroughly evaluate pedestrian motion reconstruction methods.
Yiyi Zhang 0002, Xinhao Hu, Li Niu 0002, Jianfu Zhang 0003, Yasushi Makihara, Yasushi Yagi, Wenlong Liao, Junchi Yan, Liqing Zhang 0001
ICLR2
2025 Convergence of Consistency Model with Multistep Sampling under General Data Assumptions
abstract
Diffusion models accomplish remarkable success in data generation tasks across various domains. However, the iterative sampling process is computationally expensive. Consistency models are proposed to learn consistency functions to map from noise to data directly, which allows one-step fast data generation and multistep sampling to improve sample quality. In this paper, we study the convergence of consistency models when the self-consistency property holds approximately under the training distribution. Our analysis requires only mild data assumption and applies to a family of forward processes. When the target data distribution has bounded support or has tails that decay sufficiently fast, we show that the samples generated by the consistency model are close to the target distribution in Wasserstein distance; when the target distribution satisfies some smoothness assumption, we show that with an additional perturbation step for smoothing, the generated samples are close to the target distribution in total variation distance. We provide two case studies with commonly chosen forward processes to demonstrate the benefit of multistep sampling.
Yiding Chen, Yiyi Zhang 0002, Owen Oertell, Wen Sun 0002
ICML2
2025 Scaling Offline RL via Efficient and Expressive Shortcut Models
abstract
Diffusion and flow models have emerged as powerful generative approaches capable of modeling diverse and multimodal behavior. However, applying these models to offline RL remains challenging due to the iterative nature of their noise sampling processes, making policy optimization difficult. In this paper, we introduce Scalable Offline Reinforcement Learning (SORL), a new offline RL algorithm that leverages shortcut models – a novel class of generative models – to scale both training and inference. SORL's policy can capture complex data distributions and can be trained simply and efficiently in a one-stage training procedure. At test time, SORL supports both sequential and parallel inference scaling by using the learned Q-function as a verifier. We demonstrate that SORL achieves strong performance across a range of offline RL tasks and exhibits positive scaling behavior with increased test-time compute.
Nicolas A. Espinosa Dice, Yiyi Zhang 0002, Yiding Chen, Bradley Guo, Owen Oertell, Gokul Swamy 0001, Kianté Brantley, Wen Sun 0002
NeurIPS2
2025 ID-MotionNet: Identity-Preserved 3D Skeleton Sequence Generation via Information Bottleneck Disentanglement
Jingyu Xie, Jianfu Zhang 0003, Liqing Zhang 0001, Yiyi Zhang 0002
PRCV (15)6
2024 Assessing Image Inpainting via Re-Inpainting Self-Consistency Evaluation
abstract
Image inpainting, the task of reconstructing missing segments in corrupted images using available data, faces challenges in ensuring consistency and fidelity, especially under information-scarce conditions. Traditional evaluation methods, heavily dependent on the existence of unmasked reference images, inherently favor certain inpainting outcomes, introducing biases. Addressing this issue, we introduce an innovative evaluation paradigm that utilizes a self-supervised metric based on multiple re-inpainting passes. This approach, diverging from conventional reliance on direct comparisons in pixel or feature space with original images, emphasizes the principle of self-consistency to enable the exploration of various viable inpainting solutions, effectively reducing biases. Our extensive experiments across numerous benchmarks validate the alignment of our evaluation method with human judgment.
Jianfu Zhang 0003, Yan Hong 0001, Yiyi Zhang 0002, Liqing Zhang 0001
CIKM4
2023 Critical Firms Prediction for Stemming Contagion Risk in Networked-Loans through Graph-Based Deep Reinforcement Learning
abstract
The networked-loan is major financing support for Micro, Small and Medium-sized Enterprises (MSMEs) in some developing countries. But external shocks may weaken the financial networks' robustness; an accidental default may spread across the network and collapse the whole network. Thus, predicting the critical firms in networked-loans to stem contagion risk and prevent potential systemic financial crises is of crucial significance to the long-term health of inclusive finance and sustainable economic development. Existing approaches in the banking industry dismiss the contagion risk across loan networks and need extensive knowledge with sophisticated financial expertise. Regarding the issues, we propose a novel approach to predict critical firms for stemming contagion risk in the bank industry with deep reinforcement learning integrated with high-order graph message-passing networks. We demonstrate that our approach outperforms the state-of-the-art baselines significantly on the dataset from a large commercial bank. Moreover, we also conducted empirical studies on the real-world loan dataset for risk mitigation. The proposed approach enables financial regulators and risk managers to better track and understands contagion and systemic risk in networked-loans. The superior performance also represents a paradigm shift in addressing the modern challenges in financing support of MSMEs and sustainable economic development.
Dawei Cheng, Zhibin Niu, Jianfu Zhang 0003, Yiyi Zhang 0002, Changjun Jiang 0002
AAAI4
2023 Natural Image Matting with Attended Global Context
Yiyi Zhang 0002, Li Niu 0002, Yasushi Makihara, Jianfu Zhang 0003, Weijie Zhao 0003, Yasushi Yagi, Liqing Zhang 0001
J. Comput. Sci. Technol.1
2022 Hallucinating uncertain motion and future for static image action recognition
Li Niu 0002, Shengyuan Huang, Xing Zhao 0010, Liwei Kang, Yiyi Zhang 0002, Liqing Zhang 0001
Comput. Vis. Image Underst.5
2020 Spatio-Temporal Attention-Based Neural Network for Credit Card Fraud Detection
abstract
Credit card fraud is an important issue and incurs a considerable cost for both cardholders and issuing institutions. Contemporary methods apply machine learning-based approaches to detect fraudulent behavior from transaction records. But manually generating features needs domain knowledge and may lay behind the modus operandi of fraud, which means we need to automatically focus on the most relevant patterns in fraudulent behavior. Therefore, in this work, we propose a spatial-temporal attention-based neural network (STAN) for fraud detection. In particular, transaction records are modeled by attention and 3D convolution mechanisms by integrating the corresponding information, including spatial and temporal behaviors. Attentional weights are jointly learned in an end-to-end manner with 3D convolution and detection networks. Afterward, we conduct extensive experiments on real-word fraud transaction dataset, the result shows that STAN performs better than other state-of-the-art baselines in both AUC and precision-recall curves. Moreover, we conduct empirical studies with domain experts on the proposed method for fraud post-analysis; the result demonstrates the effectiveness of our proposed method in both detecting suspicious transactions and mining fraud patterns.
Dawei Cheng, Sheng Xiang 0001, Chencheng Shang, Yiyi Zhang 0002, Fangzhou Yang, Liqing Zhang 0001
AAAI4
2020 Exploiting Motion Information from Unlabeled Videos for Static Image Action Recognition
abstract
Static image action recognition, which aims to recognize action based on a single image, usually relies on expensive human labeling effort such as adequate labeled action images and large-scale labeled image dataset. In contrast, abundant unlabeled videos can be economically obtained. Therefore, several works have explored using unlabeled videos to facilitate image action recognition, which can be categorized into the following two groups: (a) enhance visual representations of action images with a designed proxy task on unlabeled videos, which falls into the scope of self-supervised learning; (b) generate auxiliary representations for action images with the generator learned from unlabeled videos. In this paper, we integrate the above two strategies in a unified framework, which consists of Visual Representation Enhancement (VRE) module and Motion Representation Augmentation (MRA) module. Specifically, the VRE module includes a proxy task which imposes pseudo motion label constraint and temporal coherence constraint on unlabeled videos, while the MRA module could predict the motion information of a static action image by exploiting unlabeled videos. We demonstrate the superiority of our framework based on four benchmark human action datasets with limited labeled data.
Yiyi Zhang 0002, Li Niu 0002, Meichao Luo, Jianfu Zhang 0003, Dawei Cheng, Liqing Zhang 0001
AAAI1
2020 Contagious Chain Risk Rating for Networked-guarantee Loans
abstract
The small and medium-sized enterprises (SMEs) are allowed to guarantee each other and form complex loan networks to receive loans from banks during the economic expansion stage. However, external shocks may weaken the robustness, and an accidental default may spread across the network and lead to large-scale defaults, even systemic crisis. Thus, predicting and rating the default contagion chains in the guarantee network in order to reduce or prevent potential systemic financial risk, attracts a grave concern from the Regulatory Authority and the banks. Existing credit risk models in the banking industry utilize machine learning methods to generate a credit score for each customer. Such approaches dismiss the contagion risk from guarantee chains and need extensive feature engineering with deep domain expertise. To this end, we propose a novel approach to rate the risk of contagion chains in the bank industry with the deep neural network. We employed the temporal inter-chain attention network on graph-structured loan behavior data to compute risk scores for the contagion chains. We show that our approach is significantly better than the state-of-the-art baselines on the dataset from a major financial institution in Asia. Besides, we conducted empirical studies on the real-world loan dataset for risk assessment. The proposed approach enabled loan managers to monitor risks in a boarder view and avoid significant financial losses for the financial institution.
Dawei Cheng, Zhibin Niu, Yiyi Zhang 0002
KDD3
2019 A Dynamic Default Prediction Framework for Networked-guarantee Loans
abstract
Commercial banks normally require Small and Medium Enterprises (SMEs) to provide their warranties when applying for a loan. If the borrower defaults, the guarantor is obligated to repay its loan. Such a guarantee system is designed to reduce delinquent risks, but may introduce a new dimension risk if more and more SMEs involve and subsequently form complex temporal networks. Monitoring the financial status of SMEs in these networks, and preventing or reducing systematic loan risk, is an area of great concern for both the regulatory commission and the banks. To allow possible actions to be taken in advance, this paper studies the problem of predicting repayment delinquency in the networked-guarantee loans. We propose a dynamic default prediction framework (DDPF), which preserves temporal network structures and loan behavior sequences in an end-to-end model. In particular, we design a gated recursive and attention mechanism to integrate both the loan behavior and network information. Then, we uncover risky warrant patterns by the learned weights, which effectively accelerate risk evaluation process. Finally, we conduct extensive experiments in a real-world loan risk control system to evaluate its performance, the results demonstrate the effectiveness of our proposed approach compared with state-of-the-art baselines.
Dawei Cheng, Yiyi Zhang 0002, Fangzhou Yang, Zhibin Niu, Liqing Zhang 0001
CIKM2