Jiaying Zhu

dblp:278/8488 · DBLP profile ↗
← Back
13ranked-venue papers
3as first author
12since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 2 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Systems, architecture and hardware · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Skill Path: Unveiling Language Skills from Circuit Graphs
abstract
Circuit graph discovery has emerged as a fundamental approach to elucidating the skill mechanistic of language models. Despite the output faithfulness of circuit graphs, they suffer from atomic ablation, which causes the loss of causal dependencies between connected components. In addition, their discovery process, designed to preserve output faithfulness, inadvertently captures extraneous effects other than an isolated target skill. To alleviate these challenges, we introduce skill paths, which offer a more refined and compact representation by isolating individual skills within a linear chain of components. To enable skill path extracting from circuit graphs, we propose a three-step framework, consisting of decomposition, pruning, and post-hoc causal mediation. In particular, we offer a complete linear decomposition of the transformer model which leads to a disentangled computation graph. After pruning, we further adopt causal analysis techniques, including counterfactuals and interventions, to extract the final skill paths from the circuit graph. To underscore the significance of skill paths, we investigate three generic language skills—Previous Token Skill, Induction Skill, and In-Context Learning Skill—using our framework. Experiments support two crucial properties of these skills, namely stratification and inclusiveness.
Hang Chen 0002, Xinyu Yang 0001, Jiaying Zhu, Wenya Wang 0001
AAAI3
2025 A Lottery Ticket Hypothesis Approach with Sparse Fine-tuning and MAE for Image Forgery Detection and Localization
abstract
The rise in sophisticated image forgery techniques, driven by advancements in image editing and generation, has posed new security challenges. Traditional methods, designed for specific tampering artifacts, struggle with out-of-distribution image forgery detection. In this paper, we propose a shift in paradigm, placing greater emphasis on the universal characteristics of authentic images, as opposed to solely focusing on specific forgery signals. We introduce an enhancement to the Masked Autoencoder (MAE), aptly termed the Forgery MAE (FMAE). This modification retains the inherent characteristics of natural images while integrating multi-source forgery information. Our implementation involves applying the lottery ticket hypothesis during pre-training to identify forgery-sensitive parameters, followed by their sparse fine-tuning to target the forgery detection and localization task. Concurrently, we develop a ``mixture of experts'' noise extractor to compile multi-source forgery data. Our FMAE effectively extracts forgery features and shows strong resilience against unseen forgeries. Extensive experiments across multiple datasets confirm our method's superior accuracy and generalization capability over existing techniques.
Jiaying Zhu, Dong Li 0055, Xueyang Fu, Gege Shi, Jie Xiao 0002, Aiping Liu, Zhengjun Zha
AAAI1
2025 Quantifying Semantic Emergence in Language Models
abstract
Large language models (LLMs) are widely recognized for their exceptional capacity to capture semantics meaning.Yet, there remains no established metric to quantify this capability.In this work, we introduce a quantitative metric, Information Emergence (IE), designed to measure LLMs' ability to extract semantics from input tokens.We formalize "semantics" as the meaningful information abstracted from a sequence of tokens and quantify this by comparing the entropy reduction observed for a sequence of tokens (macro-level) and individual tokens (micro-level).To achieve this, we design a lightweight estimator to compute the mutual information at each transformer layer, which is agnostic to different tasks and language model architectures.We apply IE in both synthetic in-context learning (ICL) scenarios and natural sentence contexts.Experiments demonstrate informativeness and patterns about semantics.While some of these patterns confirm the conventional prior linguistic knowledge, the rest are relatively unexpected, which may provide new insights.
Hang Chen 0002, Xinyu Yang 0001, Jiaying Zhu, Wenya Wang 0001
ACL (1)3
2025 Logic Optimization Meets SAT: A Novel Framework for Circuit-SAT Solving
abstract
The Circuit Satisfiability (CSAT) problem, a variant of the Boolean Satisfiability (SAT) problem, plays a critical role in integrated circuit design and verification. However, existing SAT solvers, optimized for Conjunctive Normal Form (CNF), often struggle with the intrinsic complexity of circuit structures when directly applied to CSAT instances. To address this challenge, we propose a novel preprocessing framework that leverages advanced logic synthesis techniques and a reinforcement learning (RL) agent to optimize CSAT problem instances. The framework introduces a cost-customized Look-Up Table (LUT) mapping strategy that prioritizes solving efficiency, effectively transforming circuits into simplified forms tailored for SAT solvers. Our method achieves significant runtime reductions across diverse industrial-scale CSAT benchmarks, seamlessly integrating with state-of-the-art SAT solvers. Extensive experimental evaluations demonstrate up to $63 \%$ reduction in solving time compared to conventional approaches, highlighting the potential of EDAdriven innovations to advance SAT-solving capabilities.
Zhengyuan Shi, Tiebing Tang, Jiaying Zhu, Sadaf Khan, Hui-Ling Zhen, Mingxuan Yuan, Zhufei Chu, Qiang Xu 0001
DAC3
2025 Learnable Frequency Decomposition for Image Forgery Detection and Localization
abstract
Concern for image authenticity spurs research in image forgery detection and localization (IFDL). Most deep learning-based methods focus primarily on spatial domain modeling and have not fully explored frequency domain strategies. In this paper, we observe and analyze the frequency characteristic changes caused by image tampering. Observations indicate that manipulation traces are especially prominent in phase components and span both low and high-frequency bands. Based on these findings, we propose a forensic frequency decomposition network (F2D-Net), which incorporates deep Fourier transforms and leverages both phase information and high and low-frequency components to enhance IFDL. Specifically, F2D-Net consists of the Spectral Decomposition Subnetwork (SDSN) and the Frequency Separation Subnetwork (FSSN). The former decomposes the image into amplitude and phase, focusing on learning the semantic content in the phase spectrum to identify forged objects, thus improving forgery detection accuracy. The latter further adaptively decomposes the output of the SDSN to obtain corresponding high and low frequencies, and applies a divide-and-conquer strategy to refine each frequency band, mitigating the optimization difficulties caused by coupled forgery traces across different frequencies, thereby better capturing the pixels belonging to the forged object to improve localization accuracy. Experiments on multiple datasets demonstrate that our method outperforms state-of-the-art image forgery detection and localization techniques both qualitatively and quantitatively.
Dong Li 0055, Jiaying Zhu, Yidi Liu, Xin Lu 0008, Xueyang Fu, Jiawei Liu 0001, Aiping Liu, Zhengjun Zha
IJCAI2
2025 Rethinking Circuit Completeness in Language Models: AND, OR, and ADDER Gates
abstract
Circuit discovery has gradually become one of the prominent methods for mechanistic interpretability, and research on circuit completeness has also garnered increasing attention. Methods of circuit discovery that do not guarantee completeness not only result in circuits that are not fixed across different runs but also cause key mechanisms to be omitted. The nature of incompleteness arises from the presence of OR gates within the circuit, which are often only partially detected in standard circuit discovery methods. To this end, we systematically introduce three types of logic gates: AND, OR, and ADDER gates, and decompose the circuit into combinations of these logical gates. Through the concept of these gates, we derive the minimum requirements necessary to achieve faithfulness and completeness. Furthermore, we propose a framework that combines noising-based and denoising-based interventions, which can be easily integrated into existing circuit discovery methods without significantly increasing computational complexity. This framework is capable of fully identifying the logic gates and distinguishing them within the circuit. In addition to the extensive experimental validation of the framework's ability to restore the faithfulness, completeness, and sparsity of circuits, using this framework, we uncover fundamental properties of the three logic gates, such as their proportions and contributions to the output, and explore how they behave among the functionalities of language models.
Hang Chen 0002, Jiaying Zhu, Xinyu Yang 0001, Wenya Wang 0001
NeurIPS2
2025 A Robust Frequency MMSE Estimator Based on the Weighted Average of Multipath Signal for OTFS Systems
abstract
A robust frequency minimum mean square error (MMSE) estimator based on the weighted average of multipath signal for OTFS systems is proposed in this paper. Previous frequency estimators are based on the strongest single path signal or equally weighted average of multiple path signals so that the accuracies of estimators are limited. We generalise the weight coefficients to be taken arbitrarily, and their optimal values are determined such that the MSE of estimator is minimum. Simulation results show that the proposed method outperforms existing frequency estimators in terms of MSE performance, particularly when the normalized carrier frequency offset (CFO) is less than 0.5 and the symbol signal-to-noise ratio (SNR) of received signal is below 2 dB. Under static channel conditions, the proposed estimator can achieve a reduction in pilot energy of approximately 21.6 dB.
Yishan He, Jie Wang 0105, Jiaying Zhu, Nengtang Hua, Huilin Song
VTC2025-Spring5
2024 Learning Discriminative Noise Guidance for Image Forgery Detection and Localization
abstract
This study introduces a new method for detecting and localizing image forgery by focusing on manipulation traces within the noise domain. We posit that nearly invisible noise in RGB images carries tampering traces, useful for distinguishing and locating forgeries. However, the advancement of tampering technology complicates the direct application of noise for forgery detection, as the noise inconsistency between forged and authentic regions is not fully exploited. To tackle this, we develop a two-step discriminative noise-guided approach to explicitly enhance the representation and use of noise inconsistencies, thereby fully exploiting noise information to improve the accuracy and robustness of forgery detection. Specifically, we first enhance the noise discriminability of forged regions compared to authentic ones using a de-noising network and a statistics-based constraint. Then, we merge a model-driven guided filtering mechanism with a data-driven attention mechanism to create a learnable and differentiable noise-guided filter. This sophisticated filter allows us to maintain the edges of forged regions learned from the noise. Comprehensive experiments on multiple datasets demonstrate that our method can reliably detect and localize forgeries, surpassing existing state-of-the-art methods.
Jiaying Zhu, Dong Li 0055, Xueyang Fu, Jie Huang 0017, Aiping Liu, Zhengjun Zha
AAAI1
2024 Noise-Assisted Prompt Learning for Image Forgery Detection and Localization
Dong Li 0055, Jiaying Zhu, Xueyang Fu, Xun Guo 0001, Yidi Liu, Jiawei Liu 0001, Zhengjun Zha
ECCV (11)2
2023 Edge-aware Regional Message Passing Controller for Image Forgery Localization
abstract
Digital image authenticity has promoted research on image forgery localization. Although deep learning-based methods achieve remarkable progress, most of them usually suffer from severe feature coupling between the forged and authentic regions. In this work, we propose a two-step Edge-aware Regional Message Passing Controlling strategy to address the above issue. Specifically, the first step is to account for fully exploiting the edge information. It consists of two core designs: context-enhanced graph construction and threshold-adaptive differentiable binarization edge algorithm. The former assembles the global semantic information to distinguish the features between the forged and authentic regions, while the latter stands on the output of the former to provide the learnable edges. In the second step, guided by the learnable edges, a region message passing controller is devised to weaken the message passing between the forged and authentic regions. In this way, our ERMPC is capable of explicitly modeling the inconsistency between the forged and authentic regions and enabling it to perform well on refined forged images. Extensive experiments on several challenging benchmarks show that our method is superior to state-of-the-art image forgery localization methods qualitatively and quantitatively.
Dong Li 0055, Jiaying Zhu, Menglu Wang 0003, Jiawei Liu 0001, Xueyang Fu, Zhengjun Zha
CVPR2
2022 Distributed Event-Triggered Impulsive Consensus Control of Nonlinear Multi-Agent Systems Under Malicious Attacks
abstract
This paper addresses a secure event-triggered consensus control problem for a class of nonlinear leader-following multi-agent systems (MASs) under denial of service (DoS) attacks and deception attacks. The novelty of this study lies in the development of a secure event-triggered impulsive control (ETIC) method that can effectively deal with the effects of random DoS and deception attacks, while also achieving the desired resource-efficient consensus control performance. More specifically, in order to save the previous communication resources, an event-triggered impulsive consensus control protocol is proposed to reduce data exchanges among the agents. Our further aim is to simultaneously guarantee the mean-square exponential consensus performance of the controlled MAS and achieve attack resilience as well as resource efficiency. Sufficient conditions on the consensus performance analysis are derived, and the lower bound of impulsive sequences is further specified to exclude the Zeno behavior. Finally, a numerical example involving a group of Chua’s circuits is given to illustrate the effectiveness of the derived theoretical results.
Jiaying Zhu, Wangli He, Xiaohua Ge
IECON1
2021 LPF: A Language-Prior Feedback Objective Function for De-biased Visual Question Answering
abstract
Most existing Visual Question Answering (VQA) systems tend to overly rely on the language bias and hence fail to reason from the visual clue. To address this issue, we propose a novel Language-Prior Feedback (LPF) objective function, to re-balance the proportion of each answer's loss value in the total VQA loss. The LPF firstly calculates a modulating factor to determine the language bias using a question-only branch. Then, the LPF assigns a self-adaptive weight to each training sample in the training process. With this reweighting mechanism, the LPF ensures that the total VQA loss can be reshaped to a more balanced form. By this means, the samples that require certain visual information to predict will be efficiently used during training. Our method is simple to implement, model-agnostic, and end-to-end trainable. We conduct extensive experiments and the results show that the LPF (1) brings a significant improvement over various VQA models, (2) achieves competitive performance on the bias-sensitive VQA-CP v2 benchmark.
Zujie Liang, Haifeng Hu 0001, Jiaying Zhu
SIGIR3
2020 Learning to Contrast the Counterfactual Samples for Robust Visual Question Answering
abstract
In the task of Visual Question Answering (VQA), most state-of-the-art models tend to learn spurious correlations in the training set and achieve poor performance in out-ofdistribution test data.Some methods of generating counterfactual samples have been proposed to alleviate this problem.However, the counterfactual samples generated by most previous methods are simply added to the training data for augmentation and are not fully utilized.Therefore, we introduce a novel selfsupervised contrastive learning mechanism to learn the relationship between original samples, factual samples and counterfactual samples.With the better cross-modal joint embeddings learned from the auxiliary training objective, the reasoning capability and robustness of the VQA model are boosted significantly.We evaluate the effectiveness of our method by surpassing current state-of-the-art models on the VQA-CP dataset, a diagnostic benchmark for assessing the VQA model's robustness.
Zujie Liang, Weitao Jiang, Haifeng Hu 0001, Jiaying Zhu
EMNLP (1)4