Houcheng Jiang

dblp:389/5053 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
0009-0008-6599-3572ORCID · reported

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
8 papers
Language models and text generation · 57% Vision and language · 9% Generative modeling · 9%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
knowledge editing
5.672026
LPEdit: Locality-Preserving Knowledge Editing for MultiModal Large Language Models · WWW 2026
Explainable and Efficient Editing for Large Language Models · WWW 2025
AnyEdit: Edit Any Knowledge Encoded in Language Models · ICML 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
1.822026
LPEdit: Locality-Preserving Knowledge Editing for MultiModal Large Language Models · WWW 2026
Towards Neuron Attributions in Multi-Modal Large Language Models · NeurIPS 2024
Natural language and speech › Language models and text generation › knowledge editing
multimodal knowledge editing
1.012026
LPEdit: Locality-Preserving Knowledge Editing for MultiModal Large Language Models · WWW 2026
Machine learning › Efficient and distributed learning › inference acceleration
caching
0.912025
Accelerating Diffusion Transformer via Error-Optimized Cache · ACM Multimedia 2025
Machine learning › Generative modeling
diffusion model
0.912025
Accelerating Diffusion Transformer via Error-Optimized Cache · ACM Multimedia 2025
Machine learning › Generative modeling › diffusion model
diffusion transformer
0.912025
Accelerating Diffusion Transformer via Error-Optimized Cache · ACM Multimedia 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models · ICLR 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Neuron-Level Sequential Editing for Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation › knowledge editing
lifelong model editing
0.912025
Reinforced Lifelong Editing for Language Models · ICML 2025
Machine learning › Efficient and distributed learning
model acceleration
0.912025
Accelerating Diffusion Transformer via Error-Optimized Cache · ACM Multimedia 2025
Natural language and speech › Language models and text generation › knowledge editing
neuron-level editing
0.912025
Neuron-Level Sequential Editing for Large Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation › knowledge editing
sequential model editing
0.912025
Neuron-Level Sequential Editing for Large Language Models · ACL (1) 2025
Machine learning › Trustworthy machine learning
interpretability
0.812024
Towards Neuron Attributions in Multi-Modal Large Language Models · NeurIPS 2024
Machine learning › Trustworthy machine learning › interpretability › attribution methods
neuron attribution
0.812024
Towards Neuron Attributions in Multi-Modal Large Language Models · NeurIPS 2024
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge base
factual knowledge storage
0.312025
Explainable and Efficient Editing for Large Language Models · WWW 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › belief change
knowledge update
0.312025
Reinforced Lifelong Editing for Language Models · ICML 2025

Methods — techniques the papers use, named apart from their topics

policy optimization · 0.9null space projection · 0.9mutual information · 0.9locating-then-editing · 0.9hypernetwork · 0.9hidden state optimization · 0.9error-optimized cache · 0.9clustering · 0.9autoregressive editing · 0.9activation-based neuron selection · 0.9
YearPublicationVenuePosition
2026 LPEdit: Locality-Preserving Knowledge Editing for MultiModal Large Language Models
Junfeng Fang, Houcheng Jiang, Xiang Wang 0010, Xiangnan He 0001
WWW3
2025 Neuron-Level Sequential Editing for Large Language Models
abstract
This work explores sequential model editing in large language models (LLMs), a critical task that involves modifying internal knowledge within LLMs continuously through multi-round editing, each incorporating updates or corrections to adjust the model’s outputs without the need for costly retraining. Existing model editing methods, especially those that alter model parameters, typically focus on single-round editing and often face significant challenges in sequential model editing-most notably issues of model forgetting and failure. To address these challenges, we introduce a new model editing method, namely Neuron-level Sequential Editing (NSE), tailored for supporting sequential model editing. Specifically, we optimize the target layer’s hidden states using the model’s original weights to prevent model failure. Furthermore, we iteratively select neurons in multiple layers for editing based on their activation values to mitigate model forgetting. Our empirical experiments demonstrate that NSE significantly outperforms current modifying parameters model editing methods, marking a substantial advancement in the field of sequential model editing. Our code is released on https://anonymous.4open.science/r/NSE-0A8D/.
Houcheng Jiang, Junfeng Fang, Baolong Bi, An Zhang 0003, Xiang Wang 0010
ACL (1)1
2025 AlphaEdit: Null-Space Constrained Knowledge Editing for Language Models
abstract
Large language models (LLMs) often exhibit hallucinations, producing incorrect or outdated knowledge. Hence, model editing methods have emerged to enable targeted knowledge updates. To achieve this, a prevailing paradigm is the locating-then-editing approach, which first locates influential parameters and then edits them by introducing a perturbation. While effective, current studies have demonstrated that this perturbation inevitably disrupt the originally preserved knowledge within LLMs, especially in sequential editing scenarios. To address this, we introduce AlphaEdit, a novel solution that projects perturbation onto the null space of the preserved knowledge before applying it to the parameters. We theoretically prove that this projection ensures the output of post-edited LLMs remains unchanged when queried about the preserved knowledge, thereby mitigating the issue of disruption. Extensive experiments on various LLMs, including LLaMA3, GPT2-XL, and GPT-J, show that AlphaEdit boosts the performance of most locating-then-editing methods by an average of 36.7% with a single line of additional code for projection solely.
Junfeng Fang, Houcheng Jiang, Kun Wang 0056, Yunshan Ma 0002, Jie Shi 0005, Xiang Wang 0010, Xiangnan He 0001, Tat-Seng Chua
ICLR2
2025 Reinforced Lifelong Editing for Language Models
abstract
Large language models (LLMs) acquire information from pre-training corpora, but their stored knowledge can become inaccurate or outdated over time. Model editing addresses this challenge by modifying model parameters without retraining, and prevalent approaches leverage hypernetworks to generate these parameter updates. However, they face significant challenges in lifelong editing due to their incompatibility with LLM parameters that dynamically change during the editing process. To address this, we observed that hypernetwork-based lifelong editing aligns with reinforcement learning modeling and proposed **RLEdit**, an RL-based editing method. By treating editing losses as rewards and optimizing hypernetwork parameters at the full knowledge sequence level, we enable it to precisely capture LLM changes and generate appropriate parameter updates. Our extensive empirical evaluation across several LLMs demonstrates that RLEdit outperforms existing methods in lifelong editing with superior effectiveness and efficiency, achieving a **59.24%** improvement while requiring only **2.11%** of the time compared to most approaches.
Zherui Li 0001, Houcheng Jiang, Baolong Bi, Zhenhong Zhou, Fei Sun 0001, Junfeng Fang, Xiang Wang 0010
ICML2
2025 AnyEdit: Edit Any Knowledge Encoded in Language Models
abstract
Large language models (LLMs) often produce incorrect or outdated information, necessitating efficient and precise knowledge updates. Current model editing methods, however, struggle with long-form knowledge in diverse formats, such as poetry, code snippets, and mathematical derivations. These limitations arise from their reliance on editing a single token’s hidden state, a limitation we term as ``efficacy barrier''. To solve this, we propose \textbf{AnyEdit}, a new autoregressive editing paradigm. It decomposes long-form knowledge into sequential chunks and iteratively edits the key token in each chunk, ensuring consistent and accurate outputs. Theoretically, we ground AnyEdit in the Chain Rule of Mutual Information, showing its ability to update any knowledge within LLMs. Empirically, it outperforms strong baselines by 21.5\% on benchmarks including UnKEBench, AKEW, and our new \textbf{EditEverything} dataset for long-form diverse-formatted knowledge. Additionally, AnyEdit serves as a plug-and-play framework, enabling current editing methods to update knowledge with arbitrary length and format, significantly advancing the scope and practicality of LLM knowledge editing. Our code is available at: \url{https://github.com/jianghoucheng/AnyEdit}.
Houcheng Jiang, Junfeng Fang, Ningyu Zhang 0001, Mingyang Wan, Guojun Ma, Xiang Wang 0010, Xiangnan He 0001, Tat-Seng Chua
ICML1
2025 Accelerating Diffusion Transformer via Error-Optimized Cache
abstract
Diffusion Transformer (DiT) is a crucial method for content generation. However, it needs a lot of time to sample. Many studies have attempted to use caching to reduce the time consumption of sampling. Existing caching methods accelerate generation by reusing DiT features from the previous time step and skipping calculations in the next, but they tend to locate and cache low-error modules without focusing on reducing caching-induced errors, resulting in a sharp decline in generated content quality when increasing caching intensity. To solve this problem, we propose the Error-Optimized Cache (EOC). This method introduces three key improvements: (1) Prior knowledge extraction: Extract and process the caching differences; (2) A judgment method for cache optimization: Determine whether certain caching steps need to be optimized; (3) Cache optimization: reduce caching errors. Experiments show that this algorithm significantly reduces the error accumulation caused by caching, especially excessive caching. On the ImageNet dataset, without substantially increasing the computational load, this method improves the FID↓ of the generated images when the rule-based model FORA has a caching level of 75%, 50%, and 25%, and the training-based model Learning-to-cache has a caching level of 22%. Specifically, the FID↓ values change from 30.454 to 21.690 (28.8%), from 6.857 to 5.821 (15.1%), from 3.870 to 3.692 (4.6%), and from 3.539 to 3.451 (2.5%) respectively. Code is available at https://github.com/qiujx0520/EOC_MM2025.git.
Junxiang Qiu, Shuo Wang 0008, Jinda Lu, Houcheng Jiang, Yanbin Hao
ACM Multimedia5
2025 Explainable and Efficient Editing for Large Language Models
abstract
Large Language Models (LLMs) exhibit remarkable capabilities in storing and retrieving vast amounts of factual knowledge. However, they retain outdated or incorrect information from Web corpora. Since full retraining is costly, locate-and-edit model editing methods offer a feasible alternative. Current methods typically follow a two-stage paradigm: (1) identifying critical layers that store knowledge and (2) updating their parameters to store new knowledge. However, both phases have their inherent limitations. Firstly, layer identification is independent of the knowledge being updated, ignoring the differences in knowledge storage patterns. Secondly, parameter updating suffers from high computational overhead due to gradient descent. To solve these, we propose an Explainable and effiCient model Editing method, termed ECE. Specifically, we integrate LLM explainability into the editing process, enabling the adaptive identification of the crucial neurons. Through clustering similar knowledge, we enable batch optimization in a single gradient step, significantly reducing computational time without compromising effectiveness. Extensive experiments demonstrate that ECE can achieve superior performance, showcasing the potential of explainability-driven editing methods for LLMs. Code is available at https://github.com/tianyuzhangterry/ECE.
Junfeng Fang, Houcheng Jiang, Baolong Bi, Xiang Wang 0010, Xiangnan He 0001
WWW3
2024 Towards Neuron Attributions in Multi-Modal Large Language Models
abstract
As Large Language Models (LLMs) demonstrate impressive capabilities, demystifying their internal mechanisms becomes increasingly vital. Neuron attribution, which attributes LLM outputs to specific neurons to reveal the semantic properties they learn, has emerged as a key interpretability approach. However, while neuron attribution has made significant progress in deciphering text-only LLMs, its application to Multimodal LLMs (MLLMs) remains less explored. To address this gap, we propose a novel Neuron Attribution method tailored for MLLMs, termed NAM. Specifically, NAM not only reveals the modality-specific semantic knowledge learned by neurons within MLLMs, but also highlights several intriguing properties of neurons, such as cross-modal invariance and semantic sensitivity. These properties collectively elucidate the inner workings mechanism of MLLMs, providing a deeper understanding of how MLLMs process and generate multi-modal content. Through theoretical analysis and empirical validation, we demonstrate the efficacy of NAM and the valuable insights it offers. Furthermore, leveraging NAM, we introduce a multi-modal knowledge editing paradigm, underscoring the practical significance of our approach for downstream applications of MLLMs.
Junfeng Fang, Zac Bi, Houcheng Jiang, Yuan Gao 0020, Kun Wang 0056, An Zhang 0003, Jie Shi 0005, Xiang Wang 0010, Tat-Seng Chua
NeurIPS4