Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Haitian Zhong

dblp:354/9723 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2025
0009-0002-4741-3411ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 46% Trustworthy machine learning · 32% Vision and language · 22%
Databases, data mining, and information retrieval
1 paper
Knowledge graphs · 100%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
knowledge editing
1.622025
REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing · EMNLP 2025
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark · NeurIPS 2024
Machine learning › Trustworthy machine learning › robustness
overfitting mitigation
0.912025
REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing · EMNLP 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.812024
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark · NeurIPS 2024
Machine learning › Trustworthy machine learning
interpretability
0.312025
REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing · EMNLP 2025
Knowledge graphs
multimodal knowledge graph
0.212024
VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

knowledge editing · 1.5benchmark construction · 1.5principal component analysis · 0.9linear transformation · 0.9hidden-state perturbation · 0.9gated classifiers · 0.9
YearPublicationVenuePosition
2025 REACT: Representation Extraction And Controllable Tuning to Overcome Overfitting in LLM Knowledge Editing
abstract
Large language model editing methods frequently suffer from overfitting, wherein factual updates can propagate beyond their intended scope, overemphasizing the edited target even when it's contextually inappropriate.To address this challenge, we introduce REACT (Representation Extraction And Controllable Tuning), a unified two-phase framework designed for precise and controllable knowledge editing.In the initial phase, we utilize tailored stimuli to extract latent factual representations and apply Principal Component Analysis with a simple learnbale linear transformation to compute a directional "belief shift" vector for each instance.In the second phase, we apply controllable perturbations to hidden states using the obtained vector with a magnitude scalar, gated by a pre-trained classifier that permits edits only when contextually necessary.Relevant experiments on EVOKE benchmarks demonstrate that REACT significantly reduces overfitting across nearly all evaluation metrics, and experiments on COUNTERFACT and MQuAKE shows that our method preserves balanced basic editing performance (reliability, locality, and generality) under diverse editing scenarios.
Haitian Zhong, Yuhuan Liu, Guofan Liu, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan
EMNLP1
2025 PhosF3C: a feature fusion architecture with fine-tuned protein language model and conformer for prediction of general phosphorylation site
abstract
Protein phosphorylation, a key post-translational modification, provides essential insight into protein properties, making its prediction highly significant. Using the emerging capabilities of large language models (LLMs), we apply Low-Rank Adaptation (LoRA) fine-tuning to ESM2, a powerful protein large language model, to efficiently extract features with minimal computational resources, optimizing task-specific text alignment. Additionally, we integrate the conformer architecture with the feature coupling unit to enhance local and global feature exchange, further improving prediction accuracy. Our model achieves state-of-the-art performance, obtaining area under the curve scores of 79.5%, 76.3%, and 71.4% at the S, T, and Y sites of the general data sets. Based on the powerful feature extraction capabilities of LLMs, we conduct a series of analyses on protein representations, including studies on their structure, sequence, and various chemical properties [such as hydrophobicity (GRAVY), surface charge, and isoelectric point]. We propose a test method called linear regression tomography which is a top-down method using representation to explore the model's feature extraction capabilities. Our resources, including data and code, are publicly accessible at https://github.com/SkywalkerLuke/PhosF3C.
Yuhuan Liu, Haitian Zhong, Jixiu Zhai, Xiaojuan Gong, Tianchi Lu
Briefings Bioinform.3
2024 VLKEB: A Large Vision-Language Model Knowledge Editing Benchmark
abstract
Recently, knowledge editing on large language models (LLMs) has received considerable attention. Compared to this, editing Large Vision-Language Models (LVLMs) faces extra challenges from diverse data modalities and complicated model components, and data for LVLMs editing are limited. The existing LVLM editing benchmark, which comprises three metrics (Reliability, Locality, and Generality), falls short in the quality of synthesized evaluation images and cannot assess whether models apply edited knowledge in relevant content. Therefore, we employ more reliable data collection methods to construct a new Large $\textbf{V}$ision-$\textbf{L}$anguage Model $\textbf{K}$nowledge $\textbf{E}$diting $\textbf{B}$enchmark, $\textbf{VLKEB}$, and extend the Portability metric for more comprehensive evaluation. Leveraging a multi-modal knowledge graph, our image data are bound with knowledge entities. This can be further used to extract entity-related knowledge, which constitutes the base of editing data. We conduct experiments of different editing methods on five LVLMs, and thoroughly analyze how do they impact the models. The results reveal strengths and deficiencies of these methods and hopefully provide insights for future research. The codes and dataset are available at: https://github.com/VLKEB/VLKEB.
Haitian Zhong, Qiang Liu 0006, Liang Wang 0001, Tieniu Tan
NeurIPS2
2024 PTransIPs: Identification of Phosphorylation Sites Enhanced by Protein PLM Embeddings
abstract
Phosphorylation is pivotal in numerous fundamental cellular processes and plays a significant role in the onset and progression of various diseases. The accurate identification of these phosphorylation sites is crucial for unraveling the molecular mechanisms within cells and during viral infections, potentially leading to the discovery of novel therapeutic targets. In this study, we develop PTransIPs, a new deep learning framework for the identification of phosphorylation sites. Independent testing results demonstrate that PTransIPs outperforms existing state-of-the-art (SOTA) methods, achieving AUCs of 0.9232 and 0.9660 for the identification of phosphorylated S/T and Y sites, respectively. PTransIPs contributes from three aspects. 1) PTransIPs is the first to apply protein pre-trained language model (PLM) embeddings to this task. It utilizes ProtTrans and EMBER2 to extract sequence and structure embeddings, respectively, as additional inputs into the model, effectively addressing issues of dataset size and overfitting, thus enhancing model performance; 2) PTransIPs is based on Transformer architecture, optimized through the integration of convolutional neural networks and TIM loss function, providing practical insights for model design and training; 3) The encoding of amino acids in PTransIPs enables it to serve as a universal framework for other peptide bioactivity tasks, with its excellent performance shown in extended experiments of this paper.
Haitian Zhong, Bingrui He, Tianchi Lu
IEEE J. Biomed. Health Informatics2