Xinzhong Wang

dblp:36/9545 · DBLP profile ↗
← Back
10ranked-venue papers
1as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 4 · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Keep the General, Inject the Specific: Structured Dialogue Fine-Tuning for Knowledge Injection without Catastrophic Forgetting
abstract
Large Vision-Language Models (LVLMs) demonstrate impressive general-purpose capabilities but often suffer from catastrophic forgetting when incorporating specialized knowledge. To address this plasticity-stability dilemma, we introduce Structured Dialogue Fine-Tuning (SDFT), a data-centric approach that injects domain-specific concepts while preserving foundational abilities. Distinct from parameter-constrained continual learning methods, SDFT leverages a three-phase dialogue structure: Foundation Preservation reinforces pre-trained visual-linguistic alignment through captioning tasks; Contrastive Disambiguation uses carefully designed counterfactual examples to establish precise semantic boundaries; and Knowledge Specialization embeds specialized information via chain-of-thought reasoning. Evaluations across personalized entity recognition, abstract concept understanding, and biomedical domains show that SDFT significantly outperforms state-of-the-art baselines, including EWC-LoRA and O-LoRA. The results confirm SDFT’s effectiveness in balancing specialized knowledge acquisition and general capability retention.
Yijie Hong, Xiaofei Yin, Xinzhong Wang, Huijia Zhu, Sufeng Duan
ICMR3
2026 Semantic similarity guided contrastive hashing for unsupervised cross-modal retrieval
Limeng Gao, Xinzhong Wang, Zhen Zheng, Mingzhe Yang
J. Vis. Commun. Image Represent.3
2026 Relative semantic relationship preserving hashing for unsupervised cross-modal retrieval
Limeng Gao, Xinzhong Wang, Zhen Zheng
Multim. Syst.3
2025 Can Knowledge be Transferred from Unimodal to Multimodal? Investigating the Transitivity of Multimodal Knowledge Editing
Lingyong Fang, Xinzhong Wang, Depeng Wang, Zongru Wu, Huijia Zhu, Zhuosheng Zhang 0001, Gongshen Liu
ICCV2
2025 KMoP: Knowledge-injected Mixture-of-Prefix for Joint Multimodal Aspect-Based Sentiment Analysis
abstract
Multimodal Aspect-based Sentiment Analysis (MABSA) aims to extract aspect-sentiment pairs from a combination of text and images. However, images often contain content that is either irrelevant to the textual information or not related to the sentiment prediction, which can adversely affect the accuracy of the model predictions. Furthermore, existing models neglect precise regional information beyond global image features, which could also assist in enhancing aspect-based sentiment prediction, but may also introduce noise that damages the model. To address these issues, this study proposes Knowledge-injected Mixture-of-Prefix (KMoP) to inject various types of knowledge into the language model and reduce external noise. Specifically, external knowledge is injected in the form of prefixes into the language model, which minimizes catastrophic forgetting issue and generate noise-insensitive representations. Additionally, to allow different layers of the language model to automatically select the required knowledge, we differentiate the aggregation of prefixes from different knowledge sources for each layer through Mixture-of-Prefix. This paper simultaneously divides the training process into two parts, with the first phase training on the original clean dataset and the second phase fine-tuning on the original dataset with added noise. KMoP achieves state-of-the-art performance on MABSA task, with extensive supplementary experiments demonstrating its enhanced robustness to noise.
Xinzhong Wang, Lingyong Fang, Jidong Li, Gongshen Liu
ICME1
2024 Gicnet: global information capture network for visual place recognition
Shaoqi Hou, Zebang Qin, Guangqiang Yin, Xinzhong Wang, Zhiguo Wang 0004
Multim. Syst.5
2024 Fuzzy Dynamic Event-Triggered Containment Control for Human-in-the-Loop MASs With Error Constraints
abstract
In this article, the fuzzy dynamic event-triggered (DET) containment control problem for human-in-the-loop (HiTL) multiagent systems (MASs) with error constraints is investigated. Through utilization of fuzzy logic systems (FLSs), a high-gain state observer is presented to estimate unavailable states. To improve the transient performance, a fixed-time prescribed performance (FTPP) function is presented to restrict the containment errors and virtual errors. Under the backstepping control framework, a barrier Lyapunov function is designed to guarantee that the transform errors do not transgress the constraint bounds. Meanwhile, a nonlinear filter is introduced to obtain reduction calculation and enhance the control performance. Furthermore, the DET mechanism is presented to decrease the network communication burden. By directly controlling multiple leaders to form a dynamic convex hull, the proposed control strategy ensures that output signals of followers can converge to this convex hull and the containment errors can converge to the prescribed bounds within fixed time. Herein, a simulation example is presented to assess the effectiveness of the proposed control scheme.
Guohuai Lin, Hongru Ren, Qi Zhou 0002, Xinzhong Wang
IEEE Trans. Fuzzy Syst.4
2023 Semantic Information Mining and Fusion Method for Bot Detection
Lijia Liang, Xinzhong Wang, Gongshen Liu
ICANN (9)2
2023 A multitask joint framework for real-time person search
Ye Li 0024, Kangning Yin, Zhuofu Tan, Xinzhong Wang, Guangqiang Yin, Zhiguo Wang 0004
Multim. Syst.5
2023 Joint Detection and Association for End-to-End Multi-object Tracking
Ye Li 0024, Junyu Shi, Xinzhong Wang, Guangqiang Yin, Zhiguo Wang 0004
Neural Process. Lett.4