VLDB 2026 Research / reviewers in the wild / expert
Xi Wang 0018
dblp:08/5760-18
· DBLP profile ↗
11ranked-venue papers
3as first author
11since 2021 · last 2026
0009-0005-7668-3965ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 2 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EMSEdit: Efficient Multi-Step Meta-Learning-based Model EditingabstractLarge Language Models (LLMs) power numerous AI applications, yet updating their knowledge remains costly. Model editing provides a lightweight alternative through targeted parameter modifications, with meta-learning-based model editing (MLME) demonstrating strong effectiveness and efficiency. However, we find that MLME struggles in low-data regimes and incurs high training costs due to the use of KL divergence. To address these issues, we propose $\textbf{E}$fficient $\textbf{M}$ulti-$\textbf{S}$tep $\textbf{Edit (EMSEdit)}$, which leverages multi-step backpropagation (MSBP) to effectively capture gradient-activation mapping patterns within editing samples, performs multi-step edits per sample to enhance editing performance under limited data, and introduces norm-based regularization to preserve unedited knowledge while improving training efficiency. Experiments on two datasets and three LLMs show that EMSEdit consistently outperforms state-of-the-art methods in both sequential and batch editing. Moreover, MSBP can be seamlessly integrated into existing approaches to yield additional performance gains. Further experiments on a multi-hop reasoning editing task demonstrate EMSEdit's robustness in handling complex edits, while ablation studies validate the contribution of each design component. Our code is available at https://github.com/xpq-tech/emsedit. Xiaopeng Li 0006, Shasha Li 0001, Xi Wang 0018, Shezheng Song, Bin Ji 0002, Shangwen Wang, Jun Ma 0015, Xiaodong Liu 0004, Mina Liu, Jie Yu 0008 |
WWW | 3 |
| 2025 | SWEA: Updating Factual Knowledge in Large Language Models via Subject Word Embedding AlteringabstractThe general capabilities of large language models (LLMs) make them the infrastructure for various AI applications, but updating their inner knowledge requires significant resources. Recent model editing is a promising technique for efficiently updating a small amount of knowledge of LLMs and has attracted much attention. In particular, local editing methods, which directly update model parameters, are proven suitable for updating small amounts of knowledge. Local editing methods update weights by computing least squares closed-form solutions and identify edited knowledge by vector-level matching in inference, which achieve promising results. However, these methods still require a lot of time and resources to complete the computation. Moreover, vector-level matching lacks reliability, and such updates disrupt the original organization of the model's parameters. To address these issues, we propose a detachable and expandable Subject Word Embedding Altering (SWEA) framework, which finds the editing embeddings through token-level matching and adds them to the subject word embeddings in Transformer input. To get these editing embeddings, we propose optimizing then suppressing fusion method, which first optimizes learnable embedding vectors for the editing target and then suppresses the Knowledge Embedding Dimensions (KEDs) to obtain final editing embeddings. We thus propose SWEAOS method for editing factual knowledge in LLMs. We demonstrate the overall state-of-the-art (SOTA) performance of SWEAOS on the CounterFact and zsRE datasets. To further validate the reasoning ability of SWEAOS in editing knowledge, we evaluate it on the more complex RippleEdits benchmark. The results demonstrate that SWEAOS possesses SOTA reasoning ability. Xiaopeng Li 0006, Shasha Li 0001, Shezheng Song, Huijun Liu 0003, Bin Ji 0002, Xi Wang 0018, Jun Ma 0015, Jie Yu 0008, Xiaodong Liu 0004 |
AAAI | 6 |
| 2025 | DCTMamba: Advancing JPEG Image Restoration Through Long-Sequence Modeling and Adaptive Frequency StrategyabstractDespite the advanced long-sequence modeling of Mamba, which has expanded its applications in image restoration, there remains a lack of exploration combining its strengths with the specific characteristics of JPEG image restoration, where high-frequency components are lost after the Discrete Cosine Transform (DCT). To address this, we introduce DCTMamba, a new framework designed to apply Mamba more effectively to JPEG image restoration. Specifically, our method integrates the Discrete Cosine Transform (DCT) into the Mamba to establish the sequential scanning from lower to higher frequencies, enabling the network to initially reconstruct coarse structures and progressively refine the image with more intricate details. Furthermore, recognizing the variable frequency distributions that arise from DCT transformations across different image sizes, we have developed Scale-Adaptive Normalization to manage these variations adeptly. Comprehensive experiments confirm that DCTMamba significantly outperforms existing solutions, achieving high fidelity in both coarse structures and fine details.CTMamba significantly outperforms existing solutions, achieving high fidelity in both coarse structures and fine details. Xi Wang 0018, Xueyang Fu, Liang Li 0003, Zhengjun Zha |
AAAI | 1 |
| 2025 | IntentVCNet: Bridging Spatio-Temporal Gaps for Intention-Oriented Controllable Video CaptioningabstractIntent-oriented controlled video captioning aims to generate targeted descriptions for specific targets in a video based on customized user intent. Current Large Visual Language Models (LVLMs) have gained strong instruction following and visual comprehension capabilities. Although the LVLMs demonstrated proficiency in spatial and temporal understanding respectively, it was not able to perform fine-grained spatial control in time sequences in direct response to instructions. This substantial spatio-temporal gap complicates efforts to achieve fine-grained intention-oriented control in video. Towards this end, we propose a novel IntentVCNet that unifies the temporal and spatial understanding knowledge inherent in LVLMs to bridge the spatio-temporal gap from both prompting and model perspectives. Specifically, we first propose a prompt combination strategy designed to enable LLM to model the implicit relationship between prompts that characterize user intent and video sequences. We then propose a parameter efficient box adapter that augments the object semantic information in the global visual context so that the visual token has a priori information about the user intent. The final experiment proves that the combination of the two strategies can further enhance the LVLM's ability to model spatial details in video sequences, and facilitate the LVLMs to accurately generate controlled intent-oriented captions. Our proposed method achieved state-of-the-art results in several open source LVLMs and was the runner-up in the IntentVC challenge. Our code is available on https://github.com/thqiu0419/IntentVCNet. Tianheng Qiu, Jingchun Gao, Huiyi Leong, Xi Wang 0018, Xiaocheng Zhang, Kele Xu, Lan Zhang 0002 |
ACM Multimedia | 6 |
| 2025 | Deep Unfolding Network for Image Desnowing With Snow Shape PriorabstractEffectively leveraging snow image formulation, which accounts for atmospheric light and snow masks, is crucial for enhancing image desnowing performance and improving interpretability. However, current direct-learning approaches often neglect this formulation, while model-based methods use it in overly simplistic ways. To address this, we propose a novel unfolding network that iteratively refines the desnowing process for more thorough optimization. Additionally, model-based techniques usually rely on real-world snow masks for supervision, a requirement that is impractical in many real-world applications. To overcome this limitation, we introduce a snow shape prior as a surrogate supervision signal. We further integrate the physical properties of atmospheric light and heavy snow by decomposing the optimization task into manageable sub-problems within our unfolding network. Extensive evaluations on multiple benchmark datasets confirm that our method outperforms current state-of-the-art techniques. Xin Guo 0018, Xi Wang 0018, Xueyang Fu, Zhengjun Zha |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | DDCNet: Advanced Decoupling of Degradation and Content for Adverse Weather Image RestorationabstractAdverse weather image restoration aims to recover clear images from those affected by weather conditions such as rain, haze, and snow. Different weather types affect images in distinct ways, necessitating specific degradation removal strategies, while content reconstruction generally benefits from a consistent approach since the underlying image structure remains largely consistent. Previous methods, despite their ability to handle multiple weather conditions within a single framework, often failed to adequately separate these two critical processes, thereby adversely affecting image restoration quality. In this article, we present DDCNet, a novel framework designed to explicitly decouple degradation removal and content reconstruction when processing various adverse weather conditions within a unified network. We achieve this by separating tailored degradation removal from uniform content reconstruction at the feature level, based on channel statistics. Additionally, we utilize the Fourier transform to enhance both processes. Furthermore, to address the differing optimization directions required by different adverse weather types, we propose a novel degradation mapping (DM) loss function to constrain their respective optimization paths. Extensive experiments show that DDCNet establishes new performance standards across multiple adverse weather scenarios. Xi Wang 0018, Xueyang Fu, Yurui Zhu, Zhengjun Zha |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2024 | Offline Textual Adversarial Attacks against Large Language ModelsabstractThis work centers on textual adversarial attacks against large language models (LLMs) and proposes a new reproducible benchmark for future study. Unlike pre-trained language models (PLMs) which can output predicted class probabilities as feedback to instruct the generation of adversarial examples, LLMs cannot accurately provide such feedback due to their generative nature, making existing attack modes unsuitable. To address this issue, we propose Offline-Attack, an offline method tailored for LLMs that contains a novel Transformer-based Adversarial Machine Translation (AMT) framework. AMT is trained on one self-constructed large-scale adversarial dataset and used to translate original texts to adversarial examples. To mitigate training bias, we induce LLMs to generate stable prediction confidence and incorporate it into AMT training process. The evaluation, spanning four text classification datasets against LLaMA-2-13b-chat, showcases Offline-Attack’s robust performance, particularly achieving 44.3% attack success rate on average. Moreover, Offline-Attack exhibits promising attack ability to other LLMs like Vicuna-33b and ChatGPT. Our study paves the way for future study by presenting strong and reproducible baselines for textual adversarial attacks against LLMs. Huijun Liu 0003, Bin Ji 0002, Jie Yu 0008, Shasha Li 0001, Jun Ma 0015, Miaomiao Li 0001, Xi Wang 0018 |
IJCNN | 7 |
| 2024 | FEAttack: A Fast and Efficient Hard-Label Textual Attack Framework
Miaomiao Li 0001, Jun Ma 0015, Jie Yu 0008, Shasha Li 0001, Huijun Liu 0003, Xi Wang 0018 |
WASA (2) | 8 |
| 2022 | JPEG Artifacts Removal via Contrastive Representation Learning
Xi Wang 0018, Xueyang Fu, Yurui Zhu, Zhengjun Zha |
ECCV (17) | 1 |
| 2022 | Single Image Shadow Detection via Complementary MechanismabstractIn this paper, we present a novel shadow detection framework by investigating the mutual complementary mechanisms contained in this specific task. Our method is based on a key observation: in a single shadow image, shadow regions and non-shadow counterparts are complementary to each other in nature, thus a better estimation on one side leads to an improved estimation on the other, and vice versa. Motivated by this observation, we first leverage two parallel interactive branches to jointly produce shadow and non-shadow masks. The interaction between two parallel branches is to retain the deactivated intermediate features of one branch by introducing the negative activation technique, which could serve as complementary features to the other branch. Besides, we also apply identity reconstruction loss as complementary training guidance at the image level. Finally, we design two discriminative losses to satisfy the complementary requirements of shadow detection, i.e., neither missing any shadow regions nor falsely detecting non-shadow regions. By fully exploring and exploiting the complementary mechanism of shadow detection, our method can confidently predict more accurate shadow detection results. Extensive experiments on the three widely-used benchmarks demonstrate our proposed method achieves superior shadow detection performance against state-of-the-art methods with a relatively low computational cost. Yurui Zhu, Xueyang Fu, Chengzhi Cao, Xi Wang 0018, Qibin Sun, Zhengjun Zha |
ACM Multimedia | 4 |
| 2021 | Learning Dual Priors for JPEG Compression Artifacts RemovalabstractDeep learning (DL)-based methods have achieved great success in solving the ill-posed JPEG compression artifacts removal problem. However, as most DL architectures are designed to directly learn pixel-level mapping relationship-s, they largely ignore semantic-level information and lack sufficient interpretability. To address the above issues, in this work, we propose an interpretable deep network to learn both pixel-level regressive prior and semantic-level discriminative prior. Specifically, we design a variational model to formulate the image de-blocking problem and propose two prior terms for the image content and gradient, respectively. The content-relevant prior is formulated as a DL-based image-to-image regressor to perform as a de-blocker from the pixel-level. The gradient-relevant prior serves as a DL-based classifier to distinguish whether the image is compressed from the semantic-level. To effectively solve the variational model, we design an alternating minimization algorithm and unfold it into a deep network architecture. In this way, not only the interpretability of the deep network is increased, but also the dual priors can be well estimated from training samples. By integrating the two priors into a single framework, the image de-blocking problem can be well-constrained, leading to a better performance. Experiments on benchmarks and real-world use cases demonstrate the superiority of our method to the existing state-of-the-art approaches. Xueyang Fu, Xi Wang 0018, Aiping Liu, Junwei Han 0001, Zhengjun Zha |
ICCV | 2 |