Yuji Wang

dblp:24/9868 · DBLP profile ↗
← Back
18ranked-venue papers
8as first author
16since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorSecurity and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2026 SH-ETRs:Learning Soft-Hard rules with entity type constraints for document-level relation extraction
Haisong Chen, Nisuo Du, Qing He 0007, Yuji Wang
Expert Syst. Appl.4
2026 Multimodal summarization via coarse-and-fine granularity synergy and region counterfactual reasoning filter
Rulong Liu, Qing He 0007, Yuji Wang, Nisuo Du
Knowl. Based Syst.3
2026 Armor: Shielding Unlearnable Examples Against Data Augmentation
abstract
Private data, when published online, may be collected by unauthorized parties to train deep neural networks (DNNs). To protect privacy, defensive noises can be added to original samples to degrade their learnability by DNNs. Recently, unlearnable examples (Huang et al., 2021) are proposed to minimize the training loss such that the model learns almost nothing. However, raw data are often pre-processed before being used for training, which may restore the private information of protected data. In this paper, we reveal the data privacy violation induced by data augmentation, a commonly used data pre-processing technique to improve model generalization capability, which is the first of its kind as far as we are concerned. We demonstrate that data augmentation can significantly raise the accuracy of the model trained on unlearnable examples from 21.3% to 66.1%. To address this issue, we propose a defense framework, dubbed Armor, to protect data privacy from potential breaches of data augmentation. To overcome the difficulty of having no access to the model training process, we design a non-local module-assisted surrogate model that better captures the effect of data augmentation. In addition, we design a surrogate augmentation selection strategy that maximizes distribution alignment between augmented and non-augmented samples, to choose the optimal augmentation strategy for each class. We also use a dynamic step size adjustment algorithm to enhance the defensive noise generation process. Extensive experiments are conducted on 4 datasets and 5 data augmentation methods to verify the performance of Armor. Comparisons with 6 state-of-the-art defense methods have demonstrated that Armor can preserve the unlearnability of protected private data under data augmentation. Armor reduces the test accuracy of the model trained on augmented protected samples by as much as 60% more than baselines. We also show that Armor is robust to adversarial training. We will open-source our codes upon publication.
Xueluan Gong, Yuji Wang, Yanjiao Chen, Haocheng Dong, Yiming Li 0004, Mengyuan Sun 0001, Shuaike Li, Qian Wang 0002
IEEE Trans. Pattern Anal. Mach. Intell.2
2026 Sleight: Hidden Data Privacy Breaches in Federated Learning
abstract
Federated Learning (FL) has emerged as a paradigm for conducting machine learning across broad and decentralized datasets, promising enhanced privacy by obviating the need for direct data sharing. However, recent studies show that attackers can steal private data through model manipulation or gradient analysis. Existing attacks are constrained by low theft quantity or low-resolution data, and they are often easily detected through anomaly monitoring in gradients or weights. In this paper, we propose Sleight, a novel data-reconstruction attack, supported by two key techniques, i.e., distinctive and sparse encoding design and block partitioning. Unlike conventional methods that require detectable changes to the model, Sleight stealthily embeds a hidden model using parameter sharing to systematically extract sensitive data. The Fibonacci-based index design ensures efficient, structured retrieval of memorized data, while the block partitioning method enhances Sleight's capability to handle high-resolution images by dividing them into smaller, manageable units. Extensive experiments on 4 datasets confirmed that Sleight is superior to 5 state-of-the-art data-reconstruction attacks under 5 respective detection methods. Sleight can handle large-scale and high-resolution data without being detected or mitigated by state-of-the-art data reconstruction defense methods. In contrast to baselines, Sleight can be directly applied to both FedAvg and FedSGD scenarios, underscoring the need for developers to devise new defenses against such vulnerabilities. We will open-source our code upon acceptance.
Xueluan Gong, Yuji Wang, Shuike Li, Mengyuan Sun 0001, Chen Chen 0115, Qian Wang 0002, Kwok-Yan Lam
IEEE Trans. Dependable Secur. Comput.2
2025 IteRPrimE: Zero-shot Referring Image Segmentation with Iterative Grad-CAM Refinement and Primary Word Emphasis
abstract
Zero-shot Referring Image Segmentation (RIS) identifies the instance mask that best aligns with a specified referring expression without training and fine-tuning, significantly reducing the labor-intensive annotation process. Despite achieving commendable results, previous CLIP-based models have a critical drawback: the models exhibit a notable reduction in their capacity to discern relative spatial relationships of objects. This is because they generate all possible masks on an image and evaluate each masked region for similarity to the given expression, often resulting in decreased sensitivity to direct positional clues in text inputs. Moreover, most methods have weak abilities to manage relationships between primary words and their contexts, causing confusion and reduced accuracy in identifying the correct target region. To address these challenges, we propose IteRPrimE (Iterative Grad-CAM Refinement and Primary word Emphasis), which leverages a saliency heatmap through Grad-CAM from a Vision-Language Pre-trained (VLP) model for image-text matching. An iterative Grad-CAM refinement strategy is introduced to progressively enhance the model's focus on the target region and overcome positional insensitivity, creating a self-correcting effect. Additionally, we design the Primary Word Emphasis module to help the model handle complex semantic relations, enhancing its ability to attend to the intended object. Extensive experiments conducted on the RefCOCO/+/g, and PhraseCut benchmarks demonstrate that IteRPrimE outperforms previous SOTA zero-shot methods, particularly excelling in out-of-domain scenarios.
Yuji Wang, Jingchen Ni, Yong Liu 0033, Chun Yuan 0003, Yansong Tang
AAAI1
2025 SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes
abstract
Reference Audio-Visual Segmentation (Ref-AVS) aims to provide a pixel-wise scene understanding in Language-aided Audio-Visual Scenes (LAVS). This task requires the model to continuously segment objects referred to by text and audio from a video. Previous dual-modality methods always fail due to the lack of a third modality and the existing triple-modality method struggles with spatio-temporal consistency, leading to the target shift of different frames. In this work, we introduce a novel framework, termed SAM2-LOVE, which integrates textual, audio, and visual representations into a learnable token to prompt and align SAM2 for achieving Ref-AVS in the LAVS. Technically, our approach includes a multimodal fusion module aimed at improving multimodal understanding of SAM2, as well as token propagation and accumulation strategies designed to enhance spatio-temporal consistency without forgetting historical information. We conducted extensive experiments to demonstrate that SAM2-LOVE outperforms the SOTA by 8.5% in $\mathcal{J}\& \mathcal{F}$ on the Ref-AVS benchmark and showcase the simplicity and effectiveness of the components. Our code will be available here.
Yuji Wang, Yong Liu 0033, Yansong Tang
CVPR1
2025 FrameBridge: Improving Image-to-Video Generation with Bridge Models
abstract
Diffusion models have achieved remarkable progress on image-to-video (I2V) generation, while their noise-to-data generation process is inherently mismatched with this task, which may lead to suboptimal synthesis quality. In this work, we present FrameBridge. By modeling the frame-to-frames generation process with a bridge model based data-to-data generative process, we are able to fully exploit the information contained in the given image and improve the consistency between the generation process and I2V task. Moreover, we propose two novel techniques toward the two popular settings of training I2V models, respectively. Firstly, we propose SNR-Aligned Fine-tuning (SAF), making the first attempt to fine-tune a diffusion model to a bridge model and, therefore, allowing us to utilize the pre-trained diffusion-based text-to-video (T2V) models. Secondly, we propose neural prior, further improving the synthesis quality of FrameBridge when training from scratch. Experiments conducted on WebVid-2M and UCF-101 demonstrate the superior quality of FrameBridge in comparison with the diffusion counterpart (zero-shot FVD 95 vs. 192 on MSR-VTT and non-zero-shot FVD 122 vs. 171 on UCF-101), and the advantages of our proposed SAF and neural prior for bridge-based I2V models. The project page: https://framebridge-icml.github.io/
Yuji Wang, Zehua Chen 0005, Yixiang Wei, Jun Zhu 0001, Jianfei Chen 0001
ICML1
2025 Identity-Preserving Text-to-Video Generation Guided by Simple yet Effective Spatial-Temporal Decoupled Representations
abstract
Identity-preserving text-to-video (IPT2V) generation, which aims to create high-fidelity videos with consistent human identity, has become crucial for downstream applications. However, current end-to-end frameworks suffer a critical spatial-temporal trade-off: optimizing for spatially coherent layouts of key elements ( e.g., character identity preservation) often compromises instruction-compliant temporal smoothness, while prioritizing dynamic realism risks disrupting the spatial coherence of visual structures. To tackle this issue, we propose a simple yet effective spatial-temporal decoupled framework that decomposes representations into spatial features for layouts and temporal features for motion dynamics. Specifically, our paper proposes a semantic prompt optimization mechanism and stage-wise decoupled generation paradigm. The former module decouples the prompt into spatial and temporal components. Aligned with the subsequent stage-wise decoupled approach, the spatial prompts guide the text-to-image (T2I) stage to generate coherent spatial features, while the temporal prompts direct the sequential image-to-video (I2V) stage to ensure motion consistency. Experimental results validate that our approach achieves excellent spatiotemporal consistency, demonstrating outstanding performance in identity preservation, text relevance, and video quality. By leveraging this simple yet robust mechanism, our algorithm secures the runner-up position in 2025 ACM Multimedia Challenge. Our code is available at https://github.com/rain152/IPVG.
Yuji Wang, Moran Li, Xiaobin Hu, Ran Yi 0002, Jiangning Zhang, Weijian Cao, Yabiao Wang, Chengjie Wang 0001, Lizhuang Ma
ACM Multimedia1
2025 Dual-Branch Mamba Based Multi-label Anuran Species Classification
Yuji Wang, Longhui Zhao, Juan Gabriel Colonna, Zujie Kang, Faming Zhang, Jie Xie 0001
PRICAI1
2025 GS-3Det: Elevating 3D Gaussian Splatting for Real-Time Multi-View 3D Object Detection
abstract
3D Gaussian Splatting (3DGS) has emerged as a high-quality and efficient alternative to Neural Radiance Fields (NeRF), offering distinct advantages in scene representation and suitability for multi-view 3D object detection (MV-3DOD) tasks. However, conventional 3DGS-based detection methods typically employ a reconstruction-then-detection pipeline, which is time-consuming and unsuitable for real-time applications. This approach arises from 3DGS’s original design for reconstruction tasks, which lacks network-based training. In this paper, we propose GS-3Det, an online Gaussian detection framework for MV-3DOD, achieving real-time performance through a single forward pass for scene reconstruction and detection. Specifically, we introduce a detection-aware Gaussian grid that enables directly prediction of explicit 3D Gaussians from multi-view images, enabling efficient and robust 3D scene understanding. Additionally, we propose a Dual-Path Consistency module that leverages 3D constraints to improve the accuracy of Gaussian grid representation and detection. Experiments on the ScanNet V2 dataset demonstrate that GS-3Det surpasses state-of-the-art methods by 3.5% in [email protected] and 2.7% in [email protected], underscoring its generalization capability and real-time performance.
Yifei Han, Jixuan Fan, Sule Bai, Yuji Wang, Yansong Tang
VCIP4
2025 Aspect-based sentiment analysis with semantic and syntactic enhanced multi-layer fusion model
Qing He 0007, Yuji Wang, Nisuo Du, Wenjing Lei
Eng. Appl. Artif. Intell.3
2024 Robust Multimodal Learning via Representation Decoupling
Shicai Wei, Yang Luo 0001, Yuji Wang, Chunbo Luo
ECCV (42)3
2024 Convolution Meets Transformer: Efficient Hybrid Transformer for Semantic Segmentation with Very High Resolution Imagery
abstract
In this paper, we introduce an efficient and lightweight hybrid Transformer architecture, ingeniously integrating convolutions within Transformer blocks for semantic segmentation of remote sensing Very High Resolution (VHR) imagery. To simultaneously avoid the high computational complexity in the shallow layers and capture the local representations of the VHR images, we propose the Group-Team Convolution Modulation (GTCM) module that uses convolutions to approximate the effect of attention mechanisms and modulates features in channel dimension. Additionally, to enlarge the effective receptive field (ERF) in the decoder, based on the grouping philosophy, we adopt dilated convolutions with multiple dilated rates to further enhance the performance. The superiority and efficiency of our proposed hybrid structure are demonstrated by outperforming state-of-the-art methods on the Vaihingen and Potsdam datasets with relatively lower complexity and fewer parameters.
Yuji Wang, Ruojun Zhao, Shicai Wei, Jingchen Ni, Yang Luo 0001, Chunbo Luo
IGARSS1
2024 Fight Back Against Jailbreaking via Prompt Adversarial Tuning
abstract
While Large Language Models (LLMs) have achieved tremendous success in various applications, they are also susceptible to jailbreaking attacks. Several primary defense strategies have been proposed to protect LLMs from producing harmful information, mostly focusing on model fine-tuning or heuristical defense designs. However, how to achieve intrinsic robustness through prompt optimization remains an open problem. In this paper, motivated by adversarial training paradigms for achieving reliable robustness, we propose an approach named **Prompt Adversarial Tuning (PAT)** that trains a prompt control attached to the user prompt as a guard prefix. To achieve our defense goal whilst maintaining natural performance, we optimize the control prompt with both adversarial and benign prompts. Comprehensive experiments show that our method is effective against both grey-box and black-box attacks, reducing the success rate of advanced attacks to nearly 0, while maintaining the model's utility on the benign task and incurring only negligible computational overhead, charting a new perspective for future explorations in LLM security. Our code is available at https://github.com/PKU-ML/PAT.
Yichuan Mo, Yuji Wang, Zeming Wei, Yisen Wang 0001
NeurIPS2
2023 FEditNet: Few-Shot Editing of Latent Semantics in GAN Spaces
abstract
Generative Adversarial networks (GANs) have demonstrated their powerful capability of synthesizing high-resolution images, and great efforts have been made to interpret the semantics in the latent spaces of GANs. However, existing works still have the following limitations: (1) the majority of works rely on either pretrained attribute predictors or large-scale labeled datasets, which are difficult to collect in most cases, and (2) some other methods are only suitable for restricted cases, such as focusing on interpretation of human facial images using prior facial semantics. In this paper, we propose a GAN-based method called FEditNet, aiming to discover latent semantics using very few labeled data without any pretrained predictors or prior knowledge. Specifically, we reuse the knowledge from the pretrained GANs, and by doing so, avoid overfitting during the few-shot training of FEditNet. Moreover, our layer-wise objectives which take content consistency into account also ensure the disentanglement between attributes. Qualitative and quantitative results demonstrate that our method outperforms the state-of-the-art methods on various datasets. The code is available at https://github.com/THU-LYJ-Lab/FEditNet.
Mengfei Xia, Yezhi Shu, Yuji Wang, Yukun Lai, Qiang Li 0024, Pengfei Wan 0001, Zhongyuan Wang 0006, Yong-Jin Liu 0001
AAAI3
2023 Efficient Remote Sensing Transformer for Coastline Detection with Sentinel-2 Satellite Imagery
abstract
This paper proposes an efficient and lightweight transformer architecture to segment land and water and detect coastlines using Sentinel-2 satellite multispectral images. To address the quadratic complexity caused by long patch embedding sequence length, we propose a novel efficient self-attention that utilizes the associative properties of matrix multiplication. A feature fusion module (FFM) without additional parameters is introduced to enhance the decoder’s ability to reconstruct images. The proposed method detects coastline accurately and outperforms the baseline approach with 24 times lower complexities and 3 times fewer parameters. The high efficiency makes it suitable for resource-constrained devices.
Yuji Wang, Ruojun Zhao, Zijun Sun
IGARSS1
2016 Neural adaptive control of hypersonic aircraft with actuator fault using randomly assigned nodes
Wenxing Fu, Yuji Wang, Supeng Zhu, Yingzhou Xia
Neurocomputing2
2010 Performance improvement of force feedback in bilateral teleoperation with PD controller
abstract
This paper deals with the performance improvement of force feedback in bilateral teleoperation with PD controller. In traditional PD structures, the force feedback is simply determined by the position and velocity of the master and the slave manipulators, which may induce large resistance forces to the operator even in free motion. In this paper, a novel PD bilateral controller is proposed to tackle this problem. By incorporating a distance variable in the controller, we show that the appropriate force feedback can be obtained which still guarantees the system stability. To validate the proposed algorithm, an experiment is also carried out on our single degree of freedom teleoperation system. The results indicate that this strategy is effective for safe teleoperation missions.
Yuji Wang, Fuchun Sun 0001, Huaping Liu 0001, Haibo Min
IROS1