Yunke Wang

dblp:165/9106 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
18since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 14 · 7 first-author · 14 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 2 first-author · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Security and privacy · 1
YearPublicationVenuePosition
2026 Rethinking Visual Token Reduction in LVLMs Under Cross-Modal Misalignment
abstract
Large Vision-Language Models (LVLMs) encode visual inputs as dense sequences of patch-level tokens to capture fine-grained semantics. These visual tokens often outnumber their textual counterparts by a large margin, leading to substantial computational overhead and limiting the scalability of LVLMs in practice. Previous efforts have explored visual token reduction either prior to or within the large language models (LLMs). However, most in-LLM reduction approaches rely on text-conditioned interactions, implicitly assuming that textual tokens can reliably capture the importance of visual tokens. In this work, we revisit this assumption and reveal causal, semantic, and spatial forms of cross-modal misalignment. These misalignments undermine the effectiveness of text-guided visual token reduction. To address this, we introduce VisionDrop, a training-free, visual-only pruning framework that selects informative visual tokens based on intra-modal (visual-to-visual) attention, without relying on textual signals. To further suppress redundancy throughout the model hierarchy, we treat the visual encoder and the LLM as a unified system and design a progressive pruning pipeline. Our method performs dominant token selection and lightweight contextual merging at multiple stages, enabling fine-grained visual information to be retained even under aggressive token budgets. Extensive experiments across diverse benchmarks show that VisionDrop achieves consistent improvements over existing approaches, despite requiring no additional training or complex modifications. Notably, when integrated with LLaVA-NeXT-7B, VisionDrop achieves a 2.7x reduction in inference latency and 6x in FLOPs, while retaining 95.71% of the original performance.
Rui Xu 0031, Yunke Wang, Yong Luo 0002, Bo Du 0001
AAAI2
2026 Mitigating Legal Hallucinations via Symbolic Constraints and Analogical Precedents
abstract
With the growing potential of large language models (LLMs) in the legal domain, domain-specific finetuning and retrieval-augmented generation (RAG) methods have received widespread attention. However, current methods still suffer from hallucination risk and failing to resolve semantic drift and adapt to varying citation numbers. To address this, we propose Authoritative and Accurate Lawyer (AALawyer), a complementary dual-retriever framework based on the Legal Syllogism and the nature of different legal data. First, we introduce Symbolic Constrained Retrieval (SCR) for closed-set article retrieval, by constraining retrieval to the generative prediction. Second, we build Analogical Precedent Retrieval (APR) to retrieve open-set judicial precedents for reasoning with a newly collected large criminal dataset.Extensive experiments, including LawBench, our Hallucination Risk-Benchmark, and comprehensive ablation studies, demonstrate the effectiveness of AALawyer, which mitigates hallucinations while improving the explainability of legal reasoning.
Yanxiang Ma, Yunke Wang, Duo Shi, Chang Xu 0002
ACL (1)4
2026 MMFormer: Multi-Modality semi-Supervised vision transformer in remote sensing imagery classification
Daixun Li, Weiying Xie, Leyuan Fang, Yunke Wang, Mingxiang Cao, Jitao Ma, Yunsong Li 0001, Chang Xu 0002
Neural Networks4
2026 Marine Saliency Segmenter: Object-Focused Conditional Diffusion With Region-Level Semantic Knowledge Distillation
abstract
Marine Saliency Segmentation (MSS) plays a pivotal role in a wide range of vision-based marine exploration tasks. However, existing techniques often face the dilemma of imprecise boundaries due to the interference-rich nature of underwater environments, where suspended particles, low contrast, and color distortion hinder accurate segmentation. Although diffusion models have shown impressive performance in visual tasks, their potential to incorporate contextual semantics for enhancing feature learning of region-level salient objects remains underexplored, thereby hindering segmentation outcomes. Building on this insight, we propose DiffMSS, a novel marine saliency segmenter based on the diffusion model, which utilizes semantic knowledge distillation to guide the detection of marine salient objects. Specifically, we design the Word-level Semantic Saliency Extraction module that identifies salient terms at the word level from the captions by computing region-word similarity. These high-level semantic features are distilled into the Conditional Feature Learning Network to generate accurate and semantically informed diffusion conditions. The Object-Focused Conditional Diffusion module then leverages these conditions to iteratively generate fine-grained segmentation masks of marine instances, while a Consensus Deterministic Sampling scheme is further employed to suppress overconfident mis-segmentations and enhance structural fidelity. Extensive experiments demonstrate the superior performance of DiffMSS over state-of-the-art methods in both quantitative and qualitative evaluations. Our code and pre-trained models will be released on GitHub.
Laibin Chang, Yunke Wang, Jiaxing Huang 0001, Longxiang Deng, Bo Du 0001, Chang Xu 0002
IEEE Trans. Image Process.2
2026 Color Correction Meets Cross-Spectral Refinement: A Distribution-Aware Diffusion for Underwater Image Restoration
abstract
Underwater imaging is often plagued by significant degradation in visual quality, primarily due to the effects of light absorption and scattering in water. Although recent underwater image enhancement (UIE) methods rely on the current advances in deep neural network architecture designs, there is still considerable room for improvement in cross-scene robustness and computational efficiency. Diffusion models have shown great success in image generation, prompting us to explore their application to UIE tasks. However, directly applying them to UIE tasks will pose two challenges,i.e., high computational budget and color unbalanced perturbations. To tackle these issues, we propose DiffColor, a distribution-aware diffusion and cross-spectral refinement model for efficient UIE. Unlike single-noise image restoration tasks, underwater imaging exhibits unbalanced channel distributions due to the selective absorption of light by water. To address this, we design the Global Color Correction to balance the diverse color shifts, thereby avoiding potential global degradation disturbances during the denoising process. Instead of diffusing in the raw pixel space, we transform the image into the wavelet domain to obtain such low-frequency and high-frequency spectra. For the sacrificed image details caused by underwater scattering, we further present the Cross-Spectral Detail Refinement to enhance the high-frequency details, which are then integrated with the low-frequency signal as a dual-condition for guiding the diffusion. This strategy ensures the high-fidelity of sampled content and compensates for the sacrificed details. Extensive experiments demonstrate the superior performance of DiffColor over state-of-the-art methods in both quantitative and qualitative evaluations.
Laibin Chang, Yunke Wang, Bo Du 0001, Chang Xu 0002
IEEE Trans. Multim.2
2025 WaterDiffusion: Learning a Prior-involved Unrolling Diffusion for Joint Underwater Saliency Detection and Visual Restoration
abstract
Underwater salient object detection (USOD) plays a pivotal role in various vision-based marine exploration tasks. However, existing USOD techniques face the dilemma of object mislocalization and imprecise boundaries due to the complex underwater environment. The quality degradation of raw underwater images (caused by selective absorption and medium scattering) makes it challenging to perform instance detection directly. One conceivable approach involves initially removing visual disturbances through underwater image enhancement (UIE), followed by saliency detection. However, this two-stage approach neglects the potential positive impact of the restoration procedure on saliency detection due to it executes in a cascade. Based on this insight, we propose a generalized prior-involved diffusion model, called WaterDiffusion for collaborative underwater saliency detection and visual restoration. Specifically, we first propose a revised self-attention joint diffusion, which embeds dynamic saliency masks into the diffusive network as latent features. By extending the underwater degradation prior into the multi-scale decoder, we innovatively exploit optical transmission maps to aid in localizing underwater salient objects. Then, we further design a gate-guided binary indicator to select either normalized or raw channels for improving feature generalization. Finally, the Half-quadratic Splitting is introduced into the unfolding sampling to refine saliency masks iteratively. Comprehensive experiments demonstrate the superior performance of WaterDiffusion over state-of-the-art methods in both quantitative and qualitative evaluations.
Laibin Chang, Yunke Wang, Longxiang Deng, Bo Du 0001, Chang Xu 0002
AAAI2
2025 Masked Diffusion Models for Unsupervised Anomaly Detection in Brain Images
abstract
Unsupervised anomaly detection has gained significant attention in the field of medical imaging due to its capability of reducing the need for costly pixel-level annotation. To achieve this, existing approaches usually utilize generative models to produce healthy references of the diseased images and then identify the abnormalities by comparing healthy references and original diseased images. However, intrinsic characteristics of brain images, e.g., the low contrast and the intricate anatomical structure, make reconstruction challenging. To address those challenges, we propose a Masked Diffusion Model (MDiff), which incorporates a hierarchical patch partition strategy into the diffusion model for precise reconstruction of detailed content. Aligned with this strategy, we treat the perturbed upper-level patch as masked and introduce a masked modeling mechanism into MDiff's diffusion U-Net. This mechanism operates on sub-level patches, enhancing the model's ability to process contextual information surrounding the perturbed upper-level patch. To further improve the quality of healthy references, we integrate a memory module within the mechanism's encoder to retrieve the most relevant memory items as contextual information, while employing the learnable query embedding in its decoder to prevent the network from learning identical shortcuts. Experiments on tumor and multiple sclerosis lesion data demonstrate MDiff's effectiveness.
Rui Xu 0031, Yunke Wang, Yong Luo 0002, Shu Yang 0004, Yihui Wang 0002, Bo Du 0001, Hao Chen 0011
BIBM2
2025 FusionSAM: Visual Multi-Modal Learning with Segment Anything Model
abstract
Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance during training. While the Segment Anything Model (SAM) allows precise control during fine-tuning through its flexible prompting encoder, its potential remains largely unexplored in the context of multimodal segmentation for natural images. In this paper, we introduce SAM into multimodal image segmentation for the first time, proposing a novel framework that combines Latent Space Token Generation (LSTG) and Fusion Mask Prompting (FMP) modules. This approach transforms the training methodology for multimodal segmentation from a traditional black-box approach to a controllable, prompt-based mechanism. Specifically, we obtain latent space features for both modalities through vector quantization and embed them into a cross-attention-based inter-domain fusion module to establish long-range dependencies between modalities. We then use these comprehensive fusion features as prompts to guide precise pixel-level segmentation. Extensive experiments on multiple public datasets demonstrate that our method significantly outperforms SAM and SAM2 in multimodal autonomous driving scenarios, achieving an average improvement of 4.1% over the state-of-the-art method in segmentation mIoU, and the performance is also optimized in other multi-modal visual scenes.
Daixun Li, Weiying Xie, Mingxiang Cao, Yunke Wang, Leyuan Fang, Yunsong Li 0001, Chang Xu 0002
KDD (2)4
2025 VLA-Cache: Efficient Vision-Language-Action Manipulation via Adaptive Token Caching
abstract
Vision-Language-Action (VLA) models have demonstrated strong multi-modal reasoning capabilities, enabling direct action generation from visual perception and language instructions in an end-to-end manner. However, their substantial computational cost poses a challenge for real-time robotic control, where rapid decision-making is essential. This paper introduces VLA-Cache, a training-free inference acceleration method that reduces computational overhead by adaptively caching and reusing static visual tokens across frames. Exploiting the temporal continuity in robotic manipulation, VLA-Cache identifies minimally changed tokens between adjacent frames and reuses their cached key-value representations, thereby circumventing redundant computations. Additionally, to maintain action precision, VLA-Cache selectively re-computes task-relevant tokens that are environmentally sensitive, ensuring the fidelity of critical visual information. To further optimize efficiency, we introduce a layer adaptive token reusing strategy that dynamically adjusts the reuse ratio based on attention concentration across decoder layers, prioritizing critical tokens for recomputation. Extensive experiments on two simulation platforms (LIBERO and SIMPLER) and a real-world robotic system demonstrate that VLA-Cache achieves up to 1.7× speedup in CUDA latency and a 15\% increase in control frequency, with negligible loss on task success rate. The code and videos can be found at our project page: https://vla-cache.github.io.
Siyu Xu 0001, Yunke Wang, Chenghao Xia, Dihao Zhu, Tao Huang 0020, Chang Xu 0002
NeurIPS2
2025 Rectangling and enhancing underwater stitched image via content-aware warping and perception balancing
Laibin Chang, Yunke Wang, Bo Du 0001, Chang Xu 0002
Neural Networks2
2025 On Positive-Unlabeled Classification From Corrupted Data in GANs
abstract
This paper defines a positive and unlabeled classification problem for standard GANs, which then leads to a novel technique to stabilize the training of the discriminator in GANs and deal with corrupted data. Traditionally, real data are taken as positive while generated data are negative. This positive-negative classification criterion was kept fixed all through the learning process of the discriminator without considering the gradually improved quality of generated data, even if they could be more realistic than real data at times. In contrast, it is more reasonable to treat the generated data as unlabeled, which could be positive or negative according to their quality. The discriminator is thus a classifier for this positive and unlabeled classification problem, and we derive a new Positive-Unlabeled GAN (PUGAN). We theoretically discuss the global optimality the proposed model will achieve and the equivalent optimization goal. Empirically, we find that PUGAN can achieve comparable or even better performance than those sophisticated discriminator stabilization methods. Considering the potential corrupted data problem in real-world scenarios, we further extend our approach to PUGAN-C, which treats real data as unlabeled that accounts for both clean and corrupted instances, and generated data as positive. The samples from generator could be closer to those corrupted data within unlabeled data at first, but within the framework of adversarial training, the generator will be optimized to cheat the discriminator and produce samples that are similar to those clean data. Experimental results on image generation from several corrupted datasets demonstrate the effectiveness and generalization of PUGAN-C.
Yunke Wang, Chang Xu 0002, Tianyu Guo 0001, Bo Du 0001, Dacheng Tao
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Imitation Learning from Purified Demonstrations
abstract
Imitation learning has emerged as a promising approach for addressing sequential decision-making problems, with the assumption that expert demonstrations are optimal. However, in real-world scenarios, most demonstrations are often imperfect, leading to challenges in the effectiveness of imitation learning. While existing research has focused on optimizing with imperfect demonstrations, the training typically requires a certain proportion of optimal demonstrations to guarantee performance. To tackle these problems, we propose to purify the potential noises in imperfect demonstrations first, and subsequently conduct imitation learning from these purified demonstrations. Motivated by the success of diffusion model, we introduce a two-step purification via diffusion process. In the first step, we apply a forward diffusion process to smooth potential noises in imperfect demonstrations by introducing additional noise. Subsequently, a reverse generative process is utilized to recover the optimal demonstration from the diffused ones. We provide theoretical evidence supporting our approach, demonstrating that the distance between the purified and optimal demonstration can be bounded. Empirical results on MuJoCo and RoboSuite demonstrate the effectiveness of our method from different aspects.
Yunke Wang, Minjing Dong, Yukun Zhao, Bo Du 0001, Chang Xu 0002
ICML1
2024 Multi-tailed vision transformer for efficient inference
Yunke Wang, Bo Du 0001, Chang Xu 0002
Neural Networks1
2024 A Condition-Based Maintenance Policy Considering Batch Sizes for Warm Standby Systems With Priority to Repair
abstract
In this article, a novel condition-based maintenance (CBM) policy for a two-component warm standby system is proposed. In contrast to the traditional models of warm standby systems, joint optimization of batch sizes and CBM policy is investigated. From the just-in-time view, traditional mass production is divided into different small-batch production, and the small moving batches are applied in the presented model. In the active switching strategy, in addition to the common failure switching, preventive switching of components is also considered. The components' condition in the system can be assessed by periodic inspections, and preventive maintenance (PM) or corrective maintenance is performed once the degradation level of one component exceeds the maintenance threshold. The optimal size of small moving batches and PM threshold are derived by minimizing the long-term average total cost rate. Finally, a numerical example and sensitivity analysis are adopted to demonstrate the effectiveness of the policy.
Faqun Qi, Yunke Wang, Fengping Li
IEEE Trans. Reliab.2
2023 Unlabeled Imperfect Demonstrations in Adversarial Imitation Learning
abstract
Adversarial imitation learning has become a widely used imitation learning framework. The discriminator is often trained by taking expert demonstrations and policy trajectories as examples respectively from two categories (positive vs. negative) and the policy is then expected to produce trajectories that are indistinguishable from the expert demonstrations. But in the real world, the collected expert demonstrations are more likely to be imperfect, where only an unknown fraction of the demonstrations are optimal. Instead of treating imperfect expert demonstrations as absolutely positive or negative, we investigate unlabeled imperfect expert demonstrations as they are. A positive-unlabeled adversarial imitation learning algorithm is developed to dynamically sample expert demonstrations that can well match the trajectories from the constantly optimized agent policy. The trajectories of an initial agent policy could be closer to those non-optimal expert demonstrations, but within the framework of adversarial imitation learning, agent policy will be optimized to cheat the discriminator and produce trajectories that are similar to those optimal expert demonstrations. Theoretical analysis shows that our method learns from the imperfect demonstrations via a self-paced way. Experimental results on MuJoCo and RoboSuite platforms demonstrate the effectiveness of our method from different aspects.
Yunke Wang, Bo Du 0001, Chang Xu 0002
AAAI1
2023 Learning to Schedule in Diffusion Probabilistic Models
abstract
Recently, the field of generative models has seen a significant advancement with the introduction of Diffusion Probabilistic Models (DPMs). The Denoising Diffusion Implicit Model (DDIM) was designed to reduce computational time by skipping a number of steps in the inference process of DPMs. However, the hand-crafted sampling schedule in DDIM, which relies on human expertise, has its limitations in considering all relevant factors in the sampling process. Additionally, the assumption that all instances should have the same schedule is not always valid. To address these problems, this paper proposes a method that leverages reinforcement learning to automatically search for an optimal sampling schedule for DPMs. This is achieved by a policy network that predicts the next step to visit based on the current state of the noisy image. The optimization of the policy network is accomplished using an episodic actor-critic framework, which incorporates reinforcement learning. Empirical results demonstrate the superiority of our approach over various datasets with different timesteps. We also observe that the trained sampling schedule has a strong generalization ability across different DPM baselines.
Yunke Wang, AnhDung Dinh, Bo Du 0001, Chang Xu 0002
KDD1
2021 Learning to Weight Imperfect Demonstrations
abstract
This paper investigates how to weight imperfect expert demonstrations for generative adversarial imitation learning (GAIL). The agent is expected to perform behaviors demonstrated by experts. But in many applications, experts could also make mistakes and their demonstrations would mislead or slow the learning process of the agent. Recently, existing methods for imitation learning from imperfect demonstrations mostly focus on using the preference or confidence scores to distinguish imperfect demonstrations. However, these auxiliary information needs to be collected with the help of an oracle, which is usually hard and expensive to afford in practice. In contrast, this paper proposes a method of learning to weight imperfect demonstrations in GAIL without imposing extensive prior information. We provide a rigorous mathematical analysis, presenting that the weights of demonstrations can be exactly determined by combining the discriminator and agent policy in GAIL. Theoretical analysis suggests that with the estimated weights the agent can learn a better policy beyond those plain expert demonstrations. Experiments in the Mujoco and Atari environments demonstrate that the proposed algorithm outperforms baseline methods in handling imperfect expert demonstrations.
Yunke Wang, Chang Xu 0002, Bo Du 0001, Honglak Lee
ICML1
2021 Robust Adversarial Imitation Learning via Adaptively-Selected Demonstrations
abstract
The agent in imitation learning (IL) is expected to mimic the behavior of the expert. Its performance relies highly on the quality of given expert demonstrations. However, the assumption that collected demonstrations are optimal cannot always hold in real-world tasks, which would seriously influence the performance of the learned agent. In this paper, we propose a robust method within the framework of Generative Adversarial Imitation Learning (GAIL) to address imperfect demonstration issue, in which good demonstrations can be adaptively selected for training while bad demonstrations are abandoned. Specifically, a binary weight is assigned to each expert demonstration to indicate whether to select it for training. The reward function in GAIL is employed to determine this weight (i.e. higher reward results in higher weight). Compared to some existing solutions that require some auxiliary information about this weight, we set up the connection between weight and model so that we can jointly optimize GAIL and learn the latent weight. Besides hard binary weighting, we also propose a soft weighting scheme. Experiments in the Mujoco demonstrate the proposed method outperforms other GAIL-based methods when dealing with imperfect demonstrations.
Yunke Wang, Chang Xu 0002, Bo Du 0001
IJCAI1
2015 A statistical study of covert timing channels using network packet frequency
abstract
This paper first reviews covert timing channels with network packet frequencies as information carriers. Then, based on the study of communication and statistical models, it proposes a method to detect an enhanced covert timing channel and its use of carrier frequencies. With the help of MATLAB for simulation, several experiments have been conducted for the verification of the proposed method.
Fangyue Chen, Yunke Wang, Heng Song
ISI2