VLDB 2026 Research / reviewers in the wild / expert
Jinwu Hu
dblp:65/6584
· DBLP profile ↗
16ranked-venue papers
4as first author
16since 2021 · last 2026
0000-0002-6025-8578ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Mask-Guided Proxy Mining Network for Few-Shot Medical Image SegmentationabstractFew-shot medical image segmentation (FSMIS) has attracted increasing attention as a promising technique for solving medical image segmentation tasks by relying on only a small amount of labeled data from new classes. Current FSMIS methods typically employ pixel-level semantic correlations between support-query image pairs to guide the segmentation of query images. However, the class information gap between support and query images may induce severe mismatches, leading to semantic ambiguity between foreground and background pixels. To address this issue, we propose a novel mask-guided proxy mining network (MPMNet), which mines a set of representative reference features (termed proxies) from support and query images to rectify foreground-background ambiguity. Specifically, to eliminate false pairwise matches caused by excessive intra-class variations, we design a mask-guided proxy mining module to adaptively learn representative proxies that can perceive visual differences between objects with different scales and shapes. Moreover, we integrate a hierarchical prior generation module and a context-aware feature enrichment module into MPMNet to obtain multi-scale information and enhance the discriminability of features. With these well-designed components and structures, our MPMNet can effectively overcome the adverse effects of false pixel matches by establishing proxy-level semantic correlations. Extensive experiments on three standard medical segmentation benchmarks demonstrate that our MPMNet significantly outperforms previous state-of-the-art methods, with a mean gain of 2.71% in DSC across all datasets. The code is available at: https://github.com/donglongzi/MPMNet. Wendong Huang, Jinwu Hu, Yongchao Wang 0004, Xiuli Bi, Yucheng Shu, Xuezong Yang, Bin Xiao 0002 |
IEEE Trans. Image Process. | 2 |
| 2026 | Dynamic Prompt Compression for Efficient Inference of Large Language ModelsabstractLarge language models (LLMs) have shown outstanding performance across a variety of tasks, partly due to advanced prompting techniques. However, these techniques often require lengthy prompts, which increase computational costs and can hinder performance because of the limited context windows of LLMs. While prompt compression is a straightforward solution, existing methods confront the challenges of retaining essential information, adapting to context changes, and remaining effective across different tasks. To tackle these issues, we propose a task-agnostic method called Dynamic Prompt Compression (LLM-DPC). Our method reduces the number of prompt tokens while minimizing any degradation in LLM performance. We model prompt compression as a Markov Decision Process (MDP), enabling the DPC-Agent to sequentially remove redundant tokens by adapting to dynamic contexts and retaining crucial content. We develop a reward function for training the DPC-Agent that balances the compression ratio, the quality of the LLM output, and the retention of key information. This allows for prompt token reduction without needing an external black-box LLM. Inspired by the progressive difficulty adjustment in curriculum learning, we introduce a Hierarchical Prompt Compression (HPC) training strategy that gradually increases the compression difficulty, enabling the DPC-Agent to learn an effective compression method that maintains information integrity. Experiments demonstrate that our method outperforms state-of-the-art techniques, especially at higher compression ratio. Jinwu Hu, Wei Zhang 0098, Yufeng Wang 0004, Yu Hu 0004, Bin Xiao 0002, Mingkui Tan |
IEEE Trans. Knowl. Data Eng. | 1 |
| 2025 | Harnessing Frequency Spectrum Insights for Image Copyright Protection Against Diffusion ModelsabstractDiffusion models have achieved remarkable success in novel view synthesis, but their reliance on large, diverse, and often untraceable Web datasets has raised pressing concerns about image copyright protection. Current methods fall short in reliably identifying unauthorized image use, as they struggle to generalize across varied generation tasks and fail when the training dataset includes images from multiple sources with few identifiable (watermarked or poisoned) samples. In this paper, we present novel evidence that diffusion-generated images faithfully preserve the statistical properties of their training data, particularly reflected in their spectral features. Leveraging this insight, we introduce CoprGuard, a robust frequency domain watermarking framework to safeguard against unauthorized image usage in diffusion model training and fine-tuning. CoprGuard demonstrates remarkable effectiveness against a wide range of models, from naive diffusion models to sophisticated text-to-image models, and is robust even when watermarked images comprise a mere 1% of the training dataset. This robust and versatile approach empowers content owners to protect their intellectual property in the era of AI-driven image generation. Zhenguang Liu, Chao Shuai, Shaojing Fan, Ziping Dong, Jinwu Hu, Zhongjie Ba, Kui Ren 0001 |
CVPR | 5 |
| 2025 | Test-Time Learning for Large Language ModelsabstractWhile Large Language Models (LLMs) have exhibited remarkable emergent capabilities through extensive pre-training, they still face critical limitations in generalizing to specialized domains and handling diverse linguistic variations, known as distribution shifts. In this paper, we propose a Test-Time Learning (TTL) paradigm for LLMs, namely TLM, which dynamically adapts LLMs to target domains using only unlabeled test data during testing. Specifically, we first provide empirical evidence and theoretical insights to reveal that more accurate predictions from LLMs can be achieved by minimizing the input perplexity of the unlabeled test data. Based on this insight, we formulate the Test-Time Learning process of LLMs as input perplexity minimization, enabling self-supervised enhancement of LLM performance. Furthermore, we observe that high-perplexity samples tend to be more informative for model optimization. Accordingly, we introduce a Sample Efficient Learning Strategy that actively selects and emphasizes these high-perplexity samples for test-time updates. Lastly, to mitigate catastrophic forgetting and ensure adaptation stability, we adopt Low-Rank Adaptation (LoRA) instead of full-parameter optimization, which allows lightweight model updates while preserving more original knowledge from the model. We introduce the AdaptEval benchmark for TTL and demonstrate through experiments that TLM improves performance by at least 20% compared to original LLMs on domain knowledge adaptation. Jinwu Hu, Zitian Zhang, Xutao Wen, Chao Shuai, Wei Luo 0006, Bin Xiao 0002, Yuanqing Li 0001, Mingkui Tan |
ICML | 1 |
| 2025 | Enhancing User-Oriented Proactivity in Open-Domain Dialogues with Critic GuidanceabstractOpen-domain dialogue systems aim to generate natural and engaging conversations, providing significant practical value in real applications such as social robotics and personal assistants. The advent of large language models (LLMs) has greatly advanced this field by improving context understanding and conversational fluency. However, existing LLM-based dialogue systems often fall short in proactively understanding the user's chatting preferences and guiding conversations toward user-centered topics. This lack of user-oriented proactivity can lead users to feel unappreciated, reducing their satisfaction and willingness to continue the conversation in human-computer interactions. To address this issue, we propose a User-oriented Proactive Chatbot (UPC) to enhance the user-oriented proactivity. Specifically, we first construct a critic to evaluate this proactivity inspired by the LLM-as-a-judge strategy. Given the scarcity of high-quality training data, we then employ the critic to guide dialogues between the chatbot and user agents, generating a corpus with enhanced user-oriented proactivity. To ensure the diversity of the user backgrounds, we introduce the ISCO-800, a diverse user background dataset for constructing user agents. Moreover, considering the communication difficulty varies among users, we propose an iterative curriculum learning method that trains the chatbot from easy-to-communicate users to more challenging ones, thereby gradually enhancing its performance. Experiments demonstrate that our proposed training method is applicable to different LLMs, improving user-oriented proactivity and attractiveness in open-domain dialogues. Code and appendix are available at github.com/wang678/LLM-UPC. Yufeng Wang 0004, Jinwu Hu, Ziteng Huang, Kunyang Lin, Zitian Zhang, Peihao Chen, Yu Hu 0004, Qianyue Wang, Zhu Liang Yu, Bin Sun 0001, Xiaofen Xing, Mingkui Tan |
IJCAI | 2 |
| 2025 | Efficient Dynamic Ensembling for Multiple LLM ExpertsabstractLLMs have demonstrated impressive performance across various language tasks. However, the strengths of LLMs can vary due to different architectures, model sizes, areas of training data, etc. Therefore, ensemble reasoning for the strengths of different LLM experts is critical to achieving consistent and satisfactory performance on diverse inputs across a wide range of tasks. However, existing LLM ensemble methods are either computationally intensive or incapable of leveraging complementary knowledge among LLM experts for various inputs. In this paper, we propose an efficient Dynamic Ensemble Reasoning paradigm, called DER to integrate the strengths of multiple LLM experts conditioned on dynamic inputs. Specifically, we model the LLM ensemble reasoning problem as a Markov Decision Process, wherein an agent sequentially takes inputs to request knowledge from an LLM candidate and passes the output to a subsequent LLM candidate. Moreover, we devise a reward function to train a DER-Agent to dynamically select an optimal answering route given the input questions, aiming to achieve the highest performance with as few computational resources as possible. Last, to fully transfer the expert knowledge from the prior LLMs, we develop a Knowledge Transfer Prompt that enables the subsequent LLM candidates to transfer complementary knowledge effectively. Experiments demonstrate that our method uses fewer computational resources to achieve better performance compared to state-of-the-art baselines. Code and appendix are available at https://github.com/Fhujinwu/DER. Jinwu Hu, Yufeng Wang 0004, Shuhai Zhang, Yu Hu 0004, Bin Xiao 0004, Mingkui Tan |
IJCAI | 1 |
| 2025 | Generating Long-form Story Using Dynamic Hierarchical Outlining with Memory-EnhancementabstractQianyue Wang, Jinwu Hu, Zhengping Li, Yufeng Wang, Daiyuan Li, Yu Hu, Mingkui Tan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Qianyue Wang, Jinwu Hu, Zhengping Li, Yufeng Wang 0004, Daiyuan Li, Yu Hu 0004, Mingkui Tan |
NAACL (Long Papers) | 2 |
| 2025 | Continual Knowledge Adaptation for Reinforcement LearningabstractReinforcement Learning enables agents to learn optimal behaviors through interactions with environments. However, real-world environments are typically non-stationary, requiring agents to continuously adapt to new tasks and changing conditions. Although Continual Reinforcement Learning facilitates learning across multiple tasks, existing methods often suffer from catastrophic forgetting and inefficient knowledge utilization. To address these challenges, we propose Continual Knowledge Adaptation for Reinforcement Learning (CKA-RL), which enables the accumulation and effective utilization of historical knowledge. Specifically, we introduce a Continual Knowledge Adaptation strategy, which involves maintaining a task-specific knowledge vector pool and dynamically using historical knowledge to adapt the agent to new tasks. This process mitigates catastrophic forgetting and enables efficient knowledge transfer across tasks by preserving and adapting critical model parameters. Additionally, we propose an Adaptive Knowledge Merging mechanism that combines similar knowledge vectors to address scalability challenges, reducing memory requirements while ensuring the retention of essential knowledge. Experiments on three benchmarks demonstrate that the proposed CKA-RL outperforms state-of-the-art methods, achieving an improvement of 4.20% in overall performance and 8.02% in forward transfer. The source code is available at https://github.com/Fhujinwu/CKA-RL. Jinwu Hu, Zihao Lian, Zhiquan Wen, Xutao Wen, Bin Xiao 0002, Mingkui Tan |
NeurIPS | 1 |
| 2025 | Open-World Drone Active Tracking with Goal-Centered RewardsabstractDrone Visual Active Tracking aims to autonomously follow a target object by controlling the motion system based on visual observations, providing a more practical solution for effective tracking in dynamic environments. However, accurate Drone Visual Active Tracking using reinforcement learning remains challenging due to the absence of a unified benchmark and the complexity of open-world environments with frequent interference. To address these issues, we pioneer a systematic solution. First, we propose DAT, the first open-world drone active air-to-ground tracking benchmark. It encompasses 24 city-scale scenes, featuring targets with human-like behaviors and high-fidelity dynamics simulation. DAT also provides a digital twin tool for unlimited scene generation. Additionally, we propose a novel reinforcement learning method called GC-VAT, which aims to improve the performance of drone tracking targets in complex scenarios. Specifically, we design a Goal-Centered Reward to provide precise feedback across viewpoints to the agent, enabling it to expand perception and movement range through unrestricted perspectives. Inspired by curriculum learning, we introduce a Curriculum-Based Training strategy that progressively enhances the tracking performance in complex environments. Besides, experiments on simulator and real-world images demonstrate the superior performance of GC-VAT, achieving a Tracking Success Rate of approximately 72% on the simulator. The benchmark and code are available at https://github.com/SHWplus/DAT_Benchmark. Haowei Sun, Jinwu Hu, Zhirui Zhang, Haoyuan Tian, Xinze Xie, Yufeng Wang 0004, Xiaohua Xie, Zhu Liang Yu, Mingkui Tan |
NeurIPS | 2 |
| 2025 | Fine-Grained Textual Guidance for Generalized Multi-Modal Face Anti-Spoofing
Daiyuan Li, Zitong Yu, Jinwu Hu, Jinghui Zeng, Mingkui Tan |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2025 | Zero-Shot Skeleton-Based Action Recognition With Prototype-Guided Feature AlignmentabstractZero-shot skeleton-based action recognition aims to classify unseen skeleton-based human actions without prior exposure to such categories during training. This task is extremely challenging due to the difficulty in generalizing from known to unknown actions. Previous studies typically use two-stage training: pre-training skeleton encoders on seen action categories using cross-entropy loss and then aligning pre-extracted skeleton and text features, enabling knowledge transfer to unseen classes through skeleton-text alignment and language models' generalization. However, their efficacy is hindered by 1) insufficient discrimination for skeleton features, as the fixed skeleton encoder fails to capture necessary alignment information for effective skeleton-text alignment; 2) the neglect of alignment bias between skeleton and unseen text features during testing. To this end, we propose a prototype-guided feature alignment paradigm for zero-shot skeleton-based action recognition, termed PGFA. Specifically, we develop an end-to-end cross-modal contrastive training framework to improve skeleton-text alignment, ensuring sufficient discrimination for skeleton features. Additionally, we introduce a prototype-guided text feature alignment strategy to mitigate the adverse impact of the distribution discrepancy during testing. We provide a theoretical analysis to support our prototype-guided text feature alignment strategy and empirically evaluate our overall PGFA on three well-known datasets. Compared with the top competitor SMIE method, our PGFA achieves absolute accuracy improvements of 22.96%, 12.53%, and 18.54% on the NTU-60, NTU-120, and PKU-MMD datasets, respectively. Shuhai Zhang, Zeng You, Jinwu Hu, Mingkui Tan, Fei Liu 0006 |
IEEE Trans. Image Process. | 4 |
| 2025 | Prototype-Guided Graph Reasoning Network for Few-Shot Medical Image SegmentationabstractFew-shot semantic segmentation (FSS) is of tremendous potential for data-scarce scenarios, particularly in medical segmentation tasks with merely a few labeled data. Most of the existing FSS methods typically distinguish query objects with the guidance of support prototypes. However, the variances in appearance and scale between support and query objects from the same anatomical class are often exceedingly considerable in practical clinical scenarios, thus resulting in undesirable query segmentation masks. To tackle the aforementioned challenge, we propose a novel prototype-guided graph reasoning network (PGRNet) to explicitly explore potential contextual relationships in structured query images. Specifically, a prototype-guided graph reasoning module is proposed to perform information interaction on the query graph under the guidance of support prototypes to fully exploit the structural properties of query images to overcome intra-class variances. Moreover, instead of fixed support prototypes, a dynamic prototype generation mechanism is devised to yield a collection of dynamic support prototypes by mining rich contextual information from support images to further boost the efficiency of information interaction between support and query branches. Equipped with the proposed two components, PGRNet can learn abundant contextual representations for query images and is therefore more resilient to object variations. We validate our method on three publicly available medical segmentation datasets, namely CHAOS-T2, MS-CMRSeg, and Synapse. Experiments indicate that the proposed PGRNet outperforms previous FSS methods by a considerable margin and establishes a new state-of-the-art performance. Wendong Huang, Jinwu Hu, Yang Wei 0002, Xiuli Bi, Bin Xiao 0002 |
IEEE Trans. Medical Imaging | 2 |
| 2024 | Anatomical Prior Guided Spatial Contrastive Learning for Few-Shot Medical Image Segmentation
Wendong Huang, Jinwu Hu, Xiuli Bi, Bin Xiao 0002 |
ACM Multimedia | 2 |
| 2024 | CTNet: Contrastive Transformer Network for Polyp SegmentationabstractSegmenting polyps from colonoscopy images is very important in clinical practice since it provides valuable information for colorectal cancer. However, polyp segmentation remains a challenging task as polyps have camouflage properties and vary greatly in size. Although many polyp segmentation methods have been recently proposed and produced remarkable results, most of them cannot yield stable results due to the lack of features with distinguishing properties and those with high-level semantic details. Therefore, we proposed a novel polyp segmentation framework called contrastive Transformer network (CTNet), with three key components of contrastive Transformer backbone, self-multiscale interaction module (SMIM), and collection information module (CIM), which has excellent learning and generalization abilities. The long-range dependence and highly structured feature map space obtained by CTNet through contrastive Transformer can effectively localize polyps with camouflage properties. CTNet benefits from the multiscale information and high-resolution feature maps with high-level semantic obtained by SMIM and CIM, respectively, and thus can obtain accurate segmentation results for polyps of different sizes. Without bells and whistles, CTNet yields significant gains of 2.3%, 3.7%, 3.7%, 18.2%, and 10.1% over classical method PraNet on Kvasir-SEG, CVC-ClinicDB, Endoscene, ETIS-LaribPolypDB, and CVC-ColonDB respectively. In addition, CTNet has advantages in camouflaged object detection and defect detection. The code is available at https://github.com/Fhujinwu/CTNet. Bin Xiao 0002, Jinwu Hu, Weisheng Li 0001, Chi-Man Pun, Xiuli Bi |
IEEE Trans. Cybern. | 2 |
| 2023 | Location-Aware Transformer Network for Few-Shot Medical Image SegmentationabstractAutomatic and precise organ segmentation plays a significant role in promoting the development of the diagnosis and treatment of the disease. Despite making enormous strides in medical image segmentation, conventional deep neural network-based methods are inherently massive data-driven techniques and are challenging to adapt to novel classes with a small number of labeled samples. Few-shot learning is a promising solution through learning novel classes from extremely limited annotated examples. However, existing few-shot segmentation methods focus excessively on targets in individual images while neglecting to model the global spatial correlation across images, which may cause severe performance degradation. To solve this issue, we propose a new Transformer-based few-shot segmentation framework for medical imaging, namely location-aware transformer network (LATNet), which establishes the spatial correlation between support and query objects, yielding location-aware prototypes, and then performs segmentation by computing the semantic similarity between query features and location-aware prototypes. Additionally, to further enhance the representativeness of the obtained location-aware prototypes in low-data regimes, we design a prediction iterative refinement module, which can iteratively exploit the query predictions output by each iteration to update the location-aware prototypes and progressively refine the query predictions. Extensive experiments on three challenging medical image datasets, i.e., Abd-MRI, Card-MRI, and Abd-CT, show that the proposed LATNet achieves remarkable improvements over current state-of-the-art methods by an average of 4.17%, 1.50%, and 4.63% in terms of the Dice Score, respectively. Wendong Huang, Bin Xiao 0002, Jinwu Hu, Xiuli Bi |
BIBM | 3 |
| 2023 | IEMask R-CNN: Information-Enhanced Mask R-CNNabstractThe instance segmentation task is relatively difficult in computer vision, which requires not only high-quality masks but also high-accuracy instance category classification. Mask R-CNN has been proven to be a feasible method. However, due to the Feature Pyramid Network (FPN) structure lack useful channel information, global information and low-level texture information, and mask branch cannot obtain useful local-global information, Mask R-CNN is prevented from obtaining high-quality masks and high-accuracy instance category classification. Therefore, we proposed the Information-enhanced Mask R-CNN, called IEMask R-CNN. In the FPN structure of IEMask R-CNN, the information-enhanced FPN will enhance the useful channel information and the global information of the feature maps to solve the issues that the high-level feature map loses useful channel information and inaccurate of instance category classification, meanwhile the bottom-up path enhancement with adaptive feature fusion will ultilize the precise positioning signal in the lower layer to enhance the feature pyramid. In the mask branch of IEMask R-CNN, an encoding-decoding mask head will strength local-global information to gain a high-quality mask. Without bells and whistles, IEMask R-CNN gains significant gains of about 2.60%, 4.00%, 3.17% over Mask R-CNN on MS COCO2017, Cityscapes and LVIS1.0 benchmarks respectively. Xiuli Bi, Jinwu Hu, Bin Xiao 0002, Weisheng Li 0001, Xinbo Gao 0001 |
IEEE Trans. Big Data | 2 |