VLDB 2026 Research / reviewers in the wild / expert
Jian Hu 0002
dblp:61/5788-2
· DBLP profile ↗
13ranked-venue papers
6as first author
10since 2021 · last 2026
0000-0001-9918-672XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 5 first-author · 6 since 2021Artificial intelligence and machine learning · 6 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Class-Aware Diversified Augmentation for Open-Set Single Domain GeneralizationabstractIn Open-Set Single Domain Generalization (OS-SDG), one only has access to a single labeled source domain for training. It assumes that the learned model generalizes well to target samples belonging to the source label space whilst classifies target samples outside the source label space into a single “unknown” class. The current method synthesizes new samples that are semantically unrelated to known classes to simulate target unknown classes. This ignores that unknown classes actually may semantically correlated to known classes, making it difficult to discriminate samples at the margins of class decision boundaries as “unknown”. In this work, we introduce a Class-Aware Diversified Augmentation (CADA) method to overcome this problem. Our key idea is to synthesize explicitly new multiple unknown target classes with diversified semantic and learn the inherent correlation among the known and unknown classes, so to both increase the coverage of multiple target unknown classes and to optimize class margin separation. CADA is optimized by enhanced diversity maximization and class-aware minimization. The former synthesizes more novel classes by considering both semantic relationships to known classes and domain shift between the source and target domains. The latter employs class-agnostic clustering with synthesized samples to simulate class correlations among target classes, maximizing class margin separation. Theoretical analysis and experiments on five benchmarks show the efficacy of our CADA. Jian Hu 0002, Shaogang Gong, Weitong Cai, Junchi Yan |
IEEE Trans. Multim. | 1 |
| 2025 | InvSeg: Test-Time Prompt Inversion for Semantic SegmentationabstractVisual-textual correlations in the attention maps derived from text-to-image diffusion models are proven beneficial to dense visual prediction tasks, e.g., semantic segmentation. However, a significant challenge arises due to the input distributional discrepancy between the context-rich sentences used for image generation and the isolated class names typically used in semantic segmentation. This discrepancy hinders diffusion models from capturing accurate visual-textual correlations. To solve this, we propose InvSeg, a test-time prompt inversion method that tackles open-vocabulary semantic segmentation by inverting image-specific visual context into text prompt embedding space, leveraging structure information derived from the diffusion model's reconstruction process to enrich text prompts so as to associate each class with a structure-consistent mask. Specifically, we introduce Contrastive Soft Clustering (CSC) to align derived masks with the image's structure information, softly selecting anchors for each class and calculating weighted distances to push inner-class pixels closer while separating inter-class pixels, thereby ensuring mask distinction and internal consistency. By incorporating sample-specific context, InvSeg learns context-rich text prompts in embedding space and achieves accurate semantic alignment across modalities. Experiments show that InvSeg achieves state-of-the-art performance on the PASCAL VOC, PASCAL Context and COCO Object datasets. Jiayi Lin 0002, Jiabo Huang, Jian Hu 0002, Shaogang Gong |
AAAI | 3 |
| 2025 | Contrastive Scenario-Aware Meta Prompting for Multi-scenario Recommendation
Ang Li 0043, Jian Hu 0002, Ke Ding 0001, Jun Zhou 0011, Yong He 0009 |
DASFAA (6) | 2 |
| 2024 | Relax Image-Specific Prompt Requirement in SAM: A Single Generic Prompt for Segmenting Camouflaged ObjectsabstractCamouflaged object detection (COD) approaches heavily rely on pixel-level annotated datasets. Weakly-supervised COD (WSCOD) approaches use sparse annotations like scribbles or points to reduce annotation efforts, but this can lead to decreased accuracy. The Segment Anything Model (SAM) shows remarkable segmentation ability with sparse prompts like points. However, manual prompt is not always feasible, as it may not be accessible in real-world application. Additionally, it only provides localization information instead of semantic one, which can intrinsically cause ambiguity in interpreting targets. In this work, we aim to eliminate the need for manual prompt. The key idea is to employ Cross-modal Chains of Thought Prompting (CCTP) to reason visual prompts using the semantic information given by a generic text prompt. To that end, we introduce a test-time instance-wise adaptation mechanism called Generalizable SAM (GenSAM) to automatically generate and optimize visual prompts from the generic task prompt for WSCOD. In particular, CCTP maps a single generic text prompt onto image-specific consensus foreground and background heatmaps using vision-language models, acquiring reliable visual prompts. Moreover, to test-time adapt the visual prompts, we further propose Progressive Mask Generation (PMG) to iteratively reweight the input image, guiding the model to focus on the targeted region in a coarse-to-fine manner. Crucially, all network parameters are fixed, avoiding the need for additional training. Experiments on three benchmarks demonstrate that GenSAM outperforms point supervision approaches and achieves comparable results to scribble supervision ones, solely relying on general task descriptions. Our codes is in https://github.com/jyLin8100/GenSAM. Jian Hu 0002, Jiayi Lin 0002, Shaogang Gong, Weitong Cai |
AAAI | 1 |
| 2024 | Feature-Distribution Perturbation and Calibration for Generalized ReidabstractPerson Re-identification (ReID) has been advanced remarkably over the last 10 years. However, the i.i.d. (independent and identically distributed) assumption is somewhat non-applicable to ReID considering its objective to identify images of the same pedestrian across cameras at different locations. In this work, we propose a Feature-Distribution Perturbation and Calibration (PECA) method to derive generic feature representations for person ReID. Specifically, we perform per-domain feature-distribution perturbation to refrain the model from overfitting to the domain-biased distribution of each source (seen) domain by enforcing feature invariance to distribution shifts caused by perturbation. Furthermore, we design a global calibration mechanism to align feature distributions across all the source domains to improve the model’s generalization capacity by eliminating domain bias. These local perturbation and global calibration are conducted simultaneously, which share the same principle to avoid models overfitting by regularization respectively on the perturbed and the original distributions. Extensive experiments were conducted and the proposed PECA model outperformed the state-of-the-art competitors by significant margins. Qilei Li, Jiabo Huang, Jian Hu 0002, Shaogang Gong |
ICASSP | 3 |
| 2024 | Leveraging Hallucinations to Reduce Manual Prompt Dependency in Promptable SegmentationabstractPromptable segmentation typically requires instance-specific manual prompts to guide the segmentation of each desired object. To minimize such a need, task-generic promptable segmentation has been introduced, which employs a single task-generic prompt to segment various images of different objects in the same task. Current methods use Multimodal Large Language Models (MLLMs) to reason detailed instance-specific prompts from a task-generic prompt for improving segmentation accuracy. The effectiveness of this segmentation heavily depends on the precision of these derived prompts. However, MLLMs often suffer hallucinations during reasoning, resulting in inaccurate prompting. While existing methods focus on eliminating hallucinations to improve a model, we argue that MLLM hallucinations can reveal valuable contextual insights when leveraged correctly, as they represent pre-trained large-scale knowledge beyond individual images. In this paper, we first utilize hallucinations to mine task-related information from images and verify its accuracy to enhance precision of the generated prompts. Specifically, we introduce an iterative \textbf{Pro}mpt-\textbf{Ma}sk \textbf{C}ycle generation framework (ProMaC) with a prompt generator and a mask generator. The prompt generator uses a multi-scale chain of thought prompting, initially leveraging hallucinations to extract extended contextual prompts on a test image. These hallucinations are then minimized to formulate precise instance-specific prompts, directing the mask generator to produce masks that are consistent with task semantics by mask semantic alignment. Iteratively the generated masks induce the prompt generator to focus more on task-relevant image areas and reduce irrelevant hallucinations, resulting jointly in better prompts and masks. Experiments on 5 benchmarks demonstrate the effectiveness of ProMaC. Code is in https://lwpyh.github.io/ProMaC/. Jian Hu 0002, Jiayi Lin 0002, Junchi Yan, Shaogang Gong |
NeurIPS | 1 |
| 2023 | Global-Aware Model-Free Self-distillation for Recommendation System
Ang Li 0043, Jian Hu 0002, Wei Lu 0011, Ke Ding 0001, Jun Zhou 0011, Yong He 0009, Liang Zhang 0045, Lihong Gu |
DASFAA (4) | 2 |
| 2023 | Uncertainty-based Heterogeneous Privileged Knowledge Distillation for Recommendation SystemabstractIn industrial recommendation systems, both data sizes and computational resources vary across different scenarios. For scenarios with limited data, data sparsity can lead to a decrease in model performance. Heterogeneous knowledge distillation-based transfer learning can be used to transfer knowledge from models in data-rich domains. However, in recommendation systems, the target domain possesses specific privileged features that significantly contribute to the model. While existing knowledge distillation methods have not taken these features into consideration, leading to suboptimal transfer weights. To overcome this limitation, we propose a novel algorithm called Uncertainty-based Heterogeneous Privileged Knowledge Distillation (UHPKD). Our method aims to quantify the knowledge of both the source and target domains, which represents the uncertainty of the models. This approach allows us to derive transfer weights based on the knowledge gain, which captures the difference in knowledge between the source and target domains. Experiments conducted on both public and industrial datasets demonstrate the superiority of our UHPKD algorithm compared to other state-of-the-art methods. Ang Li 0043, Jian Hu 0002, Ke Ding 0001, Jun Zhou 0011, Yong He 0009, Xu Min |
SIGIR | 2 |
| 2022 | Learning Unbiased Transferability for Domain Adaptation by Uncertainty Modeling
Jian Hu 0002, Haowen Zhong, Shaogang Gong, Guile Wu, Junchi Yan |
ECCV (31) | 1 |
| 2022 | Attribute-Conditioned Face Swapping Network for Low-Resolution ImagesabstractDeep learning based face swapping technologies have opened new frontiers for entertainment industries while pose novel threats to identity security. Applying face swapping to real-world products, as well as defending against its misuse, rely on the capacity to generate high quality face swapped images from realistic scenarios where high resolution images are hard to come by. To this end, we need to address the drawbacks of existing methods, especially on their lacking on maintaining the detail attributes and their dependency on high resolution images as inputs. In this paper, we propose a novel Attribute-Conditioned Face Swapping Network (AFSNet) to preserve attributes and handle low resolution images. Specifically, we use an Image Enhancement Network (IEN) to restore high resolution images from low resolution images and a Face Exchange Module (FEM) to swap the faces. In the FEM, we improve the fidelities of the generated images by using a novel multi-domain feature fusion module (MDFFM) to integrate the identity feature, context feature, IEN feature, and attribute vector to obtain the final image. We also design an attribute transfer loss to promote the consistency of the attributes such as beards and youth between the source and swapped images. The experiments demonstrate our method’s superior performance compared with the state-of-the-art methods. Ang Li 0043, Jian Hu 0002, Chilin Fu, Jun Zhou 0011 |
ICASSP | 2 |
| 2020 | Discriminative Partial Domain Adversarial Network
Jian Hu 0002, Hongya Tuo, Lingfeng Qiao, Haowen Zhong, Junchi Yan, Zhongliang Jing, Henry Leung 0001 |
ECCV (27) | 1 |
| 2019 | Multi-Weight Partial Domain Adaptation
Jian Hu 0002, Hongya Tuo, Lingfeng Qiao, Haowen Zhong, Zhongliang Jing |
BMVC | 1 |
| 2019 | Source-Constraint Adversarial Domain AdaptationabstractAdversarial adaptation has made great contributions to transfer learning, while the adversarial training strategy lacks of stability when reducing the discrepancy between domains. In this paper, we propose a novel Source-constraint Adversarial Domain Adaptation (SADA) method, which jointly use adversarial adaptation and maximum mean discrepancy (MMD) so that the method can be easily optimized by gradient descent. Furthermore, motivated by metric learning, our method introduces metric loss to constrain the structure of source domain, which explicitly increases the inter-class distance and decreases the intra-class distance. As a result, SADA can not only reduce the domain discrepancy, but also make the extracted features become more domain-invariant and discriminative. We show that our model yields state of the art results on standard datasets. Haowen Zhong, Hongya Tuo, Xuanguang Ren, Jian Hu 0002, Lingfeng Qiao |
ICIP | 5 |