EDBT 2026 Demo / reviewers in the wild / expert
Kai Han 0001
dblp:51/4757-1
· DBLP profile ↗
68ranked-venue papers
10as first author
50since 2021 · last 2026
0000-0002-7995-9999ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 61 · 7 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 39 · 7 first-author · 26 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | When Deepfake Detection Meets Graph Neural Network: A Unified and Lightweight FrameworkabstractThe proliferation of generative video models has made detecting AI-generated and manipulated videos an urgent challenge. Existing detection approaches often fail to generalize across diverse manipulation types due to their reliance on isolated spatial, temporal, or spectral information, and typically require large models to perform well. This paper introduces SSTGNN, a lightweight Spatial-Spectral-Temporal Graph Neural Network framework that represents videos as structured graphs, enabling joint reasoning over spatial inconsistencies, temporal artifacts, and spectral distortions. SSTGNN incorporates learnable spectral filters and spatial-temporal differential modeling into a unified graph-based architecture, capturing subtle manipulation traces more effectively. Extensive experiments on diverse benchmark datasets demonstrate that SSTGNN not only achieves superior performance in both in-domain and cross-domain settings, but also offers strong efficiency and resource allocation. Remarkably, SSTGNN accomplishes these results with up to 42× fewer parameters than state-of-the-art models, making it highly lightweight and resource-friendly for real-world deployment. Haoyu Liu 0001, Chaoyu Gong, Mengke He, Jiate Li, Kai Han 0001, Siqiang Luo |
KDD (1) | 5 |
| 2026 | LooC: Effective Low-Dimensional Codebook for Compositional Vector QuantizationabstractVector quantization (VQ) is a prevalent and fundamental technique that discretizes continuous feature vectors by approximating them using a codebook. As the diversity and complexity of data and models continue to increase, there is an urgent need for high-capacity, yet more compact VQ methods. This paper aims to reconcile this conflict by presenting a new approach called LooC, which utilizes an effective Low-dimensional codebook for Compositional vector quantization. Firstly, LooC introduces a parameter-efficient codebook by reframing the relationship between codevectors and feature vectors, significantly expanding its solution space. Instead of individually matching codevectors with feature vectors, LooC treats them as lower-dimensional compositional units within feature vectors and combines them, resulting in a more compact codebook with improved performance. Secondly, LooC incorporates a parameter-free extrapolation-by-interpolation mechanism to enhance and smooth features during the VQ process, which allows for better preservation of details and fidelity in feature approximation. The design of LooC leads to full codebook usage, effectively utilizing the compact codebook while avoiding the problem of collapse. Thirdly, LooC can serve as a plug-and-play module for existing methods for different downstream tasks based on VQ. Finally, extensive evaluations on different tasks, datasets, and architectures demonstrate that LooC outperforms existing VQ methods, achieving state-of-the-art performance with a significantly smaller codebook. Jie Li 0040, Kwan-Yee Kenneth Wong, Kai Han 0001 |
WACV | 3 |
| 2026 | Semantic Correspondence: Unified Benchmarking and a Strong BaselineabstractEstablishing semantic correspondence is a challenging task in computer vision, aiming to match keypoints with the same semantic information across different images. Benefiting from the rapid development of deep learning, remarkable progress has been made over the past decade. However, a comprehensive review and analysis of this task remains absent. In this paper, we present the first extensive survey of semantic correspondence methods. We first propose a taxonomy to classify existing methods based on the type of their method designs. These methods are then categorized accordingly, and we provide a detailed analysis of each approach. Furthermore, we aggregate and summarize the results of methods in the literature across various benchmarks into a unified comparative table, with detailed configurations to highlight performance variations. Additionally, to provide a detailed understanding of existing methods for semantic matching, we thoroughly conduct controlled experiments to analyze the effectiveness of the components of different methods. Finally, we propose a simple yet effective baseline that achieves state-of-the-art performance on multiple benchmarks, providing a solid foundation for future research in this field. We hope this survey serves as a comprehensive reference and consolidated baseline for future development. Xinghui Li, Kai Han 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | GAMEBoT: Transparent Assessment of LLM Reasoning in GamesabstractLarge Language Models (LLMs) are increasingly deployed in real-world applications that demand complex reasoning.To track progress, robust benchmarks are required to evaluate their capabilities beyond superficial pattern recognition.However, current LLM reasoning benchmarks often face challenges such as insufficient interpretability, performance saturation or data contamination.To address these challenges, we introduce GAMEBOT (GAME Battle of Tactics), a gaming arena designed for rigorous and transparent assessment of LLM reasoning capabilities.GAMEBOT decomposes complex reasoning in games into predefined modular subproblems.This decomposition allows us to design a suite of Chain-of-Thought (CoT) prompts that leverage domain knowledge to guide LLMs in addressing these subproblems before action selection.Furthermore, we develop a suite of rule-based algorithms to generate ground truth for these subproblems, enabling rigorous validation of the LLMs' intermediate reasoning steps.This approach facilitates evaluation of both the quality of final actions and the accuracy of the underlying reasoning process.GAMEBOT also naturally alleviates the risk of data contamination through dynamic games and head-to-head LLM competitions.We benchmark 17 prominent LLMs across eight games, encompassing various strategic abilities and game characteristics.Our results suggest that GAMEBOT presents a significant challenge, even when LLMs are provided with detailed CoT prompts.Project page: https://visual-ai.github.io/gamebot Wenye Lin, Jonathan Roberts 0004, Yunhan Yang, Samuel Albanie, Zongqing Lu 0002, Kai Han 0001 |
ACL (1) | 6 |
| 2025 | ELIP: Enhanced Visual-Language Foundation Models for Image RetrievalabstractThe objective in this paper is to improve the per-formance of text-to-image retrieval. To this end, we introduce a new framework that can boost the performance of large-scale pre-trained vision-language models, so that they can be used for text-to-image re-ranking. The approach, Enhanced Language-Image Pre-training (ELIP), uses the text query, via a simple MLP mapping network, to predict a set of visual prompts to condition the ViT image encoding. ELIP can easily be applied to the commonly used CLIP, SigLIP and BLIP-2 networks. On the evaluation side, we set up two new out-of-distribution (OOD) benchmarks, Occluded COCO and ImageNet-R, to assess the zero-shot generalisation of the models to different domains. The results demonstrate that ELIP significantly boosts CLIP/SigLIP/SigLIP-2 text-to-image retrieval performance and outperforms BLIP-2 on several benchmarks, as well as providing an easy means to adapt to OOD datasets. Guanqi Zhan, Yuanpei Liu, Kai Han 0001, Weidi Xie, Andrew Zisserman |
CBMI | 3 |
| 2025 | ICE: Intrinsic Concept Extraction from a Single Image via Diffusion ModelsabstractThe inherent ambiguity in defining visual concepts poses significant challenges for modern generative models, such as the diffusion-based Text-to-Image (T2I) models, in accurately learning concepts from a single image. Existing methods lack a systematic way to reliably extract the interpretable underlying intrinsic concepts. To address this challenge, we present ICE, short for Intrinsic Concept Extraction, a novel framework that exclusively utilizes a T2I model to automatically and systematically extract intrinsic concepts from a single image. ICE consists of two pivotal stages. In the first stage, ICE devises an automatic concept localization module to pinpoint relevant text-based concepts and their corresponding masks within the image. This critical stage streamlines concept initialization and provides precise guidance for subsequent analysis. The second stage delves deeper into each identified mask, decomposing the object-level concepts into intrinsic concepts and general concepts. This decomposition allows for a more granular and interpretable breakdown of visual elements. Our framework demonstrates superior performance on intrinsic concept extraction from a single image in an unsupervised manner. Project page: https://visual-ai.github.io/ice Fernando Julio Cendra, Kai Han 0001 |
CVPR | 2 |
| 2025 | Hyperbolic Category DiscoveryabstractGeneralized Category Discovery (GCD) is an intriguing open-world problem that has garnered increasing attention. Given a dataset that includes both labelled and unlabelled images, GCD aims to categorize all images in the unlabelled subset, regardless of whether they belong to known or unknown classes. In GCD, the common practice typically involves applying a spherical projection operator at the end of the self-supervised pretrained backbone, operating within Euclidean or spherical space. However, both of these spaces have been shown to be suboptimal for encoding samples that possesses hierarchical structures. In contrast, hyperbolic space exhibits exponential volume growth relative to radius, making it inherently strong at capturing the hierarchical structure of samples from both seen and unseen categories. Therefore, we propose to tackle the category discovery challenge in the hyperbolic space. We introduce HypCD, a simple Hyperbolic framework for learning hierarchy-aware representations and classifiers for generalized Category Discovery. HypCD first transforms the Euclidean embedding space of the backbone network into hyperbolic space, facilitating subsequent representation and classification learning by considering both hyperbolic distance and the angle between samples. This approach is particularly helpful for knowledge transfer from known to unknown categories in GCD. We thoroughly evaluate HypCD on public GCD benchmarks, by applying it to various baseline and state-of-the-art methods, consistently achieving significant improvements. Project page: https://visual-ai.github.io/hypcd/ Yuanpei Liu, Zhenqi He, Kai Han 0001 |
CVPR | 3 |
| 2025 | Parallel Sequence Modeling via Generalized Spatial Propagation NetworkabstractWe present the Generalized Spatial Propagation Network (GSPN), a new attention mechanism optimized for vision tasks that inherently captures 2D spatial structures. Existing attention models, including transformers, linear attention, and state-space models like Mamba, process multidimensional data as 1D sequences, compromising spatial coherence and efficiency. GSPN overcomes these limitations by directly operating on spatially coherent image data and forming dense pairwise connections through a line-scan approach. Central to GSPN is the Stability-Context Condition, which ensures stable, long-context propagation across 2D sequences and reduces the effective sequence length to $\sqrt N $ for a square map with N elements, which significantly enhances computational efficiency. With learnable, input-dependent weights and no reliance on positional embeddings, GSPN achieves superior spatial fidelity and state-of-the-art performance in vision tasks, including ImageNet classification, class-guided image generation, and text-to-image generation. Notably, GSPN accelerates SD-XL with softmax-attention by over 84× when generating 16K images. Project page: https://whj363636.github.io/GSPN/ Wonmin Byeon, Jinwei Gu, Ka Chun Cheung, Xiaolong Wang 0004, Kai Han 0001, Jan Kautz, Sifei Liu |
CVPR | 7 |
| 2025 | Inpaint4Drag: Repurposing Inpainting Models for Drag-Based Image Editing via Bidirectional WarpingabstractDrag-based image editing has emerged as a powerful paradigm for intuitive image manipulation. However, existing approaches predominantly rely on manipulating the latent space of generative models, leading to limited precision, delayed feedback, and model-specific constraints. Accordingly, we present Inpaint4Drag, a novel framework that decomposes drag-based editing into pixel-space bidirectional warping and image inpainting. Inspired by elastic object deformation in the physical world, we treat image regions as deformable materials that maintain natural shape under user manipulation. Our method achieves real-time warping previews (0.01s) and efficient inpainting (0.3s) at 512x512 resolution, significantly improving the interaction experience compared to existing methods that require minutes per edit. By transforming drag inputs directly into standard inpainting formats, our approach serves as a universal adapter for any inpainting model without architecture modification, automatically inheriting all future improvements in inpainting technology. Extensive experiments demonstrate that our method achieves superior visual quality and precise control while maintaining real-time performance. Project page: https://visual-ai.github.io/inpaint4drag/ Kai Han 0001 |
ICCV | 2 |
| 2025 | GRAB: A Challenging Graph Analysis Benchmark for Large Multimodal ModelsabstractLarge multimodal models (LMMs) have exhibited proficiencies across many visual tasks. Although numerous well-known benchmarks exist to evaluate model performance, they increasingly have insufficient headroom. As such, there is a pressing need for a new generation of benchmarks challenging enough for the next generation of LMMs. One area that LMMs show potential is graph analysis, specifically, the tasks an analyst might typically perform when interpreting figures such as estimating the mean, intercepts or correlations of functions and data series. In this work, we introduce GRAB, a graph analysis benchmark, fit for current and future frontier LMMs. Our benchmark is predominantly synthetic, ensuring high-quality, noise-free questions. GRAB is comprised of 3284 questions, covering five tasks and 23 graph properties. We evaluate 20 LMMs on GRAB, finding it to be a challenging benchmark, with the highest performing model attaining a score of just 21.0%. Finally, we conduct various ablations to investigate where the models succeed and struggle. We release GRAB and a lightweight GRAB-Lite to encourage progress in this important, growing domain. Jonathan Roberts 0004, Kai Han 0001, Samuel Albanie |
ICCV | 2 |
| 2025 | AvatarGO: Zero-shot 4D Human-Object Interaction Generation and AnimationabstractRecent advancements in diffusion models have led to significant improvements in the generation and animation of 4D full-body human-object interactions (HOI). Nevertheless, existing methods primarily focus on SMPL-based motion generation, which is limited by the scarcity of realistic large-scale interaction data. This constraint affects their ability to create everyday HOI scenes. This paper addresses this challenge using a zero-shot approach with a pre-trained diffusion model. Despite this potential, achieving our goals is difficult due to the diffusion model's lack of understanding of ''where'' and ''how'' objects interact with the human body. To tackle these issues, we introduce **AvatarGO**, a novel framework designed to generate animatable 4D HOI scenes directly from textual inputs. Specifically, **1)** for the ''where'' challenge, we propose **LLM-guided contact retargeting**, which employs Lang-SAM to identify the contact body part from text prompts, ensuring precise representation of human-object spatial relations. **2)** For the ''how'' challenge, we introduce **correspondence-aware motion optimization** that constructs motion fields for both human and object models using the linear blend skinning function from SMPL-X. Our framework not only generates coherent compositional motions, but also exhibits greater robustness in handling penetration issues. Extensive experiments with existing methods validate AvatarGO's superior generation and animation capabilities on a variety of human-object pairs and diverse poses. As the first attempt to synthesize 4D avatars with object interactions, we hope AvatarGO could open new doors for human-centric 4D content creation. Liang Pan, Kai Han 0001, Kwan-Yee Kenneth Wong, Ziwei Liu 0002 |
ICLR | 3 |
| 2025 | BiGR: Harnessing Binary Latent Codes for Image Generation and Improved Visual Representation CapabilitiesabstractWe introduce BiGR, a novel conditional image generation model using compact binary latent codes for generative training, focusing on enhancing both generation and representation capabilities. BiGR is the first conditional generative model that unifies generation and discrimination within the same framework.
BiGR features a binary tokenizer, a masked modeling mechanism, and a binary transcoder for binary code prediction.
Additionally, we introduce a novel entropy-ordered sampling method to enable efficient image generation.
Extensive experiments validate BiGR's superior performance in generation quality, as measured by FID-50k, and representation capabilities, as evidenced by linear-probe accuracy.
Moreover, BiGR showcases zero-shot generalization across various vision tasks, enabling applications such as image inpainting, outpainting, editing, interpolation, and enrichment, without the need for structural modifications. Our findings suggest that BiGR unifies generative and discriminative tasks effectively, paving the way for further advancements in the field. We further enable BiGR to perform text-to-image generation, showcasing its potential for broader applications. Shaozhe Hao, Xuantong Liu, Xianbiao Qi, Bojia Zi, Rong Xiao 0003, Kai Han 0001, Kwan-Yee Kenneth Wong |
ICLR | 7 |
| 2025 | DebGCD: Debiased Learning with Distribution Guidance for Generalized Category DiscoveryabstractIn this paper, we tackle the problem of Generalized Category Discovery (GCD).
Given a dataset containing both labelled and unlabelled images, the objective is to categorize all images in the unlabelled subset, irrespective of whether they are from known or unknown classes.
In GCD, an inherent label bias exists between known and unknown classes due to the lack of ground-truth labels for the latter.
State-of-the-art methods in GCD leverage parametric classifiers trained through self-distillation with soft labels, leaving the bias issue unattended.
Besides, they treat all unlabelled samples uniformly, neglecting variations in certainty levels and resulting in suboptimal learning.
Moreover, the explicit identification of semantic distribution shifts between known and unknown classes, a vital aspect for effective GCD, has been neglected.
To address these challenges, we introduce DebGCD, a Debiased learning with distribution guidance framework for GCD.
Initially, DebGCD co-trains an auxiliary debiased classifier in the same feature space as the GCD classifier, progressively enhancing the GCD features. Moreover, we introduce a semantic distribution detector in a separate feature space to implicitly boost the learning efficacy of GCD. Additionally, we employ a curriculum learning strategy based on semantic distribution certainty to steer the debiased learning at an optimized pace.
Thorough evaluations on GCD benchmarks demonstrate the consistent state-of-the-art performance of our framework, highlighting its superiority. Project page: [https://visual-ai.github.io/debgcd/](https://visual-ai.github.io/debgcd/) Yuanpei Liu, Kai Han 0001 |
ICLR | 2 |
| 2025 | Needle Threading: Can LLMs Follow Threads Through Near-Million-Scale Haystacks?abstractAs the context limits of Large Language Models (LLMs) increase, the range of
possible applications and downstream functions broadens. In many real-world
tasks, decisions depend on details scattered across collections of often disparate
documents containing mostly irrelevant information. Long-context LLMs appear
well-suited to this form of complex information retrieval and reasoning, which has
traditionally proven costly and time-consuming. However, although the development of longer context models has seen rapid gains in recent years, our understanding of how effectively LLMs use their context has not kept pace. To address
this, we conduct a set of retrieval experiments designed to evaluate the capabilities
of 17 leading LLMs, such as their ability to follow threads of information through
the context window. Strikingly, we find that many models are remarkably thread-
safe: capable of simultaneously following multiple threads without significant loss
in performance. Still, for many models, we find the effective context limit is significantly shorter than the supported context length, with accuracy decreasing as
the context window grows. Our study also highlights the important point that token counts from different tokenizers should not be directly compared—they often
correspond to substantially different numbers of written characters. We release
our code and long context experimental data. Jonathan Roberts 0004, Kai Han 0001, Samuel Albanie |
ICLR | 2 |
| 2025 | HiLo: A Learning Framework for Generalized Category Discovery Robust to Domain ShiftsabstractGeneralized Category Discovery (GCD) is a challenging task in which, given a partially labelled dataset, models must categorize all unlabelled instances, regardless of whether they come from labelled categories or from new ones. In this paper, we challenge a remaining assumption in this task: that all images share the same domain. Specifically, we introduce a new task and method to handle GCD when the unlabelled data also contains images from different domains to the labelled set. Our proposed `HiLo' networks extract High-level semantic and Low-level domain features, before minimizing the mutual information between the representations. Our intuition is that the clusterings based on domain information and semantic information should be independent. We further extend our method with a specialized domain augmentation tailored for the GCD task, as well as a curriculum learning approach. Finally, we construct a benchmark from corrupted fine-grained datasets as well as a large-scale evaluation on DomainNet with real-world domain shifts, reimplementing a number of GCD baselines in this setting. We demonstrate that HiLo outperforms SoTA category discovery models by a large margin on all evaluations. Sagar Vaze, Kai Han 0001 |
ICLR | 3 |
| 2025 | VaMP: Variational Multi-Modal Prompt Learning for Vision-Language ModelsabstractVision-language models (VLMs), such as CLIP, have shown strong generalization under zero-shot settings, yet adapting them to downstream tasks with limited supervision remains a significant challenge. Existing multi-modal prompt learning methods typically rely on fixed, shared prompts and deterministic parameters, which limits their ability to capture instance-level variation or model uncertainty across diverse tasks and domains. To tackle this issue, we propose a novel Variational Multi-Modal Prompt Learning (VaMP) framework that enables sample-specific, uncertainty-aware prompt tuning in multi-modal representation learning. VaMP generates instance-conditioned prompts by sampling from a learned posterior distribution, allowing the model to personalize its behavior based on input content.
To further enhance the integration of local and global semantics, we introduce a class-aware prior derived from the instance representation and class prototype. Building upon these, we formulate prompt tuning as variational inference over latent prompt representations and train the entire framework end-to-end through reparameterized sampling. Experiments on few-shot and domain generalization benchmarks show that VaMP achieves state-of-the-art performance, highlighting the benefits of modeling both uncertainty and task structure in our method. Project page: https://visual-ai.github.io/vamp Silin Cheng 0001, Kai Han 0001 |
NeurIPS | 2 |
| 2025 | SEAL: Semantic-Aware Hierarchical Learning for Generalized Category DiscoveryabstractThis paper investigates the problem of Generalized Category Discovery (GCD). Given a partially labelled dataset, GCD aims to categorize all unlabelled images, regardless of whether they belong to known or unknown classes. Existing approaches typically depend on either single-level semantics or manually designed abstract hierarchies, which limit their generalizability and scalability. To address these limitations, we introduce a SEmantic-aware hierArchical Learning framework (SEAL), guided by naturally occurring and easily accessible hierarchical structures. Within SEAL, we propose a Hierarchical Semantic-Guided Soft Contrastive Learning approach that exploits hierarchical similarity to generate informative soft negatives, addressing the limitations of conventional contrastive losses that treat all negatives equally. Furthermore, a Cross-Granularity Consistency (CGC) module is designed to align the predictions from different levels of granularity. SEAL consistently achieves state-of-the-art performance on fine-grained benchmarks, including the SSB benchmark, Oxford-Pet, and the Herbarium19 dataset, and further demonstrates generalization on coarse-grained datasets.
Project page: https://visual-ai.github.io/seal/ Zhenqi He, Yuanpei Liu, Kai Han 0001 |
NeurIPS | 3 |
| 2025 | Panoptic Captioning: An Equivalence Bridge for Image and TextabstractThis work introduces panoptic captioning, a novel task striving to seek the minimum text equivalent of images, which has broad potential applications. We take the first step towards panoptic captioning by formulating it as a task of generating a comprehensive textual description for an image, which encapsulates all entities, their respective locations and attributes, relationships among entities, as well as global image state. Through an extensive evaluation, our work reveals that state-of-the-art Multi-modal Large Language Models (MLLMs) have limited performance in solving panoptic captioning. To address this, we propose an effective data engine named PancapEngine to produce high-quality data and a novel method named PancapChain to improve panoptic captioning. Specifically, our PancapEngine first detects diverse categories of entities in images by an elaborate detection suite, and then generates required panoptic captions using entity-aware prompts.
Additionally, our PancapChain explicitly decouples the challenging panoptic captioning task into multiple stages and generates panoptic captions step by step. More importantly, we contribute a comprehensive metric named PancapScore and a human-curated test set for reliable model evaluation. Experiments show that our PancapChain-13B model can beat state-of-the-art open-source MLLMs like InternVL-2.5-78B and even surpass proprietary models like GPT-4o and Gemini-2.0-Pro, demonstrating the effectiveness of our data engine and method.
Project page: https://visual-ai.github.io/pancap/ Kun-Yu Lin, Weining Ren, Kai Han 0001 |
NeurIPS | 4 |
| 2025 | Fin3R: Fine-tuning Feed-forward 3D Reconstruction Models via Monocular Knowledge DistillationabstractWe present Fin3R, a simple, effective, and general fine-tuning method for feed-forward 3D reconstruction models. The family of feed-forward reconstruction model regresses pointmap of all input images to a reference frame coordinate system, along with other auxiliary outputs, in a single forward pass. However, we find that current models struggle with fine geometry and robustness due to (\textit{i}) the scarcity of high-fidelity depth and pose supervision and (\textit{ii}) the inherent geometric misalignment from multi-view pointmap regression. Fin3R jointly tackles two issues with an extra lightweight fine-tuning step. We freeze the decoder, which handles view matching, and fine-tune only the image encoder—the component dedicated to feature extraction. The encoder is enriched with fine geometric details distilled from a strong monocular teacher model on large, unlabeled datasets, using a custom, lightweight LoRA adapter. We validate our method on a wide range of models, including DUSt3R, MASt3R, CUT3R, and VGGT. The fine-tuned models consistently deliver sharper boundaries, recover complex structures, and achieve higher geometric accuracy in both single- and multi-view settings, while adding only the tiny LoRA weights, which leave test-time memory and latency virtually unchanged. Project page: \href{http://visual-ai.github.io/fin3r}{https://visual-ai.github.io/fin3r} Weining Ren, Xiao Tan 0001, Kai Han 0001 |
NeurIPS | 4 |
| 2025 | GSPN-2: Efficient Parallel Sequence ModelingabstractEfficient vision transformer remains a bottleneck for high-resolution images and long-video related real-world applications. Generalized Spatial Propagation Network (GSPN) \cite{wang2025parallel} addresses this by replacing quadratic self-attention with a line-scan propagation scheme, bringing the cost close to linear in the number of rows or columns, while retaining accuracy. Despite this advancement, the existing GSPN implementation still suffers from (i) heavy overhead due to repeatedly launching GPU kernels, (ii) excessive data transfers from global GPU memory, and (iii) redundant computations caused by maintaining separate propagation weights for each channel. We introduce GSPN-2, a joint algorithm–system redesign. In particular, we eliminate thousands of micro-launches from the previous implementation into one single 2D kernel, explicitly pin one warp to each channel slice, and stage the previous column's activations in shared memory. On the model side, we introduce a set of channel-shared propagation weights that replace per-channel matrices, trimming parameters, and align naturally with the affinity map used in transformer attention. Experiments demonstrate GSPN-2's effectiveness across image classification and text-to-image synthesis tasks, matching transformer-level accuracy with significantly lower computational cost. GSPN-2 establishes a new efficiency frontier for modeling global spatial context in vision applications through its unique combination of structured matrix transformations and GPU-optimized implementation. Yitong Jiang, Collin McCarthy, David Wehr, Hanrong Ye, Ka Chun Cheung, Wonmin Byeon, Jinwei Gu, Kai Han 0001, Hongxu Yin, Pavlo Molchanov 0001, Jan Kautz, Sifei Liu |
NeurIPS | 11 |
| 2025 | Wukong's 72 Transformations: High-fidelity Textured 3D Morphing via Flow ModelsabstractWe present WUKONG, a novel training-free framework for high-fidelity textured 3D morphing that takes a pair of source and target prompts (text or images) as input. Unlike conventional methods -- which rely on manual correspondence matching and deformation trajectory estimation (limiting generalization and requiring costly preprocessing) -- WUKONG leverages the generative prior of flow-based transformers to produce high-fidelity 3D transitions with rich texture details. To ensure smooth shape transitions, we exploit the inherent continuity of flow-based generative processes and formulate morphing as an optimal transport barycenter problem. We further introduce a sequential initialization strategy to prevent abrupt geometric distortions and preserve identity coherence. For faithful texture preservation, we propose a similarity-guided semantic consistency mechanism that selectively retains high-frequency details and enables precise control over blending dynamics. This empowers WUKONG to support both global texture transitions and identity-preserving texture morphing, catering to diverse generation needs. Through extensive quantitative and qualitative evaluations, we demonstrate that WUKONG significantly outperforms state-of-the-art methods, achieving superior results across diverse geometry and texture variations. Minghao Yin, Kai Han 0001 |
NeurIPS | 3 |
| 2025 | VipDiff: Towards Coherent and Diverse Video Inpainting via Training-Free Denoising Diffusion ModelsabstractRecent video inpainting methods have achieved encouraging improvements by leveraging optical flow to guide pixel propagation from reference frames, either in the image space or feature space. However, they would produce severe artifacts when the masked area is too large and no pixel correspondences could be found. Recently, denoising diffusion models have demonstrated impressive performance in generating diverse and high-quality images, and have been exploited in a number of works for image inpainting. These methods, however, cannot be applied directly to videos to produce temporal-coherent inpainting results. In this paper, we propose a training-free framework, named VipDiff, for conditioning diffusion model on the reverse diffusion process to produce temporal-coherent inpainting results without requiring any training data or fine-tuning the pre-trained models. VipDiff takes optical flow as guidance to extract valid pixels from reference frames to serve as constraints in optimizing the randomly sampled Gaussian noise, and uses the generated results for further pixel propagation and conditional generation. VipDiff also allows for generating diverse video inpainting results over different sampled noise. Experiments demonstrate that our VipDiff outperforms state-of-the-art methods in terms of both spatial-temporal coherence and fidelity. Chaohao Xie, Kai Han 0001, Kwan-Yee Kenneth Wong |
WACV | 2 |
| 2025 | CusConcept: Customized Visual Concept Decomposition with Diffusion ModelsabstractEnabling generative models to decompose visual concepts from a single image is a complex and challenging problem. In this paper, we study a new and challenging task, customized concept decomposition, wherein the objective is to leverage diffusion models to decompose a single image and generate visual concepts from various perspectives. To address this challenge, we propose a two-stage framework, CusConcept (short for Customized Visual Concept Decomposition), to extract customized visual concept embedding vectors that can be embedded into prompts for text-to-image generation. In the first stage, CusConcept employs a vocabulary-guided concept decomposition mechanism to build vocabularies along human-specified conceptual axes. The decomposed concepts are obtained by retrieving corresponding vocabularies and learning anchor weights. In the second stage, joint concept refinement is performed to enhance the fidelity and quality of generated images. We further curate an evaluation benchmark for assessing the performance of the open-world concept decomposition task. Our approach can effectively generate high-quality images of the decomposed concepts and produce related lexical predictions as secondary outcomes. Extensive qualitative and quantitative experiments demonstrate the effectiveness of CusConcept. Our code and data are available at https://github.com/xzLcan/CusConcept. Shaozhe Hao, Kai Han 0001 |
WACV | 3 |
| 2025 | Dissecting Out-of-Distribution Detection and Open-Set Recognition: A Critical Analysis of Methods and BenchmarksabstractAbstract Detecting test-time distribution shift has emerged as a key capability for safely deployed machine learning models, with the question being tackled under various guises in recent years. In this paper, we aim to provide a consolidated view of the two largest sub-fields within the community: out-of-distribution (OOD) detection and open-set recognition (OSR). In particular, we aim to provide rigorous empirical analysis of different methods across settings and provide actionable takeaways for practitioners and researchers. Concretely, we make the following contributions: (i) We perform rigorous cross-evaluation between state-of-the-art methods in the OOD detection and OSR settings and identify a strong correlation between the performances of methods for them; (ii) We propose a new, large-scale benchmark setting which we suggest better disentangles the problem tackled by OOD detection and OSR, re-evaluating state-of-the-art OOD detection and OSR methods in this setting; (iii) We surprisingly find that the best performing method on standard benchmarks (Outlier Exposure) struggles when tested at scale, while scoring rules which are sensitive to the deep feature magnitude consistently show promise; and (iv) We conduct empirical analysis to explain these phenomena and highlight directions for future research. Code: https://github.com/Visual-AI/Dissect-OOD-OSR Sagar Vaze, Kai Han 0001 |
Int. J. Comput. Vis. | 3 |
| 2024 | DreamAvatar: Text-and-Shape Guided 3D Human Avatar Generation via Diffusion ModelsabstractWe present DreamAvatar, a text-and-shape guided framework for generating high-quality 3D human avatars with controllable poses. While encouraging results have been reported by recent methods on text-guided 3D common object generation, generating high-quality human avatars remains an open challenge due to the complexity of the human body's shape, pose, and appearance. We propose DreamAvatar to tackle this challenge, which utilizes a train-able NeRF for predicting density and color for 3D points and pretrained text-to-image diffusion models for providing 2D self-supervision. Specifically, we leverage the SMPL model to provide shape and pose guidance for the generation. We introduce a dual-observation-space design that involves the joint optimization of a canonical space and a posed space that are related by a learnable deformation field. This facilitates the generation of more complete textures and geometry faithful to the target pose. We also jointly optimize the losses computed from the full body and from the zoomed-in 3D head to alleviate the common multi-face “Janus” problem and improve facial details in the generated avatars. Extensive evaluations demonstrate that DreamAvatar significantly outperforms existing meth-ods, establishing a new state-of-the-art for text-and-shape guided 3D human avatar generation. Yan-Pei Cao 0001, Kai Han 0001, Ying Shan, Kwan-Yee Kenneth Wong |
CVPR | 3 |
| 2024 | SD4Match: Learning to Prompt Stable Diffusion Model for Semantic MatchingabstractIn this paper, we address the challenge of matching semantically similar keypoints across image pairs. Existing research indicates that the intermediate output of the UNet within the Stable Diffusion (SD) can serve as robust image feature maps for such a matching task. We demonstrate that by employing a basic prompt tuning technique, the inherent potential of Stable Diffusion can be harnessed, resulting in a significant enhancement in accuracy over previous approaches. We further introduce a novel conditional prompting module that conditions the prompt on the local details of the input image pairs, leading to a further improvement in performance. We designate our approach as SD4Match, short for Stable Diffusion for Semantic Matching. Comprehensive evaluations of SD4Match on the PF-Pascal, PF-Willow, and SPair-71k datasets show that it sets new benchmarks in accuracy across all these datasets. Particularly, SD4Match outperforms the previous state-of-the-art by a margin of 12 percentage points on the challenging SPair-71k dataset. Code is available at the project website: https://sd4match.active.vision/. Xinghui Li, Kai Han 0001, Victor Adrian Prisacariu |
CVPR | 3 |
| 2024 | IBD-SLAM: Learning Image-Based Depth Fusion for Generalizable SLAMabstractIn this paper, we address the challenging problem of visual SLAM with neural scene representations. Recently, neural scene representations have shown promise for SLAM to produce dense 3D scene reconstruction with high qual-ity. However, existing methods require scene-specific op-timization, leading to time-consuming mapping processes for each individual scene. To overcome this limitation, we propose IBD-SLAM, an Image-Based Depth fusion frame-work for generalizable SLAM. In particular, we adopt a Neural Radiance Field (NeRF) for scene representation. Inspired by multi-view image-based rendering, instead of learning a fixed-grid scene representation, we propose to learn an image-based depth fusion model that fuses depth maps of multiple reference views into a xyz-map represen-tation. Once trained, this model can be applied to new, uncalibrated monocular RGBD videos of unseen scenes, without the need for retraining, and reconstructs full 3D scenes efficiently with a light-weight pose optimization pro-cedure. We thoroughly evaluate IBD-SLAM on public visual SLAM benchmarks, outperforming the previous state-of-the-art while being 10x faster in the mapping stage. Project page:https://visual-ai.github.io/ibd-slam Minghao Yin, Shangzhe Wu, Kai Han 0001 |
CVPR | 3 |
| 2024 | PromptCCD: Learning Gaussian Mixture Prompt Pool for Continual Category Discovery
Fernando Julio Cendra, Bingchen Zhao, Kai Han 0001 |
ECCV (4) | 3 |
| 2024 | ConceptExpress: Harnessing Diffusion Models for Single-Image Unsupervised Concept Extraction
Shaozhe Hao, Kai Han 0001, Zhengyao Lv, Kwan-Yee Kenneth Wong |
ECCV (59) | 2 |
| 2024 | RegionDrag: Fast Region-Based Image Editing with Diffusion Models
Xinghui Li, Kai Han 0001 |
ECCV (17) | 3 |
| 2024 | SPTNet: An Efficient Alternative Framework for Generalized Category Discovery with Spatial Prompt TuningabstractGeneralized Category Discovery (GCD) aims to classify unlabelled images from both ‘seen’ and ‘unseen’ classes by transferring knowledge from a set of labelled ‘seen’ class images. A key theme in existing GCD approaches is adapting large-scale pre-trained models for the GCD task. An alternate perspective, however, is to adapt the data representation itself for better alignment with the pre-trained model. As such, in this paper, we introduce a two-stage adaptation approach termed SPTNet, which iteratively optimizes model parameters (i.e., model-finetuning) and data parameters (i.e., prompt learning). Furthermore, we propose a novel spatial prompt tuning method (SPT) which considers the spatial property of image data, enabling the method to better focus on object parts, which can transfer between seen and unseen classes. We thoroughly evaluate our SPTNet on standard benchmarks and demonstrate that our method outperforms existing GCD methods. Notably, we find our method achieves an average accuracy of 61.4% on the SSB, surpassing prior state-of-the-art methods by approximately 10%. The improvement is particularly remarkable as our method yields extra parameters amounting to only 0.117% of those in the backbone architecture. Project page: https://visual-ai.github.io/sptnet. Sagar Vaze, Kai Han 0001 |
ICLR | 3 |
| 2024 | SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure InterpretationabstractLarge multimodal models (LMMs) have proven flexible and generalisable across many tasks and fields. Although they have strong potential to aid scientific research, their capabilities in this domain are not well characterised. A key aspect of scientific research is the ability to understand and interpret figures, which serve as a rich, compressed source of complex information. In this work, we present SciFIBench, a scientific figure interpretation benchmark consisting of 2000 questions split between two tasks across 8 categories. The questions are curated from arXiv paper figures and captions, using adversarial filtering to find hard negatives and human verification forquality control. We evaluate 28 LMMs on SciFIBench, finding it to be a challenging benchmark. Finally, we investigate the alignment and reasoning faithfulness of the LMMs on augmented question sets from our benchmark. We release SciFIBench to encourage progress in this domain. Jonathan Roberts 0004, Kai Han 0001, Neil Houlsby, Samuel Albanie |
NeurIPS | 2 |
| 2024 | DualRC: A Dual-Resolution Learning Framework With Neighbourhood Consensus for Visual CorrespondencesabstractWe address the problem of establishing accurate correspondences between two images. We present a flexible framework that can easily adapt to both geometric and semantic matching. Our contribution consists of three parts. Firstly, we propose an end-to-end trainable framework that uses the coarse-to-fine matching strategy to accurately find the correspondences. We generate feature maps in two levels of resolution, enforce the neighbourhood consensus constraint on the coarse feature maps by 4D convolutions and use the resulting correlation map to regulate the matches from the fine feature maps. Secondly, we present three variants of the model with different focuses. Namely, a universal correspondence model named DualRC that is suitable for both geometric and semantic matching, an efficient model named DualRC-L tailored for geometric matching with a lightweight neighbourhood consensus module that significantly accelerates the pipeline for high-resolution input images, and the DualRC-D model in which we propose a novel dynamically adaptive neighbourhood consensus module (DyANC) that dynamically selects the most suitable non-isotropic 4D convolutional kernels with the proper neighbourhood size to account for the scale variation. Last, we thoroughly experiment on public benchmarks for both geometric and semantic matching, showing superior performance in both cases. Xinghui Li, Kai Han 0001, Shuda Li, Victor Adrian Prisacariu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | SeSDF: Self-Evolved Signed Distance Field for Implicit 3D Clothed Human ReconstructionabstractWe address the problem of clothed human reconstruction from a single image or uncalibrated multiview images. existing methods struggle with reconstructing detailed geometry of a clothed human and often require a calibrated setting for multiview reconstruction. We propose a flexible framework which, by leveraging the parametric SMPL-X model, can take an arbitrary number of input images to reconstruct a clothed human model under an uncalibrated setting. At the core of our framework is our novel self-evolved signed distance field (SeSDF) module which allows the framework to learn to deform the signed distance field (SDF) derived from the fitted SMPL-X model, such that detailed geometry reflecting the actual clothed human can be encoded for better reconstruction. Besides, we propose a simple method for self-calibration of multiview images via the fitted SMPL-X parameters. This lifts the requirement of tedious manual calibration and largely increases the flexibility of our method. Further, we introduce an effective occlusion-aware feature fusion strategy to account for the most useful features to reconstruct the human model. We thoroughly evaluate our framework on public benchmarks, demonstrating significant superiority over the state-of-the-arts both qualitatively and quantitatively. Kai Han 0001, Kwan-Yee Kenneth Wong |
CVPR | 2 |
| 2023 | Learning Attention as Disentangler for Compositional Zero-Shot LearningabstractCompositional zero-shot learning (CZSL) aims at learning visual concepts (i.e., attributes and objects) from seen compositions and combining concept knowledge into unseen compositions. The key to CZSL is learning the disentanglement of the attribute-object composition. To this end, we propose to exploit cross-attentions as compositional disentanglers to learn disentangled concept embeddings. For example, if we want to recognize an unseen composition “yellow flower”, we can learn the attribute concept “yellow” and object concept “flower” from different yellow objects and different flowers respectively. To further constrain the disentanglers to learn the concept of interest, we employ a regularization at the attention level. Specifically, we adapt the earth mover's distance (EMD) as a feature similarity metric in the cross-attention module. Moreover, benefiting from concept disentanglement, we improve the inference process and tune the prediction score by combining multiple concept probabilities. Comprehensive experiments on three CZSL benchmark datasets demonstrate that our method significantly outperforms previous works in both closed- and open-world settings, establishing a new state-of-the-art. Project page: https://haoosz.github.io/ade-czsl/ Shaozhe Hao, Kai Han 0001, Kwan-Yee Kenneth Wong |
CVPR | 2 |
| 2023 | Learning Semi-supervised Gaussian Mixture Models for Generalized Category DiscoveryabstractIn this paper, we address the problem of generalized category discovery (GCD), i.e., given a set of images where part of them are labelled and the rest are not, the task is to automatically cluster the images in the unlabelled data, leveraging the information from the labelled data, while the unlabelled data contain images from the labelled classes and also new ones. GCD is similar to semi-supervised learning (SSL) but is more realistic and challenging, as SSL assumes all the unlabelled images are from the same classes as the labelled ones. We also do not assume the class number in the unlabelled data is known a-priori, making the GCD problem even harder. To tackle the problem of GCD without knowing the class number, we propose an EM-like framework that alternates between representation learning and class number estimation. We propose a semi-supervised variant of the Gaussian Mixture Model (GMM) with a stochastic splitting and merging mechanism to dynamically determine the prototypes by examining the cluster compactness and separability. With these prototypes, we leverage prototypical contrastive learning for representation learning on the partially labelled data subject to the constraints imposed by the labelled data. Our framework alternates between these two steps until convergence. The cluster assignment for an unlabelled instance can then be retrieved by identifying its nearest prototype. We comprehensively evaluate our framework on both generic image classification datasets and challenging fine-grained object recognition datasets, achieving state-of-the-art performance. Our code is available at https://github.com/DTennant/GPC. Bingchen Zhao, Xin Wen 0004, Kai Han 0001 |
ICCV | 3 |
| 2023 | HeadSculpt: Crafting 3D Head Avatars with TextabstractRecently, text-guided 3D generative methods have made remarkable advancements in producing high-quality textures and geometry, capitalizing on the proliferation of large vision-language and image diffusion models.
However, existing methods still struggle to create high-fidelity 3D head avatars in two aspects:
(1) They rely mostly on a pre-trained text-to-image diffusion model whilst missing the necessary 3D awareness and head priors.
This makes them prone to inconsistency and geometric distortions in the generated avatars.
(2) They fall short in fine-grained editing. This is primarily due to the inherited limitations from the pre-trained 2D image diffusion models, which become more pronounced when it comes to 3D head avatars.
In this work, we address these challenges by introducing a versatile coarse-to-fine pipeline dubbed HeadSculpt for crafting (i.e., generating and editing) 3D head avatars from textual prompts.
Specifically, we first equip the diffusion model with 3D awareness by leveraging landmark-based control and a learned textual embedding representing the back view appearance of heads, enabling 3D-consistent head avatar generations.
We further propose a novel identity-aware editing score distillation strategy to optimize a textured mesh with a high-resolution differentiable rendering technique.
This enables identity preservation while following the editing instruction.
We showcase HeadSculpt's superior fidelity and editing capabilities through comprehensive experiments and comparisons with existing methods. Kai Han 0001, Xiatian Zhu, Jiankang Deng, Yi-Zhe Song, Tao Xiang 0002, Kwan-Yee Kenneth Wong |
NeurIPS | 3 |
| 2022 | JIFF: Jointly-aligned Implicit Face Function for High Quality Single View Clothed Human ReconstructionabstractThis paper addresses the problem of single view 3D human reconstruction. Recent implicit function based methods have shown impressive results, but they fail to recover fine face details in their reconstructions. This largely degrades user experience in applications like 3D telepresence. In this paper, we focus on improving the quality of face in the reconstruction and propose a novel Jointly-aligned Implicit Face Function (JIFF) that combines the merits of the implicit function based approach and model based approach. We employ a 3D morphable face model as our shape prior and compute space-aligned 3D features that capture detailed face geometry information. Such space-aligned 3D features are combined with pixel-aligned 2D features to jointly predict an implicit face function for high quality face reconstruction. We further extend our pipeline and introduce a coarse-to-fine architecture to predict high quality texture for our detailed face model. Extensive evaluations have been carried out on public datasets and our proposed JIFF has demonstrates superior performance (both quantitatively and qualitatively) over existing state-of-the-arts. Guanying Chen, Kai Han 0001, Wenqi Yang, Kwan-Yee Kenneth Wong |
CVPR | 3 |
| 2022 | Generalized Category DiscoveryabstractIn this paper, we consider a highly general image recognition setting wherein, given a labelled and unlabelled set of images, the task is to categorize all images in the unlabelled set. Here, the unlabelled images may come from labelled classes or from novel ones. Existing recognition methods are not able to deal with this setting, because they make several restrictive assumptions, such as the unlabelled instances only coming from known – or unknown – classes, and the number of unknown classes being known a-priori. We address the more unconstrained setting, naming it ‘Generalized Category Discovery’, and challenge all these assumptions. We first establish strong baselines by taking state-of-the-art algorithms from novel category discovery and adapting them for this task. Next, we propose the use of vision transformers with contrastive representation learning for this open-world setting. We then introduce a simple yet effective semi-supervised k-means method to cluster the unlabelled data into seen and unseen classes automatically, substantially outperforming the baselines. Finally, we also propose a new approach to estimate the number of classes in the unlabelled data. We thoroughly evaluate our approach on public datasets for generic object classification and on fine-grained datasets, leveraging the recent Semantic Shift Benchmark suite. Code: https://www.robots.ox.ac.uk/~vgg/research/gcd Sagar Vaze, Kai Han 0001, Andrea Vedaldi, Andrew Zisserman |
CVPR | 2 |
| 2022 | SharpContour: A Contour-based Boundary Refinement Approach for Efficient and Accurate Instance SegmentationabstractExcellent performance has been achieved on instance segmentation but the quality on the boundary area remains unsatisfactory, which leads to a rising attention on boundary refinement. For practical use, an ideal post-processing refinement scheme are required to be accurate, generic and efficient. However, most of existing approaches propose pixel-wise refinement, which either introduce a massive computation cost or design specifically for different backbone models. Contour-based models are efficient and generic to be incorporated with any existing segmentation methods, but they often generate over-smoothed contour and tend to fail on corner areas. In this paper, we propose an efficient contour-based boundary refinement approach, named SharpContour, to tackle the segmentation of boundary area. We design a novel contour evolution process together with an Instance-aware Point Classifier. Our method deforms the contour iteratively by updating offsets in a discrete manner. Differing from existing contour evolution methods, SharpContour estimates each offset more independently so that it predicts much sharper and accurate contours. Notably, our method is generic to seamlessly work with diverse existing models with a small computational cost. Experiments show that SharpContour achieves competitive gains whilst preserving high efficiency. Chenming Zhu, Xuanye Zhang, Yanran Li, Liangdong Qiu, Kai Han 0001, Xiaoguang Han 0001 |
CVPR | 5 |
| 2022 | Novel Class Discovery Without Forgetting
K. J. Joseph, Sujoy Paul, Gaurav Aggarwal, Soma Biswas, Piyush Rai, Kai Han 0001, Vineeth N. Balasubramanian |
ECCV (24) | 6 |
| 2022 | Open-Set Recognition: A Good Closed-Set Classifier is All You Need
Sagar Vaze, Kai Han 0001, Andrea Vedaldi, Andrew Zisserman |
ICLR | 2 |
| 2022 | Deep Photometric Stereo for Non-Lambertian SurfacesabstractThis paper addresses the problem of photometric stereo, in both calibrated and uncalibrated scenarios, for non-Lambertian surfaces based on deep learning. We first introduce a fully convolutional deep network for calibrated photometric stereo, which we call PS-FCN. Unlike traditional approaches that adopt simplified reflectance models to make the problem tractable, our method directly learns the mapping from reflectance observations to surface normal, and is able to handle surfaces with general and unknown isotropic reflectance. At test time, PS-FCN takes an arbitrary number of images and their associated light directions as input and predicts a surface normal map of the scene in a fast feed-forward pass. To deal with the uncalibrated scenario where light directions are unknown, we introduce a new convolutional network, named LCNet, to estimate light directions from input images. The estimated light directions and the input images are then fed to PS-FCN to determine the surface normals. Our method does not require a pre-defined set of light directions and can handle multiple images in an order-agnostic manner. Thorough evaluation of our approach on both synthetic and real datasets shows that it outperforms state-of-the-art methods in both calibrated and uncalibrated scenarios. Guanying Chen, Kai Han 0001, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee Kenneth Wong |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2022 | AutoNovel: Automatically Discovering and Learning Novel Visual CategoriesabstractWe tackle the problem of discovering novel classes in an image collection given labelled examples of other classes. We present a new approach called AutoNovel to address this problem by combining three ideas: (1) we suggest that the common approach of bootstrapping an image representation using the labelled data only introduces an unwanted bias, and that this can be avoided by using self-supervised learning to train the representation from scratch on the union of labelled and unlabelled data; (2) we use ranking statistics to transfer the model's knowledge of the labelled classes to the problem of clustering the unlabelled images; and, (3) we train the data representation by optimizing a joint objective function on the labelled and unlabelled subsets of the data, improving both the supervised classification of the labelled data, and the clustering of the unlabelled data. Moreover, we propose a method to estimate the number of classes for the case where the number of new categories is not known a priori. We evaluate AutoNovel on standard classification benchmarks and substantially outperform current methods for novel category discovery. In addition, we also show that AutoNovel can be used for fully unsupervised image clustering, achieving promising results. Kai Han 0001, Sylvestre-Alvise Rebuffi, Sébastien Ehrhardt, Andrea Vedaldi, Andrew Zisserman |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Anisotropic Convolutional Neural Networks for RGB-D Based Semantic Scene CompletionabstractSemantic scene completion (SSC) is a computer vision task aiming to simultaneously infer the occupancy and semantic labels for each voxel in a scene from partial information consisting of a depth image and/or a RGB image. As a voxel-wise labeling task, the key for SSC is how to effectively model the visual and geometrical variations to complete the scene. To this end, we propose the Anisotropic Network (AIC-Net), with novel convolutional modules that can model varying anisotropic receptive fields voxel-wisely in a computationally efficient manner. The basic idea to achieve such anisotropy is to decompose 3D convolution into three consecutive dimensional convolutions, and determine the dimension-wise kernels on the fly. One module, termed kernel-selection anisotropic (KSA) convolution, adaptively selects the optimal kernel sizes for each dimensional convolution from a set of candidate kernels, and the other module, termed kernel-modulation anisotropic (KMA) convolution, directly modulates a single convolutional kernel for each dimension to derive more flexible receptive field. By stacking multiple such anisotropic modules, the 3D context modeling capability and flexibility can be further enhanced. Moreover, we present a new end-to-end trainable framework to approach the SSC task avoiding the expensive TSDF pre-processing as in many existing methods. Extensive experiments on SSC benchmarks show the advantage of the proposed methods. Jie Li 0040, Peng Wang 0023, Kai Han 0001, Yu Liu 0029 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2021 | X Resolution Correspondence Networks
Georgi Tinchev, Shuda Li, Kai Han 0001, David Mitchell, Rigas Kouskouridas |
BMVC | 3 |
| 2021 | Contrastive Learning Based Hybrid Networks for Long-Tailed Image ClassificationabstractLearning discriminative image representations plays a vital role in long-tailed image classification because it can ease the classifier learning in imbalanced cases. Given the promising performance contrastive learning has shown recently in representation learning, in this work, we explore effective supervised contrastive learning strategies and tailor them to learn better image representations from imbalanced data in order to boost the classification accuracy thereon. Specifically, we propose a novel hybrid network structure being composed of a supervised contrastive loss to learn image representations and a cross-entropy loss to learn classifiers, where the learning is progressively transited from feature learning to the classifier learning to embody the idea that better features make better classifiers. We explore two variants of contrastive loss for feature learning, which vary in the forms but share a common idea of pulling the samples from the same class together in the normalized embedding space and pushing the samples from different classes apart. One of them is the recently proposed supervised contrastive (SC) loss, which is designed on top of the state-of-the-art unsupervised contrastive loss by incorporating positive samples from the same class. The other is a prototypical supervised contrastive (PSC) learning strategy which addresses the intensive memory consumption in standard SC loss and thus shows more promise under limited memory budget. Extensive experiments on three long-tailed classification datasets demonstrate the advantage of the proposed contrastive learning based hybrid networks in long-tailed classification. Peng Wang 0023, Kai Han 0001, Xiu-Shen Wei, Lei Zhang 0054, Lei Wang 0001 |
CVPR | 2 |
| 2021 | Joint Representation Learning and Novel Category Discovery on Single- and Multi-modal DataabstractThis paper studies the problem of novel category discovery on single- and multi-modal data with labels from different but relevant categories. We present a generic, end-to-end framework to jointly learn a reliable representation and assign clusters to unlabelled data. To avoid over-fitting the learnt embedding to labelled data, we take inspiration from self-supervised representation learning by noise-contrastive estimation and extend it to jointly handle labelled and unlabelled data. In particular, we propose using category discrimination on labelled data and cross-modal discrimination on multi-modal data to augment instance discrimination used in conventional contrastive learning approaches. We further employ Winner-Take-All (WTA) hashing algorithm on the shared representation space to generate pairwise pseudo labels for unlabelled data to better predict cluster assignments. We thoroughly evaluate our framework on large-scale multi-modal video benchmarks Kinetics-400 and VGG-Sound, and image benchmarks CIFAR10, CIFAR100 and ImageNet, obtaining state-of-the-art results. Xuhui Jia, Kai Han 0001, Yukun Zhu, Bradley Green |
ICCV | 2 |
| 2021 | Novel Visual Category Discovery with Dual Ranking Statistics and Mutual Knowledge DistillationabstractIn this paper, we tackle the problem of novel visual category discovery, i.e., grouping unlabelled images from new classes into different semantic partitions by leveraging a labelled dataset that contains images from other different but relevant categories. This is a more realistic and challenging setting than conventional semi-supervised learning. We propose a two-branch learning framework for this problem, with one branch focusing on local part-level information and the other branch focusing on overall characteristics. To transfer knowledge from the labelled data to the unlabelled, we propose using dual ranking statistics on both branches to generate pseudo labels for training on the unlabelled data. We further introduce a mutual knowledge distillation method to allow information exchange and encourage agreement between the two branches for discovering new categories, allowing our model to enjoy the benefits of global and local features. We comprehensively evaluate our method on public benchmarks for generic object classification, as well as the more challenging datasets for fine-grained visual recognition, achieving state-of-the-art performance. Bingchen Zhao, Kai Han 0001 |
NeurIPS | 2 |
| 2021 | Fixed Viewpoint Mirror Surface Reconstruction Under an Uncalibrated CameraabstractThis paper addresses the problem of mirror surface reconstruction, and proposes a solution based on observing the reflections of a moving reference plane on the mirror surface. Unlike previous approaches which require tedious calibration, our method can recover the camera intrinsics, the poses of the reference plane, as well as the mirror surface from the observed reflections of the reference plane under at least three unknown distinct poses. We first show that the 3D poses of the reference plane can be estimated from the reflection correspondences established between the images and the reference plane. We then form a bunch of 3D lines from the reflection correspondences, and derive an analytical solution to recover the line projection matrix. We transform the line projection matrix to its equivalent camera projection matrix, and propose a cross-ratio based formulation to optimize the camera projection matrix by minimizing reprojection errors. The mirror surface is then reconstructed based on the optimized cross-ratio constraint. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and accuracy of our method. Kai Han 0001, Miaomiao Liu 0001, Dirk Schnieders, Kwan-Yee Kenneth Wong |
IEEE Trans. Image Process. | 1 |
| 2020 | Correspondence Networks With Adaptive Neighbourhood ConsensusabstractIn this paper, we tackle the task of establishing dense visual correspondences between images containing objects of the same category. This is a challenging task due to large intra-class variations and a lack of dense pixel level annotations. We propose a convolutional neural network architecture, called adaptive neighbourhood consensus network (ANC-Net), that can be trained end-to-end with sparse key-point annotations, to handle this challenge. At the core of ANC-Net is our proposed non-isotropic 4D convolution kernel, which forms the building block for the adaptive neighbourhood consensus module for robust matching. We also introduce a simple and efficient multi-scale self-similarity module in ANC-Net to make the learned feature robust to intra-class variations. Furthermore, we propose a novel orthogonal loss that can enforce the one-to-one matching constraint. We thoroughly evaluate the effectiveness of our method on various benchmarks, where it substantially outperforms state-of-the-art methods. Shuda Li, Kai Han 0001, Theo W. Costain, Henry Howard-Jenkins, Victor Adrian Prisacariu |
CVPR | 2 |
| 2020 | Anisotropic Convolutional Networks for 3D Semantic Scene CompletionabstractAs a voxel-wise labeling task, semantic scene completion (SSC) tries to simultaneously infer the occupancy and semantic labels for a scene from a single depth and/or RGB image. The key challenge for SSC is how to effectively take advantage of the 3D context to model various objects or stuffs with severe variations in shapes, layouts, and visibility. To handle such variations, we propose a novel module called anisotropic convolution, which properties with flexibility and power impossible for the competing methods such as standard 3D convolution and some of its variations. In contrast to the standard 3D convolution that is limited to a fixed 3D receptive field, our module is capable of modeling the dimensional anisotropy voxel-wisely. The basic idea is to enable anisotropic 3D receptive field by decomposing a 3D convolution into three consecutive 1D convolutions, and the kernel size for each such 1D convolution is adaptively determined on the fly. By stacking multiple such anisotropic convolution modules, the voxel-wise modeling capability can be further enhanced while maintaining a controllable amount of model parameters. Extensive experiments on two SSC benchmarks, NYU-Depth-v2 and NYUCAD, show the superior performance of the proposed method. Jie Li 0040, Kai Han 0001, Peng Wang 0023, Yu Liu 0029, Xia Yuan |
CVPR | 2 |
| 2020 | Automatically Discovering and Learning New Visual Categories with Ranking Statistics
Kai Han 0001, Sylvestre-Alvise Rebuffi, Sébastien Ehrhardt, Andrea Vedaldi, Andrew Zisserman |
ICLR | 1 |
| 2020 | Dual-Resolution Correspondence NetworksabstractWe tackle the problem of establishing dense pixel-wise correspondences between a pair of images. In this work, we introduce Dual-Resolution Correspondence Networks (DualRC-Net), to obtain pixel-wise correspondences in a coarse-to-fine manner. DualRC-Net extracts both coarse- and fine- resolution feature maps. The coarse maps are used to produce a full but coarse 4D correlation tensor, which is then refined by a learnable neighbourhood consensus module. The fine-resolution feature maps are used to obtain the final dense correspondences guided by the refined coarse 4D correlation tensor. The selected coarse-resolution matching scores allow the fine-resolution features to focus only on a limited number of possible matches with high confidence. In this way, DualRC-Net dramatically increases matching reliability and localisation accuracy, while avoiding to apply the expensive 4D convolution kernels on fine-resolution feature maps. We comprehensively evaluate our method on large-scale public benchmarks including HPatches, InLoc, and Aachen Day-Night. It achieves state-of-the-art results on all of them. Xinghui Li, Kai Han 0001, Shuda Li, Victor Adrian Prisacariu |
NeurIPS | 2 |
| 2020 | Real-time keypoints detection for autonomous recovery of the unmanned ground vehicleabstractThe combination of a small unmanned ground vehicle (UGV) and a large unmanned carrier vehicle allows more flexibility in real applications such as rescue in dangerous scenarios. The autonomous recovery system, which is used to guide the small UGV back to the carrier vehicle, is an essential component to achieve a seamless combination of the two vehicles. This study proposes a novel autonomous recovery framework with a low‐cost monocular vision system to provide accurate positioning and attitude estimation of the UGV during navigation. First, the authors introduce a light‐weight convolutional neural network called UGV‐KPNet to detect the keypoints of the small UGV form the images captured by a monocular camera. UGV‐KPNet is computationally efficient with a small number of parameters and provides pixel‐level accurate keypoints detection results in real‐time. Then, six degrees of freedom (6‐DoF) pose is estimated using the detected keypoints to obtain positioning and attitude information of the UGV. Besides, they are the first to create a large‐scale real‐world keypoints data set of the UGV. The experimental results demonstrate that the proposed system achieves state‐of‐the‐art performance in terms of both accuracy and speed on UGV keypoint detection, and can further boost the 6‐DoF pose estimation for the UGV. Jie Li 0040, Kai Han 0001, Xia Yuan, Chunxia Zhao, Yu Liu 0029 |
IET Image Process. | 3 |
| 2019 | Self-Calibrating Deep Photometric Stereo NetworksabstractThis paper proposes an uncalibrated photometric stereo method for non-Lambertian scenes based on deep learning. Unlike previous approaches that heavily rely on assumptions of specific reflectances and light source distributions, our method is able to determine both shape and light directions of a scene with unknown arbitrary reflectances observed under unknown varying light directions. To achieve this goal, we propose a two-stage deep learning architecture, called SDPS-Net, which can effectively take advantage of intermediate supervision, resulting in reduced learning difficulty compared to a single-stage model. Experiments on both synthetic and real datasets show that our proposed approach significantly outperforms previous uncalibrated photometric stereo methods. Guanying Chen, Kai Han 0001, Boxin Shi, Yasuyuki Matsushita, Kwan-Yee Kenneth Wong |
CVPR | 2 |
| 2019 | Unsupervised Image Matching and Object Discovery as OptimizationabstractLearning with complete or partial supervision is power- ful but relies on ever-growing human annotation efforts. As a way to mitigate this serious problem, as well as to serve specific applications, unsupervised learning has emerged as an important field of research. In computer vision, unsu- pervised learning comes in various guises. We focus here on the unsupervised discovery and matching of object cate- gories among images in a collection, following the work of Cho et al. [12]. We show that the original approach can be reformulated and solved as a proper optimization problem. Experiments on several benchmarks establish the merit of our approach. Huy V. Vo, Francis R. Bach, Minsu Cho, Kai Han 0001, Yann LeCun, Patrick Pérez, Jean Ponce |
CVPR | 4 |
| 2019 | Learning to Discover Novel Visual Categories via Deep Transfer ClusteringabstractWe consider the problem of discovering novel object categories in an image collection. While these images are unlabelled, we also assume prior knowledge of related but different image classes. We use such prior knowledge to reduce the ambiguity of clustering, and improve the quality of the newly discovered classes. Our contributions are twofold. The first contribution is to extend Deep Embedded Clustering to a transfer learning setting; we also improve the algorithm by introducing a representation bottleneck, temporal ensembling, and consistency. The second contribution is a method to estimate the number of classes in the unlabelled data. This also transfers knowledge from the known classes, using them as probes to diagnose different choices for the number of classes in the unlabelled subset. We thoroughly evaluate our method, substantially outperforming state-of-the-art techniques in a large number of benchmarks, including ImageNet, OmniGlot, CIFAR-100, CIFAR-10, and SVHN. Kai Han 0001, Andrea Vedaldi, Andrew Zisserman |
ICCV | 1 |
| 2019 | Learning Transparent Object Matting
Guanying Chen, Kai Han 0001, Kwan-Yee Kenneth Wong |
Int. J. Comput. Vis. | 2 |
| 2018 | TOM-Net: Learning Transparent Object Matting From a Single ImageabstractThis paper addresses the problem of transparent object matting. Existing image matting approaches for transparent objects often require tedious capturing procedures and long processing time, which limit their practical use. In this paper, we first formulate transparent object matting as a refractive flow estimation problem. We then propose a deep learning framework, called TOM-Net, for learning the refractive flow. Our framework comprises two parts, namely a multi-scale encoder-decoder network for producing a coarse prediction, and a residual network for refinement. At test time, TOM-Net takes a single image as input, and outputs a matte (consisting of an object mask, an attenuation mask and a refractive flow field) in a fast feed-forward pass. As no off-the-shelf dataset is available for transparent object matting, we create a large-scale synthetic dataset consisting of 178K images of transparent objects rendered in front of images sampled from the Microsoft COCO dataset. We also collect a real dataset consisting of 876 samples using 14 transparent objects and 60 background images. Promising experimental results have been achieved on both synthetic and real data, which clearly demonstrate the effectiveness of our approach. Guanying Chen, Kai Han 0001, Kwan-Yee Kenneth Wong |
CVPR | 2 |
| 2018 | PS-FCN: A Flexible Learning Framework for Photometric Stereo
Guanying Chen, Kai Han 0001, Kwan-Yee Kenneth Wong |
ECCV (9) | 2 |
| 2018 | Dense Reconstruction of Transparent Objects by Altering Incident Light Paths Through Refraction
Kai Han 0001, Kwan-Yee Kenneth Wong, Miaomiao Liu 0001 |
Int. J. Comput. Vis. | 1 |
| 2017 | SCNet: Learning Semantic CorrespondenceabstractThis paper addresses the problem of establishing semantic correspondences between images depicting different instances of the same object or scene category. Previous approaches focus on either combining a spatial regularizer with hand-crafted features, or learning a correspondence model for appearance only. We propose instead a convolutional neural network architecture, called SCNet, for learning a geometrically plausible model for semantic correspondence. SCNet uses region proposals as matching primitives, and explicitly incorporates geometric consistency in its loss function. It is trained on image pairs obtained from the PASCAL VOC 2007 keypoint dataset, and a comparative evaluation on several standard benchmarks demonstrates that the proposed approach substantially outperforms both recent deep learning architectures and previous methods based on hand-crafted features. Kai Han 0001, Rafael S. Rezende, Bumsub Ham, Kwan-Yee Kenneth Wong, Minsu Cho, Cordelia Schmid, Jean Ponce |
ICCV | 1 |
| 2016 | Single View 3D Reconstruction under an Uncalibrated Camera and an Unknown Mirror SphereabstractIn this paper, we develop a novel self-calibration method for single view 3D reconstruction using a mirror sphere. Unlike other mirror sphere based reconstruction methods, our method needs neither the intrinsic parameters of the camera, nor the position and radius of the sphere be known. Based on eigen decomposition of the matrix representing the conic image of the sphere and enforcing a repeated eignvalue constraint, we derive an analytical solution for recovering the focal length of the camera given its principal point. We then introduce a robust algorithm for estimating both the principal point and the focal length of the camera by minimizing the differences between focal lengths estimated from multiple images of the sphere. We also present a novel approach for estimating both the principal point and focal length of the camera in the case of just one single image of the sphere. With the estimated camera intrinsic parameters, the position(s) of the sphere can be readily retrieved from the eigen decomposition(s) and a scaled 3D reconstruction follows. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and accuracy of our approach. Kai Han 0001, Kwan-Yee Kenneth Wong, Xiao Tan 0001 |
3DV | 1 |
| 2016 | Mirror Surface Reconstruction under an Uncalibrated CameraabstractThis paper addresses the problem of mirror surface reconstruction, and a solution based on observing the reflections of a moving reference plane on the mirror surface is proposed. Unlike previous approaches which require tedious work to calibrate the camera, our method can recover both the camera intrinsics and extrinsics together with the mirror surface from reflections of the reference plane under at least three unknown distinct poses. Our previous work has demonstrated that 3D poses of the reference plane can be registered in a common coordinate system using reflection correspondences established across images. This leads to a bunch of registered 3D lines formed from the reflection correspondences. Given these lines, we first derive an analytical solution to recover the camera projection matrix through estimating the line projection matrix. We then optimize the camera projection matrix by minimizing reprojection errors computed based on a cross-ratio formulation. The mirror surface is finally reconstructed based on the optimized cross-ratio constraint. Experimental results on both synthetic and real data are presented, which demonstrate the feasibility and accuracy of our method. Kai Han 0001, Kwan-Yee Kenneth Wong, Dirk Schnieders, Miaomiao Liu 0001 |
CVPR | 1 |
| 2015 | A fixed viewpoint approach for dense reconstruction of transparent objectsabstractThis paper addresses the problem of reconstructing the surface shape of transparent objects. The difficulty of this problem originates from the viewpoint dependent appearance of a transparent object, which quickly makes reconstruction methods tailored for diffuse surfaces fail disgracefully. In this paper, we develop a fixed viewpoint approach for dense surface reconstruction of transparent objects based on refraction of light. We introduce a simple setup that allows us alter the incident light paths before light rays enter the object, and develop a method for recovering the object surface based on reconstructing and triangulating such incident light paths. Our proposed approach does not need to model the complex interactions of light as it travels through the object, neither does it assume any parametric form for the shape of the object nor the exact number of refractions and reflections taken place along the light paths. It can therefore handle transparent objects with a complex shape and structure, with unknown and even inhomogeneous refractive index. Experimental results on both synthetic and real data are presented which demonstrate the feasibility and accuracy of our proposed approach. Kai Han 0001, Kwan-Yee Kenneth Wong, Miaomiao Liu 0001 |
CVPR | 1 |
| 2015 | A Novel Arc Segmentation Approach for Document Image ProcessingabstractIn document image processing, arc segmentation plays an important role in vectorization and graphic recognition. Moreover, the unsatisfactory results of several recent arc segmentation contests indicate that conventional methods are inadequate. This paper proposes a new arc segmentation algorithm called SymCAve (an acronym for Symmetry axis, Circle fitting and Average distribution points). First, we locate several seed points and adopt three strategies to ensure that the seed points are proper; then we calculate the center and radius utilizing the seed points. Second, the coordinates of the center and radius are adjusted by employing symmetry axes. Third, the average distribution points method is used to verify whether the points on the circumference are all black pixels. It is a complete circle if all of the points are black pixels. Otherwise, it is a partial circle if some of the points are black pixels and are continuous. Based on this information, the start and end angles of the partial circle can be determined. Finally, these arcs are verified to ensure that the results are accurate. Images and the evaluation tool were obtained from the GREC Workshop's Arc Segmentation contests, to test the systematic performance of the SymCAve algorithm. The experiments demonstrate that the proposed method can provide promising results. However, the algorithm has some drawbacks: it cannot detect a line with width of one pixel, small angles, and any large radius arcs. It is suited for segmenting images with appropriate symmetry axes. Xuan Wang 0002, Kai Han 0001, Zoe Lin Jiang |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2013 | A novel remote eye gaze tracking approach with dynamic calibrationabstractPoint of gaze estimation is the most important part of remote eye gaze tracking techniques. Though there are various point of gaze detection and estimation methods, few of them can satisfy the expectation of widely use. One primary reason is lack of accurate calibration method. In order to deal with this problem, an adaptive calibration technique is proposed based on cross-ratio. It can compensate estimation bias much better. Also, efficient remote eye gaze tracking parameters estimation procedures are employed in this paper. The eye gaze tracking approach in this paper allows users to have free head motion and has high accuracy rate, and it is based on infrared radiation (IR) light sources reflection on cornea. Experimental results demonstrate that considerable improvement can be achieved. Kai Han 0001, Xuan Wang 0002, Hainan Zhao |
MMSP | 1 |