EDBT 2026 Demo / reviewers in the wild / expert
Tao He 0007
dblp:94/5035-7
· DBLP profile ↗
32ranked-venue papers
11as first author
24since 2021 · last 2026
0000-0001-8676-7429ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 7 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 14 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | TiCAL: Typicality-Based Consistency-Aware Learning for Multimodal Emotion RecognitionabstractMultimodal Emotion Recognition (MER) aims to accurately identify human emotional states by integrating heterogeneous modalities such as visual, auditory, and textual data. Existing approaches predominantly rely on unified emotion labels to supervise model training, often overlooking a critical challenge: inter-modal emotion conflicts, wherein different modalities within the same sample may express divergent emotional tendencies. In this work, we address this overlooked issue by proposing a novel framework, Typicality-based Consistent-aware Multimodal Emotion Recognition (TiCAL), inspired by the stage-wise nature of human emotion perception. TiCAL dynamically assesses the consistency of each training sample by leveraging pseudo unimodal emotion labels alongside a typicality estimation. To further enhance emotion representation, we embed features in a hyperbolic space, enabling the capture of fine-grained distinctions among emotional categories. By incorporating consistency estimates into the learning process, our method improves model performance, particularly on samples exhibiting high modality inconsistency. Extensive experiments on benchmark datasets, e.g, MOSEI and MER2023, validate the effectiveness of TiCAL in mitigating inter-modal emotional conflicts and enhancing overall recognition accuracy, e.g., with about 2.6% improvements over the state-of-the-art DMD. Siyu Zhan, Cencen Liu, Guiduo Duan, Xiurui Xie, Yuan-Fang Li, Tao He 0007 |
AAAI | 8 |
| 2026 | Dual-Domain Imitation with Normalization: Enhancing Feature Distillation for Object Detection
Mingdong Zhang, Jielei Wang, Xuewan He, Tao He 0007, Guoming Lu |
ICIC (18) | 5 |
| 2026 | Anchor Drift No More: Hierarchical Consistency-Guided Prompt Distillation for Incomplete Multimodal LearningabstractWeb-scale content is rich in modalities yet frequently incomplete due to device limits, transmission errors, or privacy controls, making learning with missing modalities a core challenge. Prior reconstruction and alignment strategies often fail to preserve a stable class geometry when inputs are partial, leading to anchor drift -- a shift of class prototypes between complete and incomplete views that distorts the shared representation space and degrades generalization. We introduce HiCoD (Hierarchical Consistency-Guided Pro mpt Distillation), which learns a robust, class-anchored semantic space. HiCoD combines: (1) a modality-aware semantic graph that restores cross-modal structure under partial observations; (2) dual-level anchoring that unifies large-language-model–derived global category prototypes with top-K local exemplars to balance cross-modal coherence and modality-specific detail; and (3) multi-level distillation that aligns unimodal features, fused embeddings, and prompt-completed signals within a single anchor space. Across CMU-MOSI, CMU-MOSEI, and additional benchmarks, HiCoD sets a new state of the art under both fixed-pattern and random missingness, improving Acc-2 by up to 6.4 points over MPLMM and remaining robust when key modalities are absent. Ruiting Dai, Zesen Cai, Lisi Mo, Guiduo Duan, Keren Shi, Tao He 0007 |
WWW | 6 |
| 2026 | Multimodal learning with missing modalities: Can progressive diffusion achieve distribution-consistent learning from incomplete multimodal data?
Ruiting Dai, Aoqiang Jian, Ruoxuan Zhang, Yuang Jia, Lisi Mo, Siyu Zhan, Tao He 0007 |
Knowl. Based Syst. | 7 |
| 2026 | Lifelong scene graph generation
Tao He 0007, Tongtong Wu, Dongyang Zhang 0001, Ming Li 0073, Yuan-Fang Li, F. Richard Yu |
Pattern Recognit. | 1 |
| 2025 | Knowledge-Aligned Counterfactual-Enhancement Diffusion Perception for Unsupervised Cross-Domain Visual Emotion RecognitionabstractVisual Emotion Recognition (VER) is a critical yet challenging task aimed at inferring emotional states of individuals based on visual cues. However, existing works focus on single domains, e.g., realistic images or stickers, limiting VER models’ cross-domain generalizability. To fill this gap, we introduce an Unsupervised Cross-Domain Visual Emotion Recognition (UCDVER) task, which aims to generalize visual emotion recognition from the source domain (e.g., realistic images) to the low-resource target domain (e.g., stickers) in an unsupervised manner. Compared to the conventional unsupervised domain adaptation problems, UCDVER presents two key challenges: a significant emotional expression variability and an affective distribution shift. To mitigate these issues, we propose the Knowledge-aligned Counterfactual-enhancement Diffusion Perception (KCDP) framework. Specifically, KCDP leverages a VLM to align emotional representations in a shared knowledge space and guides diffusion models for improved visual affective perception. Furthermore, a Counterfactual-Enhanced Language-image Emotional Alignment (CLIEA) method generates high-quality pseudo-labels for the target domain. Extensive experiments demonstrate that our model surpasses SOTA models in both perceptibility and generalization, e.g., gaining 12% improvements over SOTA VER model TGCA-PVT. The project page is at https://yinwen2019.github.io/ucdver/. Guiduo Duan, Dongyang Zhang 0001, Yuan-Fang Li, Tao He 0007 |
CVPR | 7 |
| 2025 | KMG-LL: Knowledge-enhanced Multimodal Graph for Dialogue GenerationabstractMultimodal dialogue generation requires a comprehensive understanding of each modality and the ability to effectively integrate these modalities to produce contextually relevant and diverse responses. However, existing approaches, which primarily rely on language models, often suffer from modality bias, stemming from an over-dependence on language data in training, and inherent biases in model structure. This over-reliance results in an imbalance across modalities and leads to suboptimal performance. To address this challenge, we propose a novel dialogue generation model with Knowledge-enhanced Multimodal Graph, dubbed KMG-LL, designed to balance multimodal information and incorporate external knowledge. Specifically, we transform multimodal information into graph-structured data and integrate them with commonsense knowledge to construct knowledge-enhanced multimodal graph. To further refine the extraction of multimodal information, we propose a modal-balanced knowledge aggregation that processes modalities cross multiple levels. Extensive experiments on two multimodal dialogue datasets demonstrate that KMG-LL significantly outperforms existing baselines in multimodal dialogue generation. Yuezhou Dong, Tao He 0007, Ke Qin |
ICASSP | 2 |
| 2025 | Unbiased Multimodal Audio-to-Intent RecognitionabstractAudio-to-intent recognition is a critical task focused on identifying a speaker’s intent from spoken language. Recently, multimodal audio-to-intent approaches have emerged as the predominant strategy for enhancing audio-based intent recognition. In this study, we conduct a series of empirical experiments that reveal a significant modality bias in current multimodal audio-to-intent recognition methods. Specifically, these methods disproportionately rely on the textual modality to determine intent, often neglecting the audio data. To address this issue, we propose a context-enhanced contrastive learning framework designed to capture rich regional- and global-audio context information, thereby enabling more balanced audio-to-intent recognition. Additionally, we introduce a prototype-based intent classification strategy that encourages different intent classes and modalities to converge toward unified prototypes, leading to smoother classification boundaries as opposed to the traditionally skewed boundaries. Extensive experiments demonstrate that our approach effectively mitigates modality bias, e.g., a performance improvement of 2.12% in intent classification compared to the state-of-the-art method GZAIR on the dataset MintRec. Yuezhou Dong, Ke Qin, Guiduo Duan, Tao He 0007 |
ICASSP | 5 |
| 2025 | Unbiased Missing-Modality Multimodal Learning
Ruiting Dai, Yandong Yan, Lisi Mo, Ke Qin, Tao He 0007 |
ICCV | 6 |
| 2025 | SPADE: Spatial-Aware Denoising Network for Open-Vocabulary Panoptic Scene Graph Generation with Long- and Local-Range Context ReasoningabstractPanoptic Scene Graph Generation (PSG) integrates instance segmentation with relation understanding to capture pixel-level structural relationships in complex scenes. Although recent approaches leveraging pre-trained vision-language models (VLMs) have significantly improved performance in the open-vocabulary setting, they commonly ignore the inherent limitations of VLMs in spatial relation reasoning, such as difficulty in distinguishing object relative positions, which results in suboptimal relation prediction. Motivated by the denoising diffusion model's inversion process in preserving the spatial structure of input images, we propose SPADE (SPatial-Aware Denoising-nEtwork) framework -- a novel approach for open-vocabulary PSG. SPADE consists of two key steps: (1) inversion-guided calibration for the UNet adaptation, and (2) spatial-aware context reasoning. In the first step, we calibrate a general pre-trained teacher diffusion model into a PSG-specific denoising network with cross-attention maps derived during inversion through a lightweight LoRA-based fine-tuning strategy. In the second step, we develop a spatial-aware relation graph transformer that captures both local and long-range contextual information, facilitating the generation of high-quality relation queries. Extensive experiments on benchmark PSG and Visual Genome datasets demonstrate that SPADE outperforms state-of-the-art methods in both closed- and open-set scenarios, particularly for spatial relationship prediction. Ke Qin, Guiduo Duan, Ming Li 0065, Yuan-Fang Li, Tao He 0007 |
ICCV | 6 |
| 2025 | RobustPT: Dynamic Disentanglement Prompt Tuning in Vision-Language Models with Missing ModalitiesabstractRecently, prompt tuning has garnered considerable attention due to its success across various Vision-Language (VL) tasks. However, unimodal prompts, coupled prompts, and joint prompts in these models often lead to suboptimal performance due to differences in information density and complexity between modalities. Particularly, in scenarios with missing modalities, these prompt-based approaches tend to exacerbate 'Channel Bias'-a phonomenon where models overly rely on specific feature (such as unmissing-modal feature) channels from the base tasks, thereby undermining the model's ability to capture crucial shared knowledge applicable to new tasks and affecting its generalizability. To address this challenge, we propose RobustPT, a dynamic disentanglement prompt tuning model designed to enhance the robustness of VL models under modality missing conditions. RobustPT utilizes a multi-channel prompting mechanism to dynamically disentangle and align prompts. Specifically, RobustPT is divided into single-channel tuning and alignment-channel tuning, where prompts for each modality run independently in sequence to delve deeply into their intrinsic characteristics, followed by an integration through a non-strong coupling strategy to effectively balance information contributions and enhance overall performance. Extensive experiments demonstrate that our RobustPT achieve significant improvements over the current state-of-the-art across all benchmark datasets. Our codes are available at https://github.com/Trae1ounG/RobustPT. Ruiting Dai, Yuqiao Tan, Lisi Mo, Tao He 0007, Ke Qin, Shuang Liang 0002 |
ICMR | 4 |
| 2025 | Fine-grained Block Pruning with Tiny Sets for Vision TransformersabstractVision Transformers (ViTs) and their variants have achieved remarkable success across a broad spectrum of computer vision tasks. However, their high computational cost and significant data requirements present challenges for deployment in resource-constrained environments. Current pruning methods for ViTs predominantly focus on reducing token counts, which often disrupt the inherent spatial structure of ViTs, hindering their adaptability to hardware. Furthermore, how to compress ViTs efficiently in few-shot scenarios remains an open question. Hence, we introduce a fine-grained block pruning framework for ViTs, named FBP-ViT. Unlike traditional block pruning techniques that indiscriminately remove entire blocks, FBP-ViT selectively eliminates Multi-Head Self-Attention (MSA) or Multi-Layer Perceptron (MLP) blocks, offering enhanced flexibility and efficiency. We unify pruning and finetuning, ensuring practicality in resource-constrained environments for real-world applications. We evaluate the proposed FBP-ViT across ViTs of varying sizes and architectures, demonstrating its effectiveness in improving computational efficiency while maintaining high performance. Specifically, with a speedup factor of 1.34, FBP-ViT preserves 80.39% top-1 accuracy on ImageNet-1k using DeiT-Base, achieving more precise and efficient pruning with tiny sets. Yilin Wang 0028, Qiang Dong, Dongyang Zhang 0001, Tao He 0007 |
ICMR | 5 |
| 2025 | Unbiased multimodal intent recognition with auxiliary rationale generation
Ruiting Dai, Guiduo Duan, Ke Qin, Tao He 0007 |
Neurocomputing | 6 |
| 2025 | VQA and Visual Reasoning: An overview of approaches, datasets, and future direction
Rufai Yusuf Zakari, Jim Wilson Owusu, Ke Qin, Zaharaddeen Karami Lawal, Tao He 0007 |
Neurocomputing | 6 |
| 2025 | A unified cross-source context enhancement model for multi-source fake news detection
Ruiting Dai, Haoran Meng, Zhengdao Yuan, Lisi Mo, Wenwei Zhu, Tao He 0007 |
Knowl. Based Syst. | 6 |
| 2024 | Towards Effective Data-Free Knowledge Distillation via Diverse Diffusion AugmentationabstractData-free knowledge distillation (DFKD) has emerged as a pivotal technique in the domain of model compression, substantially reducing the dependency on the original training data. Nonetheless, conventional DFKD methods that employ synthesized training data are prone to the limitations of inadequate diversity and discrepancies in distribution between the synthesized and original datasets. To address these challenges, this paper introduces an innovative approach to DFKD through diverse diffusion augmentation (DDA). Specifically, we revise the paradigm of common data synthesis in DFKD to a composite process through leveraging diffusion models subsequent to data synthesis for self-supervised augmentation, which generates a spectrum of data samples with similar distributions while retaining controlled variations. Furthermore, to mitigate excessive deviation in the embedding space, we introduce an image filtering technique grounded in cosine similarity to maintain fidelity during the knowledge distillation process. Comprehensive experiments conducted on CIFAR-10, CIFAR-100, and Tiny-ImageNet datasets showcase the superior performance of our method across various teacher-student network configurations, outperforming the contemporary state-of-the-art DFKD methods. Code will be available at: https://github.com/SLGSP/DDA. Muquan Li, Dongyang Zhang 0001, Tao He 0007, Xiurui Xie, Yuan-Fang Li, Ke Qin |
ACM Multimedia | 3 |
| 2024 | Towards Open-vocabulary HOI Detection with Calibrated Vision-language Models and Locality-aware QueriesabstractThe open-vocabulary human-object interaction (Ov-HOI) detection aims to identify both base and novel categories of human-object interactions while only base categories are available during training. Existing Ov-HOI methods commonly leverage knowledge distilled from CLIP to extend their ability to detect previously unseen interaction categories. However, our empirical observations indicate that the inherent noise present in CLIP has a detrimental effect on HOI prediction. Moreover, the absence of novel human-object position distributions often leads to overfitting on the base categories within their learned queries. To address these issues, we propose a two-step framework named, CaM-LQ, Calibrating visual-language Models, (e.g., CLIP) for open-vocabulary HOI detection with Locality-aware Queries. By injecting the fine-grained HOI supervision from the calibrated CLIP into the HOI decoder, our model can achieve the goal of predicting novel interactions. Extensive experimental results demonstrate that our approach performs well in open-vocabulary human-object interaction detection, surpassing state-of-the-art methods across multiple metrics on mainstream datasets and showing superior open-vocabulary HOI detection performance, e.g., with 4.54 points improvement on the HICO-DET dataset over the SoTA CLIP4HOI on the UV task with the same backbone ResNet-50. Zhenhao Yang, Deqiang Ouyang, Guiduo Duan, Dongyang Zhang 0001, Tao He 0007, Yuan-Fang Li |
ACM Multimedia | 6 |
| 2023 | Transferable and differentiable discrete network embedding for multi-domains with hierarchical knowledge distillation
Tao He 0007, Lianli Gao, Jingkuan Song, Yuan-Fang Li |
Inf. Sci. | 1 |
| 2023 | State-Aware Compositional Learning Toward Unbiased Training for Scene Graph GenerationabstractHow to avoid biased predictions is an important and active research question in scene graph generation (SGG). Current state-of-the-art methods employ debiasing techniques such as resampling and causality analysis. However, the role of intrinsic cues in the features causing biased training has remained under-explored. In this paper, for the first time, we make the surprising observation that object identity information, in the form of object label embeddings (e.g. GLOVE), is principally responsible for biased predictions. We empirically observe that, even without any visual features, a number of recent SGG models can produce comparable or even better results solely from object label embeddings. Motivated by this insight, we propose to leverage a conditional variational auto-encoder to decouple the entangled visual features into two meaningful components: the object's intrinsic identity features and the extrinsic, relation-dependent state feature. We further develop two compositional learning strategies on the relation and object levels to mitigate the data scarcity issue of rare relations. On the two benchmark datasets Visual Genome and GQA, we conduct extensive experiments on the three scenarios, i.e., conventional, few-shot and zero-shot SGG. Results consistently demonstrate that our proposed Decomposition and Composition (DeC) method effectively alleviates the biases in the relation prediction. Moreover, DeC is model-free, and it significantly improves the performance of recent SGG models, establishing new state-of-the-art performance. Tao He 0007, Lianli Gao, Jingkuan Song, Yuan-Fang Li |
IEEE Trans. Image Process. | 1 |
| 2023 | Toward a Unified Transformer-Based Framework for Scene Graph Generation and Human-Object Interaction DetectionabstractScene graph generation (SGG) and human-object interaction (HOI) detection are two important visual tasks aiming at localising and recognising relationships between objects, and interactions between humans and objects, respectively. Prevailing works treat these tasks as distinct tasks, leading to the development of task-specific models tailored to individual datasets. However, we posit that the presence of visual relationships can furnish crucial contextual and intricate relational cues that significantly augment the inference of human-object interactions. This motivates us to think if there is a natural intrinsic relationship between the two tasks, where scene graphs can serve as a source for inferring human-object interactions. In light of this, we introduce SG2HOI+, a unified one-step model based on the Transformer architecture. Our approach employs two interactive hierarchical Transformers to seamlessly unify the tasks of SGG and HOI detection. Concretely, we initiate a relation Transformer tasked with generating relation triples from a suite of visual features. Subsequently, we employ another transformer-based decoder to predict human-object interactions based on the generated relation triples. A comprehensive series of experiments conducted across established benchmark datasets including Visual Genome, V-COCO, and HICO-DET demonstrates the compelling performance of our SG2HOI+ model in comparison to prevalent one-stage SGG models. Remarkably, our approach achieves competitive performance when compared to state-of-the-art HOI methods. Additionally, we observe that our SG2HOI+ jointly trained on both SGG and HOI tasks in an end-to-end manner yields substantial improvements for both tasks compared to individualized training paradigms. Tao He 0007, Lianli Gao, Jingkuan Song, Yuan-Fang Li |
IEEE Trans. Image Process. | 1 |
| 2023 | Semisupervised Network Embedding With Differentiable Deep QuantizationabstractLearning accurate low-dimensional embeddings for a network is a crucial task as it facilitates many downstream network analytics tasks. For large networks, the trained embeddings often require a significant amount of space to store, making storage and processing a challenge. Building on our previous work on semisupervised network embedding, we develop d-SNEQ, a differentiable DNN-based quantization method for network embedding. d-SNEQ incorporates a rank loss to equip the learned quantization codes with rich high-order information and is able to substantially compress the size of trained embeddings, thus reducing storage footprint and accelerating retrieval speed. We also propose a new evaluation metric, path prediction, to fairly and more directly evaluate the model performance on the preservation of high-order information. Our evaluation on four real-world networks of diverse characteristics shows that d-SNEQ outperforms a number of state-of-the-art embedding methods in link prediction, path prediction, node classification, and node recommendation while being far more space- and time-efficient. Tao He 0007, Lianli Gao, Jingkuan Song, Yuan-Fang Li |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2022 | Towards Open-Vocabulary Scene Graph Generation with Prompt-Based Finetuning
Tao He 0007, Lianli Gao, Jingkuan Song, Yuan-Fang Li |
ECCV (28) | 1 |
| 2021 | Exploiting Scene Graphs for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection is a fundamental visual task aiming at localizing and recognizing interactions between humans and objects. Existing works focus on the visual and linguistic features of the humans and objects. However, they do not capitalise on the high-level and semantic relationships present in the image, which provides crucial contextual and detailed relational knowledge for HOI inference. We propose a novel method to exploit this information, through the scene graph, for the Human-Object Interaction (SG2HOI) detection task. Our method, SG2HOI, incorporates the SG information in two ways: (1) we embed a scene graph into a global context clue, serving as the scene-specific environmental context; and (2) we build a relation-aware message-passing module to gather relationships from objects' neighborhood and transfer them into interactions. Empirical evaluation shows that our SG2HOI method outperforms the state-of-the-art methods on two benchmark HOI datasets: V-COCO and HICO-DET. Code will be available at https://github.com/ht014/SG2HOI. Tao He 0007, Lianli Gao, Jingkuan Song, Yuan-Fang Li |
ICCV | 1 |
| 2021 | Exploring students' digital informal learning: the roles of digital competence and DTPB factorsabstractLearning approaches enhanced by digital technologies have gained considerable attention in higher education. Prior research has mainly focused on digital technology adoption in formal learning settings of higher education. Yet, empirical research into discerning what influences a student’s digital informal learning has not been well investigated. This paper proposes that individuals’ digital competence affects their digital informal learning (DIL) intention and actual behaviour. To understand better learners’ DIL and the effects of digital competence, we integrated digital competence into the decomposed theory of planned behaviour (DTPB) model and tested the model by using survey data from university students in Belgium. Here, we have explored different aspects of their DIL behaviours of cognitive, meta-cognitive, social and motivational learning. The findings showed both attitudinal factors in DTPB and digital competence adequately explained students’ DIL. Finally, the roles of digital competence and other DPTB factors are discussed in students’ DIL. Tao He 0007, Qionghao Huang, Shihua Li 0004 |
Behav. Inf. Technol. | 1 |
| 2020 | SNEQ: Semi-Supervised Attributed Network Embedding with Attention-Based QuantisationabstractLearning accurate low-dimensional embeddings for a network is a crucial task as it facilitates many network analytics tasks. Moreover, the trained embeddings often require a significant amount of space to store, making storage and processing a challenge, especially as large-scale networks become more prevalent. In this paper, we present a novel semi-supervised network embedding and compression method, SNEQ, that is competitive with state-of-art embedding methods while being far more space- and time-efficient. SNEQ incorporates a novel quantisation method based on a self-attention layer that is trained in an end-to-end fashion, which is able to dramatically compress the size of the trained embeddings, thus reduces storage footprint and accelerates retrieval speed. Our evaluation on four real-world networks of diverse characteristics shows that SNEQ outperforms a number of state-of-the-art embedding methods in link prediction, node classification and node recommendation. Moreover, the quantised embedding shows a great advantage in terms of storage and time compared with continuous embeddings as well as hashing methods. Tao He 0007, Lianli Gao, Jingkuan Song, Xin Wang 0019, Kejie Huang, Yuanfang Li |
AAAI | 1 |
| 2020 | Learning from the Scene and Borrowing from the Rich: Tackling the Long Tail in Scene Graph GenerationabstractDespite the huge progress in scene graph generation in recent years, its long-tail distribution in object relationships remains a challenging and pestering issue. Existing methods largely rely on either external knowledge or statistical bias information to alleviate this problem. In this paper, we tackle this issue from another two aspects: (1) scene-object interaction aiming at learning specific knowledge from a scene via an additive attention mechanism; and (2) long-tail knowledge transfer which tries to transfer the rich knowledge learned from the head into the tail. Extensive experiments on the benchmark dataset Visual Genome on three tasks demonstrate that our method outperforms current state-of-the-art competitors. Our source code is available at https://github.com/htlsn/issg. Tao He 0007, Lianli Gao, Jingkuan Song, Jianfei Cai 0001, Yuan-Fang Li |
IJCAI | 1 |
| 2020 | Unified Binary Generative Adversarial Network for Image Retrieval and Compression
Jingkuan Song, Tao He 0007, Lianli Gao, Xing Xu 0001, Alan Hanjalic, Heng Tao Shen |
Int. J. Comput. Vis. | 2 |
| 2020 | Exam paper generation based on performance prediction of student group
Zhengyang Wu 0001, Tao He 0007, Chenjie Mao, Changqin Huang |
Inf. Sci. | 2 |
| 2019 | One Network for Multi-Domains: Domain Adaptive Hashing with Intersectant Generative Adversarial NetworksabstractWith the recent explosive increase of digital data, image recognition and retrieval become a critical practical application. Hashing is an effective solution to this problem, due to its low storage requirement and high query speed. However, most of past works focus on hashing in a single (source) domain. Thus, the learned hash function may not adapt well in a new (target) domain that has a large distributional difference with the source domain. In this paper, we explore an end-to-end domain adaptive learning framework that simultaneously and precisely generates discriminative hash codes and classifies target domain images. Our method encodes two domains images into a semantic common space, followed by two independent generative adversarial networks arming at crosswise reconstructing two domains’ images, reducing domain disparity and improving alignment in the shared space. We evaluate our framework on four public benchmark datasets, all of which show that our method is superior to the other state-of-the-art methods on the tasks of object recognition and image retrieval. Tao He 0007, Yuan-Fang Li, Lianli Gao, Dongxiang Zhang, Jingkuan Song |
IJCAI | 1 |
| 2018 | Binary Generative Adversarial Networks for Image RetrievalabstractThe most striking successes in image retrieval using deep hashing have mostly involved discriminative models, which require labels. In this paper, we use binary generative adversarial networks (BGAN) to embed images to binary codes in an unsupervised way. By restricting the input noise variable of generative adversarial networks (GAN) to be binary and conditioned on the features of each input image, BGAN can simultaneously learn a binary representation per image, and generate an image plausibly similar to the original one. In the proposed framework, we address two main problems: 1) how to directly generate binary codes without relaxation? 2) how to equip the binary representation with the ability of accurate image retrieval? We resolve these problems by proposing new sign-activation strategy and a loss function steering the learning process, which consists of new models for adversarial loss, a content loss, and a neighborhood structure loss. Experimental results on standard datasets (CIFAR-10, NUSWIDE, and Flickr) demonstrate that our BGAN significantly outperforms existing hashing methods by up to 107% in terms of mAP (See Table 2). Jingkuan Song, Tao He 0007, Lianli Gao, Xing Xu 0001, Alan Hanjalic, Heng Tao Shen |
AAAI | 2 |
| 2018 | Deep Region Hashing for Generic Instance Search from ImagesabstractInstance Search (INS) is a fundamental problem for many applications, while it is more challenging comparing to traditional image search since the relevancy is defined at the instance level. Existing works have demonstrated the success of many complex ensemble systems that are typically conducted by firstly generating object proposals, and then extracting handcrafted and/or CNN features of each proposal for matching. However, object bounding box proposals and feature extraction are often conducted in two separated steps, thus the effectiveness of these methods collapses. Also, due to the large amount of generated proposals, matching speed becomes the bottleneck that limits its application to large-scale datasets. To tackle these issues, in this paper we propose an effective and efficient Deep Region Hashing (DRH) approach for large-scale INS using an image patch as the query. Specifically, DRH is an end-to-end deep neural network which consists of object proposal, feature extraction, and hash code generation. DRH shares full-image convolutional feature map with the region proposal network, thus enabling nearly cost-free region proposals. Also, each high-dimensional, real-valued region features are mapped onto a low-dimensional, compact binary codes for the efficient object region level matching on large-scale dataset. Experimental results on four datasets show that our DRH can achieve even better performance than the state-of-the-arts in terms of mAP, while the efficiency is improved by nearly 100 times. Jingkuan Song, Tao He 0007, Lianli Gao, Xing Xu 0001, Heng Tao Shen |
AAAI | 2 |
| 2017 | Deep Discrete Hashing with Self-supervised Pairwise Labels
Jingkuan Song, Tao He 0007, Hangbo Fan, Lianli Gao |
ECML/PKDD (1) | 2 |