Yanbin Hao

dblp:96/1538 · DBLP profile ↗
← Back
82ranked-venue papers
12as first author
70since 2021 · last 2026
0000-0002-0695-1566ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 64 · 10 first-author · 58 since 2021Artificial intelligence and machine learning · 32 · 2 first-author · 29 since 2021Computer networks · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Res-Bench: Benchmarking the Robustness of Multimodal Large Language Models to Dynamic Resolution Input
abstract
Multimodal Large Language Models (MLLMs) increasingly support dynamic image resolutions. However, current evaluation paradigms primarily assess semantic performance, overlooking the critical question of resolution robustness - whether performance remains stable across varying input resolutions. To address this gap, we introduce Res-Bench, a comprehensive benchmark comprising 14,400 samples across 12 resolution levels and six core capability dimensions. We designed a novel evaluation framework that goes beyond traditional accuracy metrics to capture performance stability. This framework introduces multiple robustness metrics: Spearman's correlation for assessing resolution-performance trends, and Absolute/Relative Continuous Error (ACE/RCE) for measuring performance volatility. Using these metrics, we conducted a large-scale evaluation of leading MLLMs. Our analysis encompasses: (1) model-centric and task-centric robustness examination, (2) investigation of preprocessing strategies including padding and super-resolution, and (3) exploration of fine-tuning for stability enhancement.
Chenxu Li, Zhicai Wang, Yuan Sheng, Yanbin Hao, Xiang Wang 0010
AAAI5
2026 Accelerating Controllable Generation via Hybrid-grained Cache
abstract
Controllable generative models have been widely used to improve the realism of synthetic visual content. However, such models must handle control conditions and content generation computational requirements, resulting in generally low generation efficiency. To address this issue, we propose a Hybrid-Grained Cache (HGC) approach that reduces computational overhead by adopting cache strategies with different granularities at different computational stages. Specifically, (1) we use a coarse-grained cache (block-level) based on feature reuse to dynamically bypass redundant computations in encoder-decoder blocks between each step of model reasoning. (2) We design a fine-grained cache (prompt-level) that acts within a module, where the fine-grained cache reuses cross-attention maps within consecutive reasoning steps and extends them to the corresponding module computations of adjacent steps. These caches of different granularities can be seamlessly integrated into each computational link of the controllable generation process. We verify the effectiveness of HGC on four benchmark datasets, especially its advantages in balancing generation efficiency and visual quality. For example, on the COCO-Stuff segmentation benchmark, our HGC significantly reduces the computational cost (MACs) by 63% (from 18.22T → 6.70T↓), while keeping the loss of semantic fidelity (quantized performance degradation) within 1.5%.
Huixia Ben, Shuo Wang 0008, Jinda Lu, Junxiang Qiu, Shengeng Tang, Yanbin Hao
AAAI7
2026 SNS-Grasp: Semantic-guided Noise Scaling for Grasp Generation
abstract
While diffusion models show promise for intent-based grasp generation, their isotropic noise schedules struggle with joint-specific sensitivity and task-aware variability. This limitation leads to grasps with suboptimal semantic alignment or physical feasibility. To address this challenge, we propose Semantic-guided Noise Scaling for grasp generation (SNS-Grasp), a novel framework that integrates two key innovations. First, the Semantic-guided Noise Scaling Diffusion (SNS-Diff) module generates intent-aware grasps by replacing isotropic noise with anisotropic modulation, dynamically adapting to task semantics and joint-specific sensitivity. Specifically, SNS-Diff leverages a pretrained Intent Recognizer to extract task-aware confidence scores and joint-specific gradient sensitivities from the interaction context. These signals adjust the noise scaling during denoising, downweighting perturbations for semantically critical joints to ensure semantic alignment. Second, the Fine-grained Grasp Refinement (FGR) module establishes dynamic joint-vertex coupling through fine-grained hand-object spatial relationships, enabling iterative optimization of physically executable grasps. Extensive experiments on OakInk and GRAB demonstrate SNS-Grasp's superior performance in semantic accuracy and physical feasibility, with robust generalization to unseen objects.
Zhenhua Tang 0001, Yudian Zheng, Yuzhang Zhong, Haolun Li 0001, Yanbin Hao, Chi-Man Pun
AAAI5
2026 Selective Volume Mixup for Video Action Recognition
Yi Tan 0001, Zhaofan Qiu, Yanbin Hao, Ting Yao 0003, Tao Mei 0001
Int. J. Comput. Vis.3
2026 Modeling Long-Term Emotional Support Through Causal World Modeling With Imitation Learning
abstract
Emotional support conversation systems have emerged as a promising complement to traditional mental health consultations, offering context-aware dialogue to support seekers’ emotional well-being. Despite their potential, two fundamental challenges remain unresolved: 1) modeling long-term emotional trajectory beyond short-term relief; and 2) adapting support strategies to context in a psychologically coherent manner. To address these challenges, we propose CAIWO, a novel framework that integrates world modeling and causality-enhanced imitation learning to systematically support seekers, alleviate psychological stress, and restore emotional balance. Specifically, CAIWO comprises two core components. The first is an emotional world model, which captures long-term emotional trajectories from historical interactions to inform anticipatory guidance. The second is a causality-enhanced imitation learning module, which infers latent causal dependencies to facilitate coherent strategy transitions and mitigate compounding errors typical of conventional imitation learning. By incorporating the final latent variables into the response decoder, CAIWO dynamically adjusts the strategies and generates emotionally resonant responses. Extensive experiments on the ESConv benchmark demonstrate that CAIWO outperforms state-of-the-art baselines by 8.6%, significantly improving the generation of responses that align with seekers’ emotional development and psychological needs.
Mingzheng Li, Fei Wang 0073, Kun Li 0008, Yanyan Wei, Yiqi Nie, Yanbin Hao, Xun Yang 0001, Meng Wang 0001
IEEE Trans. Comput. Soc. Syst.6
2026 Hybrid Granularity Distribution Estimation for Few-Shot Learning: Statistics Transfer From Categories and Instances
abstract
Distribution estimation is a pivotal strategy in few-shot learning (FSL) to mitigate data scarcity by sampling from estimated distributions, utilizing statistical properties (mean and variance) transferred from related base categories. However, category-level estimation alone often fails to generate representative samples due to significant dissimilarities between base and novel categories, leading to suboptimal performance. To address this limitation, we propose Hybrid Granularity Distribution Estimation (HGDE), which integrates both coarse-grained category-level statistics and fine-grained instance-level statistics. By leveraging instance statistics from the nearest base samples, HGDE enhances the characterization of novel categories, capturing subtle features that category-level estimation overlooks. These statistics are fused through linear interpolation to form a robust distribution for novel categories, ensuring both diversity and representativeness in generated samples. Additionally, HGDE employs refined estimation techniques, such as weighted summation for mean calculation and principal component retention for covariance, to further improve accuracy. Empirical evaluations on four FSL benchmarks, including Mini-ImageNet, Tiered-ImageNet, CUB and CIFAR-FS, demonstrate that HGDE offers effective distribution estimation capabilities and leads to notable accuracy gains, with improvements of more than 1.8% in 1-shot tasks on CUB. These results highlight HGDE's ability to balance mean precision and variance diversity, making it a versatile and effective solution for FSL.
Shuo Wang 0008, Tianyu Qi, Yanbin Hao, Beier Zhu, Hanwang Zhang, Meng Wang 0001
IEEE Trans. Image Process.4
2026 PointTFA$^{m}$: Multi-Modal, Training-Free Adaptation for Point Cloud Understanding
Jinmeng Wu, Youxiang Hu, Hao Zhang 0047, Basura Fernando, Yanbin Hao, Hanyu Hong
IEEE Trans. Multim.6
2026 CookingDiffusion: Cooking Procedural Image Generation with Stable Diffusion
abstract
Recent advancements in text-to-image generation models have excelled in creating diverse and realistic images. This success extends to food imagery, where various conditional inputs like cooking styles, ingredients, and recipes are utilized. However, a yet-unexplored challenge is generating a sequence of procedural images based on cooking steps from a recipe. This could enhance the cooking experience with visual guidance and possibly lead to an intelligent cooking simulation system. To fill this gap, we introduce a novel task called cooking procedural image generation . This task is inherently demanding, as it strives to create photo-realistic images that align with cooking steps while preserving sequential consistency. To collectively tackle these challenges, we present CookingDiffusion , a novel approach that leverages Stable Diffusion and three innovative Memory Nets to model procedural prompts. These prompts encompass text prompts (representing cooking steps), image prompts (corresponding to cooking images), and multi-modal prompts (mixing cooking steps and images), ensuring the consistent generation of cooking procedural images. To validate the effectiveness of our approach, we pre-process the YouCookII dataset, establishing a new benchmark. Our experimental results demonstrate that our model excels at generating high-quality cooking procedural images with remarkable consistency across sequential cooking steps, as measured by both the FID and the proposed Average Procedure Consistency metrics. Furthermore, CookingDiffusion demonstrates the ability to manipulate ingredients and cooking methods in a recipe. We will make our code, models, and dataset publicly accessible.
Bin Zhu 0006, Yanbin Hao, Chong-Wah Ngo, Yi Tan 0001, Xiang Wang 0010
ACM Trans. Multim. Comput. Commun. Appl.3
2025 RAGG: Retrieval-Augmented Grasp Generation Model
abstract
Intent-based grasp generation inherently involves challenges such as manipulation ambiguity and modality gaps. To address these, we propose a novel Retrieval-Augmented Grasp Generation model (RAGG). Our key insight is that when humans manipulate new objects, they initially mimic the interaction patterns observed in similar objects, then progressively adjust hand-object contact. Consequently, we develop RAGG as a two-stage approach, encompassing retrieval-guided generation and structurally stable grasp refinement. In the first stage, we propose a Retrieval-Augmented Diffusion Model (ReDim), which identifies the most relevant interaction instance from a knowledge base to explicitly guide grasp generation, thereby mitigating ambiguity and bridging modality gaps to ensure semantically correct manipulation. In the second stage, we introduce a Progressive Refinement Network (PRN) with Kolmogorov-Arnold Network (KAN) layers to refine the generated coarse grasp, employing a Structural Similarity Index loss to constrain the spatial relationship between the hand and the object, thus ensuring the stability of the grasp. Extensive experiments on the OakInk and GRAB benchmarks demonstrate that RAGG achieves superior results compared to state-of-the-art approach, indicating not only better physical feasibility and controllability but also strong generalization and interpretability for unseen objects.
Zhenhua Tang 0001, Bin Zhu 0006, Yanbin Hao, Chong-Wah Ngo, Richang Hong
AAAI3
2025 Hand1000: Generating Realistic Hands from Text with Only 1, 000 Images
abstract
Text-to-image generation models have achieved remarkable advancements in recent years, aiming to produce realistic images from textual descriptions. However, these models often struggle with generating anatomically accurate representations of human hands. The resulting images frequently exhibit issues such as incorrect numbers of fingers, unnatural twisting or interlacing of fingers, or blurred and indistinct hands. These issues stem from the inherent complexity of hand structures and the difficulty in aligning textual descriptions with precise visual depictions of hands. To address these challenges, we propose a novel approach named Hand1000 that enables the generation of realistic hand images with target gesture using only 1,000 training samples. The training of Hand1000 is divided into three stages with the first stage aiming to enhance the model’s understanding of hand anatomy by using a pre-trained hand gesture recognition model to extract gesture representation. The second stage further optimizes text embedding by incorporating the extracted hand gesture representation, to improve alignment between the textual descriptions and the generated hand images. The third stage utilizes the optimized embedding to fine-tune the Stable Diffusion model to generate realistic hand images. In addition, we construct the first publicly available dataset specifically designed for text-to-hand image generation. Based on the existing hand gesture recognition dataset, we adopt advanced image captioning models and LLaMA3 to generate high-quality textual descriptions enriched with detailed gesture information. Extensive experiments demonstrate that Hand1000 significantly outperforms existing models in producing anatomically correct hand images while faithfully representing other details in the text, such as faces, clothing and colors.
Haozhuo Zhang, Yanbin Hao
AAAI4
2025 Precise, Fast, and Low-cost Concept Erasure in Value Space: Orthogonal Complement Matters
abstract
Recent success of text-to-image (T2I) generation and its increasing practical applications, enabled by diffusion models, require urgent consideration of erasing unwanted concepts, e.g., copyrighted, offensive, and unsafe ones, from the pre-trained models in a precise, timely, and low-cost manner. The twofold demand of concept erasure includes not only a precise removal of the target concept (i.e., erasure efficacy) but also a minimal change on non-target content (i.e., prior preservation), during generation. Existing methods face challenges in maintaining an effective balance between erasure efficacy and prior preservation, and they can be computationally costly. To improve, we propose a precise, fast, and low-cost concept erasure method, called Adaptive Vaule Decomposer (AdaVD), which is training-free. Our method is grounded in a classical linear algebraic operation of computing orthogonal complement, implemented in the value space of each cross-attention layer within the UNet of diffusion models. We design a shift factor to adaptively navigate the erasure strength, enhancing effectively prior preservation without sacrificing erasure efficacy. Extensive comparative experiments with both training-based and training-free state of the arts demonstrate that the proposed AdaVD excels in both single and multiple concept erasure, showing 2 to 10 times of improvement in prior preservation than the second best, meanwhile achieving the best or near best erasure efficacy. AdaVD supports a series of diffusion models and downstream image generation tasks, with code available on: https://github.com/WYuan1001/AdaVD.
Ouxiang Li, Tingting Mu, Yanbin Hao, Kuien Liu, Xiang Wang 0010, Xiangnan He 0001
CVPR4
2025 Improving Open-vocabulary Video Visual Relation Detection with Decomposed Prompt Learning and Relation Adjustment
abstract
Open-vocabulary video visual relation detection (VidVRD) expands the scope of detecting object relations in videos to include unseen categories. It marks considerable advancement in recognizing novel relations solely by training on a base set, thus extending the frontiers of automated video understanding. However, the performance of current methods on novel predicates remains significantly inferior to that on base categories. We attribute this discrepancy to two primary factors: (1) A significant task misalignment between the Visual Relation Detection (VRD) task and the pre-trained models’ visual feature extractors, which are often designed for tasks like video-text retrieval and image-text retrieval, resulting in poor generalization to the novel set. (2) The relatively small size and limited vocabulary of open-vocabulary datasets, which create a substantial gap between base and novel predicates. Consequently, text prompts trained on the base set fail to generalize effectively to the novel set. To address these issues, we propose two improvement measures: (1) We decompose base and novel relations into actional and spatial patterns and introduce an innovative text prompt learning method that leverages the shared patterns between base and novel relations. (2) We develop a relation probability adjustment mechanism that utilizes reliable base relation predictions to adjust the probabilities of relations in novel classes by considering their overlaps in either actional or spatial contents. Experimental results on the benchmark dataset demonstrate significant performance improvements.
Ming Pei, Yi Tan 0001, Yanbin Hao, Hao Zhang 0047, Jinmeng Wu, Basura Fernando, Xun Yang 0001
ICASSP3
2025 Accelerating Diffusion Transformer via Gradient-Optimized Cache
Junxiang Qiu, Shuo Wang 0008, Jinda Lu, Kezhou Chen, Yanbin Hao
ICCV6
2025 A Sanity Check for AI-generated Image Detection
abstract
With the rapid development of generative models, discerning AI-generated content has evoked increasing attention from both industry and academia. In this paper, we conduct a sanity check on whether the task of AI-generated image detection has been solved. To start with, we present Chameleon dataset, consisting of AI-generated images that are genuinely challenging for human perception. To quantify the generalization of existing methods, we evaluate 9 off-the-shelf AI-generated image detectors on Chameleon dataset. Upon analysis, almost all models misclassify AI-generated images as real ones. Later, we propose AIDE AI-generated Image DEtector with Hybrid Features, which leverages multiple experts to simultaneously extract visual artifacts and noise patterns. Specifically, to capture the high-level semantics, we utilize CLIP to compute the visual embedding. This effectively enables the model to discern AI-generated images based on semantics and contextual information. Secondly, we select the highest and lowest frequency patches in the image, and compute the low-level patchwise features, aiming to detect AI-generated images by low-level artifacts, for example, noise patterns, anti-aliasing effects. While evaluating on existing benchmarks, for example, AIGCDetectBenchmark and GenImage, AIDE achieves +3.5% and +4.6% improvements to state-of-the-art methods, and on our proposed challenging Chameleon benchmarks, it also achieves promising results, despite the problem of detecting AI-generated images remains far from being solved.
Shilin Yan, Ouxiang Li, Jiayin Cai, Yanbin Hao, Yao Hu 0002, Weidi Xie
ICLR4
2025 Improving Synthetic Image Detection Towards Generalization: An Image Transformation Perspective
abstract
With recent generative models facilitating photo-realistic image synthesis, the proliferation of synthetic images has also engendered certain negative impacts on social platforms, thereby raising an urgent imperative to develop effective detectors. Current synthetic image detection (SID) pipelines are primarily dedicated to crafting universal artifact features, accompanied by an oversight about SID training paradigm. In this paper, we re-examine the SID problem and identify two prevalent biases in current training paradigms, i.e., weakened artifact features and overfitted artifact features. Meanwhile, we discover that the imaging mechanism of synthetic images contributes to heightened local correlations among pixels, suggesting that detectors should be equipped with local awareness. In this light, we propose SAFE, a lightweight and effective detector with three simple image transformations. Firstly, for weakened artifact features, we substitute the down-sampling operator with the crop operator in image pre-processing to help circumvent artifact distortion. Secondly, for overfitted artifact features, we include ColorJitter and RandomRotation as additional data augmentations, to help alleviate irrelevant biases from color discrepancies and semantic differences in limited training samples. Thirdly, for local awareness, we propose a patch-based random masking strategy tailored for SID, forcing the detector to focus on local regions at training. Comparative experiments are conducted on an open-world dataset, comprising synthetic images generated by 26 distinct generative models. Our pipeline achieves a new state-of-the-art performance, with remarkable improvements of 4.5% in accuracy and 2.9% in average precision against existing methods. Our code is available at: https://github.com/Ouxiang-Li/SAFE.
Ouxiang Li, Jiayin Cai, Yanbin Hao, Yao Hu 0002, Fuli Feng
KDD (1)3
2025 UniSVG: A Unified Dataset for Vector Graphic Understanding and Generation with Multimodal Large Language Models
abstract
Unlike bitmap images, scalable vector graphics (SVG) maintain quality when scaled, frequently employed in computer vision and artistic design in the representation of SVG code. In this era of proliferating AI-powered systems, enabling AI to understand and generate SVG has become increasingly urgent. However, AI-driven SVG understanding and generation (U&G) remain significant challenges. SVG code, equivalent to a set of curves and lines controlled by floating-point parameters, demands high precision in SVG U&G. Besides, SVG generation operates under diverse conditional constraints, including textual prompts and visual references, which requires powerful multi-modal processing for condition-to-SVG transformation. Recently, the rapid growth of Multi-modal Large Language Models (MLLMs) have demonstrated capabilities to process multi-modal inputs and generate complex vector controlling parameters, suggesting the potential to address SVG U&G tasks within a unified model. To unlock MLLM's capabilities in the SVG area, we propose an SVG-centric dataset called UniSVG, comprising 525k data items, tailored for MLLM training and evaluation. To our best knowledge, it is the first comprehensive dataset designed for unified SVG generation (from textual prompts and images) and SVG understanding (color, category, usage, etc.). As expected, learning on the proposed dataset boosts open-source MLLMs' performance on various SVG U&G tasks, surpassing SOTA close-source MLLMs like GPT-4V. We release dataset, benchmark, weights, codes and experiment details on https://ryanlijinke.github.io/.
Jiarui Yu, Chenxing Wei, Hande Dong, Liangjing Yang, Zhicai Wang, Yanbin Hao
ACM Multimedia8
2025 Accelerating Diffusion Transformer via Error-Optimized Cache
abstract
Diffusion Transformer (DiT) is a crucial method for content generation. However, it needs a lot of time to sample. Many studies have attempted to use caching to reduce the time consumption of sampling. Existing caching methods accelerate generation by reusing DiT features from the previous time step and skipping calculations in the next, but they tend to locate and cache low-error modules without focusing on reducing caching-induced errors, resulting in a sharp decline in generated content quality when increasing caching intensity. To solve this problem, we propose the Error-Optimized Cache (EOC). This method introduces three key improvements: (1) Prior knowledge extraction: Extract and process the caching differences; (2) A judgment method for cache optimization: Determine whether certain caching steps need to be optimized; (3) Cache optimization: reduce caching errors. Experiments show that this algorithm significantly reduces the error accumulation caused by caching, especially excessive caching. On the ImageNet dataset, without substantially increasing the computational load, this method improves the FID↓ of the generated images when the rule-based model FORA has a caching level of 75%, 50%, and 25%, and the training-based model Learning-to-cache has a caching level of 22%. Specifically, the FID↓ values change from 30.454 to 21.690 (28.8%), from 6.857 to 5.821 (15.1%), from 3.870 to 3.692 (4.6%), and from 3.539 to 3.451 (2.5%) respectively. Code is available at https://github.com/qiujx0520/EOC_MM2025.git.
Junxiang Qiu, Shuo Wang 0008, Jinda Lu, Houcheng Jiang, Yanbin Hao
ACM Multimedia7
2025 Mixture of Multimodal Adapters for Sentiment Analysis
abstract
Kezhou Chen, Shuo Wang, Huixia Ben, Shengeng Tang, Yanbin Hao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kezhou Chen, Shuo Wang 0008, Huixia Ben, Shengeng Tang, Yanbin Hao
NAACL (Long Papers)5
2025 Cross-modal Feature Enhancement and Contrastive Alignment for Micro-gesture Recognition
Tuyun Shang, Yanbin Hao, Ming Pei, Huixia Ben
PRCV (7)2
2025 Prior Preserved Text-to-Image Personalization Without Image Regularization
abstract
The current state-of-the-art text-to-image (T2I) models have found numerous applications, driven by their ability to produce photorealistic images. Concept learning, as one notable application, aims to enable T2I models to generate personalized content and better enable users to create images according to their interests. Nevertheless, the process of concept learning often involves model fine-tuning, which in turn brings the potential risk of overfitting. Such overfitting causes the T2I model to have reduced output diversity and results in poor editability. To mitigate the overfitting problem, we introduce two simple yet effective designs, namely masked textual inversion (MaskTI) and text regularization (TextReg). MaskTI is a variant of vanilla textual inversion that forces the learnable identifier to only attend to the class descriptor. This modification can effectively reduce the overfitting to those uninterested backgrounds. TextReg regulates the fine-tuning of cross-attention modules with simple text prompts without identifiers, which avoids the usage of real images as the regularization prior. Our extensive experiments demonstrate that not only does our approach effectively protect prior knowledge but also has high editability for the personalized model.
Zhicai Wang, Ouxiang Li, Longhui Wei, Yanbin Hao, Xiang Wang 0010, Qi Tian 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 Cross-Modal Hashing via Diverse Instances Matching
abstract
Cross-modal hashing is a highly effective technique for searching relevant data across different modalities, owing to its low storage costs and fast similarity retrieval capability. While significant progress has been achieved in this area, prior investigations predominantly concentrate on a one-to-one feature alignment approach, where a singular feature is derived for similarity retrieval. However, the singular feature in these methods fails to adequately capture the varied multi-instance information inherent in the original data across disparate modalities. Consequently, the conventional one-to-one methodology is plagued by a semantic mismatch issue, as the rigid one-to-one alignment inhibits effective multi-instance matching. To address this issue, we propose a novel Diverse Instances Matching for Cross-modal Hashing (DIMCH), which explores the relevance between multiple instances in different modalities using a multi-instance learning algorithm. Specifically, we design a novel diverse instances learning module to extract a multi-feature set, which enables our model to capture detailed multi-instance semantics. To evaluate the similarity between two multi-feature sets, we adopt the smooth chamfer distance function, which enables our model to incorporate the conventional similarity retrieval structure. Moreover, to sufficiently exploit the supervised information from the semantic label, we adopt the weight cosine triplet loss as the objective function, which incorporates the multilevel similarity among the multi-labels into the training procedure and enables the model to mine the multi-label correlation effectively. Extensive experiments demonstrate that our diverse hashing embedding method achieves state-of-the-art performance in supervised cross-modal hashing retrieval tasks.
Junfeng Tu, Xueliang Liu, Zhen Huang 0006, Yanbin Hao, Richang Hong, Meng Wang 0001
IEEE Trans. Image Process.4
2025 CVLP-NaVD: Contrastive Visual-language Pre-training Models for Non-annotated Visual Description
abstract
Non-annotated visual description (NaVD) aims to describe generic visuals without human-annotated pairwise data. The generic visuals refer to images and videos. Existing works mainly focus on one specific visual modality, i.e., image or video. In this article, we propose a new framework for this task, which can directly be applied to both image and video with the pipeline unchanged. Essentially, it is a unified framework that flexibly adapts to images and videos. Recently, contrastive visual-language pre-training models (CVLPs) have experienced rapid development, demonstrating powerful abilities to align vision and language. To continuously leverage advanced CVLPs, our framework is designed to work well with general CVLPs. It can easily use image-language CVLPs for image input and switch to video-language CVLPs for video input. Specifically, we propose a CVLP-based framework for NaVD, named CVLP-NaVD. It follows the paradigm of adversarial learning, containing a generator and a discriminator. The generator takes an image or a video as input and produces a corresponding language description, while the discriminator evaluates the generated sentence for its naturalness in human-like language. Apart from the naturalness, CVLPs play a crucial role in enhancing the alignment between visual and language signals during generation. Particularly, we explore three rewarding strategies to compute the alignment score, including directly calculating cosine similarity (i.e., VL-cross), projecting visual embeddings into the textual domain (i.e., VL-project), and their combination (i.e., VL-mix). The three strategies are fully examined in different scenarios. Finally, we conduct extensive experiments with various unpaired and unsupervised setups in both image and video captioning tasks. The experimental results demonstrate that our CVLP-NaVD outperforms the state-of-the-art methods significantly.
Yanbin Hao, Jiarui Yu, Bin Zhu 0006, Shuo Wang 0008, Tong Xu 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2025 Mixed Attention and Channel Shift Transformer for Efficient Action Recognition
abstract
The practical use of the Transformer-based methods for processing videos is constrained by the high computing complexity. Although previous approaches adopt the spatiotemporal decomposition of 3D attention to mitigate the issue, they suffer from the drawback of neglecting the majority of visual tokens. This article presents a novel mixed attention operation that subtly fuses the random, spatial, and temporal attention mechanisms. The proposed random attention stochastically samples video tokens in a simple yet effective way, complementing other attention methods. Furthermore, since the attention operation concentrates on learning long-distance relationships, we employ the channel shift operation to encode short-term temporal characteristics. Our model can provide more comprehensive motion representations thanks to the amalgamation of these techniques. Experimental results show that the proposed method produces competitive action recognition results with low computational overhead on both large-scale and small-scale public video datasets.
Xiusheng Lu, Yanbin Hao, Lechao Cheng, Sicheng Zhao, Yutao Liu 0002, Mingli Song
ACM Trans. Multim. Comput. Commun. Appl.2
2025 A Unified Generative Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing is a highly effective and efficient method for information retrieval, enabling the search for correlated data across different modality databases using compact hash codes. Conventional cross-modal hashing typically uses separate model structures for each modality and aligns approximate continuous representations of the final hash codes. These approaches not only require specialized models for each modality but also introduce a gap between the discrete hash codes and their continuous features, yielding only approximate alignment. To address these issues, we propose a unified generative cross-modal hashing method that leverages a single Uniform Mixture-of-Expert Decoder (UMoED) for both image and text modalities. UMoED streamlines cross-modal hash learning by integrating two key design elements: (1) a cross-modal representation unification module that employs unified queries to consolidate modality-specific features into a common space, and (2) an adaptive expert enhancement module that adaptively enhances feature modeling based on the input modality. Furthermore, our decoder-based hashing method outputs hash codes in a generative manner, producing precise representations of discrete codes to bridge the gap between the discrete and continuous space, thus ensuring precise alignment during similarity learning. Extensive experiments on three benchmark datasets demonstrate that the proposed method achieves the state-of-the-art performance in cross-modal hashing retrieval.
Junfeng Tu, Xueliang Liu, Yanbin Hao, Richang Hong
ACM Trans. Multim. Comput. Commun. Appl.3
2025 Interventional Feature Generation for Few-shot Learning
abstract
Few-shot learning (FSL) aims to classify a novel object into a specific category under limited training samples. This is a challenging task since (1) the features expressed by pre-trained knowledge introduce perceived bias and then constrain the classification space, and (2) the use of general hallucination techniques based on global features fails to escape the limited classification space, resulting in sub-optimal improvements. To solve these issues, this article proposes an interventional feature generation (IFG) method. Specifically, we first use the relations of the categories or instances as interventional operations to implicitly constrain the feature representations (pre-trained knowledge) into different classification subsets. Then, we employ a parameter-free feature generation strategy to enrich each subset’s training samples of the support category. In other words, IFG provides a multi-subsets learning strategy to reduce the influence of perceived bias, enrich the diversity of generated features, and improve the robustness of the few-shot classifier. We apply our method to four benchmark datasets and observe state-of-the-art performance across all experiments. Specifically, compared to the baseline on the Mini-ImageNet dataset, our approach yields accuracy improvements of 6.03% and 3.46% for 1 and 5 support training samples, respectively. Furthermore, the proposed interventional feature generation technique can improve classifier performance in other FSL methods, demonstrating its versatility and potential for broader applications. The code is available at https://github.com/ShuoWangCS/IFG-FSL/ .
Shuo Wang 0008, Jinda Lu, Huixia Ben, Yanbin Hao, Xingyu Gao 0001, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 Boosting Few-Shot Learning via Attentive Feature Regularization
abstract
Few-shot learning (FSL) based on manifold regularization aims to improve the recognition capacity of novel objects with limited training samples by mixing two samples from different categories with a blending factor. However, this mixing operation weakens the feature representation due to the linear interpolation and the overlooking of the importance of specific channels. To solve these issues, this paper proposes attentive feature regularization (AFR) which aims to improve the feature representativeness and discriminability. In our approach, we first calculate the relations between different categories of semantic labels to pick out the related features used for regularization. Then, we design two attention-based calculations at both the instance and channel levels. These calculations enable the regularization procedure to focus on two crucial aspects: the feature complementarity through adaptive interpolation in related categories and the emphasis on specific feature channels. Finally, we combine these regularization strategies to significantly improve the classifier performance. Empirical studies on several popular FSL benchmarks demonstrate the effectiveness of AFR, which improves the recognition accuracy of novel categories without the need to retrain any feature extractor, especially in the 1-shot setting. Furthermore, the proposed AFR can seamlessly integrate into other FSL methods to improve classification performance.
Shuo Wang 0008, Jinda Lu, Yanbin Hao, Xiangnan He 0001
AAAI4
2024 GLCM-Adapter: Global-Local Content Matching for Few-shot CLIP Adaptation
Shuo Wang 0008, Xieenlong, Jinda Lu, Jinghan Li, Yanbin Hao
BMVC5
2024 Enhance Image Classification via Inter-Class Image Mixup with Diffusion Model
abstract
Text-to-image (T2I) generative models have recently emerged as a powerful tool, enabling the creation of photo-realistic images and giving rise to a multitude of applications. However, the effective integration of T2I models into fundamental image classification tasks remains an open question. A prevalent strategy to bolster image classification performance is through augmenting the training set with synthetic images generated by T2I models. In this study, we scrutinize the shortcomings of both current gener-ative and conventional data augmentation techniques. Our analysis reveals that these methods struggle to produce images that are both faithful (in terms of foreground objects) and diverse (in terms of background contexts) for domain-specific concepts. To tackle this challenge, we introduce an innovative inter-class data augmentation method known as Diff-Mix11https://github.com/Zhicaiwww/Diff-Mix, which enriches the dataset by performing image translations between classes. Our empirical results demonstrate that Diff-Mix achieves a better balance between faith-fulness and diversity, leading to a marked improvement in performance across diverse image classification scenarios, including few-shot, conventional, and long-tail classifications for domain-specific datasets.
Zhicai Wang, Longhui Wei, Heyu Chen, Yanbin Hao, Xiang Wang 0010, Xiangnan He 0001, Qi Tian 0001
CVPR5
2024 3D-GOI: 3D GAN Omni-Inversion for Multifaceted and Multi-object Editing
Haoran Li 0020, Haolin Shi, Yanbin Hao, Yong Liao 0003, Lechao Cheng, Peng Yuan Zhou
ECCV (62)4
2024 Enhancing Recipe Retrieval with Foundation Models: A Data Augmentation Perspective
Fangzhou Song, Bin Zhu 0006, Yanbin Hao, Shuo Wang 0008
ECCV (51)3
2024 Noise-NeRF: Hide Information in Neural Radiance Field Using Trainable Noise
abstract
Neural Radiance Field (NeRF) has been proposed as an innovative advancement in 3D reconstruction techniques. However, little research has been conducted on the issues of information confidentiality and security to NeRF, such as steganography. Existing NeRF steganography solutions have shortcomings in low steganography quality, model weight damage, and limited amount of steganographic information. This paper proposes Noise-NeRF, a novel NeRF steganography method employing Adaptive Pixel Selection strategy and Pixel Perturbation strategy to improve the quality and efficiency of steganography via trainable noise. Extensive experiments validate the state-of-the-art performances of Noise-NeRF on both steganography quality and rendering quality, as well as effectiveness in super-resolution image steganography.
Qinglong Huang, Haoran Li 0020, Yong Liao 0003, Yanbin Hao, Peng Yuan Zhou
ICANN (2)4
2024 PointTFA: Training-Free Clustering Adaption for Large 3D Point Cloud Models
Jinmeng Wu, Hao Zhang 0047, Basura Fernando, Yanbin Hao, Hanyu Hong
IJCAI5
2024 Selective Vision-Language Subspace Projection for Few-shot CLIP
abstract
Vision-language models such as CLIP are capable of mapping the different modality data into a unified feature space, enabling zero/few-shot inference by measuring the similarity of given images and texts. However, most existing methods overlook modality gaps in CLIP's encoded features, which is shown as the text and image features lie far apart from each other, resulting in limited classification performance. To tackle this issue, we introduce a method called Selective Vision-Language Subspace Projection (SSP), which incorporates local image features and utilizes them as a bridge to enhance the alignment between image-text pairs. Specifically, our SSP framework comprises two parallel modules: a vision projector and a language projector. Both projectors utilize local image features to span the respective subspaces for image and texts, thereby projecting the image and text features into their respective subspaces to achieve alignment. Moreover, our approach entails only training-free matrix calculations and can be seamlessly integrated into advanced CLIP-based few-shot learning frameworks. Extensive experiments on 11 datasets have demonstrated SSP's superior text-image alignment capabilities, outperforming the state-of-the-art alignment methods. The code is available at https://github.com/zhuhsingyuu/SSP
Beier Zhu, Yi Tan 0001, Shuo Wang 0008, Yanbin Hao, Hanwang Zhang
ACM Multimedia5
2024 Hierarchical Supervised Contrastive Learning for Multimodal Sentiment Analysis
Kezhou Chen, Shuo Wang 0008, Yanbin Hao
MMM (2)3
2024 Enhancing Zero-Shot Vision Models by Label-Free Prompt Distribution Learning and Bias Correcting
abstract
Vision-language models, such as CLIP, have shown impressive generalization capacities when using appropriate text descriptions. While optimizing prompts on downstream labeled data has proven effective in improving performance, these methods entail labor costs for annotations and are limited by their quality. Additionally, since CLIP is pre-trained on highly imbalanced Web-scale data, it suffers from inherent label bias that leads to suboptimal performance. To tackle the above challenges, we propose a label-**F**ree p**ro**mpt distribution **l**earning and b**i**as **c**orrection framework, dubbed as **Frolic**, which boosts zero-shot performance without the need for labeled data. Specifically, our Frolic learns distributions over prompt prototypes to capture diverse visual representations and adaptively fuses these with the original CLIP through confidence matching. This fused model is further enhanced by correcting label bias via a label-free logit adjustment. Notably, our method is not only training-free but also circumvents the necessity for hyper-parameter tuning. Extensive experimental results across 16 datasets demonstrate the efficacy of our approach, particularly outperforming the state-of-the-art by an average of $2.6\%$ on 10 datasets with CLIP ViT-B/16 and achieving an average margin of $1.5\%$ on ImageNet and its five distribution shifts with CLIP ViT-B/16. Codes are available in [https://github.com/zhuhsingyuu/Frolic](https://github.com/zhuhsingyuu/Frolic).
Beier Zhu, Yi Tan 0001, Shuo Wang 0008, Yanbin Hao, Hanwang Zhang
NeurIPS5
2024 JPA: A Joint-Part Attention for Mitigating Overfocusing on 3D Human Pose Estimation
Dengqing Yang, Zhenhua Tang 0001, Jinmeng Wu, Shuo Wang 0008, Lechao Cheng, Yanbin Hao
PRCV (6)6
2024 Space-View Decoupled 3D Gaussians for Novel-View Synthesis of Mirror Reflections
Zhenwu Wang, Zhuopeng Li, Zhenhua Tang 0001, Yanbin Hao, Huasen He
PRICAI (4)4
2024 Masked Collaborative Contrast for Weakly Supervised Semantic Segmentation
abstract
This study introduces an efficacious approach, Masked Collaborative Contrast (MCC), to highlight semantic regions in weakly supervised semantic segmentation. MCC adroitly draws inspiration from masked image modeling and contrastive learning to devise a novel framework that induces keys to contract toward semantic regions. Unlike prevalent techniques that directly eradicate patch regions in the input image when generating masks, we scrutinize the neighborhood relations of patch tokens by exploring masks considering keys on the affinity matrix. Moreover, we generate positive and negative samples in contrastive learning by utilizing the masked local output and contrasting it with the global output. Elaborate experiments on commonly employed datasets evidences that the proposed MCC mechanism effectively aligns global and local perspectives within the image, attaining impressive performance. The source code is available at https://github.com/fwu11/MCC.
Fangwen Wu, Jingxuan He 0001, Yufei Yin, Yanbin Hao, Gang Huang 0004, Lechao Cheng
WACV4
2024 PosMLP-Video: Spatial and Temporal Relative Position Encoding for Efficient Video Recognition
Yanbin Hao, Diansong Zhou, Zhicai Wang, Chong-Wah Ngo, Meng Wang 0001
Int. J. Comput. Vis.1
2024 Iterative Semantic Transformer by Greedy Distillation for Community Question Answering
abstract
The semantic matching problem consists of recognizing if the candidate text is relevant to a particular input text. Semantic similarities can be determined from human-curated knowledge, but such knowledge may not be available in every language. Instead, statistical learning techniques have been applied, but these techniques circumvent the need for manual feature engineering by using large datasets to train models to perform semantic similarity scoring between portions of text or words. The pre-trained transformer provides a further mechanism to consolidate the information throughout a sentence into single sentence-level representations, but these representations may not be optimal for the matching task. As an alternative, we propose an interactive semantic transformer based on a greedy layer-wise framework to learn a distributed similarity representation for sentence pairs. The novelty of the architecture lies in an abstract representation of the semantic similarities created by three-stage learning strategies. Model training is accomplished through a greedy layer-wise training scheme, that incorporates both supervised and unsupervised learning. The proposed model is experimentally compared to state-of-the-art approaches on three different dataset types: the library TREC, the Yahoo!, and Stack Exchange community question datasets, and results show the proposed model outperforming other approaches.
Jinmeng Wu, Tingting Mu, Jeyan Thiyagalingam, Hanyu Hong, Yanbin Hao, Tianxu Zhang, John Yannis Goulermas
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 When I Fall in Love: Capturing Video-Oriented Social Relationship Evolution via Attentive GNN
abstract
With the booming of streaming media platforms, viewers now get used to watching dramas and movies via online platforms with more intelligent services. Usually, character relationships may dynamically evolve with stories promoting in long videos. Therefore, automatic tools to capture the social relation evolution among characters are urgently required to enrich the viewing experience. However, most existing works mainly focus on shorter isolated video clips. Considering the development of the plot, they may fail to effectively summarize relationships as holistic semantic representations for the whole video. To deal with these challenges, in this paper, we propose a novel Dynamic-Evolutionary Graph Attention Network (DE-GAT) framework to generate the evolving social relation graph among characters and capture the characters’ relation evolutionary trajectory throughout the entire video. DE-GAT first integrates the multimodal cues, including visual and textual information in each video clip via the graph attention network (GAT). Expanding the temporal receptive field from clip-level to scenario-level, the most relevant factors of the evolution of social relationships can be explored. Eventually, all the scenario-level social graphs are merged to obtain the evolving global social graph for the entire movie. Extensive evaluations on the real-world MovieGraphs dataset have validated the positive impact of temporal receptive field expansion and multimodal cues on capturing evolving social relations.
Penggang Qin, Tong Xu 0001, Yanbin Hao, Fuli Feng, Chen Zhu 0003, Enhong Chen
IEEE Trans. Circuits Syst. Video Technol.4
2024 FTCM: Frequency-Temporal Collaborative Module for Efficient 3D Human Pose Estimation in Video
abstract
Capturing cross-pose correlation from a sequence of frame-level 2D poses is essential for 3D human pose estimation (3D-HPE) in the video. Recent studies have shown the promising potential of modeling the pose relation with feature-mixing operations on the temporal domain. However, they seldom consider the interaction across poses in the frequency domain. This paper studies a Frequency-Temporal Collaborative Module (FTCM) to explore the feasibility of encoding the cross-pose correlations in both frequency and temporal domains. FTCM aims to jointly capture the global and local cross-pose correlations with a more lightweight network model. Specifically, FTCM splits the pose features into two groups along the channel dimension and separately models the frequency and temporal interactions across poses with different feature-mixing operations in parallel. To achieve this goal, we purposely design two pose-mixing units, i.e., the frequency pose-mixing (FPM) and the temporal pose-mixing (TPM). Particularly, FPM is designed to reap the global correlations among different pose frequencies with the representation obtained by converting the original pose signals with Fast Fourier transform (FFT). Unlike the pose-mixing used by previous methods like Transformers that influences an individual pose with all other poses, TPM locally calibrates the pose with dynamics aggregated within several adjacent poses in the temporal domain, explicitly weighting neighboring poses more with respect to the far-away ones so as to enforce a strict locality constraint. Besides, the group strategy significantly reduces the model complexity. To verify the effectiveness of FTCM, we conduct extensive experiments on two benchmarks (i.e., Human3.6M and MPI-INF-3DHP). Experimental results not only exhibit favorable accuracy/complexity trade-offs of our FTCM but also show superior or comparable performance to state-of-the-art methods on both datasets. The code and model are publicly available at:https://github.com/zhenhuat/FTCM.
Zhenhua Tang 0001, Yanbin Hao, Jia Li 0013, Richang Hong
IEEE Trans. Circuits Syst. Video Technol.2
2024 Feature Mixture on Pre-Trained Model for Few-Shot Learning
abstract
Few-shot learning (FSL) aims at recognizing a novel object under limited training samples. A robust feature extractor (backbone) can significantly improve the recognition performance of the FSL model. However, training an effective backbone is a challenging issue since 1) designing and validating structures of backbones are time-consuming and expensive processes, and 2) a backbone trained on the known (base) categories is more inclined to focus on the textures of the objects it learns, which is hard to describe the novel samples. To solve these problems, we propose a feature mixture operation on the pre-trained (fixed) features: 1) We replace a part of the values of the feature map from a novel category with the content of other feature maps to increase the generalizability and diversity of training samples, which avoids retraining a complex backbone with high computational costs. 2) We use the similarities between the features to constrain the mixture operation, which helps the classifier focus on the representations of the novel object where these representations are hidden in the features from the pre-trained backbone with biased training. Experimental studies on five benchmark datasets in both inductive and transductive settings demonstrate the effectiveness of our feature mixture (FM). Specifically, compared with the baseline on the Mini-ImageNet dataset, it achieves 3.8% and 4.2% accuracy improvements for 1 and 5 training samples, respectively. Additionally, the proposed mixture operation can be used to improve other existing FSL methods based on backbone training.
Shuo Wang 0008, Jinda Lu, Haiyang Xu 0002, Yanbin Hao, Xiangnan He 0001
IEEE Trans. Image Process.4
2024 Efficient Unsupervised Video Hashing With Contextual Modeling and Structural Controlling
abstract
The most important effect of the video hashing technique is to support fast retrieval, which is benefiting from the high efficiency of binary calculation. Current video hash approaches are thus mainly targeted at learning compact binary codes to represent video content accurately. However, they may overlook the generation efficiency for hash codes, i.e., designing lightweight neural networks. This paper proposes anEfficientUnsupervisedVideoHashing (EUVH)method, which is not only for computing compact hash codes but also for designing a lightweight deep model. Specifically, we present an MLP-based model, where the video tensor is split into several groups and multiple axial contexts are explored to separately refine them in parallel. The axial contexts are referred to as the dynamics aggregated from different axial scales, including long/middle/short-range dependencies. The group operation significantly reduces the computational cost of the MLP backbone. Moreover, to achieve compact video hash codes, three structural losses are utilized. As demonstrated by the experiment, the three structures are highly complementary for approximating the real data structure. We conduct extensive experiments on three benchmark datasets for the unsupervised video hashing task and show the superior trade-off between performance and computational cost of our EUVH to the state of the arts.
Jingru Duan, Yanbin Hao, Bin Zhu 0006, Lechao Cheng, Peng Yuan Zhou, Xiang Wang 0010
IEEE Trans. Multim.2
2024 Two-Step Discrete Hashing for Cross-Modal Retrieval
abstract
Cross-modal hashing is an effective approach for information retrieval from large and heterogeneous cross-modal datasets, owing to its low storage cost and high computational speed. However, conventional cross-modal hashing techniques for generating hashing codes rely on cross-space dimensional compression, which results in two types of information loss: quantization information loss and dimension reduction loss. To address these limitations, we propose a novel method that decouples the one-step hashing (Fig.1a) strategy into two sub-steps (Fig.1b). Specifically, in the first step, we introduce a novel differentiable hash method, which utilizes a smooth hash module for binary quantization. This method allows our model to reduce the quantization information loss and make the model optimized by gradient descent. In the second step, we design a long-short Hamming space transformation approach to project the long code into a short one, which is effective in preserving the dimension information between long and short and mitigating the dimension reduction loss. We demonstrate the effectiveness of our approach through extensive experiments on several popular cross-modal datasets, achieving a significant improvement in cross-modal retrieval performance.
Junfeng Tu, Xueliang Liu, Yanbin Hao, Richang Hong, Meng Wang 0001
IEEE Trans. Multim.3
2023 How Can Contrastive Pre-training Benefit Audio-Visual Segmentation? A Study from Supervised and Zero-shot Perspectives
Jiarui Yu, Yanbin Hao, Jinmeng Wu, Tong Xu 0001, Shuo Wang 0008, Xiangnan He 0001
BMVC3
2023 3D Human Pose Estimation with Spatio-Temporal Criss-Cross Attention
abstract
Recent transformer-based solutions have shown great success in 3D human pose estimation. Nevertheless, to calculate the joint-to-joint affinity matrix, the computational cost has a quadratic growth with the increasing number of joints. Such drawback becomes even worse especially for pose estimation in a video sequence, which necessitates spatio-temporal correlation spanning over the entire video. In this paper, we facilitate the issue by decomposing correlation learning into space and time, and present a novel Spatio-Temporal Criss-cross attention (STC) block. Technically, STC first slices its input feature into two partitions evenly along the channel dimension, followed by performing spatial and temporal attention respectively on each partition. STC then models the interactions between joints in an identical frame and joints in an identical trajectory simultaneously by concatenating the outputs from attention layers. On this basis, we devise STCFormer by stacking multiple STC blocks and further integrate a new Structure-enhanced Positional Embedding (SPE) into STCFormer to take the structure of human body into consideration. The embedding function consists of two components: spatio-temporal convolution around neighboring joints to capture local structure, and part-aware embedding to indicate which part each joint belongs to. Extensive experiments are conducted on Human3.6M and MPI-INF-3DHP benchmarks, and superior results are reported when comparing to the state-of-the-art approaches. More remarkably, STCFormer achieves to-date the best published performance: 40.5mm P1 error on the challenging Human3.6M dataset.
Zhenhua Tang 0001, Zhaofan Qiu, Yanbin Hao, Richang Hong, Ting Yao 0003
CVPR3
2023 Bi-Directional Distribution Alignment for Transductive Zero-Shot Learning
abstract
Zero-shot learning (ZSL) suffers intensely from the domain shift issue, i.e., the mismatch (or misalignment) between the true and learned data distributions for classes without training data (unseen classes). By learning additionally from unlabelled data collected for the unseen classes, transductive ZSL (TZSL) could reduce the shift but only to a certain extent. To improve TZSL, we propose a novel approach Bi-VAEGAN which strengthens the distribution alignment between the visual space and an auxiliary space. As a result, it can reduce largely the domain shift. The proposed key designs include (1) a bi-directional distribution alignment, (2) a simple but effective L2-norm based feature normalization approach, and (3) a more sophisticated unseen class prior estimation. Evaluated by four benchmark datasets, Bi-VAEGAN11Code is available at https://github.com/Zhicaiwww/Bi-VAEGAN achieves the new state of the art under both the standard and generalized TZSL settings.
Zhicai Wang, Yanbin Hao, Tingting Mu, Ouxiang Li, Shuo Wang 0008, Xiangnan He 0001
CVPR2
2023 Semantic-based Selection, Synthesis, and Supervision for Few-shot Learning
abstract
Few-shot learning (FSL) is designed to explore the distribution of novel categories from a few samples. It is a challenging task since the classifier is usually susceptible to over-fitting when learning from limited training samples. To alleviate this phenomenon, a common solution is to achieve more training samples using a generic generation strategy in visual space. However, there are some limitations to this solution. It is because a feature extractor trained on base samples (known knowledge) tends to focus on the textures and structures of the objects it learns, which is inadequate for describing novel samples. To solve these issues, we introduce semantics and propose a Semantic-based Selection, Synthesis, and S upervision (4S) method, where semantics provide more diverse and informative supervision for recognizing novel objects. Specifically, we first utilize semantic knowledge to explore the correlation of categories in the textual space and select base categories related to the given novel category. This process can improve the efficiency of subsequent operations (synthesis and supervision). Then, we analyze the semantic knowledge to hallucinate the training samples by selectively synthesizing the contents from base and support samples. This operation not only increases the number of training samples but also takes advantage of the contents of the base categories to enhance the description of support samples. Finally, we also employ semantic knowledge as both soft and hard supervision to enrich the supervision for the fine-tuning procedure. Empirical studies on four FSL benchmarks demonstrate the effectiveness of 4S.
Jinda Lu, Shuo Wang 0008, Xinyu Zhang 0022, Yanbin Hao, Xiangnan He 0001
ACM Multimedia4
2023 CgT-GAN: CLIP-guided Text GAN for Image Captioning
abstract
The large-scale visual-language pre-trained model, Contrastive Language-Image Pre-training (CLIP), has significantly improved image captioning for scenarios without human-annotated image-caption pairs. Recent advanced CLIP-based image captioning without human annotations follows a text-only training paradigm, i.e., reconstructing text from shared embedding space. Nevertheless, these approaches are limited by the training/inference gap or huge storage requirements for text embeddings. Given that it is trivial to obtain images in the real world, we propose CLIP-guided text GAN (CgT-GAN), which incorporates images into the training process to enable the model to "see" real visual modality. Particularly, we use adversarial training to teach CgT-GAN to mimic the phrases of an external text corpus and CLIP-based reward to provide semantic guidance. The caption generator is jointly rewarded based on the caption naturalness to human language calculated from the GAN's discriminator and the semantic guidance reward computed by the CLIP-based reward module. In addition to the cosine similarity as the semantic guidance reward (i.e., CLIP-cos), we further introduce a novel semantic guidance reward called CLIP-agg, which aligns the generated caption with a weighted text embedding by attentively aggregating the entire corpus. Experimental results on three subtasks (ZS-IC, In-UIC and Cross-UIC) show that CgT-GAN outperforms state-of-the-art methods significantly across all metrics. Code is available at https://github.com/Lihr747/CgtGAN.
Jiarui Yu, Yanbin Hao, Bin Zhu 0006, Tong Xu 0001, Xiangnan He 0001
ACM Multimedia3
2023 Question-aware dynamic scene graph of local semantic representation learning for visual question answering
Jinmeng Wu, Fulin Ge, Hanyu Hong, Yu Shi 0004, Yanbin Hao, Lei Ma 0004
Pattern Recognit. Lett.5
2023 MLP-JCG: Multi-Layer Perceptron With Joint-Coordinate Gating for Efficient 3D Human Pose Estimation
abstract
Various structural relations/dependencies exist among human body joints, which makes it possible to estimate 3D poses from 2D sources. The current research on 3D human pose estimation (3D-HPE for short) mainly focuses on structural information from a specific perspective. However, this information cannot facilitate 2D-to-3D pose lifting. This paper presents a novel and efficient multi-layer perceptron with a joint-coordinate gating (MLP-JCG) model, exploring and utilizing both the local and global structural information to perform 3D pose estimations. Specifically, MLP-JCG contains two independent MLP blocks, i.e., joint-mixing MLP and coordinate-mixing MLP, which solely act on the joint and coordinate axes in modelling their local structural information. For the global structural information, we first explore two kinds of global statistics from the pose matrix embeddings, which are referred to as the dynamics aggregated along the joint/coordinate axis. Then, we propose two kinds of gating units to elementwisely contextualize the features learned from MLP blocks. All the model components are designed based on MLP, making the MLP-JCG easy to implement and train. We conduct experiments on three 3D-HPE benchmarks, and the results demonstrate the superior effectiveness and efficiency of the proposed approach.
Zhenhua Tang 0001, Jia Li 0013, Yanbin Hao, Richang Hong
IEEE Trans. Multim.3
2023 Boosting Hyperspectral Image Classification with Dual Hierarchical Learning
abstract
Hyperspectral image (HSI) classification aims at predicting the pixel-wise labels in an image, where there are only a few labeled pixel samples (hard labels) for training. It is a challenging task since the classification process is susceptible to over-fitting under training with limited samples. To relieve this problem, we propose a method based on dual hierarchical learning. First, we employ a connectionist hyperspectral convolution (HC) network to capture the representations of the pixels from different receptive fields. Specifically, an HC is designed to learn the correlation among adjacent pixels and is further extended to a connectionist hierarchical structure. These operations use the correlation to enhance one-pixel learning from multiple receptive fields. Second, we analyze the properties in the hyperspectral image and introduce a hierarchical pseudo label generation algorithm to enrich the supervision of the label information. Finally, we design a dual hierarchical learning strategy to help all HC layers learn from both the hard labels and the hierarchical pseudo labels. In other words, it addresses the HSI classification problem from different views. For inference, we employ two fusion strategies to find a better prediction. The experimental results on four popular HSI benchmarks, i.e., Salinas-A, IndianPines, PaviaU, and PaviaC, demonstrate the effectiveness of the proposed method. Our code is publicly available on GitHub: https://github.com/ShuoWangCS/HSI-DHL.
Shuo Wang 0008, Huixia Ben, Yanbin Hao, Xiangnan He 0001, Meng Wang 0001
ACM Trans. Multim. Comput. Commun. Appl.3
2022 Group Contextualization for Video Recognition
abstract
Learning discriminative representation from the complex spatio-temporal dynamic space is essential for video recognition. On top of those stylized spatio-temporal computational units, further refining the learnt feature with axial contexts is demonstrated to be promising in achieving this goal. However, previous works generally focus on utilizing a single kind of contexts to calibrate entire feature channels and could hardly apply to deal with diverse video activities. The problem can be tackled by using pair-wise spatio-temporal attentions to recompute feature response with cross-axis contexts at the expense of heavy computations. In this paper, we propose an efficient feature refinement method that decomposes the feature channels into several groups and separately refines them with different axial contexts in parallel. We refer this lightweight feature calibration as group contextualization (GC). Specifically, we design a family of efficient element-wise calibrators, i.e., ECal-G/S/T/L, where their axial contexts are information dynamics aggregated from other axes either globally or locally, to contextualize feature channel groups. The GC module can be densely plugged into each residual layer of the off-the-shelf video networks. With little computational overhead, consistent improvement is observed when plugging in GC on different networks. By utilizing calibrators to embed feature with four different kinds of contexts in parallel, the learnt representation is expected to be more resilient to diverse types of activities. On videos with rich temporal variations, empirically GC can boost the performance of 2D-CNN (e.g., TSN and TSM) to a level comparable to the state-of-the-art video networks. Code is available at https://github.com/haoyanbin918/Group-Contextualization.
Yanbin Hao, Hao Zhang 0047, Chong-Wah Ngo, Xiangnan He 0001
CVPR1
2022 Multi-directional Knowledge Transfer for Few-Shot Learning
abstract
Knowledge transfer-based few-shot learning (FSL) aims at improving the recognition ability of a novel object under limited training samples by transferring relevant potential knowledge from other data. Most related methods calculate such knowledge to refine the representation of a novel sample or enrich the supervision to a classifier during a transfer procedure. However, it is easy to introduce new noise during the transfer calculations since: (1) the unbalanced quantity of samples between the known (base) and the novel categories biases the contents capturing of the novel objects, and (2) the semantic gaps existing in different modalities weakens the knowledge interaction during the training.
Shuo Wang 0008, Xinyu Zhang 0022, Yanbin Hao, Chengbing Wang, Xiangnan He 0001
ACM Multimedia3
2022 Unsupervised Video Hashing with Multi-granularity Contextualization and Multi-structure Preservation
abstract
Unsupervised video hashing typically aims to learn a compact binary vector to represent complex video content without using manual annotations. Existing unsupervised hashing methods generally suffer from incomplete exploration of various perspective dependencies (e.g., long-range and short-range) and data structures that exist in visual contents, resulting in less discriminative hash codes. In this paper, we propose aMulti-granularity Contextualized and Multi-Structure preserved Hashing (MCMSH) method, exploring multiple axial contexts for discriminative video representation generation and various structural information for unsupervised learning simultaneously. Specifically, we delicately design three self-gating modules to separately model three granularities of dependencies (i.e., long/middle/short-range dependencies) and densely integrate them into MLP-Mixer for feature contextualization, leading to a novel model MC-MLP. To facilitate unsupervised learning, we investigate three kinds of data structures, including clusters, local neighborhood similarity structure, and inter/intra-class variations, and design a multi-objective task to train MC-MLP. These data structures show high complementarities in hash code learning. We conduct extensive experiments using three video retrieval benchmark datasets, demonstrating that our MCMSH not only boosts the performance of the backbone MLP-Mixer significantly but also outperforms the competing methods notably. Code is available at: https://github.com/haoyanbin918/MCMSH.
Yanbin Hao, Jingru Duan, Hao Zhang 0047, Bin Zhu 0006, Peng Yuan Zhou, Xiangnan He 0001
ACM Multimedia1
2022 Unified QA-aware Knowledge Graph Generation Based on Multi-modal Modeling
abstract
Understanding the long duration videos' storyline is often considered a major challenge in the field of video understanding. To promote research on understanding longer videos in the community, the deep video understanding (DVU) task is suggested for recognizing interactions at the scene level and relationships at the movie level, as well as answering questions at these two levels. In this work, we propose a unified QA-aware knowledge graph generation approach, which consists of the relation-centric graph and interaction-centric graph and demonstrates the powerful performance of multimodal pre-training models in solving such problems. Extensive validations on the HLVU dataset demonstrate the effectiveness of our proposed method.
Penggang Qin, Jiarui Yu, Yan Gao 0017, Derong Xu, Yunkai Chen, Tong Xu 0001, Enhong Chen, Yanbin Hao
ACM Multimedia9
2022 Hierarchical Hourglass Convolutional Network for Efficient Video Classification
abstract
Videos naturally contain dynamic variation over the temporal axis, which will result in the same visual clues (e.g., semantics, objects) changing their scale, position, and perspective patterns between adjacent frames. A primary trend in video CNN is adopting spatial-2D convolution for spatial semantics and temporal-1D convolution for temporal dynamics. Though the direction achieves a favorable balance between efficiency and efficacy, it suffers from misalignment of visual clues with large displacements. Particularly, rigid temporal convolution would fail to capture correct motions when a specific target moves out of the reception field of temporal convolution between adjacent frames.
Yi Tan 0001, Yanbin Hao, Hao Zhang 0047, Shuo Wang 0008, Xiangnan He 0001
ACM Multimedia2
2022 Parameterization of Cross-token Relations with Relative Positional Encoding for Vision MLP
abstract
Vision multi-layer perceptrons (MLPs) have shown promising performance in computer vision tasks, and become the main competitor of CNNs and vision Transformers. They use token-mixing layers to capture cross-token interactions, as opposed to the multi-head self-attention mechanism used by Transformers. However, the heavily parameterized token-mixing layers naturally lack mechanisms to capture local information and multi-granular non-local relations, thus their discriminative power is restrained. To tackle this issue, we propose a new positional spacial gating unit (PoSGU). It exploits the attention formulations used in the classical relative positional encoding (RPE), to efficiently encode the cross-token relations for token mixing. It can successfully reduce the current quadratic parameter complexity O(N2) of vision MLPs to $O(N)$ and O(1). We experiment with two RPE mechanisms, and further propose a group-wise extension to improve their expressive power with the accomplishment of multi-granular contexts. These then serve as the key building blocks of a new type of vision MLP, referred to as PosMLP. We evaluate the effectiveness of the proposed approach by conducting thorough experiments, demonstrating an improved or comparable performance with reduced parameter complexity. For instance, for a model trained on ImageNet1K, we achieve a performance improvement from 72.14% to 74.02% and a learnable parameter reduction from 19.4M to 18.2M. Code could be found at https://github.com/Zhicaiwww/PosMLP https://github.com/Zhicaiwww/PosMLP.
Zhicai Wang, Yanbin Hao, Xingyu Gao 0001, Hao Zhang 0047, Shuo Wang 0008, Tingting Mu, Xiangnan He 0001
ACM Multimedia2
2022 Long-term Leap Attention, Short-term Periodic Shift for Video Classification
abstract
Video transformer naturally incurs a heavier computation burden than a static vision transformer, as the former processes T times longer sequence than the latter under the current attention of quadratic complexity (T2N2). The existing works treat the temporal axis as a simple extension of spatial axes, focusing on shortening the spatio-temporal sequence by either generic pooling or local windowing without utilizing temporal redundancy.
Hao Zhang 0047, Lechao Cheng, Yanbin Hao, Chong-Wah Ngo
ACM Multimedia3
2022 MF-GAN: Multi-conditional Fusion Generative Adversarial Network for Text-to-Image Synthesis
Yuyan Yang, Yanbin Hao, Yifeng Liu 0002, Haiyong Xie 0001
MMM (1)3
2022 Attention in Attention: Modeling Context Correlation for Efficient Video Classification
abstract
Attention mechanisms have significantly boosted the performance of video classification neural networks thanks to the utilization of perspective contexts. However, the current research on video attention generally focuses on adopting a specific aspect of contexts (e.g., channel, spatial/temporal, or global context) to refine the features and neglects their underlying correlation when computing attentions. This leads to incomplete context utilization and hence bears the weakness of limited performance improvement. To tackle the problem, this paper proposes an efficient attention-in-attention (AIA) method for element-wise feature refinement, which investigates the feasibility of inserting the channel context into the spatio-temporal attention learning module, referred to as CinST, and also its reverse variant, referred to as STinC. Specifically, we instantiate the video feature contexts as dynamics aggregated along a specific axis with global average and max pooling operations. The workflow of an AIA module is that the first attention block uses one kind of context information to guide the gating weights calculation of the second attention that targets at the other context. Moreover, all the computational operations in attention units act on the pooled dimension, which results in quite few computational cost increase (https://github.com/haoyanbin918/Attention-in-Attention.
Yanbin Hao, Shuo Wang 0008, Pei Cao 0001, Xinjian Gao, Tong Xu 0001, Jinmeng Wu, Xiangnan He 0001
IEEE Trans. Circuits Syst. Video Technol.1
2022 Spatio-Temporal Collaborative Module for Efficient Action Recognition
abstract
Efficient action recognition aims to classify a video clip into a specific action category with a low computational cost. It is challenging since the integrated spatial-temporal calculation (e.g., 3D convolution) introduces intensive operations and increases complexity. This paper explores the feasibility of the integration of channel splitting and filter decoupling for efficient architecture design and feature refinement by proposing a novel spatio-temporal collaborative (STC) module. STC splits the video feature channels into two groups and separately learns spatio-temporal representations in parallel with decoupled convolutional operators. Particularly, STC consists of two computation-efficient blocks,i.e., STand TS, where they extract either spatial (S.) or temporal (T.) features and further refine their features with either temporal (∙T) or spatial (∙S) contexts globally. The spatial/temporal context refers to information dynamics aggregated from temporal/spatial axis. To thoroughly examine our method’s performance in video action recognition tasks, we conduct extensive experiments using five video benchmark datasets requiring temporal reasoning. Experimental results show that the proposed STC networks achieve a competitive trade-off between model efficiency and effectiveness.
Yanbin Hao, Shuo Wang 0008, Yi Tan 0001, Xiangnan He 0001, Zhenguang Liu, Meng Wang 0001
IEEE Trans. Image Process.1
2022 Social Context-aware Person Search in Videos via Multi-modal Cues
abstract
Person search has long been treated as a crucial and challenging task to support deeper insight in personalized summarization and personality discovery. Traditional methods, e.g., person re-identification and face recognition techniques, which profile video characters based on visual information, are often limited by relatively fixed poses or small variation of viewpoints and suffer from more realistic scenes with high motion complexity (e.g., movies). At the same time, long videos such as movies often have logical story lines and are composed of continuously developmental plots. In this situation, different persons usually meet on a specific occasion, in which informative social cues are performed. We notice that these social cues could semantically profile their personality and benefit person search task in two aspects. First, persons with certain relationships usually co-occur in short intervals; in case one of them is easier to be identified, the social relation cues extracted from their co-occurrences could further benefit the identification for the harder ones. Second, social relations could reveal the association between certain scenes and characters (e.g., classmate relationship may only exist among students), which could narrow down candidates into certain persons with a specific relationship. In this way, high-level social relation cues could improve the effectiveness of person search. Along this line, in this article, we propose a social context-aware framework, which fuses visual and social contexts to profile persons in more semantic perspectives and better deal with person search task in complex scenarios. Specifically, we first segment videos into several independent scene units and abstract out social contexts within these scene units. Then, we construct inner-personal links through a graph formulation operation for each scene unit, in which both visual cues and relation cues are considered. Finally, we perform a relation-aware label propagation to identify characters’ occurrences, combining low-level semantic cues (i.e., visual cues) and high-level semantic cues (i.e., relation cues) to further enhance the accuracy. Experiments on real-world datasets validate that our solution outperforms several competitive baselines.
Tong Xu 0001, Peilun Zhou, Weidong He, Yanbin Hao, Yi Zheng 0007, Enhong Chen
ACM Trans. Inf. Syst.5
2021 Aggregated Multi-GANs for Controlled 3D Human Motion Prediction
abstract
Human motion prediction from historical pose sequence is at the core of many applications in machine intelligence. However, in current state-of-the-art methods, the predicted future motion is confined within the same activity. One can neither generate predictions that differ from the current activity, nor manipulate the body parts to explore various future possibilities. Undoubtedly, this greatly limits the usefulness and applicability of motion prediction. In this paper, we propose a generalization of the human motion prediction task in which control parameters can be readily incorporated to adjust the forecasted motion. Our method is compelling in that it enables manipulable motion prediction across activity types and allows customization of the human movement in a variety of fine-grained ways. To this aim, a simple yet effective composite GAN structure, consisting of local GANs for different body parts and aggregated via a global GAN is presented. The local GANs game in lower dimensions, while the global GAN adjusts in high dimensional space to avoid mode collapse. Extensive experiments show that our method outperforms state-of-the-art. The codes are available at https://github.com/herolvkd/AM-GAN.
Zhenguang Liu, Kedi Lyu, Shuang Wu 0002, Haipeng Chen 0002, Yanbin Hao, Shouling Ji
AAAI5
2021 Motion Prediction using Trajectory Cues
abstract
Predicting human motion from a historical pose sequence is at the core of many applications in computer vision. Current state-of-the-art methods concentrate on learning motion contexts in the pose space, however, the high dimensionality and complex nature of human pose invoke inherent difficulties in extracting such contexts. In this paper, we instead advocate to model motion contexts in the joint trajectory space, as the trajectory of a joint is smooth, vectorial, and gives sufficient information to the model. Moreover, most existing methods consider only the dependencies between skeletal connected joints, disregarding prior knowledge and the hidden connections between geometrically separated joints. Motivated by this, we present a semi-constrained graph to explicitly encode skeletal connections and prior knowledge, while adaptively learn implicit dependencies between joints.We also explore the applications of our approach to a range of objects including human, fish, and mouse. Surprisingly, our method sets the new state-of-the-art performance on 4 different benchmark datasets, a remarkable highlight is that it achieves a 19.1% accuracy improvement over current state-of-the-art in average. To facilitate future research, we have released our code at https://github.com/Pose-Group/MPT.
Zhenguang Liu, Pengxiang Su, Shuang Wu 0002, Xuanjing Shen, Haipeng Chen 0002, Yanbin Hao, Meng Wang 0001
ICCV6
2021 NASTER: Non-local Attentional Scene Text Recognizer
abstract
Scene text recognition has been widely investigated in computer vision. In the literature, the encoder-decoder based framework, which first encodes image into feature map and then decodes them into corresponding text sequences, have achieved great success. However, this solution fails in low-quality images, as the local visual features extracted from curved or blurred images are difficult to decode into corresponding text. To address this issue, we propose a new framework for Scene Text Recognition (STR), named Non-Local Attentional Scene Text Recognizer (NASTER). We use ResNet with Global Context Block (GC block) to extract global visual features. The global context information is then captured in parallel using the self-attention module and finally decoded by a multi-layer attention decoder with an intermediate supervision module. The proposed method achieves the state-of-the-art performances on seven benchmark datasets, demonstrating the effectiveness of our approach.
Xueliang Liu, Yanbin Hao, Yunjie Ma, Richang Hong
ICMR3
2021 Selective Dependency Aggregation for Action Classification
abstract
Video data are distinct from images for the extra temporal dimension, which results in more content dependencies from various perspectives. It increases the difficulty of learning representation for various video actions. Existing methods mainly focus on the dependency under a specific perspective, which cannot facilitate the categorization of complex video actions. This paper proposes a novel selective dependency aggregation (SDA) module, which adaptively exploits multiple types of video dependencies to refine the features. Specifically, we empirically investigate various long-range and short-range dependencies achieved by the multi-direction multi-scale feature squeeze and the dependency excitation. Query structured attention is then adopted to fuse them selectively, fully considering the diversity of videos' dependency preferences. Moreover, the channel reduction mechanism is involved in SDA for controlling the additional computation cost to be lightweight. Finally, we show that the SDA module can be easily plugged into different backbones to form SDA-Nets and demonstrate its effectiveness, efficiency and robustness by conducting extensive experiments on several video benchmarks for action classification. The code and models will be available at https://github.com/ty-97/SDA.
Yi Tan 0001, Yanbin Hao, Xiangnan He 0001, Yinwei Wei, Xun Yang 0001
ACM Multimedia2
2021 Token Shift Transformer for Video Classification
abstract
Transformer achieves remarkable successes in understanding 1 and 2-dimensional signals (e.g., NLP and Image Content Understanding). As a potential alternative to convolutional neural networks, it shares merits of strong interpretability, high discriminative power on hyper-scale data, and flexibility in processing varying length inputs. However, its encoders naturally contain computational intensive operations such as pair-wise self-attention, incurring heavy computational burden when being applied on the complex 3-dimensional video signals. This paper presents Token Shift Module (i.e., TokShift), a novel, zero-parameter, zero-FLOPs operator, for modeling temporal relations within each transformer encoder. Specifically, the TokShift barely temporally shifts partial [Class] token features back-and-forth across adjacent frames. Then, we densely plug the module into each encoder of a plain 2D vision transformer for learning 3D video representation. It is worth noticing that our TokShift transformer is a pure convolutional-free video transformer pilot with computational efficiency for video understanding. Experiments on standard benchmarks verify its robustness, effectiveness, and efficiency. Particularly, with input clips of 8/12 frames, the TokShift transformer achieves SOTA precision: 79.83%/80.40% on the Kinetics-400, 66.56% on EGTEA-Gaze+, and 96.80% on UCF-101 datasets, comparable or better than existing SOTA convolutional counterparts. Our code is open-sourced in: https://github.com/VideoNetworks/TokShift-Transformer.
Hao Zhang 0047, Yanbin Hao, Chong-Wah Ngo
ACM Multimedia2
2021 Learning to Match Anchor-Target Video Pairs With Dual Attentional Holographic Networks
abstract
Video hyperlinking is the task of linking two video fragments/clips based on their multi-modal contents. Specifically, given an anchor video as a query, machine techniques automatically generate links between the anchor and target videos by modeling and comparing their content aboutness. The term "aboutness" specifically refers to contextually relevant multimedia content, i.e., a fragment is on or of something. Since video contents are multi-modal (e.g., audio and vision), the content aboutness may be reflected across different modalities. Existing approaches regard hyperlinking as a retrieval task, by embedding multi-modal video contents into one or multiple common video representation space(s) for cross-modal comparison. As a result, the aboutness between videos is scored by computing the vector-distance based similarity in the learnt common feature space. However, these methods suffer from two main limitations: (1) the video modality descriptors/features are treated equally in representation learning, which hinders the effective modeling of their respective capabilities in linking; and (2) directly using the vector-distance based similarity to measure aboutness bears the risk of returning more duplicates. This paper focuses on addressing these two problems. Specifically, we firstly build attentional neural networks to learn a compact fragment-level representation, assigning different importance weights to different descriptor/feature contents by an attention mechanism. We believe that the potentially interesting content(s) should be highlighted in the representation. Furthermore, instead of directly computing the similarity of two representation embeddings, we secondly build a holographic composition network to model the aboutness for link establishment, with the core use of circular correlation. The two networks string together to form the final hyperlinking matching system. The entire model is trained in an end-to-end fashion. We examine its effectiveness by creating four train/validate/test partitioning schemes on the Blip10000 dataset and employing two video fragmentation methods.
Yanbin Hao, Chong-Wah Ngo, Bin Zhu 0006
IEEE Trans. Image Process.1
2020 Cross-sentence Pre-trained Model for Interactive QA matching
abstract
Semantic matching measures the dependencies between query and answer representations, it is an important criterion for evaluating whether the matching is successful. In fact, such matching does not examine each sentence individually, context information outside a sentence should be considered equally important to the syntactic context inside a sentence. We proposed a new QA matching model, built upon a cross-sentence context-aware architecture. An interactive attention mechanism with a pre-trained language model is proposed to automatically select salient positional answer representations that contribute more significantly to the answer relevance of a given question. In addition to the context information captured at each word position, we incorporate a new quantity of context information jump to facilitate the attention weight formulation. This reflects the amount of new information brought by the next word and is computed by modeling the joint probability between two adjacent word states. The proposed method is compared to multiple state-of-the-art ones evaluated using the TREC library, WikiQA, and the Yahoo! community question datasets. Experimental results show that the proposed method outperforms satisfactorily the competing ones.
Jinmeng Wu, Yanbin Hao
LREC2
2020 Person-level Action Recognition in Complex Events via TSD-TSM Networks
abstract
The task of person-level action recognition in complex events aims to densely detect pedestrians and individually predict their actions from surveillance videos. In this paper, we present a simple yet efficient pipeline for this task, referred to as TSD-TSM networks. Firstly, we adopt the TSD detector for the pedestrian localization on each single keyframe. Secondly, we generate the sequential ROIs for a person proposal by replicating the adjusted bounding box coordinates around the keyframe. Particularly, we propose to conduct straddling expansion and region squaring on the original bounding box of a person proposal to widen the potential space of motion and interaction and lead to a square box for ROI detection. Finally, we adapt the TSM classifier on the generated ROI sequences to perform action classification and further adopt late fusion to promote the prediction. Our proposed pipeline achieved the 3rd place in the ACM-MM 2020 grand challenge, i.e., Large-scale Human-centric Video Analysis in Complex Events (Track-4), obtaining final 15.31% [email protected] and 20.63% [email protected] on the testing set.
Yanbin Hao, Zi-Niu Liu, Hao Zhang 0047, Bin Zhu 0006, Jingjing Chen 0001, Yu-Gang Jiang 0001, Chong-Wah Ngo
ACM Multimedia1
2020 Compact Bilinear Augmented Query Structured Attention for Sport Highlights Classification
abstract
Understanding fine-grained activities, such as sport highlights, is a problem being overlooked and receives considerably less research attention. Potential reasons include absences of specific fine-grained action benchmark datasets, research preferences to general super-categorical activities classification, and challenges of large visual similarities between fine-grained actions. To tackle these, we collect and manually annotate two sport highlights datasets, i.e., Basketball-8 & Soccer-10, for fine-grained action classification. Sample clips in the datasets are annotated with professional sub-categorical actions like "dunk", "goalkeeping" and etc. We also propose a Compact Bilinear Augmented Query Structured Attention (CBA-QSA) module and stack it on top of general three-dimensional neural networks in a plug-and-play manner to emphasize important spatio-temporal clues in highlight clips. Specifically, we adapt the hierarchical attention neural networks, which contain learnable query-scheme, on the video to identify discriminative spatial/temporal visual clues within highlight clips. We name this altered attention which separately learns a query for spatial/temporal feature as query structured attention (QSA). Furthermore, we inflate bilinear mapping, which is a mature technique to represent local pairwise interactions for image-level fine-grained classification, on video understanding. In detail, we extend its compact version (i.e., compact bilinear mapping (CBM) based on TensorSketch) to deal with the three-dimensional video signal for modeling local pairwise motion information. We eventually incorporate CBM and QSA together to form CBA-QSA neural networks for fine-grained sport highlights classifications. Experimental results demonstrate that CBA-QSA improves the general state-of-the-arts on Basketball-8 and Soccer-10 datasets.
Yanbin Hao, Hao Zhang 0047, Chong-Wah Ngo, Xiaojun Hu
ACM Multimedia1
2020 Advance on large scale near-duplicate video retrieval
Richang Hong, Yanbin Hao
Frontiers Comput. Sci.3
2020 Cross-Domain Sentiment Encoding through Stochastic Word Embedding
abstract
Sentiment analysis is an important topic concerning identification of feelings, attitudes, emotions and opinions from text. To automate such analysis, a large amount of example text needs to be manually annotated for model training. This is laborious and expensive, but the cross-domain technique is a key solution to reducing the cost by reusing annotated reviews across domains. However, its success largely relies on the learning of a robust common representation space across domains. In the recent years, significant effort has been invested to improve the cross-domain representation learning by designing increasingly more complex and elaborate model inputs and architectures. We support that it is not necessary to increase design complexity as this inevitably consumes more time in model training. Instead, we propose to explore the word polarity and occurrence information through a simple mapping and encode such information more accurately whilst managing lower computational costs. The proposed approach is unique and takes advantage of the stochastic embedding technique to tackle cross-domain sentiment alignment. Its effectiveness is benchmarked with over ten data tasks constructed from two review corpora and it is compared against ten classical and state-of-the-art methods.
Yanbin Hao, Tingting Mu, Richang Hong, Meng Wang 0001, Xueliang Liu, John Yannis Goulermas
IEEE Trans. Knowl. Data Eng.1
2020 Neighbourhood Structure Preserving Cross-Modal Embedding for Video Hyperlinking
abstract
Video hyperlinking is a task aiming to enhance the accessibility of large archives, by establishing links between fragments of videos. The links model the aboutness between fragments for efficient traversal of video content. This paper addresses the problem of link construction from the perspective of cross-modal embedding. To this end, a generalized multi-modal auto-encoder is proposed. The encoder learns two embeddings from visual and speech modalities, respectively, whereas each of the embeddings performs self-modal and cross-modal translation of modalities. Furthermore, to preserve the neighbourhood structure of fragments, which is important for video hyperlinking, the auto-encoder is devised to model data distribution of fragments in a dataset. Experiments are conducted on Blip10000 dataset using the anchor fragments provided by TRECVid Video Hyperlinking (LNK) task over the years of 2016 and 2017. This paper shares the empirical insights on a number of issues in cross-modal learning, including the preservation of neighbourhood structure in embedding, model fine-tuning and issue of missing modality, for video hyperlinking.
Yanbin Hao, Chong-Wah Ngo, Benoit Huet
IEEE Trans. Multim.1
2019 R2GAN: Cross-Modal Recipe Retrieval With Generative Adversarial Network
abstract
Representing procedure text such as recipe for crossmodal retrieval is inherently a difficult problem, not mentioning to generate image from recipe for visualization. This paper studies a new version of GAN, named Recipe Retrieval Generative Adversarial Network (R2GAN), to explore the feasibility of generating image from procedure text for retrieval problem. The motivation of using GAN is twofold: learning compatible cross-modal features in an adversarial way, and explanation of search results by showing the images generated from recipes. The novelty of R2GAN comes from architecture design, specifically a GAN with one generator and dual discriminators is used, which makes the generation of image from recipe a feasible idea. Furthermore, empowered by the generated images, a two-level ranking loss in both embedding and image spaces are considered. These add-ons not only result in excellent retrieval performance, but also generate close-to-realistic food images useful for explaining ranking of recipes. On recipe1M dataset, R2GAN demonstrates high scalability to data size, outperforms all the existing approaches, and generates images intuitive for human to interpret the search results.
Bin Zhu 0006, Chong-Wah Ngo, Jingjing Chen 0001, Yanbin Hao
CVPR4
2019 3D human pose estimation via human structure-aware fully connected network
Xiaoyan Zhang 0002, Zhenhua Tang 0001, Junhui Hou, Yanbin Hao
Pattern Recognit. Lett.4
2017 Unsupervised t-Distributed Video Hashing and Its Deep Hashing Extension
abstract
In this paper, a novel unsupervised hashing algorithm, referred to as t-USMVH, and its extension to unsupervised deep hashing, referred to as t-UDH, are proposed to support large-scale video-to-video retrieval. To improve robustness of the unsupervised learning, the t-USMVH combines multiple types of feature representations and effectively fuses them by examining a continuous relevance score based on a Gaussian estimation over pairwise distances, and also a discrete neighbor score based on the cardinality of reciprocal neighbors. To reduce sensitivity to scale changes for mapping objects that are far apart from each other, Student t-distribution is used to estimate the similarity between the relaxed hash code vectors for keyframes. This results in more accurate preservation of the desired unsupervised similarity structure in the hash code space. By adapting the corresponding optimization objective and constructing the hash mapping function via a deep neural network, we develop a robust unsupervised training strategy for a deep hashing network. The efficiency and effectiveness of the proposed methods are evaluated on two public video collections via comparisons against multiple classical and the state-of-the-art methods.
Yanbin Hao, Tingting Mu, John Yannis Goulermas, Richang Hong, Meng Wang 0001
IEEE Trans. Image Process.1
2017 Stochastic Multiview Hashing for Large-Scale Near-Duplicate Video Retrieval
abstract
Near-duplicate video retrieval (NDVR) has been a significant research task in multimedia given its high impact in applications, such as video search, recommendation, and copyright protection. In addition to accurate retrieval performance, the exponential growth of online videos has imposed heavy demands on the efficiency and scalability of the existing systems. Aiming at improving both the retrieval accuracy and speed, we propose a novel stochastic multiview hashing algorithm to facilitate the construction of a large-scale NDVR system. Reliable mapping functions, which convert multiple types of keyframe features, enhanced by auxiliary information such as video-keyframe association and ground truth relevance to binary hash code strings, are learned by maximizing a mixture of the generalized retrieval precision and recall scores. A composite Kullback-Leibler divergence measure is used to approximate the retrieval scores, which aligns stochastically the neighborhood structures between the original feature and the relaxed hash code spaces. The efficiency and effectiveness of the proposed method are examined using two public near-duplicate video collections and are compared against various classical and state-of-the-art NDVR systems.
Yanbin Hao, Tingting Mu, Richang Hong, Meng Wang 0001, Ning An 0001, John Yannis Goulermas
IEEE Trans. Multim.1
2014 On improving behavior subtraction
abstract
With the popularity of monitoring devices, huge amount of surveillance data is generated in every minute. The technique for automatic analysis of monitoring videos is in urgent demand. As an extension of background subtraction, behavior subtraction succeeds in detecting the changes of scenes dynamics instead of its photometric properties. In this paper, we first propose a new algorithm in improving behavior subtraction by maximum likelihood estimate and interval estimate methods. After that we apply the improved approach to the framework of video summarization in which the goal is to condense hours of video data into a few short segments. The compressed video clips allow human to catch their interested information quickly. We finally conduct extensive experiments on real-world surveillance videos. The experimental results demonstrate its superior performance to other state-of-the-art methods.
Yanbin Hao, Xueliang Liu, Richang Hong
SMC2
2006 TV Program Recommendation for Multiple Viewers Based on user Profile Merging
Zhiwen Yu 0001, Xingshe Zhou 0001, Yanbin Hao, Jianhua Gu
User Model. User Adapt. Interact.3