Yang Long 0001

dblp:82/10183-1 · DBLP profile ↗
← Back
89ranked-venue papers
11as first author
54since 2021 · last 2026
0000-0002-2445-6112ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 55 · 7 first-author · 36 since 2021Graphics, computer vision, multimedia, augmented reality and games · 38 · 7 first-author · 17 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 5 since 2021Databases, data management, data science and information retrieval · 5 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
YearPublicationVenuePosition
2026 vMFCoOp: Towards Equilibrium on a Unified Hyperspherical Manifold for Prompting Biomedical VLMs
abstract
Recent advances in context optimization (CoOp) guided by large language model (LLM)–distilled medical semantic priors offer a scalable alternative to manual prompt engineering and full fine-tuning for adapting biomedical CLIP-based vision-language models (VLMs). However, prompt learning in this context is challenged by semantic misalignment between LLMs and CLIP variants due to divergent training corpora and model architectures; it further lacks scalability across continuously evolving families of foundation models. More critically, pairwise multimodal alignment via conventional Euclidean-space optimization lacks the capacity to model unified representations or apply localized geometric constraints, which tends to amplify modality gaps in complex biomedical imaging and destabilize few-shot adaptation. To address these challenges, we propose vMFCoOp, a framework that inversely estimates von Mises–Fisher (vMF) distributions on a shared Hyperspherical Manifold, aligning semantic biases between arbitrary LLMs and CLIP backbones via Unified Semantic Anchors to achieve robust biomedical prompting and superior few-shot classification. Grounded in three complementary constraints, vMFCoOp demonstrates consistent improvements across 14 medical datasets, 12 medical imaging modalities, and 13 anatomical regions, outperforming state-of-the-art methods in accuracy, generalization, and clinical applicability.
Minye Shao, Sihan Guo, Xinrun Li, Xingyu Miao, Haoran Duan 0001, Yang Long 0001
AAAI6
2026 Few-shot Medical Image Segmentation via Boundary-extended Prototypes and Momentum Inference
Yazhou Zhu 0001, Yang Long 0001, Haofeng Zhang 0001
Comput. Vis. Image Underst.4
2026 Decoding visual neural representations by multimodal with dynamic balancing
Kaili Sun, Xingyu Miao, Bing Zhai, Haoran Duan 0001, Yang Long 0001
Expert Syst. Appl.5
2026 Cross-modal progressive modeling for neuro-visual representation learning
abstract
Neural decoding from scalp signals requires models that respect spatial, temporal, and spectral structure while leveraging strong visual priors. In this paper, we introduce CFT-NET for disentangled neural visual representation together with a progressive visual–semantic adaptation (PVSA) framework that aligns EEG embeddings to pretrained visual backbones under a contrastive objective followed by pairwise matching. CFT-NET integrates Frequency-Separated Weights (FSW), Spatial-Context Aggregation (SCA), and Adaptive Temporal Filtering (ATF) to explicitly extract spectral, spatial, and temporal factors. PVSA consists of an instance-guided visual encoder and a visual-guided semantic decoder linked by cross attention, enabling fine-grained neuro–image interaction. On THINGS-EEG and THINGS-MEG dataset, the approach consistently outperforms state-of-the-art baselines in both subject-dependent and subject-independent zero-shot classification. By aligning model architecture with visual cognition principles and coupling it to strong visual priors, our methods narrows the gap between neural activity and visual cognition, provides novel cross-modal neural decoding method that achieves competitive performance against recent state-of-the-art baselines.
Jiyao Pu, Kaili Sun, Zeyu Fu, Haoran Duan 0001, Yang Long 0001
Neurocomputing6
2026 Federated learning with prototype-based adaptive domain adjustment
abstract
Federated learning can struggle when clients come from different domains. We present FedPAD , a prototype-based approach that treats prototype aggregation and model aggregation as two different problems. For prototypes, the server learns simplex weights to emphasize domains that deviate more from the current global prototypes, so the aggregated prototypes better cover the full range of domains. For model aggregation, we convert these prototype weights into temperature-controlled model weights that downweight the most shifted clients, improving training stability. We analyze FedPAD under time-varying aggregation and show how heterogeneity, local updates, and weight changes across rounds enter the convergence bound. Experiments on Digits, PACS, Office-Caltech, and Office-Home demonstrate improved average accuracy and stronger worst-domain performance under the same training budget.
Leyuan Zhang, Fan Wan, Yang Long 0001
Neurocomputing3
2026 Semi-supervised crowd counting from unlabeled data
Haoran Duan 0001, Yawen Huang, Yang Long 0001, Xian Wu 0001, Feiyue Huang, Shaoxin Li 0001
Pattern Recognit.3
2026 Learning to transport for open set domain generalization
Chunming Li, Yang Long 0001, Haofeng Zhang 0001
Pattern Recognit.3
2026 ConRF: Zero-shot stylization of 3D scenes with conditioned radiation fields
abstract
• We propose a novel method that leverages CLIP for zero-shot 3D scene artistic style transfer by a single condition (i.e. image or text). • We introduce a mapping network to alleviate the ambiguity in CLIP features related to style. • We present a 3D selection volume that allows for localized style manipulation within 3D scenes, expanding the possibilities in scene stylization and manipulation. Most of the existing works on arbitrary 3D NeRF style transfer required retraining on each single style condition. This work aims to achieve zero-shot controlled stylization in 3D scenes utilizing text or visual input as conditioning factors. We introduce ConRF, a novel method of zero-shot stylization. Specifically, due to the ambiguity of CLIP features, we employ a conversion process that maps the CLIP feature space to the style space of a pre-trained VGG network and then refine the CLIP multi-modal knowledge into a style transfer neural radiation field. Additionally, we use a 3D volumetric representation to perform local style transfer. By combining these operations, ConRF offers the capability to utilize either text or images as references, resulting in the generation of sequences with novel views enhanced by global or local stylization. Our experiment demonstrates that ConRF outperforms other existing methods for 3D scene and single-text stylization in terms of visual quality. Code is available: https://xingy038.github.io/ConRF/ .
Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Fan Wan, Yawen Huang, Yang Long 0001, Yefeng Zheng 0001
Pattern Recognit.6
2026 M3GStyler: Enhancing consistency across multi-view in multi-modality 3D Gaussian style transfer
abstract
• Unified text- and image-guided 3D style transfer with consistent views across a scene, while preserving fine details. • Flow-based feature matching aligns text and image cues for faithful style. • Parallel frequency branches balance structure (low frequency) and detail (high frequency). • Cross-dimension and texture losses improve color and texture fidelity.
Xueqi Qiu, Xingyu Miao, Haoran Duan 0001, Minye Shao, Bing Zhai, Gongjie Zhang, Jingjing Deng 0001, Yang Long 0001
Pattern Recognit.8
2026 A2D2C: Adaptive attention-driven dynamic convolution for local feature adaptation
abstract
• Introduces A 2 D 2 C: attention-driven dynamic convolution for local adaptation. • Uses multi-point random sampling to route and fuse k base kernels efficiently. • Presents A 2 D 2 C + that fuses kernels once, cutting redundancy and MAdds at parity. • Shows consistent gains on ImageNet, CIFAR-100 and COCO with statistical reports.
Fan Wan, Xingyu Miao, Jingjing Deng 0001, Xianghua Xie, Yang Long 0001
Pattern Recognit.6
2026 The uncertainty advantage: Enhancing large language models' reliability through chain of uncertainty reasoning
Zirong Peng, Xiaoming Liu 0020, Guan Yang, Jie Liu 0022, Xueping Peng, Yang Long 0001
Pattern Recognit. Lett.6
2026 Leveraging language to generalize natural images to few-shot medical image segmentation
Feifan Song 0004, Yuntian Bo, Yang Long 0001, Haofeng Zhang 0001
Pattern Recognit. Lett.4
2026 Rethinking Multi-Focus Image Fusion: An Input Space Optimization View
abstract
Multi-focus image fusion (MFIF) addresses the challenge of partial focus by integrating multiple source images taken at different focal depths. Unlike most existing methods that rely on complex loss functions or large-scale synthetic datasets, this study approaches MFIF from a novel perspective: optimizing the input space. The core idea is to construct a high-quality MFIF input space in a cost-effective manner by using intermediate features from well-trained, non-MFIF networks. To this end, we propose a cascaded framework comprising two feature extractors, a Feature Distillation and Fusion Module (FDFM), and a focus segmentation network Y ${}^{U}$ Net. Based on our observation that discrepancy and edge features are essential for MFIF, we select a image deblurring network and a salient object detection network as feature extractors. To transform these extracted features into an MFIF-suitable input space, we propose FDFM as a training-free feature adapter. To make FDFM compatible with high-dimensional feature maps, we extend the manifold theory from the edge-preserving field and design a novel isometric domain transformation. Extensive experiments on six benchmark datasets show that 1) our model consistently outperforms 13 state-of-the-art methods in both qualitative and quantitative evaluations, and 2) the constructed input space can directly enhance the performance of many MFIF models without additional requirements.
Zeyu Wang 0009, Haoran Duan 0001, Yang Long 0001, Ling Shao 0001
IEEE Trans. Image Process.5
2025 Towards Scalable Spatial Intelligence Via 2D-To-3D Data Lifting
abstract
Spatial intelligence is emerging as a transformative frontier in AI, yet it remains constrained by the scarcity of largescale 3D datasets. Unlike the abundant 2D imagery, acquiring 3D data typically requires specialized sensors and laborious annotation. In this work, we present a scalable pipeline that converts single-view images into comprehensive, scale- and appearance-realistic 3D representations - including point clouds, camera poses, depth maps, and pseudo-RGBD - via integrated depth estimation, camera calibration, and scale calibration. Our method bridges the gap between the vast repository of imagery and the increasing demand for spatial scene understanding. By automatically generating authentic, scale-aware 3D data from images, we significantly reduce data collection costs and open new avenues for advancing spatial intelligence. We release two generated spatial datasets, i.e., COCO-3D and Objects365-v2-3D, and demonstrate through extensive experiments that our generated data can benefit various 3D tasks, ranging from fundamental perception to MLLMbased reasoning. These results validate our pipeline as an effective solution for developing AI systems capable of perceiving, understanding, and interacting with physical environments.
Xingyu Miao, Haoran Duan 0001, Quanhao Qian, Jiuniu Wang, Yang Long 0001, Ling Shao 0001, Deli Zhao, Gongjie Zhang
ICCV5
2025 Rethinking Score Distilling Sampling for 3D Editing and Generation
abstract
Score Distillation Sampling (SDS) has emerged as a prominent method for text-to-3D generation by leveraging the strengths of 2D diffusion models. However, SDS is limited to generation tasks and lacks the capability to edit existing 3D assets. Conversely, variants of SDS that introduce editing capabilities often can not generate new 3D assets effectively. In this work, we observe that the processes of generation and editing within SDS and its variants have unified underlying gradient terms. Building on this insight, we propose Unified Distillation Sampling (UDS), a method that seamlessly integrates both the generation and editing of 3D assets. Essentially, UDS refines the gradient terms used in vanilla SDS methods, unifying them to support both tasks. Extensive experiments demonstrate that UDS not only outperforms baseline methods in generating 3D assets with richer details but also excels in editing tasks, thereby bridging the gap between 3D generation and editing.
Xingyu Miao, Haoran Duan 0001, Yang Long 0001, Jungong Han
ICML3
2025 TRACE: Temporally Reliable Anatomically-Conditioned 3D CT Generation with Enhanced Efficiency
Minye Shao, Xingyu Miao, Haoran Duan 0001, Zeyu Wang 0009, Jingkun Chen, Yawen Huang, Xian Wu 0001, Jingjing Deng 0001, Yang Long 0001, Yefeng Zheng 0001
MICCAI (4)9
2025 Synthesizing Spreading-out features for generative zero-shot image classification
Jingren Liu, Zheng Zhang 0006, Yang Long 0001, Wankou Yang, Yunyang Yan, Haofeng Zhang 0001
Eng. Appl. Artif. Intell.4
2025 Towards stereoscopic vision: Attention-guided gaze estimation with EEG in 3D space
abstract
Since traditional gaze-tracking methods rely on line-of-sight estimation, spatial attention modeling from neural activity offers an alternative perspective to gaze estimation. This paper presents a proof-of-concept study on attention-guided gaze estimation with Electroencephalography (EEG), investigating whether brain signals can be leveraged to estimate attentional focus within a controlled 3D environment. We first conducted a preliminary survey to gather public opinions, revealing a generally positive attitude towards EEG-driven gaze tracking. Building on this insight, we collected an EEG dataset in VR, where participants engaged with stimuli presented at predefined spatial locations. We introduce a deep learning model that estimates the relative saliency of candidate positions, enabling gaze estimation through optimization within the learned representation. Our results demonstrate that attentional focus was successfully mapped in a 3D coordinate space from 5 participants, and low-frequency oscillations contributed more significantly to predictive performance. The model achieved robust accuracy in distinguishing gaze locations, highlighting the potential of EEG-based gaze estimation for attention tracking in 3D environments.
Dantong Qin, Yang Long 0001, Zhibin Zhou 0002, Yuting Jin, Pan Wang 0005
Neurocomputing2
2025 DFAN++: Enhanced triple-branch network for generalized zero-shot image classification
Yuan Zhou 0023, Haoran Duan 0001, Yang Long 0001
Neurocomputing5
2025 FMDConv: Fast multi-attention dynamic convolution via speed-accuracy trade-off
abstract
Spatial convolution is fundamental in constructing deep Convolutional Neural Networks (CNNs) for visual recognition. While dynamic convolution enhances model accuracy by adaptively combining static kernels, it incurs significant computational overhead, limiting its deployment in resource-constrained environments such as federated edge computing. To address this, we propose Fast Multi-Attention Dynamic Convolution (FMDConv), which integrates input attention, temperature-degraded kernel attention, and output attention to optimize the speed-accuracy trade-off. FMDConv achieves a better balance between accuracy and efficiency by selectively enhancing feature extraction with lower complexity. Furthermore, we introduce two novel quantitative metrics, the Inverse Efficiency Score and Rate-Correct Score, to systematically evaluate this trade-off. Extensive experiments on CIFAR-10, CIFAR-100, and ImageNet demonstrate that FMDConv reduces the computational cost by up to 49.8% on ResNet-18 and 42.2% on ResNet-50 compared to prior multi-attention dynamic convolution methods while maintaining competitive accuracy. These advantages make FMDConv highly suitable for real-world, resource-constrained applications. This figure presents an overview of the proposed Fast Multi-Attention Dynamic Convolution (FMDConv) framework, which integrates Input Attention, Temperature-Degraded Kernel Attention, and Output Attention to optimize the speed-accuracy trade-off in convolutional neural networks. The diagram illustrates how these mechanisms enhance feature selection at different stages, significantly reducing computational cost while maintaining competitive accuracy, making FMDConv suitable for resource-constrained applications. • Introduces IES and RCS to quantify speed-accuracy trade-off in CNNs. • Evaluates channel, kernel, and filter attention for effective structures. • Develops a new CNN with kernel attention to enhance efficiency and accuracy. • Demonstrates FMDConv’s advantages via standard benchmark testing.
Fan Wan, Haoran Duan 0001, Kevin W. Tong, Jingjing Deng 0001, Yang Long 0001
Knowl. Based Syst.6
2025 Multi-domain feature-enhanced attribute updater for generalized zero-shot learning
Yuyan Shi, Chenyi Jiang, Feifan Song 0004, Qiaolin Ye, Yang Long 0001, Haofeng Zhang 0001
Neural Comput. Appl.5
2025 Imaginary-Connected Embedding in Complex Space for Unseen Attribute-Object Discrimination
abstract
Compositional Zero-Shot Learning (CZSL) aims to recognize novel compositions of seen primitives. Prior studies have attempted to either learn primitives individually (non-connected) or establish dependencies among them in the composition (fully-connected). In contrast, human comprehension of composition diverges from the aforementioned methods as humans possess the ability to make composition-aware adaptation for these primitives, instead of inferring them rigidly through the aforementioned methods. However, developing a comprehension of compositions akin to human cognition proves challenging within the confines of real space. This arises from the limitation of real-space-based methods, which often categorize attributes, objects, and compositions using three independent measures, without establishing a direct dynamic connection. To tackle this challenge, we expand the CZSL distance metric scheme to encompass complex spaces to unify the independent measures, and we establish an imaginary-connected embedding in complex space to model human understanding of attributes. To achieve this representation, we introduce an innovative visual bias-based attribute extraction module that selectively extracts attributes based on object prototypes. As a result, we are able to incorporate phase information in training and inference, serving as a metric for attribute-object dependencies while preserving the independent acquisition of primitives. We evaluate the effectiveness of our proposed approach on three benchmark datasets, illustrating its superiority compared to baseline methods.
Chenyi Jiang, Yang Long 0001, Zechao Li, Haofeng Zhang 0001, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 Laser: Efficient Language-Guided Segmentation in Neural Radiance Fields
abstract
In this work, we propose a method that leverages CLIP feature distillation, achieving efficient 3D segmentation through language guidance. Unlike previous methods that rely on multi-scale CLIP features and are limited by processing speed and storage requirements, our approach aims to streamline the workflow by directly and effectively distilling dense CLIP features, thereby achieving precise segmentation of 3D scenes using text. To achieve this, we introduce an adapter module and mitigate the noise issue in the dense CLIP feature distillation process through a self-cross-training strategy. Moreover, to enhance the accuracy of segmentation edges, this work presents a low-rank transient query attention mechanism. To ensure the consistency of segmentation for similar colors under different viewpoints, we convert the segmentation task into a classification task through label volume, which significantly improves the consistency of segmentation in color-similar areas. We also propose a simplified text augmentation strategy to alleviate the issue of ambiguity in the correspondence between CLIP features and text. Extensive experimental results show that our method surpasses current state-of-the-art technologies in both training speed and performance.
Xingyu Miao, Haoran Duan 0001, Yang Bai 0011, Tejal Shah, Jun Song 0003, Yang Long 0001, Rajiv Ranjan 0001, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Dual Interspersion and Flexible Deployment for Few-Shot Medical Image Segmentation
abstract
Acquiring a large volume of annotated medical data is impractical due to time, financial, and legal constraints. Consequently, few-shot medical image segmentation is increasingly emerging as a prominent research direction. Nowadays, Medical scenarios pose two major challenges: 1) intra-class variation caused by diversity among support and query sets; 2) inter-class extreme imbalance resulting from background heterogeneity. However, existing prototypical networks struggle to tackle these obstacles effectively. To this end, we propose a Dual Interspersion and Flexible Deployment (DIFD) model. Drawing inspiration from military interspersion tactics, we design the dual Interspersion module to generate representative basis prototypes from support features. These basis prototypes are then deeply interacted with query features. Furthermore, we introduce a fusion factor to fuse and refine the basis prototypes. Ultimately, we seamlessly integrate and flexibly deploy the basis prototypes to facilitate correct matching between the query features and basis prototypes, thus conducive to improving the segmentation accuracy of the model. Extensive experiments on three publicly available medical image datasets demonstrate that our model significantly outshines other SoTAs (2.78% higher dice score on average across all datasets), achieving a new level of performance. The code is available at: https://github.com/zmcheng9/DIFD.
Ziming Cheng, Yang Long 0001, Tao Zhou 0002, Haofeng Zhang 0001, Ling Shao 0001
IEEE Trans. Medical Imaging3
2025 Rethinking Brain Tumor Segmentation From the Frequency Domain Perspective
abstract
Precise segmentation of brain tumors, particularly contrast-enhancing regions visible in post-contrast MRI (areas highlighted by contrast agent injection), is crucial for accurate clinical diagnosis and treatment planning but remains challenging. However, current methods exhibit notable performance degradation in segmenting these enhancing brain tumor areas, largely due to insufficient consideration of MRI-specific tumor features such as complex textures and directional variations. To address this, we propose the Harmonized Frequency Fusion Network (HFF-Net), which rethinks brain tumor segmentation from a frequency-domain perspective. To comprehensively characterize tumor regions, we develop a Frequency Domain Decomposition (FDD) module that separates MRI images into low-frequency components, capturing smooth tumor contours and high-frequency components, highlighting detailed textures and directional edges. To further enhance sensitivity to tumor boundaries, we introduce an Adaptive Laplacian Convolution (ALC) module that adaptively emphasizes critical high-frequency details using dynamically updated convolution kernels. To effectively fuse tumor features across multiple scales, we design a Frequency Domain Cross-Attention (FDCA) integrating semantic, positional, and slice-specific information. We further validate and interpret frequency-domain improvements through visualization, theoretical reasoning, and experimental analyses. Extensive experiments on four public datasets demonstrate that HFF-Net achieves an average relative improvement of 4.48% (ranging from 2.39% to 7.72%) in the mean Dice scores across the three major subregions, and an average relative improvement of 7.33% (ranging from 5.96% to 8.64%) in the segmentation of contrast-enhancing tumor regions, while maintaining favorable computational efficiency and clinical applicability. Our code is available at: https://github.com/VinyehShaw/HFF.
Minye Shao, Zeyu Wang 0009, Haoran Duan 0001, Yawen Huang, Bing Zhai, Shizheng Wang, Yang Long 0001, Yefeng Zheng 0001
IEEE Trans. Medical Imaging7
2025 Contextual Interaction via Primitive-based Adversarial Training for Compositional Zero-shot Learning
abstract
Compositional Zero-shot Learning (CZSL) aims to identify novel compositions via known attribute–object pairs. The primary challenge in CZSL tasks lies in the significant discrepancies introduced by the complex interaction between the visual primitives of attribute and object, consequently decreasing the classification performance toward novel compositions. Previous remarkable works primarily addressed this issue by focusing on disentangling strategy or utilizing object-based conditional probabilities to constrain the selection space of attributes. Unfortunately, few studies have explored the problem from the perspective of modeling the mechanism of visual primitive interactions. Inspired by the success of vanilla adversarial learning in Cross-Domain Few-shot Learning, we take a step further and devise a model-agnostic and Primitive-based Adversarial Training (PBadv) method to deal with this problem. Besides, the latest studies highlight the weakness of the perception of hard compositions even under data-balanced conditions. To this end, we propose a novel over-sampling strategy with object-similarity guidance to augment target compositional training data. We performed detailed quantitative analysis and retrieval experiments on well-established datasets, such as UT-Zappos50K, MIT-States, and C-GQA, to validate the effectiveness of our proposed method, and the State-of-the-Art (SOTA) performance demonstrates the superiority of our approach. The code is available at https://github.com/lisuyi/PBadv_czsl .
Suyi Li 0005, Chenyi Jiang, Yang Long 0001, Zheng Zhang 0006, Haofeng Zhang 0001
ACM Trans. Multim. Comput. Commun. Appl.4
2024 SID-NERF: Few-Shot Nerf Based on Scene Information Distribution
abstract
The novel view synthesis from a limited set of images is a significant research focus. Traditional NeRF methods, relying mainly on color supervision, struggle with accurate scene geometry reconstruction when faced with sparse input images, leading to suboptimal rendering. We propose a Few-shot NeRF Based on Scene Information Distribution(Sid-NeRF) to address this by integrating geometric and color supervision, enhancing the model’s understanding of scene geometry. We also implement a data selector during training to identify and utilize the most accurate geometric data, thus improving training efficiency. Additionally, a residual module is introduced to counteract any optimization biases from the selector. Our method was tested on three datasets and showed excellent performance in various environments with limited images. Notably, compared to other novel view synthesis methods based on fewer views, our method does not require any prior knowledge and thus does not incur additional computational and storage costs.
Fan Wan, Yang Long 0001
ICME3
2024 Dual Variational Knowledge Attention for Class Incremental Vision Transformer
abstract
Class incremental learning (CIL) strives to emulate the human cognitive process of continuously learning and adapting to new tasks while retaining knowledge from past experiences. Despite significant advancements in this field, Transformer-based models have not fully leveraged the potential of attention mechanisms to balance the transferable knowledge between tokens and the associated information. This paper addresses this gap by using a dual variational knowledge attention (DVKA) mechanism within a Transformer-based encoder-decoder framework, tailored for CIL. DVKA mechanism aims to manage the information flow through the attention maps, ensuring a balanced representation of all classes, and mitigating the risk of information dilution as new classes are incrementally introduced. This method, leverage the information bottleneck and mutual information principle, selectively filters less relevant information, directing the model’s focus towards the most significant details for each class. The DVKA is designed with two distinct attentions: one focused on the feature level and the other on the token dimension. The feature-focused attention aims to purify the complex nature of various classification tasks, ensuring a comprehensive representation of both old and new tasks. The token-focused attention mechanism highlights specific tokens, facilitating local discrimination among disparate patches and fostering global coordination for a spectrum of task tokens. Our work is a major stride towards improving transformer models for class incremental learning, presenting a theoretical rationale and effective experimental results on three widely-used datasets.
Haoran Duan 0001, Rui Sun 0010, Varun Ojha 0001, Tejal Shah, Zhuoxu Huang, Zizhou Ouyang, Yawen Huang, Yang Long 0001, Rajiv Ranjan 0001
IJCNN8
2024 Wearable-based behaviour interpolation for semi-supervised human activity recognition
abstract
While traditional feature engineering for Human Activity Recognition (HAR) involves a trial-and-error process, deep learning has emerged as a preferred method for high-level representations of sensor-based human activities. However, most deep learning-based HAR requires a large amount of labelled data and extracting HAR features from unlabelled data for effective deep learning training remains challenging. We, therefore, introduce a deep semi-supervised HAR approach, MixHAR, which concurrently uses labelled and unlabelled activities. Our MixHAR employs a linear interpolation mechanism to blend labelled and unlabelled activities while addressing both inter- and intra-activity variability. A unique challenge identified is the activity-intrusion problem during mixing, for which we propose a mixing calibration mechanism to mitigate it in the feature embedding space. Additionally, we rigorously explored and evaluated the five conventional/popular deep semi-supervised technologies on HAR, acting as the benchmark of deep semi-supervised HAR. Our results demonstrate that MixHAR significantly improves performance, underscoring the potential of deep semi-supervised techniques in HAR.
Haoran Duan 0001, Varun Ojha 0001, Shizheng Wang, Yawen Huang, Yang Long 0001, Rajiv Ranjan 0001, Yefeng Zheng 0001
Inf. Sci.6
2024 CTNeRF: Cross-time Transformer for dynamic neural radiance field from monocular video
Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Fan Wan, Yawen Huang, Yang Long 0001, Yefeng Zheng 0001
Pattern Recognit.6
2024 DS-Depth: Dynamic and Static Depth Estimation via a Fusion Cost Volume
abstract
Self-supervised monocular depth estimation methods typically rely on the reprojection error to capture geometric relationships between successive frames in static environments. However, this assumption does not hold in dynamic objects in scenarios, leading to errors during the view synthesis stage, such as feature mismatch and occlusion, which can significantly reduce the accuracy of the generated depth maps. To address this problem, we propose a novel dynamic cost volume that exploits residual optical flow to describe moving objects, improving incorrectly occluded regions in static cost volumes used in previous work. Nevertheless, the dynamic cost volume inevitably generates extra occlusions and noise, thus we alleviate this by designing a fusion module that makes static and dynamic cost volumes compensate for each other. In other words, occlusion from the static volume is refined by the dynamic volume, and incorrect information from the dynamic volume is eliminated by the static volume. Furthermore, we propose a pyramid distillation loss to reduce photometric error inaccuracy at low resolutions and an adaptive photometric error loss to alleviate the flow direction of the large gradient in the occlusion regions. We conducted extensive experiments on the KITTI and Cityscapes datasets, and the results demonstrate that our model outperforms previously published baselines for self-supervised monocular depth estimation.
Xingyu Miao, Yang Bai 0011, Haoran Duan 0001, Yawen Huang, Fan Wan, Xinxing Xu, Yang Long 0001, Yefeng Zheng 0001
IEEE Trans. Circuits Syst. Video Technol.7
2024 Rules for Expectation: Learning to Generate Rules via Social Environment Modeling
abstract
The evolution of natural life is guided by a perpetually adaptive set of rules, encompassing natural laws, human policies, and game mechanics. Automated game design, through the creation of simulated environments populated by AI agents, embodies these rules, aligning with the objectives of artificial life research that seeks to replicate the dynamics of biological life through computational models. This paper presents a comprehensive framework, the Rule Generation Networks (RGN), devised for automated rule design, evaluation, and evolution in line with controllable expectations. We refine and formalize three cardinal elements - rules, strategies, and evaluation - to elucidate the intricate relationships inherent in rule generation tasks. The RGN integrates generative neural networks for rule design and a suite of reinforcement learning models for rule evaluation. To exemplify rule evolution and adaptation across varying environments, we introduce a controllability metric to gauge game dynamics and evolve the rule designer accordingly. Furthermore, we develop two game environments, Maze Run and Trust Evolution, modelling human exploration and societal trade dynamics, to gamify and evaluate the generated rules.
Jiyao Pu, Haoran Duan 0001, Junzhe Zhao, Yang Long 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Sentinel-Guided Zero-Shot Learning: A Collaborative Paradigm Without Real Data Exposure
abstract
With increasing concerns over data privacy and model copyrights, especially in the context of collaborations between AI service providers and data owners, an innovative Sentinel-Guided Zero-Shot Learning (SG-ZSL) paradigm is proposed in this work. SG-ZSL is designed to foster efficient collaboration without the need to exchange models or sensitive data. It consists of a teacher model, a student model and a generator that links both model entities. The teacher model serves as a sentinel on behalf of the data owner, replacing real data, to guide the student model at the AI service provider’s end during training. Considering the disparity of knowledge space between the teacher and student, we introduce two variants of the teacher model: the omniscient and the quasi-omniscient teachers. Under these teachers’ guidance, the student model seeks to match the teacher model’s performance and explores domains that the teacher has not covered. To trade-off between privacy and performance, we further introduce two distinct security-level training protocols: white-box and black-box, enhancing the paradigm’s adaptability. Despite the inherent challenges of real data absence in the SG-ZSL paradigm, it consistently outperforms in ZSL and GZSL tasks, notably in the white-box protocol. Our comprehensive evaluation further attests to its robustness and efficiency across various setups, including stringent black-box training protocol.
Fan Wan, Xingyu Miao, Haoran Duan 0001, Jingjing Deng 0001, Yang Long 0001
IEEE Trans. Circuits Syst. Video Technol.6
2024 MRL-Seg: Overcoming Imbalance in Medical Image Segmentation With Multi-Step Reinforcement Learning
abstract
Medical image segmentation is a critical task for clinical diagnosis and research. However, dealing with highly imbalanced data remains a significant challenge in this domain, where the region of interest (ROI) may exhibit substantial variations across different slices. This presents a significant hurdle to medical image segmentation, as conventional segmentation methods may either overlook the minority class or overly emphasize the majority class, ultimately leading to a decrease in the overall generalization ability of the segmentation results. To overcome this, we propose a novel approach based on multi-step reinforcement learning, which integrates prior knowledge of medical images and pixel-wise segmentation difficulty into the reward function. Our method treats each pixel as an individual agent, utilizing diverse actions to evaluate its relevance for segmentation. To validate the effectiveness of our approach, we conduct experiments on four imbalanced medical datasets, and the results show that our approach surpasses other state-of-the-art methods in highly imbalanced scenarios. These findings hold substantial implications for clinical diagnosis and research.
Feiyang Yang, Haoran Duan 0001, Feilong Xu, Yawen Huang, Xiaoli Zhang 0001, Yang Long 0001, Yefeng Zheng 0001
IEEE J. Biomed. Health Informatics7
2024 Dynamic visual-guided selection for zero-shot learning
Yuan Zhou 0023, Haoran Duan 0001, Yang Long 0001
J. Supercomput.5
2024 Deconfounding Causal Inference for Zero-Shot Action Recognition
abstract
Zero-shot action recognition (ZSAR) aims to recognize unseen action categories in the test set without corresponding training examples. Most existing zero-shot methods follow the feature generation framework to transfer knowledge from seen action categories to model the feature distribution of unseen categories. However, due to the complexity and diversity of actions, it remains challenging to generate unseen feature distribution, especially for the cross-dataset scenario when there is a potentially larger domain shift. This article proposes aDeconfoundingCaUSAlGAN (DeCalGAN) for generating unseen action video features with the following technical contributions: 1) Our model unifies compositional ZSAR with traditional visual-semantic models to incorporate local object information with global semantic information for feature generation. 2) A GAN-based architecture is proposed for causal inference and unseen distribution discovery. 3) A deconfounding module is proposed to refine representations of local objects and global semantic information confounder in the training data. Action descriptions and random object features after causal inference are then used to discover unseen distributions of novel actions in different datasets. Our extensive experiments onCross-DatasetZero-ShotActionRecognition (CD-ZSAR) demonstrate substantial improvement over the UCF101 and HMDB51 standard benchmarks for this problem.
Junyan Wang 0001, Yiqi Jiang, Yang Long 0001, Xiuyu Sun, Maurice Pagnucco, Yang Song 0001
IEEE Trans. Multim.3
2023 Dual Feature Augmentation Network for Generalized Zero-shot Learning
Yuan Zhou 0023, Haoran Duan 0001, Yang Long 0001
BMVC4
2023 Privacy-Enhanced Zero-Shot Learning via Data-Free Knowledge Transfer
abstract
Considering the increasing concerns about data copyright and sensitivity issues, we present a novel Privacy-Enhanced Zero-Shot Learning (PE-ZSL) paradigm. The key innovation is to involve a teacher model as the data safeguard to guide the PE-ZSL model training without data sharing. The PE-ZSL model consists of a generator and student network, which can achieve data-free knowledge transfer while maintaining the performance of teacher model. We investigate ‘black-’ and ‘white-box’ scenarios in PE-ZSL task as different levels of framework privacy. Besides, we provide the discussion of teacher model in both omniscient and quasi-omniscient settings according to the knowledge space. Despite simple implementations and data-missing disadvantages, our PE-ZSL framework can retain state-of-the-art ZSL and GZSL performance under the ‘white-box’ scenario. Extensive qualitative and quantitative analysis also demonstrates promising results when deploying the model under ‘black-box’ scenario.
Fan Wan, Daniel Organisciak, Jiyao Pu, Haoran Duan 0001, Peng Zhang 0058, Xingsong Hou, Yang Long 0001
ICME8
2023 Community-Aware Federated Video Summarization
abstract
Video summarization aims to extract representative frames to retain high-level information. Increasing concerns about privacy issues have been raised because conventional large-scale training requires users to upload video samples that may inevitably release sensitive information. In this paper, we thoroughly discuss the Federated Video Summarization problem, i.e., how to obtain a robust video summarization model when video data is distributed on private data islands. Our key contribution includes 1) We propose a fundamental Frame-Based aggregation method to video-related tasks, which differs from the sample-based aggregation in conventional FedAvg. 2) To mitigate the heterogeneous distribution due to community diversity, we propose the Community-Aware Clustering Federated Video Summarization Framework (CFed-VS) that clusters clients via a novel data-driven clustering approach. 3) We further tackle the challenging non-IID setting with a proposed Mixture Transformer, which manifests state-of-the-art performance via extensive quantitative and qualitative experiments on TVSum and SumMe datasets.
Fan Wan, Junyan Wang 0001, Haoran Duan 0001, Yang Song 0001, Maurice Pagnucco, Yang Long 0001
IJCNN6
2023 Towards Few-shot Image Captioning with Cycle-based Compositional Semantic Enhancement Framework
abstract
Many efforts paid attention to the multi-modal task, of which image captioning is a classic work. Especially the Clip model improves the performance of image captioning; meantime, its few-shot and zero-shot problems have become a significant research project. In this work, aiming at the image captioning task, we design the new few-shot and zero-shot settings different from popular directions. The direction focuses on the impact of the exited dataset for captioning model ability. According to analysis, we discover the frequency of the word combination can directly influence the performance of the captioning model. Based on this, we define the new few-shot and zero-shot settings. In terms of this, a Cycle-based captioning framework based on data augmentation is proposed to overcome this problem, of which the novelty switcher module is the critical component. Finally, experiments demonstrate that our framework can achieve state-of-the-art performance on both traditional, few-shot and zero-shot settings.
Peng Zhang 0058, Yang Bai 0011, Jie Su 0001, Yan Huang 0008, Yang Long 0001
IJCNN5
2023 Improving Health Mention Classification Through Emphasising Literal Meanings: A Study Towards Diversity and Generalisation for Public Health Surveillance
abstract
People often use disease or symptom terms on social media and online forums in ways other than to describe their health. Thus the NLP health mention classification (HMC) task aims to identify posts where users are discussing health conditions literally, not figuratively. Existing computational research typically only studies health mentions within well-represented groups in developed nations. Developing countries with limited health surveillance abilities fail to benefit from such data to manage public health crises. To advance the HMC research and benefit more diverse populations, we present the Nairaland health mention dataset (NHMD), a new dataset collected from a dedicated web forum for Nigerians. NHMD consists of 7,763 manually labelled posts extracted based on four prevalent diseases (HIV/AIDS, Malaria, Stroke and Tuberculosis) in Nigeria. With NHMD, we conduct extensive experiments using current state-of-the-art models for HMC and identify that, compared to existing public datasets, NHMD contains out-of-distribution examples. Hence, it is well suited for domain adaptation studies. The introduction of the NHMD dataset imposes better diversity coverage of vulnerable populations and generalisation for HMC tasks in a global public health surveillance setting. Additionally, we present a novel multi-task learning approach for HMC tasks by combining literal word meaning prediction as an auxiliary task. Experimental results demonstrate that the proposed approach outperforms state-of-the-art methods statistically significantly (p < 0.01, Wilcoxon test) in terms of F1 score over the state-of-the-art and shows that our new dataset poses a strong challenge to the existing HMC methods.
Olanrewaju Tahir Aduragba, Jialin Yu 0001, Alexandra I. Cristea, Yang Long 0001
WWW4
2023 Feature fine-tuning and attribute representation transformation for zero-shot learning
Shanmin Pang, Wenyu Hao, Yang Long 0001
Comput. Vis. Image Underst.4
2023 Data driven recurrent generative adversarial network for generalized zero shot image classification
Jie Zhang 0005, Shengbin Liao, Haofeng Zhang 0001, Yang Long 0001, Zheng Zhang 0006, Li Liu 0004
Inf. Sci.4
2023 Dynamic Unary Convolution in Transformers
abstract
It is uncertain whether the power of transformer architectures can complement existing convolutional neural networks. A few recent attempts have combined convolution with transformer design through a range of structures in series, where the main contribution of this paper is to explore a parallel design approach. While previous transformed-based approaches need to segment the image into patch-wise tokens, we observe that the multi-head self-attention conducted on convolutional features is mainly sensitive to global correlations and that the performance degrades when these correlations are not exhibited. We propose two parallel modules along with multi-head self-attention to enhance the transformer. For local information, a dynamic local enhancement module leverages convolution to dynamically and explicitly enhance positive local patches and suppress the response to less informative ones. For mid-level structure, a novel unary co-occurrence excitation module utilizes convolution to actively search the local co-occurrence between patches. The parallel-designed Dynamic Unary Convolution in Transformer (DUCT) blocks are aggregated into a deep architecture, which is comprehensively evaluated across essential computer vision tasks in image-based classification, segmentation, retrieval and density estimation. Both qualitative and quantitative results show our parallel convolutional-transformer approach with dynamic and unary convolution outperforms existing series-designed structures.
Haoran Duan 0001, Yang Long 0001, Haofeng Zhang 0001, Chris G. Willcocks, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2023 The Importance of Expert Knowledge for Automatic Modulation Open Set Recognition
abstract
Automatic modulation classification (AMC) is an important technology for the monitoring, management, and control of communication systems. In recent years, machine learning approaches are becoming popular to improve the effectiveness of AMC for radio signals. However, the automatic modulation open-set recognition (AMOSR) scheme that aims to identify the known modulation types and recognize the unknown modulation signals is not well studied. Therefore, in this paper, we propose a novel multi-modal marginal prototype framework for radio frequency (RF) signals (MMPRF) to improve AMOSR performance. First, MMPRF addresses the problem of simultaneous recognition of closed and open sets by partitioning the feature space in the way of one versus other and marginal restrictions. Second, we exploit the wireless signal domain knowledge to extract a series of signal-related features to enhance the AMOSR capability. In addition, we propose a GAN-based unknown sample generation strategy to allow the model to understand the unknown world. Finally, we conduct extensive experiments on several publicly available radio modulation data, and experimental results show that our proposed MMPRF outperforms the state-of-the-art AMOSR methods.
Taotao Li, Zhenyu Wen, Yang Long 0001, Zhen Hong, Shilian Zheng, Li Yu 0001, Bo Chen 0003, Xiaoniu Yang, Ling Shao 0001
IEEE Trans. Pattern Anal. Mach. Intell.3
2023 Visual-Semantic Aligned Bidirectional Network for Zero-Shot Learning
abstract
Zero-shot learning (ZSL) aims to recognize unknown categories that are unavailable during training. Recently, generative models have shown the potential to address this challenging problem by synthesizing unseen features conditioned on semantic embeddings such as attributes. However, unidirectional generative models cannot guarantee the effective coupling between visual and semantic spaces. To this end, we propose a visual-semantic aligned bidirectional network with cycle consistency to alleviate the gap between these two spaces, generating unseen features of high quality. More importantly, we incorporate two carefully designed strategies into our bidirectional framework to improve the overall ZSL performance. Specifically, we enhance the intra-domain class divergence in both visual and semantic spaces, and in the meantime, mitigate the inter-domain shift to preserve seen-unseen domain discrimination. Experimental results on four standard benchmarks show the superiority of our framework over existing state-of-the-art methods under both conventional and generalized ZSL settings.
Xingsong Hou, Jie Qin 0004, Yuming Shen, Yang Long 0001, Li Liu 0004, Zhao Zhang 0001, Ling Shao 0001
IEEE Trans. Multim.5
2022 Towards Unified Multi-Excitation for Unsupervised Video Prediction
Junyan Wang 0001, Likun Qin, Peng Zhang 0058, Yang Long 0001, Bingzhang Hu, Maurice Pagnucco, Shizheng Wang, Yang Song 0001
BMVC4
2022 Action Quality Assessment with Temporal Parsing Transformer
Yang Bai 0011, Desen Zhou, Songyang Zhang 0001, Jian Wang 0066, Errui Ding, Yu Guan 0001, Yang Long 0001, Jingdong Wang 0001
ECCV (4)7
2022 Deep Generative Modelling: A Comparative Review of VAEs, GANs, Normalizing Flows, Energy-Based and Autoregressive Models
abstract
Deep generative models are a class of techniques that train deep neural networks to model the distribution of training samples. Research has fragmented into various interconnected approaches, each of which make trade-offs including run-time, diversity, and architectural restrictions. In particular, this compendium covers energy-based models, variational autoencoders, generative adversarial networks, autoregressive models, normalizing flows, in addition to numerous hybrid approaches. These techniques are compared and contrasted, explaining the premises behind each and how they are interrelated, while reviewing current state-of-the-art advances and implementations.
Sam Bond-Taylor, Adam Leach, Yang Long 0001, Chris G. Willcocks
IEEE Trans. Pattern Anal. Mach. Intell.3
2022 EfficientTDNN: Efficient Architecture Search for Speaker Recognition
abstract
Convolutional neural networks (CNNs), such as the time-delay neural network (TDNN), have shown their remarkable capability in learning speaker embedding. However, they meanwhile bring a huge computational cost in storage size, processing, and memory. Discovering the specialized CNN that meets a specific constraint requires a substantial effort of human experts. Compared with hand-designed approaches, neural architecture search (NAS) appears as a practical technique in automating the manual architecture design process and has attracted increasing interest in spoken language processing tasks such as speaker recognition. In this paper, we propose EfficientTDNN, an efficient architecture search framework consisting of a TDNN-based supernet and a TDNN-NAS algorithm. The proposed supernet introduces temporal convolution of different ranges of the receptive field and feature aggregation of various resolutions from different layers to TDNN. On top of it, the TDNN-NAS algorithm quickly searches for the desired TDNN architecture via weight-sharing subnets, which surprisingly reduces computation while handling the vast number of devices with various resources requirements. Experimental results on the VoxCeleb dataset show the proposed EfficientTDNN enables approximate$10^{13}$architectures concerning depth, kernel, and width. Considering different computation constraints, it achieves a 2.20% equal error rate (EER) with 204 M multiply-accumulate operations (MACs), 1.41% EER with 571 M MACs as well as 0.94% EER with 1.45 G MACs. Comprehensive investigations suggest that the trained supernet generalizes subnets not sampled during training and obtains a favorable trade-off between accuracy and efficiency.
Rui Wang 0073, Zhihua Wei 0001, Haoran Duan 0001, Shouling Ji, Yang Long 0001, Zhen Hong
IEEE ACM Trans. Audio Speech Lang. Process.5
2021 Dynamic Graph Warping Transformer for Video Alignment
Junyan Wang 0001, Yang Long 0001, Maurice Pagnucco, Yang Song 0001
BMVC2
2021 Discriminative Latent Semantic Graph for Video Captioning
abstract
Video captioning aims to automatically generate natural language sentences that can describe the visual contents of a given video. Existing generative models like encoder-decoder frameworks cannot explicitly explore the object-level interactions and frame-level information from complex spatio-temporal data to generate semantic-rich captions. Our main contribution is to identify three key problems in a joint framework for future video summarization tasks. 1) Enhanced Object Proposal: we propose a novel Conditional Graph that can fuse spatio-temporal information into latent object proposal. 2) Visual Knowledge: Latent Proposal Aggregation is proposed to dynamically extract visual words with higher semantic levels. 3) Sentence Validation: A novel Discriminative Language Validator is proposed to verify generated captions so that key semantic concepts can be effectively preserved. Our experiments on two public datasets (MVSD and MSR-VTT) manifest significant improvements over state-of-the-art approaches on all metrics, especially for BLEU-4 and CIDEr. Our code is available at https://github.com/baiyang4/D-LSG-Video-Caption.
Yang Bai 0011, Junyan Wang 0001, Yang Long 0001, Bingzhang Hu, Yang Song 0001, Maurice Pagnucco, Yu Guan 0001
ACM Multimedia3
2021 Modality independent adversarial network for generalized zero shot image classification
Haofeng Zhang 0001, Yinduo Wang, Yang Long 0001, Longzhi Yang, Ling Shao 0001
Neural Networks3
2021 A plug-in attribute correction module for generalized zero-shot learning
Haofeng Zhang 0001, Haoyue Bai 0003, Yang Long 0001, Li Liu 0004, Ling Shao 0001
Pattern Recognit.3
2020 Query Twice: Dual Mixture Attention Meta Learning for Video Summarization
abstract
Video summarization aims to select representative frames to retain high-level information, which is usually solved by predicting the segment-wise importance score via a softmax function. However, softmax function suffers in retaining high-rank representations for complex visual or sequential information, which is known as the Softmax Bottleneck problem. In this paper, we propose a novel framework named Dual Mixture Attention (DMASum) model with Meta Learning for video summarization that tackles the softmax bottleneck problem, where the Mixture of Attention layer (MoA) effectively increases the model capacity by employing twice self-query attention that can capture the second-order changes in addition to the initial query-key attention, and a novel Single Frame Meta Learning rule is then introduced to achieve more generalization to small datasets with limited training sources. Furthermore, the DMASum significantly exploits both visual and sequential attention that connects local key-frame and global attention in an accumulative way. We adopt the new evaluation protocol on two public datasets, SumMe, and TVSum. Both qualitative and quantitative experiments manifest significant improvements over the state-of-the-art methods.
Junyan Wang 0001, Yang Bai 0011, Yang Long 0001, Bingzhang Hu, Zhenhua Chai, Yu Guan 0001, Xiaolin Wei
ACM Multimedia3
2020 Semantic combined network for zero-shot scene parsing
abstract
Recently, image‐based scene parsing has attracted increasing attention due to its wide application. However, conventional models can only be valid on images with the same domain of the training set and are typically trained using discrete and meaningless labels. Inspired by the traditional zero‐shot learning methods which employ auxiliary side information to bridge the source and target domains, the authors propose a novel framework called semantic combined network (SCN), which aims at learning a scene parsing model only from the images of the seen classes while targeting on the unseen ones. In addition, with the assistance of semantic embeddings of classes, the proposed SCN can further improve the performances of traditional fully supervised scene parsing methods. Extensive experiments are conducted on the data set Cityscapes, and the results show that the proposed SCN can perform well on both zero‐shot scene parsing (ZSSP) and generalised ZSSP settings based on several state‐of‐the‐art scenes parsing architectures. Furthermore, the authors test the proposed model under the traditional fully supervised setting and the results show that the proposed SCN can also significantly improve the performances of the original network models.
Yinduo Wang, Haofeng Zhang 0001, Yang Long 0001, Longzhi Yang
IET Image Process.4
2020 Learning discriminative domain-invariant prototypes for generalized zero shot learning
Yinduo Wang, Haofeng Zhang 0001, Zheng Zhang 0006, Yang Long 0001, Ling Shao 0001
Knowl. Based Syst.4
2020 Asymmetric graph based zero shot learning
Yinduo Wang, Haofeng Zhang 0001, Zheng Zhang 0006, Yang Long 0001
Multim. Tools Appl.4
2020 Deep transductive network for generalized zero shot learning
Haofeng Zhang 0001, Li Liu 0004, Yang Long 0001, Zheng Zhang 0006, Ling Shao 0001
Pattern Recognit.3
2020 Pseudo distribution on unseen classes for generalized zero shot learning
Haofeng Zhang 0001, Jingren Liu, Yazhou Yao, Yang Long 0001
Pattern Recognit. Lett.4
2020 A Joint Label Space for Generalized Zero-Shot Classification
abstract
The fundamental problem of Zero-Shot Learning (ZSL) is that the one-hot label space is discrete, which leads to a complete loss of the relationships between seen and unseen classes. Conventional approaches rely on using semantic auxiliary information, e.g. attributes, to re-encode each class so as to preserve the inter-class associations. However, existing learning algorithms only focus on unifying visual and semantic spaces without jointly considering the label space. More importantly, because the final classification is conducted in the label space through a compatibility function, the gap between attribute and label spaces leads to significant performance degradation. Therefore, this paper proposes a novel pathway that uses the label space to jointly reconcile visual and semantic spaces directly, which is named Attributing Label Space (ALS). In the training phase, one-hot labels of seen classes are directly used as prototypes in a common space, where both images and attributes are mapped. Since mappings can be optimized independently, the computational complexity is extremely low. In addition, the correlation between semantic attributes has less influence on visual embedding training because features are mapped into labels instead of attributes. In the testing phase, the discrete condition of label space is removed, and priori one-hot labels are used to denote seen classes and further compose labels of unseen classes. Therefore, the label space is very discriminative for the Generalized ZSL (GZSL), which is more reasonable and challenging for real-world applications. Extensive experiments on five benchmarks manifest improved performance over all of compared state-of-the-art methods.
Jin Li 0011, Xuguang Lan, Yang Long 0001, Yang Liu 0069, Xingyu Chen 0001, Ling Shao 0001, Nanning Zheng 0001
IEEE Trans. Image Process.3
2020 2D Pose-Based Real-Time Human Action Recognition With Occlusion-Handling
abstract
Human Action Recognition (HAR) for CCTV-oriented applications is still a challenging problem. Real-world scenarios HAR implementations is difficult because of the gap between Deep Learning data requirements and what the CCTV-based frameworks can offer in terms of data recording equipments. We propose to reduce this gap by exploiting human poses provided by the OpenPose, which has been already proven to be an effective detector in CCTV-like recordings for tracking applications. Therefore, in this work, we first propose ActionXPose: a novel 2D pose-based approach for pose-level HAR. ActionXPose extracts low- and high-level features from body poses which are provided to a Long Short-Term Memory Neural Network and a 1D Convolutional Neural Network for the classification. We also provide a new dataset, named ISLD, for realistic pose-level HAR in a CCTV-like environment, recorded in the Intelligent Sensing Lab. ActionXPose is extensively tested on ISLD under multiple experimental settings, e.g. Dataset Augmentation and Cross-Dataset setting, as well as revising other existing datasets for HAR. ActionXPose achieves state-of-the-art performance in terms of accuracy, very high robustness to occlusions and missing data, and promising results for practical implementation in real-world applications.
Federico Angelini, Zeyu Fu, Yang Long 0001, Ling Shao 0001, Syed M. Naqvi
IEEE Trans. Multim.3
2020 A Probabilistic Zero-Shot Learning Method via Latent Nonnegative Prototype Synthesis of Unseen Classes
abstract
Zero-shot learning (ZSL), a type of structured multioutput learning, has attracted much attention due to its requirement of no training data for target classes. Conventional ZSL methods usually project visual features into semantic space and assign labels by finding their nearest prototypes. However, this type of nearest neighbor search (NNS)-based method often suffers from great performance degradation because of the nonuniform variances between different categories. In this article, we propose a probabilistic framework by taking covariance into account to deal with the above-mentioned problem. In this framework, we define a new latent space, which has two characteristics. The first is that the features in this space should gather within the classes and scatter between the classes, which is implemented by triplet learning; the second is that the prototypes of unseen classes are synthesized with nonnegative coefficients, which are generated by nonnegative matrix factorization (NMF) of relations between the seen classes and the unseen classes in attribute space. During training, the learned parameters are the projection model for triplet network and the nonnegative coefficients between the unseen classes and the seen classes. In the testing phase, visual features are projected into latent space and assigned with the labels that have the maximum probability among unseen classes for classic ZSL or within all classes for generalized ZSL. Extensive experiments are conducted on four popular data sets, and the results show that the proposed method can outperform the state-of-the-art methods in most circumstances.
Haofeng Zhang 0001, Huaqi Mao, Yang Long 0001, Wankou Yang, Ling Shao 0001
IEEE Trans. Neural Networks Learn. Syst.3
2019 Few-Shot Image and Sentence Matching via Gated Visual-Semantic Embedding
abstract
Although image and sentence matching has been widely studied, its intrinsic few-shot problem is commonly ignored, which has become a bottleneck for further performance improvement. In this work, we focus on this challenging problem of few-shot image and sentence matching, and propose a Gated Visual-Semantic Embedding (GVSE) model to deal with it. The model consists of three corporative modules in terms of uncommon VSE, common VSE, and gated metric fusion. The uncommon VSE exploits external auxiliary resources to extract generic features for representing uncommon instances and words in images and sentences, and then integrates them by modeling their semantic relation to obtain global representations for association analysis. To better model other common instances and words in rest content of images and sentences, the common VSE learns their discriminative representations directly from scratch. After obtaining two similarity metrics from the two VSE modules with different advantages, the gated metric fusion module adaptively fuses them by automatically balancing their relative importance. Based on the fused metric, we perform extensive experiments in terms of few-shot and conventional image and sentence matching, and demonstrate the effectiveness of the proposed model by achieving the state-of-the-art results on two public benchmark datasets.
Yan Huang 0008, Yang Long 0001, Liang Wang 0001
AAAI2
2019 A General Transductive Regularizer for Zero-Shot Learning
Huaqi Mao, Haofeng Zhang 0001, Yang Long 0001, Longzhi Yang
BMVC3
2019 Order Matters: Shuffling Sequence Generation for Video Prediction
Junyan Wang 0001, Bingzhang Hu, Yang Long 0001, Yu Guan 0001
BMVC3
2019 Adversarial unseen visual feature synthesis for Zero-shot Learning
Haofeng Zhang 0001, Yang Long 0001, Li Liu 0004, Ling Shao 0001
Neurocomputing2
2019 Dual-verification network for zero-shot learning
Haofeng Zhang 0001, Yang Long 0001, Wankou Yang, Ling Shao 0001
Inf. Sci.2
2019 Zero-shot leaning and hashing with binary visual similes
Haofeng Zhang 0001, Yang Long 0001, Ling Shao 0001
Multim. Tools Appl.2
2019 Classification complexity assessment for hyper-parameter optimization
Ziyun Cai, Yang Long 0001, Ling Shao 0001
Pattern Recognit. Lett.2
2019 Generic compact representation through visual-semantic ambiguity removal
Yang Long 0001, Yu Guan 0001, Ling Shao 0001
Pattern Recognit. Lett.1
2019 Zero-shot Hashing with orthogonal projection for image retrieval
Haofeng Zhang 0001, Yang Long 0001, Ling Shao 0001
Pattern Recognit. Lett.2
2019 Triple Verification Network for Generalized Zero-Shot Learning
abstract
Conventional Zero-shot Learning approaches often suffer from severe performance degradation in the Generalised Zero-shot Learning (GZSL) scenario, i.e. to recognise test images that are from both seen and unseen classes. This paper studies the Class-level Over-fitting (CO) and empirically shows its effects to GZSL. We then address ZSL as a Triple Verification problem and propose a unified optimisation of regression and compatibility functions, i.e. two main streams of existing ZSL approaches. The complementary losses mutually regularise the same model to mitigate the CO problem. Furthermore, we implement a deep extension paradigm to linear models and significantly outperforms state-of-the-art methods in both GZSL and ZSL scenarios on the four standard benchmarks.
Haofeng Zhang 0001, Yang Long 0001, Yu Guan 0001, Ling Shao 0001
IEEE Trans. Image Process.2
2019 Depth Embedded Recurrent Predictive Parsing Network for Video Scenes
abstract
Semantic segmentation-based scene parsing plays an important role in automatic driving and autonomous navigation. However, most of the previous models only consider static images, and fail to parse sequential images because they do not take the spatial-temporal continuity between consecutive frames in a video into account. In this paper, we propose a depth embedded recurrent predictive parsing network (RPPNet), which analyzes preceding consecutive stereo pairs for parsing result. In this way, RPPNet effectively learns the dynamic information from historical stereo pairs, so as to correctly predict the representations of the next frame. The other contribution of this paper is to systematically study the video scene parsing (VSP) task, in which we use the RPPNet to facilitate conventional image paring features by adding spatial-temporal information. The experimental results show that our proposed method RPPNet can achieve fine predictive parsing results on cityscapes and the predictive features of RPPNet can significantly improve conventional image parsing networks in VSP task.
Lingli Zhou, Haofeng Zhang 0001, Yang Long 0001, Ling Shao 0001, Jing-Yu Yang 0001
IEEE Trans. Intell. Transp. Syst.3
2018 Towards Affordable Semantic Searching: Zero-Shot Retrieval via Dominant Attributes
abstract
Instance-level retrieval has become an essential paradigm to index and retrieves images from large-scale databases. Conventional instance search requires at least an example of the query image to retrieve images that contain the same object instance. Existing semantic retrieval can only search semantically-related images, such as those sharing the same category or a set of tags, not the exact instances. Meanwhile, the unrealistic assumption is that all categories or tags are known beforehand. Training models for these semantic concepts highly rely on instance-level attributes or human captions which are expensive to acquire. Given the above challenges, this paper studies the Zero-shot Retrieval problem that aims for instance-level image search using only a few dominant attributes. The contributions are: 1) we utilise automatic word embedding to infer class-level attributes to circumvent expensive human labelling; 2) the inferred class-attributes can be extended into discriminative instance attributes through our proposed Latent Instance Attributes Discovery (LIAD) algorithm; 3) our method is not restricted to complete attribute signatures, query of dominant attributes can also be dealt with. On two benchmarks, CUB and SUN, extensive experiments demonstrate that our method can achieve promising performance for the problem. Moreover, our approach can also benefit conventional ZSL tasks.
Yang Long 0001, Li Liu 0004, Yuming Shen, Ling Shao 0001
AAAI1
2018 Adaptive Visual-Depth Fusion Transfer
Ziyun Cai, Yang Long 0001, Xiaoyuan Jing, Ling Shao 0001
ACCV (4)2
2018 Towards Light-weight Annotations: Fuzzy Interpolative Reasoning for Zero-shot Image Classificaiton
Yang Long 0001, Yao Tan, Daniel Organisciak, Longzhi Yang, Ling Shao 0001
BMVC1
2018 Towards Universal Representation for Unseen Action Recognition
abstract
Unseen Action Recognition (UAR) aims to recognise novel action categories without training examples. While previous methods focus on inner-dataset seen/unseen splits, this paper proposes a pipeline using a large-scale training source to achieve a Universal Representation (UR) that can generalise to a more realistic Cross-Dataset UAR (CDUAR) scenario. We first address UAR as a Generalised Multiple-Instance Learning (GMIL) problem and discover 'building-blocks' from the large-scale ActivityNet dataset using distribution kernels. Essential visual and semantic components are preserved in a shared space to achieve the UR that can efficiently generalise to new datasets. Predicted UR exemplars can be improved by a simple semantic adaptation, and then an unseen action can be directly recognised using UR during the test. Without further training, extensive experiments manifest significant improvements over the UCF101 and HMDB51 benchmarks.
Yi Zhu 0001, Yang Long 0001, Yu Guan 0001, Shawn D. Newsam, Ling Shao 0001
CVPR2
2018 Enhancing Apparel Data Based on Fashion Theory for Developing a Novel Apparel Style Recommendation System
Congying Guan, Sheng Feng Qin, Wessie Ling, Yang Long 0001
WorldCIST (3)4
2018 Face recognition with a small occluded training set using spatial and statistical pooling
Yang Long 0001, Fan Zhu 0001, Ling Shao 0001, Junwei Han 0001
Inf. Sci.1
2018 Zero-Shot Learning Using Synthesised Unseen Visual Data with Diffusion Regularisation
abstract
Sufficient training examples are the fundamental requirement for most of the learning tasks. However, collecting well-labelled training examples is costly. Inspired by Zero-shot Learning (ZSL) that can make use of visual attributes or natural language semantics as an intermediate level clue to associate low-level features with high-level classes, in a novel extension of this idea, we aim to synthesise training data for novel classes using only semantic attributes. Despite the simplicity of this idea, there are several challenges. First, how to prevent the synthesised data from over-fitting to training classes? Second, how to guarantee the synthesised data is discriminative for ZSL tasks? Third, we observe that only a few dimensions of the learnt features gain high variances whereas most of the remaining dimensions are not informative. Thus, the question is how to make the concentrated information diffuse to most of the dimensions of synthesised data. To address the above issues, we propose a novel embedding algorithm named Unseen Visual Data Synthesis (UVDS) that projects semantic features to the high-dimensional visual feature space. Two main techniques are introduced in our proposed algorithm. (1) We introduce a latent embedding space which aims to reconcile the structural difference between the visual and semantic spaces, meanwhile preserve the local structure. (2) We propose a novel Diffusion Regularisation (DR) that explicitly forces the variances to diffuse over most dimensions of the synthesised data. By an orthogonal rotation (more precisely, an orthogonal transformation), DR can remove the redundant correlated attributes and further alleviate the over-fitting problem. On four benchmark datasets, we demonstrate the benefit of using synthesised unseen data for zero-shot learning. Extensive experimental results suggest that our proposed approach significantly outperforms the state-of-the-art methods.
Yang Long 0001, Li Liu 0004, Fumin Shen, Ling Shao 0001, Xuelong Li 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 Adaptive RGB Image Recognition by Visual-Depth Embedding
abstract
Recognizing RGB images from RGB-D data is a promising application, which significantly reduces the cost while can still retain high recognition rates. However, existing methods still suffer from the domain shifting problem due to conventional surveillance cameras and depth sensors are using different mechanisms. In this paper, we aim to simultaneously solve the above two challenges: 1) how to take advantage of the additional depth information in the source domain? 2) how to reduce the data distribution mismatch between the source and target domains? We propose a novel method called adaptive Visual- Depth Embedding (aVDE) which learns the compact shared latent space between two representations of labeled RGB and depth modalities in the source domain first. Then the shared latent space can help the transfer of the depth information to the unlabeled target dataset. At last, aVDE models two separate learning strategies for domain adaptation (feature matching and instance reweighting) in a unified optimization problem, which matches features and reweights instances jointly across the shared latent space and the projected target domain for an adaptive classifier. We test our method on five pairs of datasets for object recognition and scene classification, the results of which demonstrates the effectiveness of our proposed method.
Ziyun Cai, Yang Long 0001, Ling Shao 0001
IEEE Trans. Image Process.2
2018 Unsupervised Deep Hashing With Pseudo Labels for Scalable Image Retrieval
abstract
In order to achieve efficient similarity searching, hash functions are designed to encode images into low-dimensional binary codes with the constraint that similar features will have a short distance in the projected Hamming space. Recently, deep learning-based methods have become more popular, and outperform traditional non-deep methods. However, without label information, most state-of-the-art unsupervised deep hashing (DH) algorithms suffer from severe performance degradation for unsupervised scenarios. One of the main reasons is that the ad-hoc encoding process cannot properly capture the visual feature distribution. In this paper, we propose a novel unsupervised framework that has two main contributions: 1) we convert the unsupervised DH model into supervised by discovering pseudo labels; 2) the framework unifies likelihood maximization, mutual information maximization, and quantization error minimization so that the pseudo labels can maximumly preserve the distribution of visual features. Extensive experiments on three popular data sets demonstrate the advantages of the proposed method, which leads to significant performance improvement over the state-of-the-art unsupervised hashing algorithms.
Haofeng Zhang 0001, Li Liu 0004, Yang Long 0001, Ling Shao 0001
IEEE Trans. Image Process.3
2017 From Zero-Shot Learning to Conventional Supervised Classification: Unseen Visual Data Synthesis
abstract
Robust object recognition systems usually rely on powerful feature extraction mechanisms from a large number of real images. However, in many realistic applications, collecting sufficient images for ever-growing new classes is unattainable. In this paper, we propose a new Zero-shot learning (ZSL) framework that can synthesise visual features for unseen classes without acquiring real images. Using the proposed Unseen Visual Data Synthesis (UVDS) algorithm, semantic attributes are effectively utilised as an intermediate clue to synthesise unseen visual features at the training stage. Hereafter, ZSL recognition is converted into the conventional supervised problem, i.e. the synthesised visual features can be straightforwardly fed to typical classifiers such as SVM. On four benchmark datasets, we demonstrate the benefit of using synthesised unseen data. Extensive experimental results manifest that our proposed approach significantly improve the state-of-the-art results.
Yang Long 0001, Li Liu 0004, Ling Shao 0001, Fumin Shen, Guiguang Ding, Jungong Han
CVPR1
2017 Learning to Recognise Unseen Classes by A Few Similes
abstract
Existing image classification systems often suffer from re-training models for novel unseen classes. Zero-shot learning (ZSL) aims to recognise these unseen classes directly using trained models with a further inference procedure. However, existing approaches highly rely on human-defined class-attribute associations to achieve the inference, which significantly increases the annotation cost. This paper aims to address ZSL on non-attribute tasks, i.e. only training images with labels are used as most of the supervised settings. Our main contributions are: 1) to circumvent expensive attributes, we propose to use semantic similes that directly indicate the unseen-to-seen associations; 2) a novel similarity-based representation is proposed to represent both visual images and semantic similes in a unified embedding space; 3) in order to reduce the annotation cost, we use only a few similes to infer a class-level prototype for each unseen class. On two popular benchmarks, AwA and aPY, extensive experiments manifest that our method can significantly improve the state-of-the-art results using only two similes for each unseen class. Furthermore, we revisit the Caltech 101 dataset without attributes. Our ZSL results can exceed that of previous supervised methods.
Yang Long 0001, Ling Shao 0001
ACM Multimedia1
2017 Towards Fine-Grained Open Zero-Shot Learning: Inferring Unseen Visual Features from Attributes
abstract
Zero-shot Learning (ZSL) can leverage attributes to recognise unseen instances. However, the training data is limited and cannot adequately discriminate fine-grained classes with similar attributes. In this paper, we propose a complementary procedure that inversely makes use of attributes to infer discriminative visual features for unseen classes. In this way, ZSL is fully converted into conventional supervised classification, where robust classifiers can be employed to address the fine-grained problem. To infer high-quality unseen data, we propose a novel algorithm named Orthogonal Semantic-Visual Embedding (OSVE) that can discover the tiny visual differences between different instances under the same attribute by an orthogonal embedding space. On two fine-grained benchmarks, CUB and SUN, our method remarkably improves the state-of-the-art results under standard ZSL settings. We further challenge the Open ZSL problem where the number of seen classes is significantly smaller than that of unseen classes. Substantial experiments manifest that the inferred visual features can be successfully fed to SVM which can effectively discriminate unseen classes from fine-grained open candidates.
Yang Long 0001, Li Liu 0004, Ling Shao 0001
WACV1
2017 Describing Unseen Classes by Exemplars: Zero-Shot Learning Using Grouped Simile Ensemble
abstract
Learning visual attributes is an effective approach for zero-shot recognition. However, existing methods are restricted to learning explicitly nameable attributes and cannot tell which attributes are more important to the recognition task. In this paper, we propose a unified framework named Grouped Simile Ensemble (GSE). We claim our contributions as follows. 1) We propose to substitute explicit attribute annotation by similes, which are more natural expressions that can describe complex unseen classes. Similes do not involve extra concepts of attributes, i.e. only exemplars of seen classes are needed. We provide an efficient scenario to annotate similes for two benchmark datasets, AwA and aPY. 2) We propose a graph-cut-based class clustering algorithm to effectively discover implicit attributes from the similes. 3) Our GSE can automatically find the most effective simile groups to make the prediction. On both datasets, extensive experimental results manifest that our approach can significantly improve the performance over the state-of-the-art methods.
Yang Long 0001, Ling Shao 0001
WACV1
2016 Attribute Embedding with Visual-Semantic Ambiguity Removal for Zero-shot Learning
Yang Long 0001, Li Liu 0004, Ling Shao 0001
BMVC1
2016 Recognising occluded multi-view actions using local nearest neighbour embedding
Yang Long 0001, Fan Zhu 0001, Ling Shao 0001
Comput. Vis. Image Underst.1