Zhaoquan Yuan

dblp:135/5072 · DBLP profile ↗
← Back
31ranked-venue papers
9as first author
27since 2021 · last 2026
0000-0002-4083-5155ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 9 first-author · 26 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 2 since 2021
YearPublicationVenuePosition
2026 Monocular Vehicle Pose and Shape Reconstruction via Dynamic Context Adaptation and Progressive Geometry Refinement
abstract
Accurate reconstruction of 3D vehicle pose and shape from monocular images is challenging, particularly for distant objects in autonomous driving. Existing methods often suffer from geometric ambiguity in depth estimation and structural hollowness in shape recovery, primarily due to inadequate multi-scale feature aggregation and unflexible prior modeling. To overcome these limitations, MonoVPR is proposed, a novel framework integrating dynamic context adaptation and progressive geometry refinement. Specifically, a Hierarchical Dual-Context Attention (HDCA) module is introduced to resolve scale-dependent degradation through gated cross-attention across multi-resolution feature maps, dynamically fusing object-centric geometric cues with scene-centric semantics. For shape refinement, the Bounded Iterative Mesh Refiner (BIMR) progressively optimizes template-guided deformations via multi-head attention and a tanh-bounded correction loop, ensuring physically plausible reconstructions.Extensive experiments on the ApolloCar3D benchmark demonstrate MonoVPR achieves state-of-the-art performance, showing exceptional capability in reconstructing geometrically consistent shapes and precise poses for challenging long-range scenarios.
Wei Li 0110, Long Ji, Xiao Wu 0001, Zhaoquan Yuan, Penglin Dai
AAAI5
2026 I2CD: An Invertible Causal Framework for Compositional Zero-Shot Learning via Disentangle-Compose-Disentangle
abstract
Compositional Zero-Shot Learning (CZSL) addresses the challenge of recognizing unseen attribute-object compositions in images, representing a fundamental challenge in artificial intelligence. Current approaches, which primarily focus on semantic alignment or distribution independence of primitives, have not achieved effective state-object decoupling and causal interventional invariance, limiting their performance on unseen compositions. To tackle this challenge, this study introduces I2CD (Invertible Causal framework via Disentangle-Compose-Disentangle), a novel framework that integrates invertible neural networks with causal intervention techniques to achieve state-object disentanglement. The framework employs a disentangle-compose-disentangle mechanism for counterfactual generation within the disentangled representation space, ensuring that modifications to one primitive (attribute or object) maintain independence from the other, thus enabling robust causal disentanglement. Representational consistency is maintained through semantic alignment between initial disentangled representations and their recomposed-then-disentangled counterparts with corresponding textual concepts. Comprehensive evaluations on three benchmark datasets—MIT-States, UT-Zappos, and C-GQA—demonstrate the framework's effectiveness in achieving both disentanglement and compositional generalization in CZSL tasks.
Zhaoquan Yuan, Yuankang Pan, Ao Luo, Wei Li 0110, Xiao Wu 0001, Changsheng Xu
AAAI1
2026 Prohibited Item Detection in X-ray images based on refined surface perception
Wei Li 0110, Zhaoquan Yuan
J. Vis. Commun. Image Represent.4
2026 Energy-based causal disentanglement for compositional zero-shot learning
Yuankang Pan, Zhaoquan Yuan, Wei Li 0110
Multim. Syst.3
2026 Learning Unknowns Without Forgetting Knowns: Compositional and Bidirectional Low-Rank Adaptive Open-World Detection Transformer
abstract
Open-World Object Detection (OWOD) aims to detect unseen objects as “unknown” while incrementally learning them without catastrophic forgetting. This problem presents two major challenges: (1) the lack of annotations for unknown objects during training, and (2) the risk of catastrophic forgetting during model updates. To address these issues, we propose the COmpositional and Bidirectional low-Rank Adaptive open-world detection transformer (COBRA)-a novel framework built upon a pre-trained Deformable DETR model. Specifically, COBRA first employs an attentional filtering mechanism that prunes previously known (P-Known) and currently known (C-Known) objects, yielding a purified set of candidateunknowns. To system-atically pseudo-label theseunknowns, we introduce a Primitive Composition Recognition (PCR) module, which evaluates set-level similarity between candidate objects and learned primitives, enabling accurate labeling ofpseudo-unknowns. To mitigate catastrophic forgetting during incremental updates, COBRA leverages Bidirectional Low-Rank Adaptation (Bi-LoRA)-a parameter-efficient mechanism that supports forward knowledge transfer and stable backward integration. Together, these components form a synergistic pipeline for continual object discovery and knowledge consolidation. Extensive experiments on MS COCO and PASCAL VOC demonstrate that our rehearsal-free COBRA framework outperforms SAM-powered methods in unknown recall while achieving lower forgetting compared to rehearsal-based competitors.
Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Wei Li 0110, Ao Luo, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2026 Class-Specific Knowledge-Guided Multimodal Prompt Tuning for Few-Shot Class-Incremental Learning
abstract
Few-shot class-incremental learning (FSCIL) requires a model to learn the knowledge of new categories incrementally, using only a few samples, after being trained on a base session with ample categories and sample sizes. This task presents two major challenges: catastrophic forgetting and overfitting. Current approaches primarily enhance the model’s ability to extract knowledge during the base stage to improve adaptability to new tasks. Large-scale pre-trained models, known for their high robustness and zero-shot transfer capabilities, have demonstrated promising performance in FSCIL. The key to solving FSCIL lies in effectively fine-tuning such large models to balance the learning of new knowledge and the retention of old knowledge. Inspired by human-like knowledge retrieval mechanisms, we propose Class-specific Knowledge-Guided Prompt Tuning (CKGPT), which leverages class-specific prompts to guide the model in learning targeted knowledge reuse and integration effectively. When faced with novel tasks, the model selectively activates previously learned knowledge that is the most relevant, improving performance on new tasks while minimizing updates to irrelevant knowledge to reduce forgetting. By incorporating mechanisms that balance knowledge retention and transfer, CKGPT ensures a more robust adaptation to sequential tasks. Extensive experiments on multiple benchmarks validate the effectiveness of our method in achieving superior performance.
Fangying Xiong, Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu
IEEE Trans. Circuits Syst. Video Technol.2
2026 THMM-CLIP: Task-Guided Hierarchical Multi-Modal Alignment for Rehearsal-Free Class Incremental Learning
abstract
Class incremental learning (CIL) requires models to acquire knowledge from sequential tasks containing non-overlapping classes while avoiding catastrophic forgetting. While vision-language foundation models like CLIP demonstrate remarkable potential for CIL through their pre-trained cross-modal alignment capabilities, existing CLIP-based approaches critically overlook the progressive degradation of visual representations in incremental scenarios . Through feature space analysis, we identify a crucial dichotomy : textual embeddings maintain stable discriminative power across sequential tasks, whereas visual features exhibit progressive deterioration manifested by intra-task confusion (ambiguous decision boundaries between co-occurring classes) and inter-task interference (semantic collision between historical and novel categories). To address these dual challenges, we propose task-guided hierarchical multi-modal alignment (THMM-CLIP), a framework that establishes persistent visual-textual coherence through hierarchical multi-modal alignment (HMA) and robust prompt selection (RPS). HMA adapts lightweight task-specific prompt vectors to dynamically recalibrate the CLIP image encoder, thereby achieving: (i) intra-task alignment, (ii) inter-task discriminability alignment, and (iii) global structural alignment with textual features. RPS incorporates a dual-level task identifier that integrates class-level and task-level representative features to ensure precise prompt retrieval during inference. Ablation studies validate all components’ contributions, while t-SNE visualizations, confusion matrices, and Grad-CAM analyses confirm strengthened cross-modal alignment.
Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Zechao Li, Changsheng Xu
ACM Trans. Multim. Comput. Commun. Appl.2
2026 PSGNet: Pure Smoke Image Generation With Gradient and Style Learning
abstract
The realistic and controllable generation of pure smoke is critical for smoke image editing, smoke visual special effects generation, and smoke data synthesizing within security scenarios. It is a relatively underexplored topic and continues to present significant challenges. Existing methods face challenges in the generation of smoke with intricate details and the regulation of various smoke styles. In this paper, a Pure Smoke image Generation Network (PSGNet) is proposed with a gradient and style learning approach to generate realistic and controllable smoke images. To achieve flexibility in control across the spatial dimension, the smoke shape mask is used to encode spatial details, such as the location and contour of the smoke, along with other related properties. To enhance the physical realism of synthesized smoke, a novel gradient-based learning framework is proposed to generate smoke gradient features, highlighting a special focus on explicitly encoding and exploiting gradient information. This framework uses a smoke gradient learning architecture that captures the subtle structures and patterns characteristic of real smoke, enabling the generation of highly realistic smoke with rich, fine-scale detail. In addition, a spatially aware style learning strategy is proposed to provide fine-grained control over smoke attributes such as density, color, and overall look. It is able to effectively model style features across both channel and spatial dimensions, thereby enabling spatially aware style manipulation. By combining the gradient module with this style learning framework, the method produces smoke that exhibits rich visual details and customizable image styles. Experiments conducted on six benchmark datasets demonstrate that the proposed PSGNet significantly outperforms the state-of-the-art approaches.
Jian-Jun Qiao, Xiao Wu 0001, Zhi-Qi Cheng, Wei Li 0110, Zhaoquan Yuan
IEEE Trans. Vis. Comput. Graph.5
2025 Contrastive Invariant Risk Minimization for Grounded Situation Recognition
abstract
Grounded situation recognition (GSR) is a comprehensive structured scene understanding task that predicts the salient activity (verb), entities (nouns) involved in the activity with their roles, as well as the corresponding bounding-box groundings of the entities from the given image. Existing I.I.D.-based methods for GSR are limited in their ability to recognize novel verb-noun combinations. To address this problem, in this paper, we novelly consider GSR as a Non-I.I.D. task and focus on learning verb-invariant and role-specific representations for verb and noun predictions. Based on the causality, a novel Contrastive Invariant Risk Minimization (CIRM) model for GSR is proposed. In the proposed CIRM, invariant risk minimization is integrated into a transformer architecture to learn invariant representations for verb prediction. To enhance the intra-verb compactness and the inter-verb separability, contrastive learning is utilized to learn discriminative features. As far as we know, this is the first work that regards the task of GSR as a problem of out-of-distribution generalization. Extensive experiments on the benchmark SWiG dataset demonstrate the effectiveness of our proposed CIRM over other state-of-the-art methods in all evaluation metrics.
Zhaoquan Yuan, Chengbin Zhao, Yuting Tang, Lishu Guo, Xiao Wu 0001, Changsheng Xu
ICME1
2025 HOOI Detection: Cascade-Clue Integrated Modeling over Multiple Temporal Segments
abstract
To fully comprehend a visual scene, recognizing and localizing interaction actions are essential components. Recently, significant advances have been made in detecting human-object interaction actions, which aim to capture pairwise relations between entities in the scene. Although these methods have made significant progress, they ignore the human-object-object interaction (HOOI) actions that frequently occur between a human and two objects in the real world. To advance related research, a new task named HOOI detection is introduced. It aims to accurately localize the humans in each video frame and identify the HOOI actions they perform. For this purpose, two novel HOOI datasets oriented to industrial production and daily life are constructed. These new datasets provide essential data support for in-depth research of HOOI detection. Furthermore, a cutting-edge method named Cascade-Clue Integrated Modeling over Multiple Temporal Segments (C2TS) is proposed for effectively detecting HOOI actions. Specifically, considering the phased characteristics of the action, C2TS comprehensively considers the HOOI information in the preceding, neighborhood, and subsequent temporal segments. For each temporal segment, the Cascaded Modeling and Clue Augmentation methods are applied to extract the corresponding HOOI features. The final detection result is obtained by classifying the aggregated HOOI features from the three temporal segments. Experiments conducted on the two proposed HOOI-related datasets vividly demonstrate that our method outperforms state-of-the-art approaches by achieving remarkable improvements of approximately 3% and 5%, which powerfully validates its efficacy in tackling the given challenge.
Mingxuan Zhang 0001, Qi He 0007, Zhaoquan Yuan, Tingquan He
ICMR3
2025 Latent Interactiveness Field for Non-Contact Human Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection serves a broad spectrum of applications. Despite significant progress, current approaches encounter difficulties in effectively handling Non-Contact Human-Object Interaction (NCHOI) scenarios, where humans and objects remain physically apart. To address these challenges, this paper proposes a novel approach, named Latent Interactiveness Field Modeling (LIFM), which enhances HOI detection by capturing long-range contextual dependencies. Specifically, the Latent Interactiveness Field (LIF) is introduced to define potential interactive relationships between humans and objects. To complement this, the LIF Fusion Encoder is designed to adaptively fuse visual features with LIF, resulting in more informative and discriminative feature representations. The Mobile Scanning HOI Dataset (MSHD) is introduced as a comprehensive benchmark to systematically assess the robustness of existing methods on both common HOI and NCHOI in real-world applications. Extensive experimentation indicates that the proposed approach outperforms existing state-of-the-art techniques. It offers substantial improvements, particularly in NCHOI scenarios, which highlight its effectiveness in resolving issues related to long-range interactions.
Xiang Huang 0004, Ao Luo, Xiao Wu 0001, Zhaoquan Yuan
ACM Multimedia4
2025 HOPNet: Learning Hand-Object-Person Interaction Network for Hand Contact State Detection
abstract
The detection of hand contact states, which involves identifying interactions between hands and objects or other entities, is essential for the development of human-computer interaction systems and the comprehension of social dynamics. Previous approaches have made progress in modeling hand-object interactions. Nonetheless, they neglect critical cues between their hands and bodies, as well as those of others, thus constraining their ability to accurately detect interpersonal contact. The task remains challenging due to frequent occlusions, especially in crowded multi-person scenarios with complex contexts. In this paper, a novel hand-object-person interaction network, called HOPNet, is proposed to model contextual information between hands and objects, as well as between hands and bodies. Specifically, HOPNet consists of two components: (i) the Hand-Object Relation (HOR) module analyzes interaction patterns between hands and objects, capturing spatial and semantic relationships; (ii) the Contrastive Spatial Refinement (CSR) module learns hand-body interactions through contrastive geometric embedding and relative spatial enhancement, improving interpersonal contact recognition in crowded scenarios. Experiments on ContactHands and 100DOH datasets demonstrate that HOPNet outperforms state-of-the-art methods.
Wei Li 0110, Yizhao Wan, Xiao Wu 0001, Jianshuai Wang, Penglin Dai, Zhaoquan Yuan
ACM Multimedia6
2025 DualEnhance: External Multimodal Foundation Models Guidance and Internal Fast-Slow Teacher Regulation
abstract
Source-Free Domain Adaptive Object Detection addresses cross-domain detection on an unlabeled target domain without accessing source data. Existing methods implement self-training with Mean Teacher but are bottlenecked by error accumulation from noisy pseudo-labels generated via recursive teacher-student updates. This issue is handled through the proposed dual enhancements: (1) External Guidance via Multimodal Foundation Models (FMs); (2) Internal Regulation through Fast-Slow Teacher. First, despite FMs' multimodal comprehension, their semantic misalignment with a specific task introduces noise during adaptation. Bidirectional Distillation mitigates this by calibrating the FM using task-specific knowledge transferred from the source detector. The aligned cross-modal knowledge then propagates through high-quality pseudo-label generation. Second, the conventional Mean Teacher suffers from plasticity-stability dilemma, where rapid adaptation corrupts historical knowledge. Fast-Slow Teacher introduces dual-velocity knowledge consolidation: The Fast Teacher dynamically captures emerging domain features, while the Slow Teacher preserves stable historical knowledge and periodically resets the Fast Teacher, establishing an error-correcting dynamic equilibrium. Experiments show our method achieves significant improvements over SOTA.
Qi He 0007, Xiao Wu 0001, Jun-Yan He, Wei Li 0110, Zhaoquan Yuan
ACM Multimedia5
2025 Disentanglement-Based Equivariant Learning for Compositional VQA
abstract
Compositional visual question answering (VQA) represents a challenging yet fundamental task that requires models to comprehend novel combinations of previously learned concepts. The current methods often overlook the disentanglement of underlying concepts and are restricted in terms of their ability to effectively capture the compositional variation mechanism. Moreover, the state-of-the-art techniques depend on additional clues for training, which is not feasible in real-world VQA scenarios. To address these issues, in this paper, we introduce a novelDisentanglement-basedEquivAriantLearning (DEAL) framework for compositional VQA, which is guided exclusively by ground-truth answers. In DEAL, we employ causality-inspired interventions to disentangle concepts derived from visual and textual inputs within a re-encoding framework. Based on the principle of equivariance, we subsequently perform a compositional transformation on the inference input and impose the equivariant constraint on the output to augment the compositional reasoning capacity of the model. Comprehensive experiments conducted on the benchmark CLEVR-CoGenT and GQA-SGL datasets validate the superiority of our proposed DEAL approach over the existing state-of-the-art methods for compositional VQA tasks in both visual and linguistic generalization settings.
Zhou Du, Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu
IEEE Trans. Multim.2
2025 Active Cross-Modal Domain Adaptation
abstract
Most cross-modal methods assume that training and testing data come from the same domain, which is often not the case in real-world scenarios due to cross-modal domain shifts and potential unknown concepts. Moreover, cross-modal shifts hinder the capture of unknown concepts, and the presence of unknown concepts can in turn exacerbate the cross-modal shifts. To address these challenges, this paper proposes a new paradigm called Active Cross-Modal Domain Adaptation (ACM-DA), wherein only cross-modal data from the source domain and uni-modal data from the target domain are utilized. To concurrently mitigate the adverse effects of both cross-modal domain shifts and unknown concepts, we propose a Curiosity-Driven Active Adaptation Network (CD-A2N), selectively annotating samples to maximize performance gain. First, we present Curiosity Arousal within Cross-modal Domain Adaptation (CA-CDA) to explore the complexity and novelty characteristics of target samples, while reducing cross-modal discrepancy and aligning source and target domains. Second, Curiosity-driven Active Learning (CAL) is devised to strategically select a subset of target samples for annotation, aiming to achieve more valuable data selection at a small labeling cost. Finally, we jointly train CA-CDA and CAL with the newly labeled target domain sub-dataset to alleviate the above issues. Extensive experiments demonstrate that CD-A2N provides an effective solution for achieving ACM-DA. Code will be available athttps://github.com/Feliciaxyao/ACM-DA.
Xuan Yao 0001, Junyu Gao 0002, Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu
IEEE Trans. Multim.4
2025 Person in Uniforms Re-Identification
abstract
Person in Uniforms Re-identification (PU-ReID) is an emerging computer vision task for various intelligent video surveillance applications. PU-ReID is much understudied due to the absence of large-scale annotated datasets, also this task is extremely challenging because many individuals captured in surveillance videos wear same clothing, introducing significant interference for retrieval tasks owing to the high visual similarity of outfits and subtle differences among individuals. This research initiates the exploration of person in uniforms re-identification, a novel and challenging task tailored for real industrial scenarios. To address these issues, a novel framework is proposed for PU-ReID, which aims to reduce the visual impact of similar uniforms and learn the unique cues derived from human parts and detailed visual features. Specifically, several novel techniques are built in this study: first, a uniform feature separation method with orthogonal constraints is proposed to extract non-uniform features. Second, multi-view subspace feature alignment is introduced to integrate soft-biometrics including optics-related visual features, contextual information of human parts, and cloth-invariant biometric features. In addition, to close the gap between academic research and real-world settings, a new person in uniforms ReID dataset named PU-151 is constructed, which consists of 151 gas station employees in uniforms from 1,488 videos. At last, extensive experiments conducted on five datasets demonstrate that the proposed approach significantly outperforms the state-of-the-art methods. This advancement can drive further developments in re-identification and person search technologies.
Chong-Yang Xiang, Xiao Wu 0001, Jun-Yan He, Zhaoquan Yuan, Tingquan He
ACM Trans. Multim. Comput. Commun. Appl.4
2024 MagicCartoon: 3D Pose and Shape Estimation for Bipedal Cartoon Characters
abstract
The 3D model can be estimated by regressing the pose and shape parameters from the image data of the digital model. The reconstruction of 3D cartoon characters poses a challenging task due to diverse visual representations and postural variations. This paper proposes a dual-branch structure named MagicCartoon for 3D bipedal cartoon character estimation, which models pose and shape independently through feature decoupling. Considering the correlation between category difference and shape parameters, a hybrid feature fusion technique is introduced, which integrates the global features of the original image with the corresponding local features expressed by the puzzle image, reducing the abstractness of understanding shape parameter differences. To semantically align image and geometric between feature space, a geometric-guided feedback loop is proposed in an iterative way, so that the pose of modeling results can be expressed consistently with the image. Moreover, a feature consistency loss is designed to augment the training data by incorporating the same character with different postures and the same posture of different characters. It enhances the correlation between the features extracted by the backbone network and the specific task. Experiments conducted on the 3DBiCar dataset demonstrate that MagicCartoon outperforms the state-of-the-art methods.
Yu-Pei Song, Yuantong Liu, Xiao Wu 0001, Qi He 0007, Zhaoquan Yuan, Ao Luo
ACM Multimedia5
2024 TMM-CLIP: Task-guided Multi-Modal Alignment for Rehearsal-Free Class Incremental Learning
Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Zechao Li, Changsheng Xu
MMAsia2
2023 Human-Object-Object Interaction: Towards Human-Centric Complex Interaction Detection
abstract
Localizing and recognizing interactive actions in videos is a pivotal yet intricate task that paves the way towards profound video comprehension. Recent advancements in Human-Object Interaction (HOI) detection, which involve detecting and localizing the interactions between human and object pairs, have undeniably marked significant progress. However, the realm of human-object-object interaction, an essential aspect of real-world industrial applications, remains largely uncharted. In this paper, we introduce a novel task referred to as Human-Object-Object Interaction (HOOI) detection and present a cutting-edge method named the Human-Object-Object Interaction Network (H2O-Net). The proposed H2O-Net is comprised of two principal modules: sequential motion feature extraction and HOOI modeling. The former module delves into the gradually evolving visual characteristics of entities throughout the HOOI process, harnessing spatial-temporal features across multiple fine-grained partitions. Conversely, the latter module aspires to encapsulate HOOI actions through intricate interactions between entities. It commences by capturing and amalgamating two sub-interaction features to extract comprehensive HOOI features, subsequently refining them using the interaction cues embedded within the long-term global context. Furthermore, we contribute to the research community by constructing a new video dataset, dubbed the HOOI dataset. The actions encompassed within this dataset pertain to pivotal operational behaviors in industrial manufacturing, imbuing it with substantial application potential and serving as a valuable addition to the existing repertoire of interaction action detection datasets. Experimental evaluations conducted on the proposed HOOI and widely-used AVA datasets demonstrate that our method outperforms existing state-of-the-art techniques by margins of 6.16 mAP and 1.9 mAP, respectively, thus substantiating its effectiveness.
Mingxuan Zhang 0001, Xiao Wu 0001, Zhaoquan Yuan, Qi He 0007, Xiang Huang 0004
ACM Multimedia3
2023 Learning Surface-awareness Network for X-Ray Prohibited Item Detection
abstract
X-ray image security detection is a crucial method used to identify various types of prohibited items in luggage. However, the unique characteristics of X-ray imaging can result in the loss of intricate surface details, leading to subpar detection of prohibited items within X-ray images. In this paper, a Surface-aware Prohibited Item X-ray Detection Network (SPIXDet) is proposed to address this issue, which incorporates two key components: the Boundary Aggregation Module (BAM) and the Global Cross-Feature Downsampling layer (GCFD). The BAM module effectively mines image edge information while minimizing the number of parameters involved. Meanwhile, the GCFD module is introduced to mitigate chaotic interference caused by undifferentiated boundary boosting. The surface-aware capability of the model can be enhanced through the BAM and GCFD module. Furthermore, the Focal-SIoU loss function is introduced to increase positioning accuracy and optimize the model training process. To validate the effectiveness of our model, extensive experiments are conducted on the SIXray100 dataset, and the results demonstrate the advantages of SPIXDet compared to other X-ray prohibited item detection methods.
Wei Li 0110, Zhaoquan Yuan, Xiao Wu 0001
MMAsia3
2022 Learning Graph-based Residual Aggregation Network for Group Activity Recognition
abstract
Group activity recognition aims to understand the overall behavior performed by a group of people. Recently, some graph-based methods have made progress by learning the relation graphs among multiple persons. However, the differences between an individual and others play an important role in identifying confusable group activities, which have not been elaborately explored by previous methods. In this paper, a novel Graph-based Residual AggregatIon Network (GRAIN) is proposed to model the differences among all persons of the whole group, which is end-to-end trainable. Specifically, a new local residual relation module is explicitly proposed to capture the local spatiotemporal differences of relevant persons, which is further combined with the multi-graph relation networks. Moreover, a weighted aggregation strategy is devised to adaptively select multi-level spatiotemporal features from the appearance-level information to high level relations. Finally, our model is capable of extracting a comprehensive representation and inferring the group activity in an end-to-end manner. The experimental results on two popular benchmarks for group activity recognition clearly demonstrate the superior performance of our method in comparison with the state-of-the-art methods.
Wei Li 0110, Tianzhao Yang, Xiao Wu 0001, Zhaoquan Yuan
IJCAI4
2022 Domain-Specific Conditional Jigsaw Adaptation for Enhancing transferability and Discriminability
abstract
Unsupervised Domain Adaptation (UDA) aims to transfer knowledge from a label-rich source domain to a target domain where the label is unavailable. Existing approaches tend to reduce the distribution discrepancy between the source and target domains or assign the pseudo target labels to implement a self-training strategy. However, the transferability or discriminability lackage of the traditional methods results in the limited ability to generalize on the target domain. To remedy this issue, a novel unsupervised domain adaptation framework called Domain-specific Conditional Jigsaw Adaptation Network (DCJAN) is proposed for UDA, which simultaneously encourages the network to extract transferable and discriminative features. To improve the discriminability, a conditional jigsaw module is presented to reconstruct class-aware features of the original images by reconstructing that of corresponding shuffled images. Moreover, in order to enhance the transferability, a domain-specific jigsaw adaptation is proposed to deal with the domain gaps, which utilizes the prior knowledge of jigsaw puzzles to reduce mismatching. It trains conditional jigsaw modules for each domain and updates the shared feature extractor to make the domain-specific conditional jigsaw modules could perform well not only on the corresponding domain but also on the other domain. A consistent conditioning strategy is proposed to ensure the safe training of conditional jigsaw. Experiments conducted on the widely-used Office-31, Office-Home, VisDA-2017, and DomainNet datasets demonstrate the effectiveness of the proposed approach, which outperforms the state-of-the-art methods.
Qi He 0007, Zhaoquan Yuan, Xiao Wu 0001, Jun-Yan He
ACM Multimedia2
2021 Meta-Learning Causal Feature Selection for Stable Prediction
abstract
Conventional predictive models in machine learning are based on I.I.D. hypothesis between training and testing data. However, such a hypothesis is fragile in the real world, and the model minimizing empirical errors on training data does not perform well on testing data, which makes the prediction unstable. This instability can be found widely in domain generalization, active learning, and transfer learning, etc. In this paper, we propose a novel Meta-learning Causal Feature Selection (MCFS) model for general Non-I.I.D. image classification. In MCFS, we jointly optimize a convolutional network and a causal parameter for identifying causal variables on meta-training and meta-testing data which simulate the distribution shifts in Non-I.I.D. problems. Extensive experiments conducted on public VLCS and NICO datasets demonstrate the effectiveness of the proposed MCFS, which outperforms the state-of-the-art methods.
Zhaoquan Yuan, Xiao Wu 0001, Bing-Kun Bao, Changsheng Xu
ICME1
2021 Hierarchical Multi-Task Learning for Diagram Question Answering with Multi-Modal Transformer
abstract
Diagram question answering (DQA) is an effective way to evaluate the reasoning ability for diagram semantic understanding, which is a very challenging task and largely understudied compared with natural images. Existing separate two-stage methods for DQA are limited in ineffective feedback mechanisms. To address this problem, in this paper, we propose a novel structural parsing-integrated Hierarchical Multi-Task Learning (HMTL) model for diagram question answering based on a multi-modal transformer framework. In the proposed paradigm of multi-task learning, the two tasks of diagram structural parsing and question answering are in the different semantic levels and equipped with different transformer blocks, which constituents a hierarchical architecture. The structural parsing module encodes the information of constituents and their relationships in diagrams, while the diagram question answering module decodes the structural signals and combines question-answers to infer correct answers. Visual diagrams and textual question-answers are interplayed in the multi-modal transformer, which achieves cross-modal semantic comprehension and reasoning. Extensive experiments on the benchmark AI2D and FOODWEBS datasets demonstrate the effectiveness of our proposed HMTL over other state-of-the-art methods.
Zhaoquan Yuan, Xiao Wu 0001, Changsheng Xu
ACM Multimedia1
2021 Contrastive Learning in Frequency Domain for Non-I.I.D. Image Classification
Huan Shao 0005, Zhaoquan Yuan, Xiao Wu 0001
MMM (1)2
2021 DB-LSTM: Densely-connected Bi-directional LSTM for human action recognition
Jun-Yan He, Xiao Wu 0001, Zhi-Qi Cheng, Zhaoquan Yuan, Yu-Gang Jiang 0001
Neurocomputing4
2021 Adversarial Multimodal Network for Movie Story Question Answering
abstract
Visual question answering by using information from multiple modalities has attracted more and more attention in recent years. However, it is a very challenging task, as the visual content and natural language have quite different statistical properties. In this work, we present a method called Adversarial Multimodal Network (AMN) to better understand video stories for question answering. In AMN, we propose to learn multimodal feature representations by finding a more coherent subspace for video clips and the corresponding texts (e.g., subtitles and questions) based on generative adversarial networks. Moreover, a self-attention mechanism is developed to enforce our newly introduced consistency constraint in order to preserve the self-correlation between the visual cues of the original video clips in the learned multimodal representations. Extensive experiments on the benchmark MovieQA and TVQA datasets show the effectiveness of our proposed AMN over other published state-of-the-art methods.
Zhaoquan Yuan, Lixin Duan, Xiao Wu 0001, Changsheng Xu
IEEE Trans. Multim.1
2015 Learning Feature Hierarchies: A Layer-Wise Tag-Embedded Approach
abstract
Feature representation learning is an important and fundamental task in multimedia and pattern recognition research. In this paper, we propose a novel framework to explore the hierarchical structure inside the images from the perspective of feature representation learning, which is applied to hierarchical image annotation. Different from the current trend in multimedia analysis of using pre-defined features or focusing on the end-task “flat” representation, we propose a novel layer-wise tag- embedded deep learning (LTDL) model to learn hierarchical features which correspond to hierarchical semantic structures in the tag hierarchy . Unlike most existing deep learning models, LTDL utilizes both the visual content of the image and the hierarchical information of associated social tags. In the training stage, the two kinds of information are fused in a bottom-up way. Supervised training and multi-modal fusion alternate in a layer-wise way to learn feature hierarchies. To validate the effectiveness of LTDL, we conduct extensive experiments for hierarchical image annotation on a large-scale public dataset. Experimental results show that the proposed LTDL can learn representative features with improved performances.
Zhaoquan Yuan, Changsheng Xu, Jitao Sang 0001, Shuicheng Yan, M. Shamim Hossain
IEEE Trans. Multim.1
2014 A Unified Framework of Latent Feature Learning in Social Media
abstract
The current trend in social media analysis and application is to use the pre-defined features and devoted to the later model development modules to meet the end tasks. Representation learning has been a fundamental problem in machine learning, and widely recognized as critical to the performance of end tasks. In this paper, we provide evidence that specially learned features will addresses the diverse, heterogeneous, and collective characteristics of social media data. Therefore, we propose to transfer the focus from the model development to latent feature learning, and present a unified framework of latent feature learning on social media. To address the noisy, diverse, heterogeneous, and interconnected characteristics of social media data, the popular deep learning is employed due to its excellent abstract abilities. In particular, we instantiate the proposed framework by (1) designing a novel relational generative deep learning model to solve the social media link analysis task, and (2) developing a multimodal deep learning to lambda rank model towards the social image retrieval task. We show that the derived latent features lead to improvement in both of the social media tasks.
Zhaoquan Yuan, Jitao Sang 0001, Changsheng Xu, Yan Liu 0004
IEEE Trans. Multim.1
2013 Tag-aware image classification via Nested Deep Belief nets
abstract
With the rising of internet photos-sharing web sites, the rich aware text information surrounding images on the sites are proved helpful to improve the image classification. This paper presents a novel nested deep learning model called Nested Deep Belief Network(NDBN) for tag-aware image classification. A multi-layer structure of Deep Belief Network(DBN) is established to learn a unified representation of visual feature and tag feature for an image, and an additional Gaussian Restricted Boltzmann Machine is built to capture the tag-tag dependency. Compared with conventional methods, the proposed model can not only find correlations across modalities, but mine the importance for different tags, and also bring about low-rank tag feature representation. We conduct experiments over the MIR Flickr dataset and the results show that the proposed NDBN model outperforms the existing image classification techniques.
Zhaoquan Yuan, Jitao Sang 0001, Changsheng Xu
ICME1
2013 Latent feature learning in social media network
abstract
The current trend in social media analysis and application is to use the pre-defined features and devoted to the later model development modules to meet the end tasks. In this work, we claim that representation is critical to the end tasks and contributes much to the model development module. We provide evidence that specially learned feature well addresses the diverse, heterogeneous and collective characteristics of social media data. Therefore, we propose to transfer the focus from the model development to latent feature learning, and present a general feature learning framework based on the popular deep architecture. In particular, following the proposed framework, we design a novel relational generative deep learning model to test the idea on link analysis tasks in the social media networks. We show that the derived latent features well embed both the media content and their observed links, leading to improvement in social media tasks of user recommendation and social image annotation.
Zhaoquan Yuan, Jitao Sang 0001, Yan Liu 0004, Changsheng Xu
ACM Multimedia1