Zhaoquan Yuan

dblp:135/5072 · DBLP profile ↗
← Back
3ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0002-4083-5155ORCID · verified

Domains — venue-derived; a paper can count in several

Other / Interdisciplinary · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 HOOI Detection: Cascade-Clue Integrated Modeling over Multiple Temporal Segments
abstract
To fully comprehend a visual scene, recognizing and localizing interaction actions are essential components. Recently, significant advances have been made in detecting human-object interaction actions, which aim to capture pairwise relations between entities in the scene. Although these methods have made significant progress, they ignore the human-object-object interaction (HOOI) actions that frequently occur between a human and two objects in the real world. To advance related research, a new task named HOOI detection is introduced. It aims to accurately localize the humans in each video frame and identify the HOOI actions they perform. For this purpose, two novel HOOI datasets oriented to industrial production and daily life are constructed. These new datasets provide essential data support for in-depth research of HOOI detection. Furthermore, a cutting-edge method named Cascade-Clue Integrated Modeling over Multiple Temporal Segments (C2TS) is proposed for effectively detecting HOOI actions. Specifically, considering the phased characteristics of the action, C2TS comprehensively considers the HOOI information in the preceding, neighborhood, and subsequent temporal segments. For each temporal segment, the Cascaded Modeling and Clue Augmentation methods are applied to extract the corresponding HOOI features. The final detection result is obtained by classifying the aggregated HOOI features from the three temporal segments. Experiments conducted on the two proposed HOOI-related datasets vividly demonstrate that our method outperforms state-of-the-art approaches by achieving remarkable improvements of approximately 3% and 5%, which powerfully validates its efficacy in tackling the given challenge.
Mingxuan Zhang 0001, Qi He 0007, Zhaoquan Yuan, Tingquan He
ICMR3
2024 TMM-CLIP: Task-guided Multi-Modal Alignment for Rehearsal-Free Class Incremental Learning
Yuankang Pan, Zhaoquan Yuan, Xiao Wu 0001, Zechao Li, Changsheng Xu
MMAsia2
2023 Learning Surface-awareness Network for X-Ray Prohibited Item Detection
abstract
X-ray image security detection is a crucial method used to identify various types of prohibited items in luggage. However, the unique characteristics of X-ray imaging can result in the loss of intricate surface details, leading to subpar detection of prohibited items within X-ray images. In this paper, a Surface-aware Prohibited Item X-ray Detection Network (SPIXDet) is proposed to address this issue, which incorporates two key components: the Boundary Aggregation Module (BAM) and the Global Cross-Feature Downsampling layer (GCFD). The BAM module effectively mines image edge information while minimizing the number of parameters involved. Meanwhile, the GCFD module is introduced to mitigate chaotic interference caused by undifferentiated boundary boosting. The surface-aware capability of the model can be enhanced through the BAM and GCFD module. Furthermore, the Focal-SIoU loss function is introduced to increase positioning accuracy and optimize the model training process. To validate the effectiveness of our model, extensive experiments are conducted on the SIXray100 dataset, and the results demonstrate the advantages of SPIXDet compared to other X-ray prohibited item detection methods.
Wei Li 0110, Zhaoquan Yuan, Xiao Wu 0001
MMAsia3