Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wolin Liang

dblp:414/9752 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
0009-0002-4569-3448ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Image recognition and object detection · 23% Video understanding and tracking · 18% 3D vision · 18%

Topics — the 6 heaviest of 8, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking
action recognition
1.012026
What-Meets-Where: Unified Learning of Action and Contact Localization in Images · AAAI 2026
Computer vision › Vision and language
cross-modal transformer
1.012026
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection · AAAI 2026
Computer vision › 3D vision › 3d scene understanding › object relation reasoning
human-object interaction
1.012026
What-Meets-Where: Unified Learning of Action and Contact Localization in Images · AAAI 2026
Computer vision › Image recognition and object detection
human-object interaction detection
1.012026
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection · AAAI 2026
Computer vision › Face, body and person analysis › human body analysis
body part detection
0.312026
What-Meets-Where: Unified Learning of Action and Contact Localization in Images · AAAI 2026
Computer vision › Image recognition and object detection › object detection › detection transformer
DETR-based detection
0.312026
QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection · AAAI 2026

Methods — techniques the papers use, named apart from their topics

transformer query initialization · 1.0prior-guided segmentation · 1.0knowledge distillation · 1.0interaction inference · 1.0cross-modal attention · 1.0contact prior aware module · 1.0
YearPublicationVenuePosition
2026 What-Meets-Where: Unified Learning of Action and Contact Localization in Images
abstract
People control their bodies to establish contact with the environment. To comprehensively understand actions across diverse visual contexts, it is essential to simultaneously consider what action is occurring and where it is happening. Current methodologies, however, often inadequately capture this duality, typically failing to jointly model both action semantics and their spatial contextualization within scenes. To bridge this gap, we introduce a novel vision task that simultaneously predicts high-level action semantics and fine-grained body-part contact regions. Our proposed framework, PaIR-Net, comprises three key components: the Contact Prior Aware Module (CPAM) for identifying contact-relevant body parts, the Prior-Guided Concat Segmenter (PGCS) for pixel-wise contact segmentation, and the Interaction Inference Module (IIM) responsible for integrating global interaction relationships. To facilitate this task, we present PaIR (Part-aware Interaction Representation), a comprehensive dataset containing 13,979 images that encompass 654 actions, 80 object categories, and 17 body parts. Experimental evaluation demonstrates that PaIR-Net significantly outperforms baseline approaches, while ablation studies confirm the efficacy of each architectural component.
Yuxiao Wang 0003, Wolin Liang, Weiying Xue, Zhenao Wei, Nan Zhuang, Qi Liu 0005
AAAI3
2026 QueryCraft: Transformer-Guided Query Initialization for Enhanced Human-Object Interaction Detection
abstract
Human-Object Interaction (HOI) detection aims to localize human-object pairs and recognize their interactions in images. Although DETR-based methods have recently emerged as the mainstream framework for HOI detection, they still suffer from a key limitation: Randomly initialized queries lack explicit semantics, leading to suboptimal detection performance. To address this challenge, we propose QueryCraft, a novel plug-and-play HOI detection framework that incorporates semantic priors and guided feature learning through transformer-based query initialization. Central to our approach is ACTOR (Action-aware Cross-modal TransfORmer), a cross-modal Transformer encoder that jointly attends to visual regions and textual prompts to extract action-relevant features. Rather than merely aligning modalities, ACTOR leverages language-guided attention to infer interaction semantics and produce semantically meaningful query representations. To further enhance object-level query quality, we introduce a Perceptual Distilled Query Decoder (PDQD), which distills object category awareness from a pre-trained detector to serve as object query initiation. This dual-branch query initialization enables the model to generate more interpretable and effective queries for HOI detection. Extensive experiments on HICO-Det and V-COCO benchmarks demonstrate that our method achieves state-of-the-art performance and strong generalization.
Yuxiao Wang 0003, Wolin Liang, Weiying Xue, Nan Zhuang, Qi Liu 0005
AAAI2