Xubin Zhong

dblp:272/4321 · DBLP profile ↗
← Back
5ranked-venue papers
4as first author
4since 2021 · last 2022
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 4 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Image recognition and object detection · 67% Efficient and distributed learning · 10% Vision and language · 9%

Topics — the 6 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
human-object interaction detection
2.652022
Towards Hard-Positive Query Mining for DETR-Based Human-Object Interaction Detection · ECCV (27) 2022
Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection · CVPR 2022
Polysemy Deciphering Network for Robust Human-Object Interaction Detection · Int. J. Comput. Vis. 2021
Computer vision › Image recognition and object detection › object detection
detection transformer
0.612022
Towards Hard-Positive Query Mining for DETR-Based Human-Object Interaction Detection · ECCV (27) 2022
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.612022
Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection · CVPR 2022
Natural language and speech › Information extraction and text analysis › word sense disambiguation
polysemy resolution
0.412020
Polysemy Deciphering Network for Human-Object Interaction Detection · ECCV (20) 2020
Machine learning › Deep learning architectures and training
transformer
0.212022
Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection · CVPR 2022
Machine learning › Deep learning architectures and training
attention mechanism
0.112021
Glance and Gaze: Inferring Action-Aware Points for One-Stage Human-Object Interaction Detection · CVPR 2021

Methods — techniques the papers use, named apart from their topics

neural network · 0.9knowledge distillation · 0.6hard-positive query mining · 0.6data augmentation · 0.6context-consistent stitching · 0.6hard negative attentive loss · 0.5action-aware point modeling · 0.5
YearPublicationVenuePosition
2022 Distillation Using Oracle Queries for Transformer-based Human-Object Interaction Detection
abstract
Transformer-based methods have achieved great success in the field of human-object interaction (HOI) detection. However, these models tend to adopt semantically ambigu-ous queries, which lowers the transformer's representation learning power. Moreover, there are a very limited num-ber of labeled human-object pairs for most images in ex-isting datasets, which constrains the transformer's set pre-diction power. To handle the first problem, we propose an efficient knowledge distillation model, named Distillation using Oracle Queries (DOQ), which shares parameters be-tween teacher and student networks. The teacher network adopts oracle queries that are semantically clear and gener-ates high-quality decoder embeddings. By mimicking both the attention maps and decoder embeddings of the teacher network, the representation learning power of the student network is significantly promoted. To address the sec-ond problem, we introduce an efficient data augmentation method, named Context-Consistent Stitching (CCS), which generates complicated images online. Each new image is obtained by stitching labeled human-object pairs cropped from multiple training images. By selecting source images with similar context, the new synthesized image is made visually realistic. Our methods significantly promote both the accuracy and training efficiency of transformer-based HOI detection models. Experimental results show that our proposed approach consistently outperforms state-of-the-art methods on three benchmarks: HICO-DET, HOI-A, and V-COCO. Code is available at ht tps: / / gi thub. com/ SherlockHolmes221/DOQ.
Xian Qu, Changxing Ding, Xingao Li, Xubin Zhong, Dacheng Tao
CVPR4
2022 Towards Hard-Positive Query Mining for DETR-Based Human-Object Interaction Detection
Xubin Zhong, Changxing Ding, Zijian Li 0011, Shaoli Huang
ECCV (27)1
2021 Glance and Gaze: Inferring Action-Aware Points for One-Stage Human-Object Interaction Detection
abstract
Modern human-object interaction (HOI) detection approaches can be divided into one-stage methods and two-stage ones. One-stage models are more efficient due to their straightforward architectures, but the two-stage models are still advantageous in accuracy. Existing one-stage models usually begin by detecting predefined interaction areas or points, and then attend to these areas only for interaction prediction; therefore, they lack reasoning steps that dynamically search for discriminative cues. In this paper, we propose a novel one-stage method, namely Glance and Gaze Network (GGNet), which adaptively models a set of action-aware points (ActPoints) via glance and gaze steps. The glance step quickly determines whether each pixel in the feature maps is an interaction point. The gaze step leverages feature maps produced by the glance step to adaptively infer ActPoints around each pixel in a progressive manner. Features of the refined ActPoints are aggregated for interaction prediction. Moreover, we design an action-aware approach that effectively matches each detected interaction with its associated human-object pair, along with a novel hard negative attentive loss to improve the optimization of GGNet. All the above operations are conducted simultaneously and efficiently for all pixels in the feature maps. Finally, GGNet outperforms state-of-the-art methods by significant margins on both V-COCO and HICO-DET benchmarks. Code of GGNet is available at https://github.com/SherlockHolmes221/GGNet.
Xubin Zhong, Xian Qu, Changxing Ding, Dacheng Tao
CVPR1
2021 Polysemy Deciphering Network for Robust Human-Object Interaction Detection
Xubin Zhong, Changxing Ding, Xian Qu, Dacheng Tao
Int. J. Comput. Vis.1
2020 Polysemy Deciphering Network for Human-Object Interaction Detection
Xubin Zhong, Changxing Ding, Xian Qu, Dacheng Tao
ECCV (20)1