Tong Yang 0005

dblp:44/7710-5 · DBLP profile ↗
← Back
10ranked-venue papers
2as first author
5since 2021 · last 2022
0000-0002-2276-8534ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2022 Anchor DETR: Query Design for Transformer-Based Detector
abstract
In this paper, we propose a novel query design for the transformer-based object detection. In previous transformer-based detectors, the object queries are a set of learned embeddings. However, each learned embedding does not have an explicit physical meaning and we cannot explain where it will focus on. It is difficult to optimize as the prediction slot of each object query does not have a specific mode. In other words, each object query will not focus on a specific region. To solve these problems, in our query design, object queries are based on anchor points, which are widely used in CNN-based detectors. So each object query focuses on the objects near the anchor point. Moreover, our query design can predict multiple objects at one position to solve the difficulty: ``one region, multiple objects''. In addition, we design an attention variant, which can reduce the memory cost while achieving similar or better performance than the standard attention in DETR. Thanks to the query design and the attention variant, the proposed detector that we called Anchor DETR, can achieve better performance and run faster than the DETR with 10x fewer training epochs. For example, it achieves 44.2 AP with 19 FPS on the MSCOCO dataset when using the ResNet50-DC5 feature for training 50 epochs. Extensive experiments on the MSCOCO benchmark prove the effectiveness of the proposed methods. Code is available at https://github.com/megvii-research/AnchorDETR.
Xiangyu Zhang 0005, Tong Yang 0005, Jian Sun 0001
AAAI3
2022 LGD: Label-Guided Self-Distillation for Object Detection
abstract
In this paper, we propose the first self-distillation framework for general object detection, termed LGD (Label-Guided self-Distillation). Previous studies rely on a strong pretrained teacher to provide instructive knowledge that could be unavailable in real-world scenarios. Instead, we generate an instructive knowledge by inter-and-intra relation modeling among objects, requiring only student representations and regular labels. Concretely, our framework involves sparse label-appearance encoding, inter-object relation adaptation and intra-object knowledge mapping to obtain the instructive knowledge. They jointly form an implicit teacher at training phase, dynamically dependent on labels and evolving student representations. Modules in LGD are trained end-to-end with student detector and are discarded in inference. Experimentally, LGD obtains decent results on various detectors, datasets, and extensive tasks like instance segmentation. For example in MS-COCO dataset, LGD improves RetinaNet with ResNet-50 under 2x single-scale training from 36.2% to 39.0% mAP (+ 2.8%). It boosts much stronger detectors like FCOS with ResNeXt-101 DCN v2 under 2x multi-scale training from 46.1% to 47.9% (+ 1.8%). Compared with a classical teacher-based method FGFI, LGD not only performs better without requiring pretrained teacher but also reduces 51% training cost beyond inherent student learning.
Peizhen Zhang, Zijian Kang, Tong Yang 0005, Xiangyu Zhang 0005, Nanning Zheng 0001, Jian Sun 0001
AAAI3
2021 Co-mining: Self-Supervised Learning for Sparsely Annotated Object Detection
abstract
Object detectors usually achieve promising results with the supervision of complete instance annotations. However, their performance is far from satisfactory with sparse instance annotations. Most existing methods for sparsely annotated object detection either re-weight the loss of hard negative samples or convert the unlabeled instances into ignored regions to reduce the interference of false negatives. We argue that these strategies are insufficient since they can at most alleviate the negative effect caused by missing annotations. In this paper, we propose a simple but effective mechanism, called Co-mining, for sparsely annotated object detection. In our Co-mining, two branches of a siamese network predict the pseudo-label sets for each other. To enhance multi-view learning and better mine unlabeled instances, the original image and corresponding augmented image are used as the inputs of two branches of the siamese network, respectively. Co-mining can serve as a general training mechanism applied to most of modern object detectors. Experiments are performed on MS COCO dataset with three different sparsely annotated settings using two typical frameworks: anchor-based detector RetinaNet and anchor-free detector FCOS. Experimental results show that our Co-mining with RetinaNet achieves 1.4%∼2.1% improvements compared with different baselines and surpasses existing methods under the same sparsely annotated setting.
Tiancai Wang, Tong Yang 0005, Jiale Cao, Xiangyu Zhang 0005
AAAI2
2021 Points As Queries: Weakly Semi-Supervised Object Detection by Points
abstract
We propose a novel point annotated setting for the weakly semi-supervised object detection task, in which the dataset comprises small fully annotated images and large weakly annotated images by points. It achieves a balance between tremendous annotation burden and detection performance. Based on this setting, we analyze existing detectors and find that these detectors have difficulty in fully exploiting the power of the annotated points. To solve this, we introduce a new detector, Point DETR, which extends DETR by adding a point encoder. Extensive experiments conducted on MS-COCO dataset in various data settings show the effectiveness of our method. In particular, when using 20% fully labeled data from COCO, our detector achieves a promising performance, 33.3 AP, which outperforms a strong baseline (FCOS) by 2.0 AP, and we demonstrate the point annotations bring over 10 points in various AR metrics.
Liangyu Chen 0002, Tong Yang 0005, Xiangyu Zhang 0005, Wei Zhang 0016, Jian Sun 0001
CVPR2
2021 You Only Look One-Level Feature
abstract
This paper revisits feature pyramids networks (FPN) for one-stage detectors and points out that the success of FPN is due to its divide-and-conquer solution to the optimization problem in object detection rather than multi-scale feature fusion. From the perspective of optimization, we introduce an alternative way to address the problem instead of adopting the complex feature pyramids - utilizing only one-level feature for detection. Based on the simple and efficient solution, we present You Only Look One-level Feature (YOLOF). In our method, two key components, Dilated Encoder and Uniform Matching, are proposed and bring considerable improvements. Extensive experiments on the COCO benchmark prove the effectiveness of the proposed model. Our YOLOF achieves comparable results with its feature pyramids counterpart RetinaNet while being 2.5× faster. Without transformer layers, YOLOF can match the performance of DETR in a single-level feature manner with 7× less training epochs. Code is available at https://github.com/megvii-model/YOLOF.
Qiang Chen 0007, Tong Yang 0005, Xiangyu Zhang 0005, Jian Cheng 0001, Jian Sun 0001
CVPR3
2020 Learning Human-Object Interaction Detection Using Interaction Points
abstract
Understanding interactions between humans and objects is one of the fundamental problems in visual classification and an essential step towards detailed scene understanding. Human-object interaction (HOI) detection strives to localize both the human and an object as well as the identification of complex interactions between them. Most existing HOI detection approaches are instance-centric where interactions between all possible human-object pairs are predicted based on appearance features and coarse spatial information. We argue that appearance features alone are insufficient to capture complex human-object interactions. In this paper, we therefore propose a novel fully-convolutional approach that directly detects the interactions between human-object pairs. Our network predicts interaction points, which directly localize and classify the inter-action. Paired with the densely predicted interaction vectors, the interactions are associated with human and object detections to obtain final predictions. To the best of our knowledge, we are the first to propose an approach where HOI detection is posed as a keypoint detection and grouping problem. Experiments are performed on two popular benchmarks: V-COCO and HICO-DET. Our approach sets a new state-of-the-art on both datasets. Code is available at https://github.com/vaesl/IP-Net.
Tiancai Wang, Tong Yang 0005, Martin Danelljan, Fahad Shahbaz Khan, Xiangyu Zhang 0005, Jian Sun 0001
CVPR2
2019 DetNAS: Backbone Search for Object Detection
abstract
Object detectors are usually equipped with backbone networks designed for image classification. It might be sub-optimal because of the gap between the tasks of image classification and object detection. In this work, we present DetNAS to use Neural Architecture Search (NAS) for the design of better backbones for object detection. It is non-trivial because detection training typically needs ImageNetpre-training while NAS systems require accuracies on the target detection task as supervisory signals. Based on the technique of one-shot supernet, which contains all possible networks in the search space, we propose a framework for backbone search on object detection. We train the supernet under the typical detector training schedule: ImageNet pre-training and detection fine-tuning. Then, the architecture search is performed on the trained supernet, using the detection task as the guidance. This framework makes NAS on backbones very efficient. In experiments, we show the effectiveness of DetNAS on various detectors, for instance, one-stage RetinaNetand the two-stage FPN. We empirically find that networks searched on object detection shows consistent superiority compared to those searched on ImageNet classification. The resulting architecture achieves superior performance than hand-crafted networks on COCO with much less FLOPs complexity.
Yukang Chen, Tong Yang 0005, Xiangyu Zhang 0005, Gaofeng Meng, Xinyu Xiao, Jian Sun 0001
NeurIPS2
2018 Multi-Label Dilated Recurrent Network for Sequential Face Alignment
abstract
Compared with detection in still image, sequential face landmark detection in video is relatively less studied. In this work, we present a novel network with a dilated residual network and a residual convolutional LSTM to preserve the detection acuity at both spatial and temporal dimensions respectively. We also introduce and compare multi-label loss with regression loss and multi-class loss for face landmark detection. We perform extensive experiments to verify the contribution of different components. Our method achieves state-of-the-art performance on benchmark datasets.
Tong Yang 0005, Shizheng Qin, Junchi Yan
ICME1
2018 MetaAnchor: Learning to Detect Objects with Customized Anchors
abstract
We propose a novel and flexible anchor mechanism named MetaAnchor for object detection frameworks. Unlike many previous detectors model anchors via a predefined manner, in MetaAnchor anchor functions could be dynamically generated from the arbitrary customized prior boxes. Taking advantage of weight prediction, MetaAnchor is able to work with most of the anchor-based object detection systems such as RetinaNet. Compared with the predefined anchor scheme, we empirically find that MetaAnchor is more robust to anchor settings and bounding box distributions; in addition, it also shows the potential on the transfer task. Our experiment on COCO detection task shows MetaAnchor consistently outperforms the counterparts in various scenarios.
Tong Yang 0005, Xiangyu Zhang 0005, Jian Sun 0001
NeurIPS1
2017 Automatic tongue image matting for remote medical diagnosis
abstract
With the rapid adoption of smartphones and tablets, more and more remote medical diagnostic applications have mushroomed. Tongue Diagnosis (TD) is a kind of noninvasive diagnostic technique, which offers significant information for health conditions. However, it is rather tough to extract the tongue from a high-quality image, in which there is a definite large area of the tongue, to say nothing of extracting the tongue from a digital image captured by photographers who often lack the necessary skills using different mobile front facing cameras. Fundamentally, automatic tongue image segmentation is difficult due to two special factors: the particularity of the tongue and the diversity of the image. Our paper first addresses these problems by proposing a new end-to-end iterative network for tongue image matting, which directly learns the alpha matte from the input image by correcting misunderstanding in intermediate steps. Neither user interaction nor initialization is required. In addition, we create a large-scale tongue image matting dataset including 7,0680 training images. Compared with other high-performance algorithms, our algorithm achieves the true sense of the pixel-wise automatic tongue segmentation.
Tong Yang 0005, Yangyang Hu, Menglong Xu, Fufeng Li
BIBM2