Xingyu Zeng

dblp:135/4927 · DBLP profile ↗
← Back
27ranked-venue papers
5as first author
10since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 23 · 5 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SCALAR: Spatial-concept alignment for robust vision in harsh open world
Xiaoyu Yang 0007, Lijian Xu, Xingyu Zeng, Xiaosong Wang 0001, Hongsheng Li 0001, Shaoting Zhang 0001
Pattern Recognit.3
2025 ARise: Towards Knowledge-Augmented Reasoning via Risk-Adaptive Search
abstract
Large language models (LLMs) have demonstrated impressive capabilities and are receiving increasing attention to enhance their reasoning through scaling test-time compute. However, their application in open-ended, knowledge-intensive, complex reasoning scenarios is still limited. Reasoning-oriented methods struggle to generalize to open-ended scenarios due to implicit assumptions of complete world knowledge. Meanwhile, knowledge-augmented reasoning (KAR) methods fails to address two core challenges: 1) error propagation, where errors in early steps cascade through the chain, and 2) verification bottleneck, where the explore–exploit trade-off arises in multi-branch decision processes. To overcome these limitations, we introduce ARise, a novel framework that integrates risk assessment of intermediate reasoning states with dynamic retrieval-augmented generation (RAG) within a Monte Carlo tree search paradigm. This approach enables effective construction and optimization of reasoning plans across multiple maintained hypothesis branches. Experimental results show that ARise significantly outperforms the state-of-the-art KAR methods by up to 23.10%, and the latest RAG-equipped large reasoning models by up to 25.37%. Our project page is at https://opencausalab.github.io/ARise.
Yize Zhang, Tianshu Wang 0002, Xingyu Zeng, Xianpei Han, Le Sun 0001, Chaochao Lu
ACL (1)5
2025 Spy Inside: Scalable Verification of Dependable Transformers for Event Time Series Systems
abstract
Event time series appear in many software scenarios and are a necessary data type in data analytics systems. Transformers are the preferred type of sequential neural network for advanced analytics on event time series, particularly due to their significant contributions to the recent surge of large language models (LLMs). Event series analytics heavily depends on the quality of input data, which may contain natural measurement errors or adversarial noises. Since the input data deviates from the true state, the opaque nature of neural networks presents a challenge in ensuring the reliability of output, which might be deemed untrustworthy. In this paper, we introduce an innovative formal verification framework for Transformer-based event series systems, leveraging sampling, linear programming, and the extreme value theorem. This framework can support the verification of the dependability of Transformers in managing inputs characterized by unpredictability and uncertainty. To exemplify its utility, we apply our verification approach to verify natural requirements from a real-world event series environments: network traffic classification. It outperforms the current state-of-the-art verifier in terms of effectiveness, providing more stringent verified bounds. Our experimental findings provide valuable benchmarks for guaranteeing reliable deployment of systems in scenarios where the credibility of event data is compromised, and for exposing specific cases in which the expected requirements are not satisfied.
Haodong Deng, Qi Qi 0001, Lu Lu 0015, Zirui Zhuang, Xingyu Zeng, Jinguang Wang, Bo He 0003, Wei Li 0119, Jingyu Wang 0001
ICASSP5
2025 PUMA: Empowering Unified MLLM with Multi-Granular Visual Generation
abstract
Recent advancements in multimodal foundation models have yielded significant progress in vision-language understanding. Initial attempts have also explored the potential of multimodal large language models (MLLMs) for visual content generation. However, existing works have insufficiently addressed the varying granularity demands of different image generation tasks within a unified MLLM paradigm - from the diversity required in text-to-image generation to the precise controllability needed in image manipulation. In this work, we propose PUMA, emPowering Unified MLLM with Multi-grAnular visual generation. PUMA unifies multi-granular visual features as both inputs and outputs of MLLMs, elegantly addressing the different granularity requirements of various image generation tasks within a unified MLLM framework. Following multimodal pretraining and task-specific instruction tuning, PUMA demonstrates proficiency in a wide range of multimodal tasks. This work represents a significant step towards a truly unified MLLM capable of adapting to the granularity demands of various visual tasks. The code and model will be released in https://github.com/rongyaofang/PUMA.
Rongyao Fang, Chengqi Duan, Kun Wang 0056, Hao Li 0069, Linjiang Huang, Hao Tian 0006, Xingyu Zeng, Rui Zhao 0001, Jifeng Dai, Hongsheng Li 0001, Xihui Liu
ICCV7
2025 GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and Editing
abstract
Current image generation and editing methods primarily process textual prompts as direct inputs without explicit reasoning about visual composition or operational steps. We present Generation Chain-of-Thought (GoT), a novel paradigm that empowers a Multimodal Large Language Model (MLLM) to first generate an explicit, structured reasoning chain in natural language—detailing semantic relationships, object attributes, and, crucially, precise spatial coordinates—before any image synthesis occurs. This intermediate reasoning output directly guides the subsequent visual generation or editing process. This approach transforms conventional text-to-image generation and editing into a reasoning-guided framework that analyzes semantic relationships and spatial arrangements. We define the formulation of GoT and construct large-scale GoT datasets containing over \textbf{9M} samples with detailed reasoning chains capturing semantic-spatial relationships. To leverage the advantages of GoT, we implement a unified framework that integrates Qwen2.5-VL for reasoning chain generation with an end-to-end diffusion model enhanced by our novel Semantic-Spatial Guidance Module. Experiments show our GoT framework achieves excellent performance on both generation and editing tasks, with significant improvements over baselines. Additionally, our approach enables interactive visual generation, allowing users to explicitly modify reasoning steps for precise image adjustments. GoT pioneers a new direction for reasoning-driven visual generation and editing, producing images that better align with human intent. We will release our datasets and models to facilitate future research.
Rongyao Fang, Chengqi Duan, Kun Wang 0056, Linjiang Huang, Hao Li 0069, Hao Tian 0006, Shilin Yan, Weihao Yu 0005, Xingyu Zeng, Jifeng Dai, Xihui Liu, Hongsheng Li 0001
NeurIPS9
2025 Robustness Verification of Deep Graph Neural Networks Tightened by Linear Approximation
abstract
Recent research indicates that adding residual connections in Graph Neural Networks (GNNs) would amplify susceptibility to anomalous nodes, consequently undermining the robustness of deep GNNs in practical settings. However, existing verification methods encounter challenges with the increasing number of parameters and computational overhead in deep GNNs. In this paper, we derive the general form of the residual connections and apply the dual backpropagation network to deep GNNs. Considering the heightened computational errors arising from the increased number of layers in deep GNNs, we propose a new method for calculating intermediate activation bounds of GNNs based on linear approximation. Experimental results show that new method can effectively enhance the verification accuracy. Notably, the maximum perturbation value of nodes correctly classified shows an average improvement of 119.5%. To showcase the the efficacy and scalability of our method, we verify robustness of deep GNNs on six different graph datasets, and our method can effectively verify the robustness of deep GNNs even with 32 layers of residual connections, i.e. verify over 87.29% of nodes in the Citeseer dataset. Furthermore, we analyse the influence of the graph structural properties on the robustness of the model.
Xingyu Zeng, Qi Qi 0001, Jingyu Wang 0001, Haodong Deng, Haifeng Sun 0001, Zirui Zhuang, Jianxin Liao
WSDM1
2024 Gradient-based Visual Explanation for Transformer-based CLIP
abstract
Significant progress has been achieved on the improvement and downstream usages of the Contrastive Language-Image Pre-training (CLIP) vision-language model, while less attention is paid to the interpretation of CLIP. We propose a Gradient-based visual Explanation method for CLIP (Grad-ECLIP), which interprets the matching result of CLIP for specific input image-text pair. By decomposing the architecture of the encoder and discovering the relationship between the matching similarity and intermediate spatial features, Grad-ECLIP produces effective heat maps that show the influence of image regions or words on the CLIP results. Different from the previous Transformer interpretation methods that focus on the utilization of self-attention maps, which are typically extremely sparse in CLIP, we produce high-quality visual explanations by applying channel and spatial weights on token features. Qualitative and quantitative evaluations verify the superiority of Grad-ECLIP compared with the state-of-the-art methods. A series of analysis are conducted based on our visual explanation results, from which we explore the working mechanism of image-text matching, and the strengths and limitations in attribution identification of CLIP. Codes are available here: https://github.com/Cyang-Zhao/Grad-Eclip.
Chenyang Zhao 0011, Xingyu Zeng, Antoni B. Chan
ICML3
2023 SeqCo-DETR: Sequence Consistency Training for Self-Supervised Object Detection with Transformers
Guoqiang Jin, Fan Yang 0089, Mingshan Sun, Ruyi Zhao, Yakun Liu, Wei Li 0314, Tianpeng Bao, Xingyu Zeng, Rui Zhao 0001
BMVC9
2022 Three-stage Training Pipeline with Patch Random Drop for Few-shot Object Detection
Shaobo Lin, Xingyu Zeng, Shilin Yan, Rui Zhao 0001
ACCV (6)2
2022 Scale-Aware Spatio-Temporal Relation Learning for Video Anomaly Detection
Guoqiu Li, Guanxiong Cai, Xingyu Zeng
ECCV (4)3
2020 Monocular 3D Object Detection with Decoupled Structured Polygon Estimation and Height-Guided Depth Estimation
abstract
Monocular 3D object detection task aims to predict the 3D bounding boxes of objects based on monocular RGB images. Since the location recovery in 3D space is quite difficult on account of absence of depth information, this paper proposes a novel unified framework which decomposes the detection problem into a structured polygon prediction task and a depth recovery task. Different from the widely studied 2D bounding boxes, the proposed novel structured polygon in the 2D image consists of several projected surfaces of the target object. Compared to the widely-used 3D bounding box proposals, it is shown to be a better representation for 3D detection. In order to inversely project the predicted 2D structured polygon to a cuboid in the 3D physical world, the following depth recovery task uses the object height prior to complete the inverse projection transformation with the given camera projection matrix. Moreover, a fine-grained 3D box refinement scheme is proposed to further rectify the 3D detection results. Experiments are conducted on the challenging KITTI benchmark, in which our method achieves state-of-the-art detection accuracy.
Yingjie Cai, Buyu Li, Zeyu Jiao, Hongsheng Li 0001, Xingyu Zeng, Xiaogang Wang 0001
AAAI5
2020 Rethinking Pseudo-LiDAR Representation
Xinzhu Ma, Shinan Liu, Zhiyi Xia, Hongwen Zhang 0001, Xingyu Zeng, Wanli Ouyang
ECCV (13)5
2020 Adapting Object Detectors with Conditional Domain Normalization
Kun Wang 0056, Xingyu Zeng, Shixiang Tang, Dapeng Chen, Di Qiu, Xiaogang Wang 0001
ECCV (11)3
2019 GS3D: An Efficient 3D Object Detection Framework for Autonomous Driving
abstract
We present an efficient 3D object detection framework based on a single RGB image in the scenario of autonomous driving. Our efforts are put on extracting the underlying 3D information in a 2D image and determining the accurate 3D bounding box of object without point cloud or stereo data. Leveraging the off-the-shelf 2D object detector, we propose an artful approach to efficiently obtain a coarse cuboid for each predicted 2D box. The coarse cuboid has enough accuracy to guide us to determine the 3D box of the object by refinement. In contrast to previous state-of-the-art methods that only use the features extracted from the 2D bounding box for box refinement, we explore the 3D structure information of the object by employing the visual features of visible surfaces. The new features from surfaces are utilized to eliminate the problem of representation ambiguity brought by only using 2D bounding box. Moreover, we investigate different methods of 3D box refinement and discover that a classification formulation with quality aware loss have much better performance than regression. Evaluated on KITTI benchmark, our approach outperforms current state-of-the-art methods for single RGB image based 3D object detection.
Buyu Li, Wanli Ouyang, Lu Sheng, Xingyu Zeng, Xiaogang Wang 0001
CVPR4
2018 Crafting GBD-Net for Object Detection
abstract
The visual cues from multiple support regions of different sizes and resolutions are complementary in classifying a candidate box in object detection. Effective integration of local and contextual visual cues from these regions has become a fundamental problem in object detection. In this paper, we propose a gated bi-directional CNN (GBD-Net) to pass messages among features from different support regions during both feature learning and feature extraction. Such message passing can be implemented through convolution between neighboring support regions in two directions and can be conducted in various layers. Therefore, local and contextual visual patterns can validate the existence of each other by learning their nonlinear relationships and their close interactions are modeled in a more complex way. It is also shown that message passing is not always helpful but dependent on individual samples. Gated functions are therefore needed to control message transmission, whose on-or-offs are controlled by extra visual evidence from the input sample. The effectiveness of GBD-Net is shown through experiments on three object detection datasets, ImageNet, Pascal VOC2007 and Microsoft COCO. Besides the GBD-Net, this paper also shows the details of our approach in winning the ImageNet object detection challenge of 2016, with source code provided on https://github.com/craftGBD/craftGBD. In this winning system, the modified GBD-Net, new pretraining scheme and better region proposal designs are provided. We also show the effectiveness of different network structures and existing techniques for object detection, such as multi-scale testing, left-right flip, bounding box voting, NMS, and context.
Xingyu Zeng, Wanli Ouyang, Hongsheng Li 0001, Tong Xiao 0003, Kun Wang 0056, Yu Liu 0015, Yucong Zhou, Bin Yang 0022, Zhe Wang 0006, Hui Zhou 0005, Xiaogang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2018 T-CNN: Tubelets With Convolutional Neural Networks for Object Detection From Videos
abstract
The state-of-the-art performance for object detection has been significantly improved over the past two years. Besides the introduction of powerful deep neural networks, such as GoogleNet and VGG, novel object detection frameworks, such as R-CNN and its successors, Fast R-CNN, and Faster R-CNN, play an essential role in improving the state of the art. Despite their effectiveness on still images, those frameworks are not specifically designed for object detection from videos. Temporal and contextual information of videos are not fully investigated and utilized. In this paper, we propose a deep learning framework that incorporates temporal and contextual information from tubelets obtained in videos, which dramatically improves the baseline performance of existing still-image detection frameworks when they are applied to videos. It is called T-CNN, i.e., tubelets with convolutional neueral networks. The proposed framework won newly introduced an object-detection-from-video task with provided data in the ImageNet Large-Scale Visual Recognition Challenge 2015. Code is publicly available athttps://github.com/myfavouritekk/T-CNN.
Kai Kang 0006, Hongsheng Li 0001, Xingyu Zeng, Bin Yang 0022, Tong Xiao 0003, Cong Zhang 0005, Zhe Wang 0006, Ruohui Wang, Xiaogang Wang 0001, Wanli Ouyang
IEEE Trans. Circuits Syst. Video Technol.4
2017 DeepID-Net: Object Detection with Deformable Part Based Convolutional Neural Networks
abstract
In this paper, we propose deformable deep convolutional neural networks for generic object detection. This new deep learning object detection framework has innovations in multiple aspects. In the proposed new deep architecture, a new deformation constrained pooling (def-pooling) layer models the deformation of object parts with geometric constraint and penalty. A new pre-training strategy is proposed to learn feature representations more suitable for the object detection task and with good generalization capability. By changing the net structures, training strategies, adding and removing some key components in the detection pipeline, a set of models with large diversity are obtained, which significantly improves the effectiveness of model averaging. The proposed approach improves the mean averaged precision obtained by RCNN [16], which was the state-of-the-art, from 31% to 50.3% on the ILSVRC2014 detection test set. It also outperforms the winner of ILSVRC2014, GoogLeNet, by 6.1%. Detailed component-wise analysis is also provided through extensive experimental evaluation, which provides a global view for people to understand the deep learning object detection pipeline.
Wanli Ouyang, Xingyu Zeng, Xiaogang Wang 0001, Ping Luo 0002, Yonglong Tian, Hongsheng Li 0001, Shuo Yang 0003, Zhe Wang 0006, Hongyang Li 0001, Kun Wang 0056, Chen Change Loy, Xiaoou Tang
IEEE Trans. Pattern Anal. Mach. Intell.2
2017 Visual Importance and Distortion Guided Deep Image Quality Assessment Framework
abstract
In this paper, we tackle the problem of no-reference image quality assessment (IQA). A learning-based IQA framework “VIDGIQA” is proposed, which extracts quality features from the input image and regresses the visual quality on these features. Since different distortions lead to different visual perceptions in the human visual system, distortion information is adopted to guide the feature learning process together with the human quality scores. Besides, a regression method is proposed to model and estimate the visual importance weights of all local regions, which can effectively improve the performance. More importantly, all these operations are integrated into one deep neural network, so that they can be jointly optimized and well cooperate with each other. Experiments were conducted to demonstrate the power of the proposed method on several datasets, including the LIVE dataset [1], the TID 2013 dataset [2], the LIVE multiply distorted IQA dataset [3], CSIQ [4] , and the LIVE in the wild image quality database [5]. The proposed method achieves 0.969 and 0.973 on the LIVE dataset [1] in terms of the spearman rank-order correlation coefficient and the Pearson linear correlation coefficient, respectively, which outperforms the state-of-the-art methods.
Jingwei Guan, Shuai Yi, Xingyu Zeng, Wai-kuen Cham, Xiaogang Wang 0001
IEEE Trans. Multim.3
2016 Gated Bi-directional CNN for Object Detection
Xingyu Zeng, Wanli Ouyang, Bin Yang 0022, Xiaogang Wang 0001
ECCV (7)1
2016 Learning Mutual Visibility Relationship for Pedestrian Detection with a Deep Model
Wanli Ouyang, Xingyu Zeng, Xiaogang Wang 0001
Int. J. Comput. Vis.2
2016 Partial Occlusion Handling in Pedestrian Detection With a Deep Model
abstract
Part-based models have demonstrated their merit in object detection. However, there is a key issue to be solved on how to integrate the inaccurate scores of part detectors when there are occlusions, abnormal deformations, appearances, or illuminations. To handle the imperfection of part detectors, this paper presents a probabilistic pedestrian detection framework. In this framework, a deformable part-based model is used to obtain the scores of part detectors and the visibilities of parts are modeled as hidden variables. Once the occluded parts are identified, their effects are properly removed from the final detection score. Unlike previous occlusion handling approaches that assumed independence among the visibility probabilities of parts or manually defined rules for the visibility relationship, a deep model is proposed in this paper for learning the visibility relationship among overlapping parts at multiple layers. The proposed approach can be viewed as a general postprocessing of part-detection results and can take detection scores of existing part-based models as input. The experimental results on three public datasets (Caltech, ETH, and Daimler) and a new CUHK occlusion dataset (http://www.ee.cuhk.edu.hk/~xgwang/CUHK_pedestrian.html), which is specially designed for the evaluation of occlusion handling approaches, show the effectiveness of the proposed approach.
Wanli Ouyang, Xingyu Zeng, Xiaogang Wang 0001
IEEE Trans. Circuits Syst. Video Technol.2
2015 DeepID-Net: Deformable deep convolutional neural networks for object detection
abstract
In this paper, we propose deformable deep convolutional neural networks for generic object detection. This new deep learning object detection framework has innovations in multiple aspects. In the proposed new deep architecture, a new deformation constrained pooling (def-pooling) layer models the deformation of object parts with geometric constraint and penalty. A new pre-training strategy is proposed to learn feature representations more suitable for the object detection task and with good generalization capability. By changing the net structures, training strategies, adding and removing some key components in the detection pipeline, a set of models with large diversity are obtained, which significantly improves the effectiveness of model averaging. The proposed approach improves the mean averaged precision obtained by RCNN [14], which was the state-of-the-art, from 31% to 50.3% on the ILSVRC2014 detection test set. It also outperforms the winner of ILSVRC2014, GoogLeNet, by 6.1%. Detailed component-wise analysis is also provided through extensive experimental evaluation, which provide a global view for people to understand the deep learning object detection pipeline.
Wanli Ouyang, Xiaogang Wang 0001, Xingyu Zeng, Ping Luo 0002, Yonglong Tian, Hongsheng Li 0001, Shuo Yang 0003, Zhe Wang 0006, Chen Change Loy, Xiaoou Tang
CVPR3
2015 Learning Deep Representation with Large-Scale Attributes
abstract
Learning strong feature representations from large scale supervision has achieved remarkable success in computer vision as the emergence of deep learning techniques. It is driven by big visual data with rich annotations. This paper contributes a large-scale object attribute database that contains rich attribute annotations (over 300 attributes) for ~180k samples and 494 object classes. Based on the ImageNet object detection dataset, it annotates the rotation, viewpoint, object part location, part occlusion, part existence, common attributes, and class-specific attributes. Then we use this dataset to train deep representations and extensively evaluate how these attributes are useful on the general object detection task. In order to make better use of the attribute annotations, a deep learning scheme is proposed by modeling the relationship of attributes and hierarchically clustering them into semantically meaningful mixture types. Experimental results show that the attributes are helpful in learning better features and improving the object detection accuracy by 2.6% in mAP on the ILSVRC 2014 object detection dataset and 2.4% in mAP on PASCAL VOC 2007 object detection dataset. Such improvement is well generalized across datasets.
Wanli Ouyang, Hongyang Li 0001, Xingyu Zeng, Xiaogang Wang 0001
ICCV3
2015 Single-Pedestrian Detection Aided by Two-Pedestrian Detection
abstract
In this paper, we address the challenging problem of detecting pedestrians who appear in groups. A new approach is proposed for single-pedestrian detection aided by two-pedestrian detection. A mixture model of two-pedestrian detectors is designed to capture the unique visual cues which are formed by nearby pedestrians but cannot be captured by single-pedestrian detectors. A probabilistic framework is proposed to model the relationship between the configurations estimated by single- and two-pedestrian detectors, and to refine the single-pedestrian detection result using two-pedestrian detection. The two-pedestrian detector can integrate with any single-pedestrian detector. Twenty-five state-of-the-art single-pedestrian detection approaches are combined with the two-pedestrian detector on three widely used public datasets: Caltech, TUD-Brussels, and ETH. Experimental results show that our framework improves all these approaches. The average improvement is 9 percent on the Caltech-Test dataset, 11 percent on the TUD-Brussels dataset and 17 percent on the ETH dataset in terms of average miss rate. The lowest average miss rate is reduced from 37 to percent on the Caltech-Test dataset, from 55 to 50 percent on the TUD-Brussels dataset and from 43 to 38 percent on the ETH dataset.
Wanli Ouyang, Xingyu Zeng, Xiaogang Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2014 Deep Learning of Scene-Specific Classifier for Pedestrian Detection
Xingyu Zeng, Wanli Ouyang, Xiaogang Wang 0001
ECCV (3)1
2013 Modeling Mutual Visibility Relationship in Pedestrian Detection
abstract
Detecting pedestrians in cluttered scenes is a challenging problem in computer vision. The difficulty is added when several pedestrians overlap in images and occlude each other. We observe, however, that the occlusion/visibility statuses of overlapping pedestrians provide useful mutual relationship for visibility estimation - the visibility estimation of one pedestrian facilitates the visibility estimation of another. In this paper, we propose a mutual visibility deep model that jointly estimates the visibility statuses of overlapping pedestrians. The visibility relationship among pedestrians is learned from the deep model for recognizing co-existing pedestrians. Experimental results show that the mutual visibility deep model effectively improves the pedestrian detection results. Compared with existing image-based pedestrian detection approaches, our approach has the lowest average miss rate on the Caltech-Train dataset, the Caltech-Test dataset and the ETH dataset. Including mutual visibility leads to 4% - 8% improvements on multiple benchmark datasets.
Wanli Ouyang, Xingyu Zeng, Xiaogang Wang 0001
CVPR2
2013 Multi-stage Contextual Deep Learning for Pedestrian Detection
abstract
Cascaded classifiers have been widely used in pedestrian detection and achieved great success. These classifiers are trained sequentially without joint optimization. In this paper, we propose a new deep model that can jointly train multi-stage classifiers through several stages of back propagation. It keeps the score map output by a classifier within a local region and uses it as contextual information to support the decision at the next stage. Through a specific design of the training strategy, this deep architecture is able to simulate the cascaded classifiers by mining hard samples to train the network stage-by-stage. Each classifier handles samples at a different difficulty level. Unsupervised pre-training and specifically designed stage-wise supervised training are used to regularize the optimization problem. Both theoretical analysis and experimental results show that the training strategy helps to avoid over fitting. Experimental results on three datasets (Caltech, ETH and TUD-Brussels) show that our approach outperforms the state-of-the-art approaches.
Xingyu Zeng, Wanli Ouyang, Xiaogang Wang 0001
ICCV1