VLDB 2026 Research / reviewers in the wild / expert
Xiao Sun 0001
dblp:151/8845
· DBLP profile ↗
31ranked-venue papers
3as first author
22since 2021 · last 2026
0000-0001-7459-804XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 3 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 3 first-author · 19 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | RacketVision: A Multiple Racket Sports Benchmark for Unified Ball and Racket AnalysisabstractWe introduce RacketVision, a novel dataset and benchmark for advancing computer vision in sports analytics, covering table tennis, tennis, and badminton. The dataset is the first to provide large-scale, fine-grained annotations for racket pose alongside traditional ball positions, enabling research into complex human-object interactions. It is designed to tackle three interconnected tasks: fine-grained ball tracking, articulated racket pose estimation, and predictive ball trajectory forecasting. Our evaluation of established baselines reveals a critical insight for multi-modal fusion: while naively concatenating racket pose features degrades performance, a Cross-Attention mechanism is essential to unlock their value, leading to trajectory prediction results that surpass strong unimodal baselines. RacketVision provides a versatile resource and a strong starting point for future research in dynamic object tracking, conditional motion forecasting, and multi-modal analysis in sports. Linfeng Dong, Yuchen Yang 0003, Wei Wang 0333, Yuenan Hou, Zhihang Zhong, Xiao Sun 0001 |
AAAI | 7 |
| 2026 | Velocity Disambiguation for Video Frame InterpolationabstractExisting video frame interpolation (VFI) methods blindly predict where each object is at a specific timestep $t$t ("time indexing"), which struggles to predict precise object movements. Given two images of a baseball, there are infinitely many possible trajectories: accelerating or decelerating, straight or curved. This often results in blurry frames as the method averages out these possibilities. Instead of forcing the network to learn this complicated time-to-location mapping implicitly together with predicting the frames, we provide the network with an explicit hint on how far the object has traveled between start and end frames, a novel approach termed "distance indexing". This method offers a clearer learning goal for models, reducing the uncertainty tied to object speeds. We further observed that, even with this extra guidance, objects can still be blurry especially when they are equally far from both input frames (i.e., halfway in-between), due to the directional ambiguity in long-range motion. To solve this, we propose an iterative reference-based estimation strategy that breaks down a long-range prediction into several short-range steps. When integrating our plug-and-play strategies into state-of-the-art learning-based models, they exhibit markedly sharper outputs and superior perceptual quality in arbitrary time interpolations, using a uniform distance indexing map in the same format as time indexing without requiring extra computation. Furthermore, we demonstrate that if additional latency is acceptable, a continuous map estimator can be employed to compute a pixel-wise dense distance indexing using multiple nearby frames. Combined with efficient multi-frame refinement, this extension can further disambiguate complex motion, thus enhancing performance both qualitatively and quantitatively. Additionally, the ability to manually specify distance indexing allows for independent temporal manipulation of each object, providing a novel tool for video editing tasks such as re-timing. Zhihang Zhong, Wei Wang 0333, Xiao Sun 0001, Yu Qiao 0001, Gurunandan Krishnan, Sizhuo Ma, Jian Wang 0100 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | OccMamba: Semantic Occupancy Prediction with State Space ModelsabstractTraining deep learning models for semantic occupancy prediction is challenging due to factors such as a large number of occupancy cells, severe occlusion, limited visual cues, complicated driving scenarios, etc. Recent methods often adopt transformer-based architectures given their strong capability in learning input-conditioned weights and long-range relationships. However, transformer-based networks are notorious for their quadratic computation complexity, seriously undermining their efficacy and deployment in semantic occupancy prediction. Inspired by the global modeling and linear computation complexity of the Mamba architecture, we present the first Mamba-based network for semantic occupancy prediction, termed OccMamba. Specifically, we first design the hierarchical Mamba module and local context processor to better aggregate global and local contextual information, respectively. Besides, to relieve the inherent domain gap between the linguistic and 3D domains, we present a simple yet effective 3D-to-1D reordering scheme, i.e., height-prioritized 2D Hilbert expansion. It can maximally retain the spatial structure of 3D voxels as well as facilitate the processing of Mamba blocks. Endowed with the aforementioned designs, our OccMamba is capable of directly and efficiently processing large volumes of dense scene grids, achieving state-of-the-art performance across three prevalent occupancy prediction benchmarks, including OpenOccupancy, SemanticKITTI, and SemanticPOSS. Notably, on OpenOc-cupancy, our OccMamba outperforms the previous state-of-the-art Co-Occ by 5.1% IoU and 4.3% mIoU, respectively. Our implementation is open-sourced and available at: https://github.com/USTCLH/OccMamba. Yuenan Hou, Xiaohan Xing, Yuexin Ma, Xiao Sun 0001, Yanyong Zhang |
CVPR | 5 |
| 2025 | MaskGaussian: Adaptive 3D Gaussian Representation from Probabilistic MasksabstractWhile 3D Gaussian Splatting (3DGS) has demonstrated remarkable performance in novel view synthesis and real-time rendering, the high memory consumption due to the use of millions of Gaussians limits its practicality. To mitigate this issue, improvements have been made by pruning unnecessary Gaussians, either through a hand-crafted criterion or by using learned masks. However, these methods deterministically remove Gaussians based on a snapshot of the pruning moment, leading to sub-optimized reconstruction performance from a long-term perspective. To address this issue, we introduce MaskGaussian, which models Gaussians as probabilistic entities rather than permanently removing them, and utilize them according to their probability of existence. To achieve this, we propose a masked-rasterization technique that enables unused yet probabilistically existing Gaussians to receive gradients, allowing for dynamic assessment of their contribution to the evolving scene and adjustment of their probability of existence. Hence, the importance of Gaussians iteratively changes and the pruned Gaussians are selected diversely. Extensive experiments demonstrate the superiority of the proposed method in achieving better rendering quality with fewer Gaussians than previous pruning methods, pruning over 60% of Gaussians on average with only a 0.02 PSNR decline. Our code can be found at: https://github.com/kaikai23/MaskGaussian Zhihang Zhong, Yifan Zhan, Xiao Sun 0001 |
CVPR | 5 |
| 2025 | Towards Explicit Exoskeleton for the Reconstruction of Complicated 3D Human Avatars
Yifan Zhan, Qingtian Zhu, Muyao Niu, Mingze Ma, Jiancheng Zhao, Zhihang Zhong, Xiao Sun 0001, Yu Qiao 0001, Yinqiang Zheng |
ICCV | 7 |
| 2025 | Towards Label-Free 3D Visual Grounding with Vision Foundation Models
Xiaopei Wu, Yuenan Hou, Binbin Lin 0001, Xinge Zhu, Yuexin Ma, Haifeng Liu 0001, Deng Cai 0001, Xiao Sun 0001 |
IROS | 8 |
| 2025 | Intrinsic Feature Rectification: Mitigating RAG Dependency by Addressing Information Loss in Image CaptioningabstractLarge Language Models (LLMs) have significantly advanced image captioning, yet their performance often degrades in limited-data settings due to overfitting in visual encoders, resulting in information loss during image encoding and poor generalization. While RAG is commonly used to boost performance in such cases, we argue that it primarily compensates for these visual encoding shortcomings rather than introducing truly novel knowledge. To address this issue at its root, we propose Intrinsic Feature Rectification (IFR), which enhances visual representations and reduces reliance on retrieval. IFR first aligns target image features with robust pre-trained CLIP representations and complementary auxiliary features from a vision-centric encoder. It then optionally fuses these features to construct a more holistic and generalizable visual representation, less prone to overfitting. Experiments show that IFR matches or exceeds the performance of RAG-based methods while eliminating retrieval overhead. Moreover, RAG offers little to no additional benefit when IFR is applied, supporting our hypothesis that RAG often serves as a fallback for weak intrinsic features. By improving visual understanding from the outset, IFR provides a more efficient solution for generalizable image captioning. Code will be available at https://github.com/haowuxc/IFR. Zhihang Zhong, Xiao Sun 0001 |
MMAsia | 3 |
| 2024 | Aleth-NeRF: Illumination Adaptive NeRF with Concealing Field AssumptionabstractThe standard Neural Radiance Fields (NeRF) paradigm employs a viewer-centered methodology, entangling the aspects of illumination and material reflectance into emission solely from 3D points. This simplified rendering approach presents challenges in accurately modeling images captured under adverse lighting conditions, such as low light or over-exposure. Motivated by the ancient Greek emission theory that posits visual perception as a result of rays emanating from the eyes, we slightly refine the conventional NeRF framework to train NeRF under challenging light conditions and generate normal-light condition novel views unsupervisedly. We introduce the concept of a ``Concealing Field," which assigns transmittance values to the surrounding air to account for illumination effects. In dark scenarios, we assume that object emissions maintain a standard lighting level but are attenuated as they traverse the air during the rendering process. Concealing Field thus compel NeRF to learn reasonable density and colour estimations for objects even in dimly lit situations. Similarly, the Concealing Field can mitigate over-exposed emissions during rendering stage. Furthermore, we present a comprehensive multi-view dataset captured under challenging illumination conditions for evaluation. Our code and proposed dataset are available at https://github.com/cuiziteng/Aleth-NeRF. Ziteng Cui, Lin Gu 0003, Xiao Sun 0001, Xianzheng Ma, Yu Qiao 0001, Tatsuya Harada |
AAAI | 3 |
| 2024 | DIBS: Enhancing Dense Video Captioning with Unlabeled Videos via Pseudo Boundary Enrichment and Online RefinementabstractWe present Dive Into the BoundarieS (DIBS), a novel pretraining framework for dense video captioning (DVC), that elaborates on improving the quality of the generated event captions and their associated pseudo event bound-aries from unlabeled videos. By leveraging the capabil-ities of diverse large language models (LLMs), we gen-erate rich DVC-oriented caption candidates and optimize the corresponding pseudo boundaries under several metic-ulously designed objectives, considering diversity, event-centricity, temporal ordering, and coherence. Moreover, we further introduce a novel online boundary refinement strat-egy that iteratively improves the quality of pseudo bound-aries during training. Comprehensive experiments have been conducted to examine the effectiveness of the pro-posed technique components. By leveraging a substantial amount of unlabeled video data, such as HowToI00M [16], we achieve a remarkable advancement on standard DVC datasets like YouCook2 [31] and ActivityNet [13]. We out-perform the previous state-of-the-art Vid2Seq [27] across a majority of metrics, achieving this with just 0.4% of the unlabeled video data used for pre-training by Vid2Seq. Huabin Liu 0001, Yu Qiao 0001, Xiao Sun 0001 |
CVPR | 4 |
| 2024 | Within the Dynamic Context: Inertia-Aware 3D Human Modeling with Pose Sequence
Yifan Zhan, Zhihang Zhong, Wei Wang 0333, Xiao Sun 0001, Yu Qiao 0001, Yinqiang Zheng |
ECCV (49) | 5 |
| 2024 | Mask as Supervision: Leveraging Unified Mask Information for Unsupervised 3D Pose Estimation
Yuchen Yang 0003, Yu Qiao 0001, Xiao Sun 0001 |
ECCV (44) | 3 |
| 2024 | Clearer Frames, Anytime: Resolving Velocity Ambiguity in Video Frame Interpolation
Zhihang Zhong, Gurunandan Krishnan, Xiao Sun 0001, Yu Qiao 0001, Sizhuo Ma, Jian Wang 0100 |
ECCV (33) | 3 |
| 2023 | Randomized Quantization: A Generic Augmentation for Data Agnostic Self-supervised LearningabstractSelf-supervised representation learning follows a paradigm of withholding some part of the data and tasking the network to predict it from the remaining part. Among many techniques, data augmentation lies at the core for creating the information gap. Towards this end, masking has emerged as a generic and powerful tool where content is withheld along the sequential dimension, e.g., spatial in images, temporal in audio, and syntactic in language. In this paper, we explore the orthogonal channel dimension for generic data augmentation by exploiting precision redundancy. The data for each channel is quantized through a non-uniform quantizer, with the quantized value sampled randomly within randomly sampled quantization bins. From another perspective, quantization is analogous to channel-wise masking, as it removes the information within each bin, but preserves the information across bins. Our approach significantly surpasses existing generic data augmentation methods, while showing on par performance against modality-specific augmentations. We comprehensively evaluate our approach on vision, audio, 3D point clouds, as well as the DABS benchmark which is comprised of various data modalities. The code is available at https://github.com/microsoft/random_quantize. Huimin Wu 0001, Chenyang Lei, Xiao Sun 0001, Peng-Shuai Wang, Qifeng Chen 0001, Kwang-Ting Cheng, Stephen Lin 0001, Zhirong Wu |
ICCV | 3 |
| 2023 | Long-Term Rhythmic Video SoundtrackerabstractWe consider the problem of generating musical soundtracks in sync with rhythmic visual cues. Most existing works rely on pre-defined music representations, leading to the incompetence of generative flexibility and complexity. Other methods directly generating video-conditioned waveforms suffer from limited scenarios, short lengths, and unstable generation quality. To this end, we present Long-Term Rhythmic Video Soundtracker (LORIS), a novel framework to synthesize long-term conditional waveforms. Specifically, our framework consists of a latent conditional diffusion probabilistic model to perform waveform synthesis. Furthermore, a series of context-aware conditioning encoders are proposed to take temporal information into consideration for a long-term generation. Notably, we extend our model’s applicability from dances to multiple sports scenarios such as floor exercise and figure skating. To perform comprehensive evaluations, we establish a benchmark for rhythmic video soundtracks including the pre-processed dataset, improved evaluation metrics, and robust generative baselines. Extensive experiments show that our model generates long-term soundtracks with state-of-the-art musical quality and rhythmic correspondence. Codes are available at https://github.com/OpenGVLab/LORIS. Jiashuo Yu, Yaohui Wang 0001, Xiao Sun 0001, Yu Qiao 0001 |
ICML | 4 |
| 2022 | A Simple Multi-Modality Transfer Learning Baseline for Sign Language TranslationabstractThis paper proposes a simple transfer learning baseline for sign language translation. Existing sign language datasets (e.g. PHOENIX-2014T, CSL-Daily) contain only about 10 K-20K pairs of sign videos, gloss annotations and texts, which are an order of magnitude smaller than typical parallel data for training spoken language translation models. Data is thus a bottleneck for training effective sign language translation models. To mitigate this problem, we propose to progressively pretrain the model from general- domain datasets that include a large amount of external supervision to within-domain datasets. Concretely, we pretrain the sign-to-gloss visual network on the general domain of human actions and the within-domain of a sign-to-gloss dataset, and pretrain the gloss-to-text translation network on the general domain of a multilingual corpus and the within-domain of a gloss-to-text corpus. The joint model is fine-tuned with an additional module named the visual-language mapper that connects the two networks. This simple baseline surpasses the previous state-of-the-art results on two sign language translation benchmarks, demonstrating the effectiveness of transfer learning. With its simplicity and strong performance, this approach can serve as a solid baseline for future research. Fangyun Wei, Xiao Sun 0001, Zhirong Wu, Stephen Lin 0001 |
CVPR | 3 |
| 2022 | Cross-Model Pseudo-Labeling for Semi-Supervised Action RecognitionabstractSemi-supervised action recognition is a challenging but important task due to the high cost of data annotation. A common approach to this problem is to assign unlabeled data with pseudo-labels, which are then used as additional supervision in training. Typically in recent work, the pseudo-labels are obtained by training a model on the labeled data, and then using confident predictions from the model to teach itself. In this work, we propose a more effective pseudo-labeling scheme, called Cross-Model Pseudo-Labeling (CMPL). Concretely, we introduce a lightweight auxiliary network in addition to the primary backbone, and ask them to predict pseudo-labels for each other. We observe that, due to their different structural biases, these two models tend to learn complementary representations from the same video clips. Each model can thus benefit from its counterpart by utilizing cross-model predictions as supervision. Experiments on different data partition protocols demonstrate the significant improvement of our framework over existing alternatives. For example, CMPL achieves 17.6% and 25.1% Top-1 accuracy on Kinetics-400 and UCF-101 using only the RGB modality and 1% labeled data, outperforming our baseline model, FixMatch [17], by 9.0% and 10.3%, respectively.11Project page is at https://justimyhxu.github.io/projects/cmpl/. Yinghao Xu 0001, Fangyun Wei, Xiao Sun 0001, Ceyuan Yang, Yujun Shen, Bo Dai 0002, Bolei Zhou, Stephen Lin 0001 |
CVPR | 3 |
| 2022 | Unsupervised Learning of Efficient Geometry-Aware Neural Articulated Representations
Atsuhiro Noguchi, Xiao Sun 0001, Stephen Lin 0001, Tatsuya Harada |
ECCV (17) | 2 |
| 2022 | Bringing Rolling Shutter Images Alive with Dual Reversed Distortion
Zhihang Zhong, Mingdeng Cao, Xiao Sun 0001, Zhirong Wu, Zhongyi Zhou, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (7) | 3 |
| 2022 | Animation from Blur: Multi-modal Blur Decomposition with Motion Guidance
Zhihang Zhong, Xiao Sun 0001, Zhirong Wu, Yinqiang Zheng, Stephen Lin 0001, Imari Sato |
ECCV (19) | 2 |
| 2021 | Neural Articulated Radiance FieldabstractWe present Neural Articulated Radiance Field (NARF), a novel deformable 3D representation for articulated objects learned from images. While recent advances in 3D implicit representation have made it possible to learn models of complex objects, learning pose-controllable representations of articulated objects remains a challenge, as current methods require 3D shape supervision and are unable to render appearance. In formulating an implicit representation of 3D articulated objects, our method considers only the rigid transformation of the most relevant object part in solving for the radiance field at each 3D location. In this way, the proposed method represents pose-dependent changes without significantly increasing the computational complexity. NARF is fully differentiable and can be trained from images with pose annotations. Moreover, through the use of an autoencoder, it can learn appearance variations over multiple instances of an object class. Experiments show that the proposed method is efficient and can generalize well to novel poses. The code is available for research purposes at https://github.com/nogu-atsu/NARF. Atsuhiro Noguchi, Xiao Sun 0001, Stephen Lin 0001, Tatsuya Harada |
ICCV | 2 |
| 2021 | Learning Skeletal Graph Neural Networks for Hard 3D Pose EstimationabstractVarious deep learning techniques have been proposed to solve the single-view 2D-to-3D pose estimation problem. While the average prediction accuracy has been improved significantly over the years, the performance on hard poses with depth ambiguity, self-occlusion, and complex or rare poses is still far from satisfactory. In this work, we target these hard poses and present a novel skeletal GNN learning solution. To be specific, we propose a hop-aware hierarchical channel-squeezing fusion layer to effectively extract relevant information from neighboring nodes while suppressing undesired noises in GNN learning. In addition, we propose a temporal-aware dynamic graph construction procedure that is robust and effective for 3D pose estimation. Experimental results on the Human3.6M dataset show that our solution achieves 10.3% average prediction accuracy improvement and greatly improves on hard poses over state-of-the-art techniques. We further apply the proposed technique on the skeleton-based action recognition task and also achieve state-of-the-art performance. Our code is available at https://github.com/ailingzengzzz/Skeletal-GNN. Ailing Zeng, Xiao Sun 0001, Nanxuan Zhao, Minhao Liu, Qiang Xu 0001 |
ICCV | 2 |
| 2021 | ACP++: Action Co-Occurrence Priors for Human-Object Interaction DetectionabstractA common problem in the task of human-object interaction (HOI) detection is that numerous HOI classes have only a small number of labeled examples, resulting in training sets with a long-tailed distribution. The lack of positive labels can lead to low classification accuracy for these classes. Towards addressing this issue, we observe that there exist natural correlations and anti-correlations among human-object interactions. In this paper, we model the correlations as action co-occurrence matrices and present techniques to learn these priors and leverage them for more effective training, especially on rare classes. The efficacy of our approach is demonstrated experimentally, where the performance of our approach consistently improves over the state-of-the-art methods on both of the two leading HOI detection benchmark datasets, HICO-Det and V-COCO. Dong-Jin Kim 0003, Xiao Sun 0001, Jinsoo Choi, Stephen Lin 0001, In-So Kweon |
IEEE Trans. Image Process. | 2 |
| 2020 | Detecting Human-Object Interactions with Action Co-occurrence Priors
Dong-Jin Kim 0003, Xiao Sun 0001, Jinsoo Choi, Stephen Lin 0001, In-So Kweon |
ECCV (21) | 2 |
| 2020 | Point-Set Anchors for Object Detection, Instance Segmentation and Pose Estimation
Fangyun Wei, Xiao Sun 0001, Hongyang Li 0001, Jingdong Wang 0001, Stephen Lin 0001 |
ECCV (10) | 2 |
| 2020 | SRNet: Improving Generalization in 3D Human Pose Estimation with a Split-and-Recombine Approach
Ailing Zeng, Xiao Sun 0001, Fuyang Huang, Minhao Liu, Qiang Xu 0001, Stephen Lin 0001 |
ECCV (14) | 2 |
| 2018 | Integral Human Pose Regression
Xiao Sun 0001, Fangyin Wei, Shuang Liang 0001 |
ECCV (6) | 1 |
| 2018 | Compositional Human Pose Regression
Shuang Liang 0001, Xiao Sun 0001 |
Comput. Vis. Image Underst. | 2 |
| 2017 | Compositional Human Pose RegressionabstractRegression based methods are not performing as well as detection based methods for human pose estimation. A central problem is that the structural information in the pose is not well exploited in the previous regression methods. In this work, we propose a structure-aware regression approach. It adopts a reparameterized pose representation using bones instead of joints. It exploits the joint connection structure to define a compositional loss function that encodes the long range interactions in the pose. It is simple, effective, and general for both 2D and 3D pose estimation in a unified setting. Comprehensive evaluation validates the effectiveness of our approach. It significantly advances the state-of-the-art on Human3.6M [20] and is competitive with state-of-the-art results on MPII [3]. Xiao Sun 0001, Jiaxiang Shang, Shuang Liang 0001 |
ICCV | 1 |
| 2017 | Towards 3D Human Pose Estimation in the Wild: A Weakly-Supervised ApproachabstractIn this paper, we study the task of 3D human pose estimation in the wild. This task is challenging due to lack of training data, as existing datasets are either in the wild images with 2D pose or in the lab images with 3D pose.,, We propose a weakly-supervised transfer learning method that uses mixed 2D and 3D labels in a unified deep neutral network that presents two-stage cascaded structure. Our network augments a state-of-the-art 2D pose estimation sub-network with a 3D depth regression sub-network. Unlike previous two stage approaches that train the two sub-networks sequentially and separately, our training is end-to-end and fully exploits the correlation between the 2D pose and depth estimation sub-tasks. The deep features are better learnt through shared representations. In doing so, the 3D pose labels in controlled lab environments are transferred to in the wild images. In addition, we introduce a 3D geometric constraint to regularize the 3D pose prediction, which is effective in the absence of ground truth depth labels. Our method achieves competitive results on both 2D and 3D benchmarks. Xingyi Zhou, Qixing Huang, Xiao Sun 0001, Xiangyang Xue 0001 |
ICCV | 3 |
| 2015 | Cascaded hand pose regressionabstractWe extends the previous 2D cascaded object pose regression work [9] in two aspects so that it works better for 3D articulated objects. Our first contribution is 3D pose-indexed features that generalize the previous 2D parameterized features and achieve better invariance to 3D transformations. Our second contribution is a principled hierarchical regression that is adapted to the articulated object structure. It is therefore more accurate and faster. Comprehensive experiments verify the state-of-the-art accuracy and efficiency of the proposed approach on the challenging 3D hand pose estimation problem, on a public dataset and our new dataset. Xiao Sun 0001, Shuang Liang 0001, Xiaoou Tang, Jian Sun 0001 |
CVPR | 1 |
| 2014 | Realtime and Robust Hand Tracking from DepthabstractWe present a realtime hand tracking system using a depth sensor. It tracks a fully articulated hand under large viewpoints in realtime (25 FPS on a desktop without using a GPU) and with high accuracy (error below 10 mm). To our knowledge, it is the first system that achieves such robustness, accuracy, and speed simultaneously, as verified on challenging real data. Our system is made of several novel techniques. We model a hand simply using a number of spheres and define a fast cost function. Those are critical for realtime performance. We propose a hybrid method that combines gradient based and stochastic optimization methods to achieve fast convergence and good accuracy. We present new finger detection and hand initialization methods that greatly enhance the robustness of tracking. Chen Qian 0006, Xiao Sun 0001, Xiaoou Tang, Jian Sun 0001 |
CVPR | 2 |