EDBT 2026 Demo / reviewers in the wild / expert
Tao Zhuo
dblp:153/2336
· DBLP profile ↗
31ranked-venue papers
6as first author
16since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 20 · 4 first-author · 10 since 2021Artificial intelligence and machine learning · 10 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SOMA: Feature Gradient Enhanced Affine-Flow Matching for SAR-Optical RegistrationabstractAchieving pixel-level registration between SAR and optical images remains a challenging task due to their fundamentally different imaging mechanisms and visual characteristics. Although deep learning has achieved great success in many cross-modal tasks, its performance on SAR-Optical registration tasks is still unsatisfactory. Gradient-based information has traditionally played a crucial role in handcrafted descriptors by highlighting structural differences. However, such gradient cues have not been effectively leveraged in deep learning frameworks for SAR-Optical image matching. To address this gap, we propose SOMA, a dense registration framework that integrates structural gradient priors into deep features and refines alignment through a hybrid matching strategy. Specifically, we introduce the Feature Gradient Enhancer (FGE), which embeds multi-scale, multi-directional gradient filters into the feature space using attention and reconstruction mechanisms to boost feature distinctiveness. Furthermore, we propose the Global-Local Affine-Flow Matcher (GLAM), which combines affine transformation and flow-based refinement within a coarse-to-fine architecture to ensure both structural consistency and local accuracy. Experimental results demonstrate that SOMA significantly improves registration precision, increasing the CMR@1px by 12.29% on the SEN1-2 dataset and 18.50% on the GFGE_SO dataset. In addition, SOMA exhibits strong robustness and generalizes well across diverse scenes and resolutions. Tao Zhuo, Xiuwei Zhang 0001, Hanlin Yin, Wencong Wu, Yanning Zhang 0001 |
AAAI | 2 |
| 2025 | Visual Object Tracking With Multi-Frame Distractor SuppressionabstractWith the rapid development of CNN or Transformer, the present mainstream approaches regard an image patch as the reference of the target to perform tracking, which is known as template matching-based trackers. However, most existing template matching-based trackers only consider the per-frame localization accuracy, neglecting the potential distractor (similar object) dependencies among multiple video frames, which poses a fundamental challenge in template matching-based tracking. In this work, we propose a novel comprehensive framework with multi-frame distractor suppression for visual object tracking (MFDSTrack), which explicitly models the temporal history of both the target object and potential distractors. Specifically, we utilize a universal target candidate generation module to detect target candidates (both target and distractors), providing a holistic view of the scene. In addition, a temporal and distractor-aware association module is designed to suppress multi-frame distractors by adopting a simple encoder-decoder Transformer architecture. The encoder accepts inputs of target candidates’ history, while the decoder takes current target candidate queries and the output of the encoder as inputs to associate current target candidate queries with historical trajectories. We extensively evaluate our trackers, MFDSTrack-SD, MFDSTrack-OS, MFDSTrack-GRM, and MFDSTrack-LT on the LaSOT,${\mathrm {LaSOT}}_{ext}$, TrackingNet, GOT-10k, UAV123, NFS, and OTB100 benchmark. Extensive experiments show that our methods outperform previous state-of-the-art trackers on seven tracking benchmarks. Mingyu Cai, Zhixuan Bai, Tao Zhuo, Hongming Zhang 0002, Yanning Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | AdaSemiCD: An Adaptive Semi-Supervised Change Detection Method Based on Pseudo-Label EvaluationabstractChange detection (CD) is an essential field in remote sensing, with a primary focus on identifying areas of change in bitemporal image pairs captured at varying intervals of the same region. The data annotation process for CD tasks is both time-consuming and labor-intensive. To better utilize the scarce labeled data and abundant unlabeled data, we introduce an adaptive semi-supervised learning (SSL) method, AdaSemiCD, to improve pseudo-label usage and optimize the training process. Initially, due to the extreme class imbalance inherent in CD, the model is more inclined to focus on the background class, and it is easy to confuse the boundary of the target object. Considering these two points, we develop a measurable evaluation metric for pseudo-labels that enhances the representation of information entropy by class rebalancing and amplification of ambiguous areas, assigning greater weights to prospective change objects. Subsequently, to enhance the reliability of sample wise pseudo-labels, we introduce the AdaFusion module, to dynamically identify the most uncertain region and substitute it with more trustworthy content. Lastly, to ensure better training stability, we introduce the AdaEMA module, which updates the teacher model using only batches of trusted samples. Experimental results on ten public CD datasets validate the efficacy and generalizability of our proposed adaptive training framework. Lingyan Ran, Wen Dongcheng, Tao Zhuo, Shizhou Zhang, Xiuwei Zhang 0001, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Hand-Centric Motion Refinement for 3D Hand-Object Interaction via Hierarchical Spatial-Temporal ModelingabstractHands are the main medium when people interact with the world. Generating proper 3D motion for hand-object interaction is vital for applications such as virtual reality and robotics. Although grasp tracking or object manipulation synthesis can produce coarse hand motion, this kind of motion is inevitably noisy and full of jitter. To address this problem, we propose a data-driven method for coarse motion refinement. First, we design a hand-centric representation to describe the dynamic spatial-temporal relation between hands and objects. Compared to the object-centric representation, our hand-centric representation is straightforward and does not require an ambiguous projection process that converts object-based prediction into hand motion. Second, to capture the dynamic clues of hand-object interaction, we propose a new architecture that models the spatial and temporal structure in a hierarchical manner. Extensive experiments demonstrate that our method outperforms previous methods by a noticeable margin. Yuze Hao, Jianrong Zhang, Tao Zhuo, Fuan Wen, Hehe Fan |
AAAI | 3 |
| 2024 | Edge-Guided Detector-Free Network for Robust and Accurate Visible-Thermal Image MatchingabstractRecent detector-free models strive to leverage both local and global context for image matching, showcasing enhanced robustness, particularly in scenarios with weak-textured scenes. Despite these advancements, automatically establishing feature correspondences between visible and thermal images still introduces additional challenges. Differences in radiation and geometry between these modalities often result in degraded performance for the majority of existing methods. To this end, we propose edge-guided detector-free model termed EDMatcher for visible-thermal image matching. Besides local and global context in the images, EDMatcher also leverages modality-robust structural information in image edges, which demonstrates promising robustness to images with distinct modalities. Moreover, an edge-masked ground-truth matrix generation strategy is introduced during the training, which helps EDMatcher to further focus on more salient regions while leaving out texture-less regions, leading to more efficient learning. Extensive experiments show that EDMatcher has strong generalization and achieves excellent matching performances. Zhaoshuai Qi, Xiuwei Zhang 0001, Tao Zhuo, Yanning Zhang 0001 |
ICME | 4 |
| 2024 | RGB-T Object Detection via Group Shuffled Multi-receptive Attention and Multi-modal Supervision
Jinzhong Wang, Xuetao Tian, Shun Dai, Tao Zhuo, Haorui Zeng, Hongjuan Liu, Xiuwei Zhang 0001, Yanning Zhang 0001 |
ICPR (17) | 4 |
| 2024 | A Semantic Perception and CNN-Transformer Hybrid Network for Occluded Person Re-IdentificationabstractThe objective of the occluded person re-identification (ReID) task is to capture the same person from different camera angles when the pedestrian’s body is partially occluded. In this task, there are two main challenges: 1) pedestrians are often occluded by other persons or objects, and 2) pedestrians change poses. Moreover, these two issues often simultaneously occur. Although many occluded person ReID algorithms have been proposed, many existing methods can often only solve one of these issues well, and the other issue is often ignored. In this work, a novel semantic perception and CNN-transformer hybrid network (abbreviated as SPH) is proposed for occluded person ReID, which consists of a CNN-based human semantic perception stream and a transformer-based pose perception stream. In the former, a human semantic auxiliary module and a human semantic perception module are designed to obtain human semantic information where multi-granularity region features of the human body are extracted to solve the issues of occlusion. In the latter, we propose a token-based pose integration module to obtain the corresponding patch for each pose key-point and the relative position information to solve the change in pedestrian pose. Moreover, these two streams are jointly optimized in a unified framework. In addition, to further solve the issue of occlusion, the human completion strategy is proposed for the query sample where the gallery samples are used to complete the missing parts of the query. Extensive experimental results on three public occluded person ReID datasets, Occluded-DukeMTMC, P-DukeMTMC-reID, and Occluded-REID, demonstrate that the proposed method can outperform all SOTA occluded person ReID methods in terms of the mAP and Rank-1. Compared with PAT (CVPR21) on the Occluded-DukeMTMC and Occluded-REID datasets, the improvements in mAP/Rank-1 reached 10.1%/7.4%, and 10%/1%, respectively. Moreover, when TransReID (ICCV21) was used, SPH achieved improvements of 4.5% (mAP) and 5.5% (Rank-1) on the Occluded-DukeMTMC dataset. Zan Gao 0001, Peng Chen 0047, Tao Zhuo, Meng Liu 0006, Lei Zhu 0002, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | DDF: A Novel Dual-Domain Image Fusion Strategy for Remote Sensing Image Semantic Segmentation With Unsupervised Domain AdaptationabstractThe semantic segmentation of remote sensing (RS) images is a challenging and hot issue due to the large amount of unlabeled data and domain variation. Unsupervised domain adaptation (UDA) has proven to be advantageous in leveraging unlabeled information from the target domain. However, traditional approaches of independently fine-tuning UDA models in the source and target domains have a limited effect on the result. In this article, we propose a hybrid training strategy that boosts self-training methods with domain fusion images. First, we introduce a novel dual-domain image fusion (DDF) strategy to effectively utilize the original image, the style-transferred image, and the intermediate-domain information. Second, to further refine the precision of pseudolabels, we present a region-specific reweighting strategy that assigns different weights to pseudolabel regions based on their spatial context. Finally, we conduct a series of extensive benchmark experiments and ablation studies on the ISPRS Vaihingen and Potsdam datasets. These results show the efficiency of our approach and establish a practical basis for implementing semantic segmentation in remote sensors. Lingyan Ran, Lushuang Wang, Tao Zhuo, Yinghui Xing, Yanning Zhang 0001 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2023 | A Novel Temporal Channel Enhancement and Contextual Excavation Network for Temporal Action LocalizationabstractThe temporal action localization (TAL) task aims to locate and classify action instances in untrimmed videos. Most previous methods use classifiers and locators to act on the same feature; thus, the classification and localization processes are relatively independent. Therefore, if the classification results and localization results are fused, there will be a problem that the classification results are correct while the localization results are wrong, resulting in inaccurate final results, and vice versa. To solve this problem, we propose a novel temporal channel enhancement and contextual excavation network (TCN) for the TAL task, which generates robust classification and localization features and refines the final localization results. Specifically, a temporal channel enhancement module is designed to enhance the temporal and channel information of the feature sequence. Then, the temporal semantic contextual excavation module is developed to establish relationships between similar frames. Finally, the features with enhanced contextual information are transferred to a classifier. While executing the classification process, we obtain powerful classification features. Most importantly, with the robust classification features, the final localization features are produced by the refine localization module, which is applied to obtain the final localization results. Extensive experiments show that TCN can outperform all the SOTA methods on the THUMOS14 dataset, and achieves a comparable performance on the ActivityNet1.3 dataset. Compared with ActionFormer (ECCV 2022) and BREM (MM 2022) on the THUMOS14 dataset, the proposed TCN can achieve improvements of 1.8% and 5.0%, respectively. Zan Gao 0001, Xinglei Cui, Yibo Zhao 0001, Tao Zhuo, Weili Guan, Meng Wang 0001 |
ACM Multimedia | 4 |
| 2023 | Automatic Network Architecture Search for RGB-D Semantic SegmentationabstractRecent RGB-D semantic segmentation networks are usually manually designed. However, due to limited human efforts and time costs, their performance might be inferior for complex scenarios. To address this issue, we propose the first Neural Architecture Search (NAS) method that designs the network automatically. Specifically, the target network consists of an encoder and a decoder. The encoder is designed with two independent branches, where each branch specializes in extracting features from RGB and depth images, respectively. The decoder fuses the features and generates the final segmentation result. Besides, for automatic network design, we design a grid-like network-level search space combined with a hierarchical cell-level search space. By further developing an effective gradient-based search strategy, the network structure with hierarchical cell architectures is discovered. Extensive results on two datasets show that the proposed method outperforms the state-of-the-art approaches, which achieves a mIoU score of 55.1% on the NYU-Depth v2 dataset and 50.3% on the SUN-RGBD dataset. Wenna Wang, Tao Zhuo, Xiuwei Zhang 0001, Mingjun Sun, Hanlin Yin, Yinghui Xing, Yanning Zhang 0001 |
ACM Multimedia | 2 |
| 2023 | A Multitemporal Scale and Spatial-Temporal Transformer Network for Temporal Action LocalizationabstractTemporal action localization plays an important role in video analysis, which aims to localize and classify actions in untrimmed videos. Previous methods often predict actions on a feature space of a single temporal scale. However, the temporal features of a low-level scale lack sufficient semantics for action classification, while a high-level scale cannot provide the rich details of the action boundaries. In addition, the long-range dependencies of video frames are often ignored. To address these issues, a novel multitemporal-scale spatial–temporal transformer (MSST) network is proposed for temporal action localization, which predicts actions on a feature space of multiple temporal scales. Specifically, we first use refined feature pyramids of different scales to pass semantics from high-level scales to low-level scales. Second, to establish the long temporal scale of the entire video, we use a spatial–temporal transformer encoder to capture the long-range dependencies of video frames. Then, the refined features with long-range dependencies are fed into a classifier for coarse action prediction. Finally, to further improve the prediction accuracy, we propose a frame-level self-attention module to refine the classification and boundaries of each action instance. Most importantly, these three modules are jointly explored in a unified framework, and MSST has an anchor-free and end-to-end architecture. Extensive experiments show that the proposed method can outperform state-of-the-art approaches on the THUMOS14 dataset and achieve comparable performance on the ActivityNet1.3 dataset. Compared with A2Net (TIP20, Avg{0.3:0.7}), Sub-Action (CSVT2022, Avg{0.1:0.5}), and AFSD (CVPR21, Avg{0.3:0.7}) on the THUMOS14 dataset, the proposed method can achieve improvements of 12.6%, 17.4%, and 2.2%, respectively. Zan Gao 0001, Xinglei Cui, Tao Zhuo, Zhiyong Cheng 0001, Anan Liu, Meng Wang 0001, Shengyong Chen |
IEEE Trans. Hum. Mach. Syst. | 3 |
| 2022 | One-shot Video Graph Generation for Explainable Action Reasoning
Tao Zhuo, Peng Zhang 0005, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001, Mohan Kankanhalli |
Neurocomputing | 2 |
| 2022 | Entropy guided attention network for weakly-supervised action localization
Ying Sun 0001, Hehe Fan, Tao Zhuo, Joo-Hwee Lim, Mohan Kankanhalli |
Pattern Recognit. | 4 |
| 2022 | Understanding Atomic Hand-Object Interaction With Human IntentionabstractHand-object interaction plays a very important role when humans manipulate objects. While existing methods focus on improving hand-object recognition with fully automatic methods, human intention has been largely neglected in the recognition process, thus leading to undesirable interaction descriptions. To better interpret human-object interaction that is aligned to human intention, we argue that a reference specifying human intention should be taken into account. Thus, we propose a new approach to represent interactions while reflecting human purpose with three key factors,i.e., hand, object and reference. Specifically, we design a pattern ofhand-object, object-reference, hand, object, reference> (HOR) to recognize intention based atomic hand-object interactions. This pattern aims to model interactions with the states of hand, object, reference and their relationships. Furthermore, we design a simple yet effective Spatially Part-based (3+1)D convolutional neural network, namely SP(3+1)D, which leverages 3D and 1D convolutions to model visual dynamics and object position changes based on our HOR, respectively. With the help of our SP(3+1)D network, the recognition results are able to indicate human purposes accurately. To evaluate the proposed method, we annotate a Something-1.3k dataset, which contains 10 atomic hand-object interactions and about 130 videos for each interaction. Experimental results on Something-1.3k demonstrate the effectiveness of our SP(3+1)D network. Hehe Fan, Tao Zhuo, Xin Yu 0002, Yi Yang 0001, Mohan Kankanhalli |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2021 | Effective Abstract Reasoning with Dual-Contrast Network
Tao Zhuo, Mohan Kankanhalli |
ICLR | 1 |
| 2021 | Unsupervised Abstract Reasoning for Raven's Problem MatricesabstractRaven's Progressive Matrices (RPM) is highly correlated with human intelligence, and it has been widely used to measure the abstract reasoning ability of humans. In this paper, to study the abstract reasoning capability of deep neural networks, we propose the first unsupervised learning method for solving RPM problems. Since the ground truth labels are not allowed, we design a pseudo target based on the prior constraints of the RPM formulation to approximate the ground-truth label, which effectively converts the unsupervised learning strategy into a supervised one. However, the correct answer is wrongly labelled by the pseudo target, and thus the noisy contrast will lead to inaccurate model training. To alleviate this issue, we propose to improve the model performance with negative answers. Moreover, we develop a decentralization method to adapt the feature representation to different RPM problems. Extensive experiments on three datasets demonstrate that our method even outperforms some of the supervised approaches. Our code is available at https://github.com/visiontao/ncd. Tao Zhuo, Mohan Kankanhalli |
IEEE Trans. Image Process. | 1 |
| 2020 | Attention feature matching for weakly-supervised video relocalizationabstractLocalizing the desired video clip for a given query in an untrimmed video has been a hot research topic for multimedia understanding. Recently, a new task named video relocalization, in which the query is a video clip, has been raised. Some methods have been developed for this task, however, these methods often require dense annotations of the temporal boundaries inside long videos for training. A more practical solution is the weakly-supervised approach, which only needs the matching information between the query and video. Haoyu Tang 0002, Jihua Zhu, Zan Gao 0001, Tao Zhuo, Zhiyong Cheng 0001 |
MMAsia | 4 |
| 2020 | Unsupervised Online Video Object Segmentation With Motion Property UnderstandingabstractUnsupervised video object segmentation aims to automatically segment moving objects over an unconstrained video without any user annotation. So far, only few unsupervised online methods have been reported in the literature, and their performance is still far from satisfactory because the complementary information from future frames cannot be processed under online setting. To solve this challenging problem, in this paper, we propose a novel unsupervised online video object segmentation (UOVOS) framework by construing the motion property to mean moving in concurrence with a generic object for segmented regions. By incorporating the salient motion detection and the object proposal, a pixel-wise fusion strategy is developed to effectively remove detection noises, such as dynamic background and stationary objects. Furthermore, by leveraging the obtained segmentation from immediately preceding frames, a forward propagation algorithm is employed to deal with unreliable motion detection and object proposals. Experimental results on several benchmark datasets demonstrate the efficacy of the proposed method. Compared to state-of-the-art unsupervised online segmentation algorithms, the proposed method achieves an absolute gain of 6.2%. Moreover, our method achieves better performance than the best unsupervised offline algorithm on the DAVIS-2016 benchmark dataset. Our code is available on the project website: https://www.github.com/visiontao/uovos. Tao Zhuo, Zhiyong Cheng 0001, Peng Zhang 0005, Yongkang Wong, Mohan Kankanhalli |
IEEE Trans. Image Process. | 1 |
| 2020 | Ensemble Tracking Based on Diverse Collaborative Framework With Multi-Cue Dynamic FusionabstractTracking with deep neural networks has been verified to arrive at a new level accuracy in many challenging scenarios, but the tracking robustness has been still challenged by model singularity and self-learning loop mechanism. As a promising solution for the limitations, to ensemble diverse tracking strategies into a highly-interactive framework has shown a potential effectiveness in recent studies. In this work, a collaborative tracking framework is proposed by exploiting both discriminative correlation filters and deep classifiers into an ensembling framework. With a multi-cue dynamic fusion scheme performed on all the ensembled members’ outputs, a robust long-term tracking can be achieved by calculating the optimal robustness scores based on a dynamic weighted sum of multi-cue metrics. Meanwhile, the obtained reliable and diverse training samples are also utilized to adaptively update the tracker in each branch with heuristic frequency, which is able to alleviate the training samples’ contamination and model corruption. Experiments on the OTB-2015, Temple color 128, UAV123, VOT2016, and VOT2018 benchmark datasets have shown superior performance in comparison to other state-of-the-art tracking approaches. Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yufei Zha, Yanning Zhang 0001 |
IEEE Trans. Multim. | 3 |
| 2019 | Explainable Video Action Reasoning via Prior Knowledge and State TransitionsabstractHuman action analysis and understanding in videos is an important and challenging task. Although substantial progress has been made in past years, the explainability of existing methods is still limited. In this work, we propose a novel action reasoning framework that uses prior knowledge to explain semantic-level observations of video state changes. Our method takes advantage of both classical reasoning and modern deep learning approaches. Specifically, prior knowledge is defined as the information of a target video domain, including a set of objects, attributes and relationships in the target video domain, as well as relevant actions defined by the temporal attribute and relationship changes (i.e. state transitions). Given a video sequence, we first generate a scene graph on each frame to represent concerned objects, attributes and relationships. Then those scene graphs are associated by tracking objects across frames to form a spatio-temporal graph (also called video graph), which represents semantic-level video states. Finally, by sequentially examining each state transition in the video graph, our method can detect and explain how those actions are executed with prior knowledge, just like the logical manner of thinking by humans. Compared to previous works, the action reasoning results of our method can be explained by both logical rules and semantic-level observations of video content changes. Besides, the proposed method can be used to detect multiple concurrent actions with detailed information, such as who (particular objects), when (time), where (object locations) and how (what kind of changes). Experiments on a re-annotated dataset CAD-120 show the effectiveness of our method. Tao Zhuo, Zhiyong Cheng 0001, Peng Zhang 0005, Yongkang Wong, Mohan Kankanhalli |
ACM Multimedia | 1 |
| 2019 | Multi-model cooperative task assignment and path planning of multiple UCAV formation
Hanqiao Huang, Tao Zhuo |
Multim. Tools Appl. | 2 |
| 2018 | Robust tracking based on H-CNN with low-resource sampling and scaling by frame-wise motion localization
Peng Zhang 0005, Tao Zhuo, Hanqiao Huang, Kangli Chen, Mohan Kankanhalli |
Multim. Tools Appl. | 2 |
| 2018 | Going deeper with two-stream ConvNets for action recognition in video surveillance
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Yanning Zhang 0001 |
Pattern Recognit. Lett. | 3 |
| 2018 | Saliency flow based video segmentation via motion guided contour refinement
Peng Zhang 0005, Tao Zhuo, Hanqiao Huang, Mohan Kankanhalli |
Signal Process. | 2 |
| 2017 | Online object tracking based on CNN with spatial-temporal saliency guided sampling
Peng Zhang 0005, Tao Zhuo, Wei Huang 0013, Kangli Chen, Mohan Kankanhalli |
Neurocomputing | 2 |
| 2016 | Deformable object tracking with spatiotemporal segmentation in big vision surveillance
Peng Zhang 0005, Tao Zhuo, Lei Xie 0001, Yanning Zhang 0001 |
Neurocomputing | 2 |
| 2016 | Online tracking based on efficient transductive learning with sample matching costs
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Dapeng Tao, Jun Cheng 0002 |
Neurocomputing | 2 |
| 2016 | Bayesian tracking fusion framework with online classifier ensemble for immersive visual applications
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Hanqiao Huang, Kangli Chen |
Multim. Tools Appl. | 2 |
| 2016 | Real-time tracking-by-learning with high-order regularization fusion for big video abstraction
Peng Zhang 0005, Tao Zhuo, Yanning Zhang 0001, Lei Xie 0001, Dapeng Tao |
Signal Process. | 2 |
| 2015 | Superframe segmentation based on content-motion correspondence for social video summarizationabstractThe goal of video summarization is to turn large volume of video data into a compact visual summary that can be easily interpreted by users in a while. Existing summarization strategies employed the point based feature correspondence for the superframe segmentation. Unfortunately, the information carried by those sparse points is far from sufficiency and stability to describe the change of interesting regions of each frame. Therefore, in order to overcome the limitations of point feature, we propose a region correspondence based superframe segmentation to achieve more effective video summarization. Instead of utilizing the motion of feature points, we calculate the similarity of content-motion to obtain the strength of change between the consecutive frames. With the help of circulant structure kernel, the proposed method is able to perform more accurate motion estimation efficiently. Experimental testing on the videos from benchmark database has demonstrate the effectiveness of the proposed method. Tao Zhuo, Peng Zhang 0005, Kangli Chen, Yanning Zhang 0001 |
ACII | 1 |
| 2014 | Object Tracking using Reformative Transductive Learning with Sample Variational CorrespondenceabstractTracking-by-learning strategies have effectively solved many challenging problems for visual tracking. When labeled samples are limited, the learning performance can be improved by exploiting unlabeled ones. Thus, a key issue for semi-supervised learning is the label assignment of the unlabeled samples, which is the principal focus of transductive learning. Unfortunately, the optimization scheme employed by the transductive learning is hard to be applied to online tracking because of its large amount of computation for sample labeling. In this paper, a reformative transductive learning was proposed with the variational correspondence between the learning samples, which are utilized to build an effective matching cost function for more efficient label assignment during the learning of representative separators. By using a weighted accumulative average to update the coefficients via a fixed budget of support vectors, the proposed tracking has been demonstrated to outperform most of the state-of-art trackers. Tao Zhuo, Peng Zhang 0005, Yanning Zhang 0001, Wei Huang 0013, Hichem Sahli |
ACM Multimedia | 1 |