VLDB 2026 Research / reviewers in the wild / expert
Yang Li 0041
dblp:37/4190-41
· DBLP profile ↗
30ranked-venue papers
3as first author
23since 2021 · last 2026
0000-0001-9427-7665ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 21 · 3 first-author · 15 since 2021Artificial intelligence and machine learning · 14 · 3 first-author · 10 since 2021Systems, architecture and hardware · 2 · 2 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ChatTracker: Enhancing Visual Tracking via LLM-Driven Iterative Description RefinementabstractVisual object tracking focuses on locating a target object within a video sequence based on an initial bounding box. Recently, Vision-Language (VL) trackers have been proposed to utilize additional natural language descriptions to enhance versatility in various applications. Despite this potential, VL trackers still underperform the State-of-the-Art (SoTA) visual trackers in terms of tracking accuracy. We find that this inferiority is primarily due to their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we identify, for the first time, that over 10% of textual annotations in existing VL tracking datasets suffer from inaccuracies through manual evaluation. To address this problem, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel Reflection-based Language Description Refinement Module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed, which can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that ChatTracker achieves comparable performance to existing SoTA tracking methods. In addition, language descriptions generated by ChatTracker enhance the performance of various VL trackers and exhibit better text-to-image alignment than annotations in the original dataset. Moreover, our proposed framework can improve the performance of various visual tasks, including Referring Expression Comprehension (REC), Referring Expression Segmentation (RES), and Referring Video Object Segmentation (R-VOS) tasks by providing more accurate language descriptions, which demonstrates the universality of ChatTracker. We release the manual evaluation results and the generated textual descriptions, aiming to drive advancements in VL tracking. Yiming Sun 0006, Mi Zhang 0001, Shaoxiang Chen 0001, Yang Li 0041, Changbo Wang, Jianke Zhu, Steven C. H. Hoi |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2025 | Motion-Zero: A Zero-Shot Trajectory Control Framework of Moving Object for Diffusion-Based Video GenerationabstractRecent large-scale pre-trained diffusion models have demonstrated a powerful generative ability to produce high-quality videos from detailed text descriptions. However, exerting control over the motion of objects in videos generated by any video diffusion model remains a challenging problem. In this paper, we propose a novel zero-shot moving object trajectory control framework, Motion-Zero, to enable arbitrary single-object-trajectory control for the text-to-video diffusion model. To this end, an initial noise prior module is designed to provide a position-based prior to improve the stability of the appearance of the moving object and the accuracy of position. In addition, based on the attention map of the U-Net, spatial constraints are directly applied to the denoising process of diffusion models, which further ensures the positional consistency of moving objects during the inference. Furthermore, temporal consistency is guaranteed with a proposed shift temporal attention mechanism. Our method can be flexibly applied to various state-of-the-art video diffusion models without any training process. Extensive experiments demonstrate our proposed method can control the motion trajectories of arbitrary objects while preserving the original ability to generate high-quality videos. Changgu Chen, Junwei Shu, Gaoqi He, Changbo Wang, Yang Li 0041 |
AAAI | 5 |
| 2025 | Open-World Reinforcement Learning over Long Short-Term ImaginationabstractTraining visual reinforcement learning agents in a high-dimensional open world presents significant challenges. While various model-based methods have improved sample efficiency by learning interactive world models, these agents tend to be “short-sighted”, as they are typically trained on short snippets of imagined experiences. We argue that the primary challenge in open-world decision-making is improving the exploration efficiency across a vast state space, especially for tasks that demand consideration of long-horizon payoffs. In this paper, we present LS-Imagine, which extends the imagination horizon within a limited number of state transition steps, enabling the agent to explore behaviors that potentially lead to promising long-term feedback. The foundation of our approach is to build a $\textit{long short-term world model}$. To achieve this, we simulate goal-conditioned jumpy state transitions and compute corresponding affordance maps by zooming in on specific areas within single images. This facilitates the integration of direct long-term values into behavior learning. Our method demonstrates significant improvements over state-of-the-art techniques in MineDojo. Jiajian Li, Qi Wang 0080, Yunbo Wang, Xin Jin 0014, Yang Li 0041, Wenjun Zeng 0001, Xiaokang Yang 0001 |
ICLR | 5 |
| 2025 | ABC-GS: Alignment-Based Controllable Style Transfer for 3D Gaussian Splattingabstract3D scene stylization approaches based on Neural Radiance Fields (NeRF) achieve promising results by optimizing with Nearest Neighbor Feature Matching (NNFM) loss. However, NNFM loss does not consider global style information. In addition, the implicit representation of NeRF limits their fine-grained control over the resulting scenes. In this paper, we introduce ABC-GS, a novel framework based on 3D Gaussian Splatting to achieve high-quality 3D style transfer. To this end, a controllable matching stage is designed to achieve precise alignment between scene content and style features through segmentation masks. Moreover, a style transfer loss function based on feature alignment is proposed to ensure that the outcomes of style transfer accurately reflect the global style of the reference image. Furthermore, the original geometric information of the scene is preserved with the depth loss and Gaussian regularization terms. Extensive experiments show that our ABC-GS provides controllability of style transfer and achieves stylization results that are more faithfully aligned with the global style of the chosen artistic reference. Our homepage is available at https://vpx-ecnu.github.io/ABC-GS-website. Zhongliang Liu, Man Sha, Yang Li 0041 |
ICME | 5 |
| 2025 | GTAD: Global Temporal Aggregation Denoising Learning for 3D Semantic Occupancy PredictionabstractAccurately perceiving dynamic environments is a fundamental task for autonomous driving and robotic systems. Existing methods inadequately utilize temporal information, relying mainly on local temporal interactions between adjacent frames and failing to leverage global sequence information effectively. To address this limitation, we investigate how to effectively aggregate global temporal features from temporal sequences, aiming to achieve occupancy representations that efficiently utilize global temporal information from historical observations. For this purpose, we propose a global temporal aggregation denoising network named GTAD, introducing a global temporal information aggregation framework as a new paradigm for holistic 3D scene understanding. Our method employs an in-model latent denoising network to aggregate local temporal features from the current moment and global temporal features from historical sequences. This approach enables the effective perception of both fine-grained temporal information from adjacent frames and global temporal patterns from historical observations. As a result, it provides a more coherent and comprehensive understanding of the environment. Extensive experiments on the nuScenes and Occ3D-nuScenes benchmark and ablation studies demonstrate the superiority of our method. Yang Li 0041, Yisheng Deng, Weifeng Ge |
IROS | 2 |
| 2025 | TimeSoccer: An End-to-End Multimodal Large Language Model for Soccer Commentary GenerationabstractSoccer is a globally popular sporting event, typically characterized by long matches and distinctive highlight moments. Recent advances in Multimodal Large Language Models (MLLMs) show promising capabilities in temporal grounding and video understanding. However, generating soccer commentary requires both precise temporal localization and semantically rich descriptions over long-form videos. Existing soccer MLLMs often rely on temporal priors for caption generation, which limits their ability to process the entire video in an end-to-end manner. Traditional approaches, on the other hand, follow a complex two-step paradigm that fails to capture the global context, leading to suboptimal performance. To solve the above issues, we present TimeSoccer, the first end-to-end soccer MLLM for Single-anchor Dense Video Captioning (SDVC) in full-match soccer videos. TimeSoccer jointly predicts timestamps and generates captions in a single pass, enabling global context modeling across 45-minute matches. To support long video understanding of soccer matches, we introduce MoFA-Select, a training-free, motion-aware frame compression module that adaptively selects representative frames via a coarse-to-fine strategy, and incorporates complementary training paradigms to strengthen the model's ability to handle long temporal sequences. Extensive experiments demonstrate that our TimeSoccer achieves State-of-The-Art (SoTA) performance on the SDVC task in an end-to-end form, generating high-quality commentary with accurate temporal alignment and strong semantic relevance. For more information, please visit: https://vpx-ecnu.github.io/TimeSoccer-Website/. Ling You, Wenxuan Huang 0001, Xinni Xie, Xiangyi Wei, Bangyan Li, Shaohui Lin, Yang Li 0041, Changbo Wang |
ACM Multimedia | 7 |
| 2025 | Wandering and feeling the Scenes: Body-Aware Diffusion for 3D Human Motion GenerationabstractAs demand for virtual digital characters grows in fields such as virtual reality, gaming, and animation, generating highly controllable human motion within scenes has become a key research focus. Existing methods for scene-aware motion generation typically rely on global alignment or latent space matching, which provides limited control over the fine-grained movements of individual body parts. This limitation often leads to rigid and unrealistic motions when interacting with complex environments. Therefore, we propose the Body-Aware Interaction Diffusion Model (BA-IDM), which enables fine-grained control of human motion within a scene by leveraging multimodal information. Text descriptions, motion scenes, and movement trajectories can all serve as inputs, allowing for precise control of each body part and facilitating the generation of a wide range of complex actions. Moreover, our approach is designed to operate on de-identified motion data, effectively protecting user privacy throughout the process, which is essential for practical and user-centric applications. Jingyu Gong, Shaohui Lin, Yang Li 0041, Zhizhong Zhang 0001 |
MMAsia | 4 |
| 2025 | Detail-preserving shape completion of point cloud models with articulated structure
Yi Quan, Chen Li 0035, Yang Li 0041, Changbo Wang, Hong Qin 0001 |
Comput. Aided Geom. Des. | 3 |
| 2025 | CeRF: Convolutional neural radiance derivative fields for new view synthesis
Ling You, Dingbo Lu, Yang Li 0041, Changbo Wang |
Comput. Graph. | 5 |
| 2025 | Warped convolutional neural networks for large homography transformation with psl(3) algebra
Xinrui Zhan, Wenyu Liu 0001, Risheng Yu, Jianke Zhu, Yang Li 0041 |
Neurocomputing | 5 |
| 2024 | Multi-Prototype Space Learning for Commonsense-Based Scene Graph GenerationabstractIn the domain of scene graph generation, modeling commonsense as a single-prototype representation has been typically employed to facilitate the recognition of infrequent predicates. However, a fundamental challenge lies in the large intra-class variations of the visual appearance of predicates, resulting in subclasses within a predicate class. Such a challenge typically leads to the problem of misclassifying diverse predicates due to the rough predicate space clustering. In this paper, inspired by cognitive science, we maintain multi-prototype representations for each predicate class, which can accurately find the multiple class centers of the predicate space. Technically, we propose a novel multi-prototype learning framework consisting of three main steps: prototype-predicate matching, prototype updating, and prototype space optimization. We first design a triple-level optimal transport to match each predicate feature within the same class to a specific prototype. In addition, the prototypes are updated using momentum updating to find the class centers according to the matching results. Finally, we enhance the inter-class separability of the prototype space through iterations of the inter-class separability loss and intra-class compactness loss. Extensive evaluations demonstrate that our approach significantly outperforms state-of-the-art methods on the Visual Genome dataset. Lianggangxu Chen, Youqi Song, Yiqing Cai, Jiale Lu, Yang Li 0041, Changbo Wang, Gaoqi He |
AAAI | 5 |
| 2024 | HeRF: A Hierarchical Framework for Efficient and Extendable New View SynthesisabstractRecently, neural radiance fields have made significant advancements in rendering new views. However, limited research has focused on dynamically loading implicit radiance fields with efficient memory utilization and extended scene representation. This paper introduces HeRF, a novel framework with a hierarchical scene representation based on layered sparse voxels. With such an adaptive design, our method is able to partition scenes into different levels for faster modeling and reduced memory cost. Furthermore, these partitioned scenes can be dynamically loaded and joined for a better immersive experience. Quantitative and qualitative analysis using objectlevel, indoor, and outdoor datasets demonstrates the effectiveness of HeRF. Remarkably, our proposed method requires only about 38% of the training rays and 45% of the GPU memory cost, yet achieves a 9% improvement in PSNR compared to NeRFusion on the ScanNet dataset. The code is available at https://github.com/Minisal/HeRF. Dingbo Lu, Ling You, Yang Li 0041, Changbo Wang |
IJCNN | 5 |
| 2024 | Channel Robust Strategies with Data Augmentation for Audio Anti-spoofing
Sardor Mamarasulov, Yang Li 0041, Changbo Wang |
ISC (2) | 2 |
| 2024 | Generative Data Augmentation with Liveness Information Preserving for Face Anti-SpoofingabstractFace anti-spoofing is a critical aspect of ensuring security in the context of human-robot interaction and collaboration. Recently, disentangled-based data augmentation methods have achieved great success in face anti-spoofing tasks. The underlying assumption of those methods is that the liveness information could be completely disentangled and the labeling of the augmented data could totally depend on the liveness-related feature branch. However, we observe that it is almost impossible to extract the liveness-related information completely, which makes the current labeling strategy inaccurate. In this paper, we rethink the disentangling process and propose a novel generative-based data augmentation framework without forcing liveness information encoded into any specific feature space. Specifically, the original images are decomposed into statistic feature space and spatial feature space with liveness information preserving. With these two feature spaces, synthesized liveness-preserving images are generated with the Cartesian product to further approach the distribution of real face anti-spoofing data. Along with the original samplings, the augmented data are fed to a ResNet-based classifier with our proposed pseudo-label strategy for liveness information augmentation. Both qualitative and quantitative experiments demonstrate promising results to show the effectiveness of our proposed method. Changgu Chen, Yang Li 0041, Jian Zhang 0079, Changbo Wang |
ICMR | 2 |
| 2024 | FIND: Fine-tuning Initial Noise Distribution with Policy Optimization for Diffusion ModelsabstractIn recent years, large-scale pre-trained diffusion models have demonstrated their outstanding capabilities in image and video generation tasks. However, existing models tend to produce visual objects commonly found in the training dataset, which diverges from user input prompts. The underlying reason behind the inaccurate generated results lies in the model's difficulty in sampling from specific intervals of the initial noise distribution corresponding to the prompt. Moreover, it is challenging to directly optimize the initial distribution, given that the diffusion process involves multiple denoising steps. In this paper, we introduce a Fine-tuning Initial Noise Distribution (FIND) framework with policy optimization, which unleashes the powerful potential of pre-trained diffusion networks by directly optimizing the initial distribution to align the generated contents with user-input prompts. To this end, we first reformulate the diffusion denoising procedure as a one-step Markov decision process and employ policy optimization to directly optimize the initial distribution. In addition, a dynamic reward calibration module is proposed to ensure training stability during optimization. Furthermore, we introduce a ratio clipping algorithm to utilize historical data for network training and prevent the optimized distribution from deviating too far from the original policy to restrain excessive optimization magnitudes. Extensive experiments demonstrate the effectiveness of our method in both text-to-image and text-to-video tasks, surpassing SOTA methods in achieving consistency between prompts and the generated content. Our method achieves 10 times faster than the SOTA approach. Changgu Chen, Libing Yang, Lianggangxu Chen, Gaoqi He, Changbo Wang, Yang Li 0041 |
ACM Multimedia | 7 |
| 2024 | ChatTracker: Enhancing Visual Tracking Performance via Chatting with Multimodal Large Language ModelabstractVisual object tracking aims to locate a targeted object in a video sequence based on an initial bounding box. Recently, Vision-Language~(VL) trackers have proposed to utilize additional natural language descriptions to enhance versatility in various applications. However, VL trackers are still inferior to State-of-The-Art (SoTA) visual trackers in terms of tracking performance. We found that this inferiority primarily results from their heavy reliance on manual textual annotations, which include the frequent provision of ambiguous language descriptions. In this paper, we propose ChatTracker to leverage the wealth of world knowledge in the Multimodal Large Language Model (MLLM) to generate high-quality language descriptions and enhance tracking performance. To this end, we propose a novel reflection-based prompt optimization module to iteratively refine the ambiguous and inaccurate descriptions of the target with tracking feedback. To further utilize semantic information produced by MLLM, a simple yet effective VL tracking framework is proposed and can be easily integrated as a plug-and-play module to boost the performance of both VL and visual trackers. Experimental results show that our proposed ChatTracker achieves a performance comparable to existing methods. Yiming Sun 0006, Shaoxiang Chen 0001, Junwei Huang, Yang Li 0041, Chenhui Li 0001, Changbo Wang |
NeurIPS | 6 |
| 2024 | InvVis: Large-Scale Data Embedding for Invertible VisualizationabstractWe present InvVis, a new approach for invertible visualization, which is reconstructing or further modifying a visualization from an image. InvVis allows the embedding of a significant amount of data, such as chart data, chart information, source code, etc., into visualization images. The encoded image is perceptually indistinguishable from the original one. We propose a new method to efficiently express chart data in the form of images, enabling large-capacity data embedding. We also outline a model based on the invertible neural network to achieve high-quality data concealing and revealing. We explore and implement a variety of application scenarios of InvVis. Additionally, we conduct a series of evaluation experiments to assess our method from multiple perspectives, including data embedding quality, data restoration accuracy, data encoding capacity, etc. The result of our experiments demonstrates the great potential of InvVis in invertible visualization. Huayuan Ye, Chenhui Li 0001, Yang Li 0041, Changbo Wang |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2023 | Multi-Source Templates Learning for Real-Time Aerial TrackingabstractAerial tracking aims at tracking an arbitrary visual object in a video captured by Unmanned Aerial Vehicles (UAV). Due to the scarce computation resources, the deployment of high-consuming state-of-the-art trackers on UAV becomes impractical. On the other hand, lightweight trackers suffer from inferior performance caused by the low sampling frequency and resolution of UAV videos. In this paper, we propose a novel multi-source templates learning method to alleviate the paradox of efficiency and effectiveness for aerial tracking. Besides conventional static and dynamic templates, our work introduces an additional general-object template to learn common feature properties of a general object during training time. To exploit all templates information, a multi-source templates fusion scheme is proposed to capture characteristics of object in low quality UAV video streams. Furthermore, a joint optimization process is employed to enforce the lightness of model while achieving comparable tracking performance. Our experimental results demonstrate an appealing performance trade-off between accuracy and speed. The proposed tracker achieves 200 FPS on GPU, 100 FPS on CPU, and 12 FPS on Nvidia Jetson Xavier NX, respectively. Our code will be released at https://github.com/vpx-ecnu/MSTL. Yiming Sun 0006, Yang Li 0041, Changbo Wang |
ICASSP | 2 |
| 2023 | AdaptMVSNet: Efficient Multi-View Stereo with adaptive convolution and attention fusion
Yuanjie Chen, Wenjie Song 0002, Yang Li 0041 |
Comput. Graph. | 5 |
| 2023 | Contact-conditioned hand-held object reconstruction from single-view images
Yang Li 0041, Adnane Boukhayma, Changbo Wang, Marc Christie |
Comput. Graph. | 2 |
| 2022 | Homography Decomposition Networks for Planar Object TrackingabstractPlanar object tracking plays an important role in AI applications, such as robotics, visual servoing, and visual SLAM. Although the previous planar trackers work well in most scenarios, it is still a challenging task due to the rapid motion and large transformation between two consecutive frames. The essential reason behind this problem is that the condition number of such a non-linear system changes unstably when the searching range of the homography parameter space becomes larger. To this end, we propose a novel Homography Decomposition Networks~(HDN) approach that drastically reduces and stabilizes the condition number by decomposing the homography transformation into two groups. Specifically, a similarity transformation estimator is designed to predict the first group robustly by a deep convolution equivariant network. By taking advantage of the scale and rotation estimation with high confidence, a residual transformation is estimated by a simple regression model. Furthermore, the proposed end-to-end network is trained in a semi-supervised fashion. Extensive experiments show that our proposed approach outperforms the state-of-the-art planar tracking methods at a large margin on the challenging POT, UCSB and POIC datasets. Codes and models are available at https://github.com/zhanxinrui/HDN. Xinrui Zhan, Yueran Liu, Jianke Zhu, Yang Li 0041 |
AAAI | 4 |
| 2021 | SuPer Deep: A Surgical Perception Framework for Robotic Tissue Manipulation using Deep Learning for Feature ExtractionabstractRobotic automation in surgery requires precise tracking of surgical tools and mapping of deformable tissue. Previous works on surgical perception frameworks require significant effort in developing features for surgical tool and tissue tracking. In this work, we overcome the challenge by exploiting deep learning methods for surgical perception. We integrated deep neural networks, capable of efficient feature extraction, into the tissue tracking and surgical tool tracking processes. By leveraging transfer learning, the deep-learning-based approach requires minimal training data and reduced feature engineering efforts to fully perceive a surgical scene. The framework was tested on three publicly available datasets, which use the da Vinci® Surgical System, for comprehensive analysis. Experimental results show that our framework achieves state-of-the-art tracking performance in a surgical environment by utilizing deep learning for feature extraction. Jingpei Lu, Ambareesh Jayakumari, Florian Richter 0002, Yang Li 0041, Michael C. Yip |
ICRA | 4 |
| 2021 | Attribute-Aware Pedestrian Detection in a CrowdabstractPedestrian detection is an initial step to perform outdoor scene analysis, which plays an essential role in many real-world applications. Although having enjoyed the merits of deep learning frameworks from the generic object detectors, pedestrian detection is still a very challenging task due to heavy occlusions, and highly crowded group. Generally, the conventional detectors are unable to differentiate individuals from each other effectively under such a dense environment. To tackle this critical problem, we propose an attribute-aware pedestrian detector to explicitly model people's semantic attributes in a high-level feature detection fashion. Besides the typical semantic features, center position, target's scale, and offset, we introduce a pedestrian-oriented attribute feature to encode the high-level semantic differences among the crowd. Moreover, a novel attribute-feature-based Non-Maximum Suppression (NMS) is proposed to distinguish the person from a highly overlapped group by adaptively rejecting the false-positive results in a very crowd settings. Furthermore, an enhanced ground truth target is designed to alleviate the difficulties caused by the attribute configuration, and to ease the class imbalance issue during training. Finally, we evaluate our proposed attribute-aware pedestrian detector on three benchmark datasets including CityPerson, CrowdHuman, and EuroCityPerson, and achieves the state-of-the-art results. Lixiang Lin, Jianke Zhu, Yang Li 0041, Yun-chen Chen, Yao Hu 0002, Steven C. H. Hoi |
IEEE Trans. Multim. | 4 |
| 2020 | DeepFacade: A Deep Learning Approach to Facade Parsing With Symmetric LossabstractParsing building facades into procedural grammars plays an important role for 3D building model generation tasks, which have been long desired in computer vision. Deep learning is a promising approach to facade parsing, however, a straightforward solution by directly applying standard deep learning approaches cannot always yield the optimal results. This is primarily due to two reasons: 1) it is nontrivial to train existing semantic segmentation networks for facade parsing, e.g., Fully-Convolutional Neural Networks (FCN) which are usually weak at predicting fine-grained shapes (J. Long et al., 2015); and 2) building facades are man-made architectures with highly regularized shape priors, and the prior knowledge plays an important role in facade parsing, for which how to integrate the prior knowledge into deep neural networks remains an open problem. In this paper, we present a novel symmetric loss function that can be used in deep neural networks for end-to-end training. This novel loss is based on the assumption that most of windows and doors have a highly symmetric rectangle shape, and it penalizes all window predictions that are non-rectangles. This prior knowledge is smoothly integrated into the end-to-end training process. Quantitative evaluation demonstrates that our method has outperformed previous state-of-art methods significantly on five popular facade parsing datasets. Qualitative results have shown that our method effectively aids deep convolutional neural networks to predict more accurate, visually pleasing, and symmetric shapes. To the best of our knowledge, we are the first to incorporate symmetry constraint into end-to-end training in deep neural networks for facade parsing. Hantang Liu, Jianke Zhu, Yang Li 0041, Steven C. H. Hoi |
IEEE Trans. Multim. | 5 |
| 2019 | Robust Estimation of Similarity Transformation for Visual Object TrackingabstractMost of existing correlation filter-based tracking approaches only estimate simple axis-aligned bounding boxes, and very few of them is capable of recovering the underlying similarity transformation. To tackle this challenging problem, in this paper, we propose a new correlation filter-based tracker with a novel robust estimation of similarity transformation on the large displacements. In order to efficiently search in such a large 4-DoF space in real-time, we formulate the problem into two 2-DoF sub-problems and apply an efficient Block Coordinates Descent solver to optimize the estimation result. Specifically, we employ an efficient phase correlation scheme to deal with both scale and rotation changes simultaneously in log-polar coordinates. Moreover, a variant of correlation filter is used to predict the translational motion individually. Our experimental results demonstrate that the proposed tracker achieves very promising prediction performance compared with the state-of-the-art visual object tracking methods while still retaining the advantages of high efficiency and simplicity in conventional correlation filter-based tracking methods. Yang Li 0041, Jianke Zhu, Steven C. H. Hoi, Wenjie Song 0002, Hantang Liu |
AAAI | 1 |
| 2018 | Temporally-adjusted correlation filter-based tracking
Wenjie Song 0002, Yang Li 0041, Jianke Zhu, Chun Chen 0001 |
Neurocomputing | 2 |
| 2017 | CFNN: Correlation Filter Neural Network for Visual Object TrackingabstractAlbeit convolutional neural network (CNN) has shown promising capacity in many computer vision tasks, applying it to visual tracking is yet far from solved. Existing methods either employ a large external dataset to undertake exhaustive pre-training or suffer from less satisfactory results in terms of accuracy and robustness. To track single target in a wide range of videos, we present a novel Correlation Filter Neural Network architecture, as well as a complete visual tracking pipeline, The proposed approach is a special case of CNN, whose initialization does not need any pre-training on the external dataset. The initialization of network enjoys the merits of cyclic sampling to achieve the appealing discriminative capability, while the network updating scheme adopts advantages from back-propagation in order to capture new appearance variations. The tracking pipeline integrates both aspects well by making them complementary to each other. We validate our tracker on OTB-2013 benchmark. The proposed tracker obtains the promising results compared to most of existing representative trackers. Yang Li 0041, Jianke Zhu |
IJCAI | 1 |
| 2016 | Image Alignment by Online Robust PCA via Stochastic Gradient DescentabstractAligning a given set of images is usually conducted in batch mode manner, which not only requires large amount of memory but also adjusts all the previous transformations to register an input image. To address this issue, we propose a novel approach to image alignment by incorporating the geometric transformation into online robust principal component analysis (PCA). Instead of calculating the warp update using noisy input samples like the conventional methods, we suggest directly linearizing the object function by performing warp update on the recovered samples, which corresponds to an efficient inverse composition algorithm. Since the basis matrix is kept constant for a given sample, both the latent vector and warp update can be very efficiently computed. Moreover, we present two basis updating methods for robust PCA, including the closed-form solution and stochastic gradient descent scheme. We have conducted the extensive experiments on the real-world tasks of background subtraction with camera motion and visual tracking on the challenging video sequences, whose promising results demonstrate the efficacy of our presented approach. Wenjie Song 0002, Jianke Zhu, Yang Li 0041, Chun Chen 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2015 | Reliable Patch Trackers: Robust visual tracking by exploiting reliable patchesabstractMost modern trackers typically employ a bounding box given in the first frame to track visual objects, where their tracking results are often sensitive to the initialization. In this paper, we propose a new tracking method, Reliable Patch Trackers (RPT), which attempts to identify and exploit the reliable patches that can be tracked effectively through the whole tracking process. Specifically, we present a tracking reliability metric to measure how reliably a patch can be tracked, where a probability model is proposed to estimate the distribution of reliable patches under a sequential Monte Carlo framework. As the reliable patches distributed over the image, we exploit the motion trajectories to distinguish them from the background. Therefore, the visual object can be defined as the clustering of homo-trajectory patches, where a Hough voting-like scheme is employed to estimate the target state. Encouraging experimental results on a large set of sequences showed that the proposed approach is very effective and in comparison to the state-of-the-art trackers. The full source code of our implementation will be publicly available. Yang Li 0041, Jianke Zhu, Steven C. H. Hoi |
CVPR | 1 |
| 2011 | Adaptive lattice-based light rendering of participating mediaabstractABSTRACT The visual world around us displays a rich set of light effects because of translucent and participating media. It is hard and time consuming to render these effects with scattering, caustic, and shaft because of the complex interaction between light and different media. This paper presents a new rendering method based on adaptive lattice for lighting participating media of translucent materials such as marble, wax, and shaft light. Firstly, on the basis of the lattice‐based photon tracing model, multi‐scale hierarchical lattice was constructed by mixed lattice types sampling combined cubic Cartesian and face‐centered cubic with view‐dependent adaptive resolution. Then, an adaptive method to trace diffuse photons and marked specular photons with different phase functions was suggested. Multiple lights and heterogeneous materials were also considered here. Further, the mixed rendering method and GPU accelerate technology were introduced to render different light effects under different participating media. Copyright © 2011 John Wiley & Sons, Ltd. Changbo Wang, Chenhui Li 0001, Jinqiu Dai, Yang Li 0041 |
Comput. Animat. Virtual Worlds | 4 |