Qiankun Liu 0001

dblp:179/8822-1 · DBLP profile ↗
← Back
24ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-7457-2090ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 15 · 5 first-author · 14 since 2021
YearPublicationVenuePosition
2026 MagicPaint: Operate Anything for Image Inpainting with Diffusion Model
abstract
Recent diffusion-based models have significantly improved inpainting quality. However, existing methods struggle with multi-task inpainting due to conflicting optimization objectives, and current datasets are typically limited to task-specific scenarios, hindering joint training. To address these challenges, we propose MagicPaint, a unified diffusion-based inpainting model that supports object addition, removal, and unconditional inpainting across both text and image modalities. MagicPaint semantically decouples operation types and target content by learnable tokens in MMToken Module, effectively reconciling conflicting optimization objectives and enabling robust multi-task, multi-modal inpainting. Besides, a novel inpainting paradigm named MagicMask, encodes operating intent directly into the mask and applies a mask loss for spatially precise supervision. In addition, existing inpainting datasets are insufficient for multi-task and multi-modal scenarios, limiting the capability of inpainting models. Thus, we further introduce a new dataset comprising 2.1M image tuples. It is dedicatedly designed to support diverse inpainting scenarios and significantly improves upon existing datasets, particularly in object removal. Through efforts from both model and data perspectives, MagicPaint enables users to operate anything—add, remove or inpaint content which is specified through either text or image modalities in a seamless and unified manner. Extensive experiments demonstrate that MagicPaint achieves state-of-the-art performance across three key tasks (i.e., text-guided addition, image-guided addition, and object removal) and produces outputs with superior visual consistency and contextual fidelity compared to existing methods.
Qinhong Yang, Dongdong Chen 0001, Qi Chu 0001, Qiankun Liu 0001, Zhentao Tan, Xulin Li, Huamin Feng, Nenghai Yu
AAAI5
2026 MSFA Image Denoising Using Physics-Based Noise Model and Noise-Decoupled Network
abstract
Multispectral filter array (MSFA) camera is increasingly used due to its compact size and fast capturing speed. However, because of its narrow-band property, it often suffers from the light-deficient problem, and images captured are easily overwhelmed by noise. As a type of commonly used denoising method, neural networks have shown their power to achieve satisfactory denoising results. However, their performance highly depends on high-quality noisy-clean image pairs. For the task of MSFA image denoising, there is currently neither a paired real dataset nor an accurate noise model capable of generating realistic noisy images. To this end, we present a physics-based noise model that is capable to match the real noise distribution and synthesize realistic noisy images. In our noise model, those different types of noise can be divided into SimpleDist component and ComplexDist component. The former contains all the types of noise that can be described using a simple probability distribution like Gaussian or Poisson distribution, and the latter contains the complicated color bias noise that cannot be modeled using a simple probability distribution. Besides, we design a noise-decoupled network consisting of a SimpleDist noise removal network (SNRNet) and a ComplexDist noise removal network (CNRNet) to sequentially remove each component. Moreover, according to the non-uniformity of color bias noise in our noise model, we introduce a learnable position embedding in CNRNet to indicate the position information. To verify the effectiveness of our physics-based noise model and noise-decoupled network, we collect a real MSFA denoising dataset with paired long-exposure clean images and short-exposure noisy images. Experiments are conducted to prove that the network trained using synthetic data generated by our noise model performs as well as trained using paired real data, and our noise-decoupled network outperforms other state-of-the-art denoising methods.
Ying Fu 0001, Qiankun Liu 0001, Jun Zhang 0007
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 UNICL-SAM: Uncertainty-Driven In-Context Segmentation with Part Prototype Discovery
abstract
Recent advancements in in-context segmentation generalists have demonstrated significant success in performing various image segmentation tasks using a limited number of labeled example images. However, real-world applications present challenges due to the variability of support examples, which often exhibit quality issues resulting from various sources and inaccurate labeling. How to extract more robust representations from these examples has always been one of the goals of in-context visual learning. In response, we propose UNICL-SAM, to better model the example distribution and extract robust representations to help in-context segmentation. We incorporate an uncertainty probabilistic module to quantify each example’s reliability during both the training and testing phases. Utilizing this uncertainty estimation, we introduce an uncertainty-guided graph augmentation and feature refinement strategy, aimed at mitigating the impact of high-uncertainty regions to enhance the learning of robust representations. Subsequently, we construct prototypes for each example by aggregating part information, thereby creating reliable in-context instruction that effectively represents fine-grained local semantics. This approach serves as a valuable complement to traditional global pooling features. Experimental results demonstrate the effectiveness of the proposed framework, underscoring its potential for real-world applications.
Dianmo Sheng, Dongdong Chen 0001, Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Wenbin Tu, Shengwei Xu, Nenghai Yu
CVPR4
2025 MonoMVSNet: Monocular Priors Guided Multi-View Stereo Network
abstract
Learning-based Multi-View Stereo (MVS) methods aim to predict depth maps for a sequence of calibrated images to recover dense point clouds. However, existing MVS methods often struggle with challenging regions, such as textureless regions and reflective surfaces, where feature matching fails. In contrast, monocular depth estimation inherently does not require feature matching, allowing it to achieve robust relative depth estimation in these regions. To bridge this gap, we propose MonoMVSNet, a novel monocular feature and depth guided MVS network that integrates powerful priors from a monocular foundation model into multi-view geometry. Firstly, the monocular feature of the reference view is integrated into source view features by the attention mechanism with a newly designed cross-view position encoding. Then, the monocular depth of the reference view is aligned to dynamically update the depth candidates for edge regions during the sampling procedure. Finally, a relative consistency loss is further designed based on the monocular depth to supervise the depth prediction. Extensive experiments demonstrate that MonoMVSNet achieves state-of-the-art performance on the DTU and Tanks-and-Temples datasets, ranking first on the Tanks-and-Temples Intermediate and Advanced benchmarks. The source code is available at https://github.com/JianfeiJ/MonoMVSNet.
Jianfei Jiang 0006, Qiankun Liu 0001, Haochen Yu, Hongyuan Liu 0007, Liyong Wang, Huimin Ma 0001
ICCV2
2025 MVSMamba: Multi-View Stereo with State Space Model
abstract
Robust feature representations are essential for learning-based Multi-View Stereo (MVS), which relies on accurate feature matching. Recent MVS methods leverage Transformers to capture long-range dependencies based on local features extracted by conventional feature pyramid networks. However, the quadratic complexity of Transformer-based MVS methods poses challenges to balance performance and efficiency. Motivated by the global modeling capability and linear complexity of the Mamba architecture, we propose MVSMamba, the first Mamba-based MVS network. MVSMamba enables efficient global feature aggregation with minimal computational overhead. To fully exploit Mamba's potential in MVS, we propose a Dynamic Mamba module (DM-module) based on a novel reference-centered dynamic scanning strategy, which enables: (1) Efficient intra- and inter-view feature interaction from the reference to source views, (2) Omnidirectional multi-view feature representations, and (3) Multi-scale global feature aggregation. Extensive experimental results demonstrate MVSMamba outperforms state-of-the-art MVS methods on the DTU dataset and the Tanks-and-Temples benchmark with both superior performance and efficiency. The source code is available at https://github.com/JianfeiJ/MVSMamba.
Jianfei Jiang 0006, Qiankun Liu 0001, Hongyuan Liu 0007, Haochen Yu, Liyong Wang, Huimin Ma 0001
NeurIPS2
2024 Infrared Small Target Detection with Scale and Location Sensitivity
abstract
Recently, infrared small target detection (IRSTD) has been dominated by deep-learning-based methods. However, these methods mainly focus on the design of complex model structures to extract discriminative features, leaving the loss functions for IRSTD under-explored. For ex-ample, the widely used Intersection over Union (IoU) and Dice losses lack sensitivity to the scales and locations of targets, limiting the detection performance of detectors. In this paper, we focus on boosting detection performance with a more effective loss but a simpler model structure. Specifically, we first propose a novel Scale and Location Sensitive (SLS) loss to handle the limitations of existing losses: 1) for scale sensitivity, we compute a weight for the IoU loss based on target scales to help the detector distinguish targets with different scales: 2) for location sensitivity, we introduce a penalty term based on the center points of targets to help the detector localize targets more precisely. Then, we design a simple Multi-Scale Head to the plain U-Net (MSHNet). By applying SLS loss to each scale of the predictions, our MSHNet outperforms existing state-of-the-art methods by a large margin. In addition, the detection performance of existing detectors can be further improved when trained with our SLS loss, demonstrating the effectiveness and generalization of our SLS loss. The code is available at https://github.com/ying-fu/MSHNet.
Qiankun Liu 0001, Rui Liu 0039, Bolun Zheng, Hongkui Wang, Ying Fu 0001
CVPR1
2024 Towards More Unified In-Context Visual Understanding
abstract
The rapid advancement of large language models (LLMs) has accelerated the emergence of in-context learning (ICL) as a cutting-edge approach in the natural language processing domain. Recently, ICL has been employed in visual understanding tasks, such as semantic segmentation and image captioning, yielding promising results. However, existing visual ICL framework can not enable producing content across multiple modalities, whicd limits their potential usage scenarios. To address this issue, we present a new ICLframeworkfor visual understanding with multi-modal output enabled. First, we quantize and embed both text and visual prompt into a unified representational space, structured as interleaved in-context sequences. Then a decoder-only sparse transformer architecture is employed to perform generative modeling on them, facilitating in-context learning. Thanks to this design, the model is capable of handling in-context vision understanding tasks with multimodal output in a unified pipeline. Experimental re-sults demonstrate that our model achieves competitive performance compared with specialized models and previous ICL baselines. Overall, our research takes a further step toward unified multimodal in-context learning.
Dianmo Sheng, Dongdong Chen 0001, Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Jianmin Bao, Bin Liu 0016, Shengwei Xu, Nenghai Yu
CVPR4
2024 Multi-Object Tracking in the Dark
abstract
Low-light scenes are prevalent in real-world applications (e.g. autonomous driving and surveillance at night). Recently, multi-object tracking in various practical use cases have received much attention, but multi-object tracking in dark scenes is rarely considered. In this paper, we focus on multi-object tracking in dark scenes. To address the lack of datasets, we first build a Low-light Multi-Object Tracking (LMOT) dataset. LMOT provides well-aligned low-light video pairs captured by our dual-camera system, and high-quality multi-object tracking annotations for all videos. Then, we propose a low-light multi-object tracking method, termed as LTrack. We introduce the adaptive low-pass downsample module to enhance low-frequency components of images outside the sensor noises. The degradation suppression learning strategy enables the model to learn invariant information under noise disturbance and image quality degradation. These components improve the robustness of multi-object tracking in dark scenes. We conducted a comprehensive analysis of our LMOT dataset and proposed LTrack. Experimental results demonstrate the superiority of the proposed method and its competitiveness in real night low-light scenes. Dataset and Code: https:/github.com/ying-fu/LMOT
Qiankun Liu 0001, Yunhao Zou, Ying Fu 0001
CVPR3
2024 Object-Aware NIR-to-Visible Translation
Yunyi Gao, Lin Gu 0003, Qiankun Liu 0001, Ying Fu 0001
ECCV (23)3
2024 Transformer Based Pluralistic Image Completion With Reduced Information Loss
abstract
Transformer based methods have achieved great success in image inpainting recently. However, we find that these solutions regard each pixel as a token, thus suffering from an information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration. 2) They quantize 2563RGB values to a small number (such as 512) of quantized color values. The indices of quantized pixels are used as tokens for the inputs and prediction targets of the transformer. To mitigate these issues, we propose a new transformer based framework called “PUT”. Specifically, to avoid input downsampling while maintaining computation efficiency, we design a patch-based auto-encoder P-VQVAE. The encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by input quantization, an Un-quantized Transformer is applied. It directly takes features from the P-VQVAE encoder as input without any quantization and only regards the quantized tokens as prediction targets.Furthermore, to make the inpainting process more controllable, we introduce semantic and structural conditions as extra guidance. Extensive experiments show that our method greatly outperforms existing transformer based methods on image fidelity and achieves much higher diversity and better fidelity than state-of-the-art pluralistic inpainting methods on complex large-scale datasets (e.g., ImageNet). Codes are available athttps://github.com/liuqk3/PUT.
Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Ying Fu 0001, Qi Chu 0001, Gang Hua 0001, Nenghai Yu
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Raw Image Based Over-Exposure Correction Using Channel-Guidance Strategy
abstract
Most existing methods for over-exposure in image correction are developed based on sRGB images, which can result in complex and non-linear degradation due to the image signal processing pipeline. By contrast, data-driven approaches based on RAW image data offer natural advantages for image processing tasks. RAW images, characterized by their near-linear correlation with scene radiance and enriched information content due to higher bit depth, demonstrate superior performance compared to sRGB-based techniques. Further, the spectral sensitivity characteristics intrinsic to digital camera sensors indicate that the blue and red channels in a Bayer pattern RAW image typically encompass more contextual information than the green channels. This property renders them less susceptible to over-exposure, thereby making them more effective for data extraction in high dynamic range scenes. In this paper, we introduce a Channel-Guidance Network (CGNet) that leverages the benefits of RAW images for over-exposure correction. The CGNet estimates the properly-exposed sRGB image directly from the over-exposed RAW image in an end-to-end manner. Specifically, we introduce a RAW-based channel-guidance branch to the U-net-based backbone, which exploits the color channel intensity prior of RAW images to achieve superior over-exposure correction performance. To further facilitate research in over-exposure correction, we present synthetic and real-world over-exposure correction benchmark datasets. These datasets comprise a large set of paired RAW and sRGB images across a variety of scenarios. Experiments on our RAW-sRGB datasets validate the advantages of our RAW-based channel guidance strategy and proposed CGNet over state-of-the-art sRGB-based methods on over-exposure correction. Our code and dataset are publicly available athttps://github.com/whiteknight-WJN/CGNet.
Ying Fu 0001, Yunhao Zou, Qiankun Liu 0001, Chenggang Yan 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Siamese-DETR for Generic Multi-Object Tracking
abstract
The ability to detect and track the dynamic objects in different scenes is fundamental to real-world applications, e.g., autonomous driving and robot navigation. However, traditional Multi-Object Tracking (MOT) is limited to track objects belonging to the pre-defined closed-set categories. Recently, Generic MOT (GMOT) is proposed to track interested objects beyond pre-defined categories and it can be divided into Open-Vocabulary MOT (OVMOT) and Template-Image-based MOT (TIMOT). Taking the consideration that the expensive well pre-trained (vision-)language model and fine-grained category annotations are required to train OVMOT models, in this paper, we focus on TIMOT and propose a simple but effective method, Siamese-DETR. Only the commonly used detection datasets (e.g., COCO) are required for training. Different from existing TIMOT methods, which train a Single Object Tracking (SOT) based detector to detect interested objects and then apply a data association based MOT tracker to get the trajectories, we leverage the inherent object queries in DETR variants. Specifically: 1) The multi-scale object queries are designed based on the given template image, which are effective for detecting different scales of objects with the same category as the template image; 2) A dynamic matching training strategy is introduced to train Siamese-DETR on commonly used detection datasets, which takes full advantage of provided annotations; 3) The online tracking pipeline is simplified through a tracking-by-query manner by incorporating the tracked boxes in the previous frame as additional query boxes. The complex data association is replaced with the much simpler Non-Maximum Suppression (NMS). Extensive experimental results show that Siamese-DETR surpasses existing MOT methods on GMOT-40 dataset by a large margin.
Qiankun Liu 0001, Ying Fu 0001
IEEE Trans. Image Process.1
2024 Exploring the Application of Large-Scale Pre-Trained Models on Adverse Weather Removal
abstract
Image restoration under adverse weather conditions (e.g., rain, snow, and haze) is a fundamental computer vision problem that has important implications for various downstream applications. Distinct from early methods that are specially designed for specific types of weather, recent works tend to simultaneously remove various adverse weather effects based on either spatial feature representation learning or semantic information embedding. Inspired by various successful applications incorporating large-scale pre-trained models (e.g., CLIP), in this paper, we explore their potential benefits for leveraging large-scale pre-trained models in this task based on both spatial feature representation learning and semantic information embedding aspects: 1) spatial feature representation learning, we design a Spatially Adaptive Residual (SAR) encoder to adaptively extract degraded areas. To facilitate training of this model, we propose a Soft Residual Distillation (CLIP-SRD) strategy to transfer spatial knowledge from CLIP between clean and adverse weather images; 2) semantic information embedding, we propose a CLIP Weather Prior (CWP) embedding module to enable the network to adaptively respond to different weather conditions. This module integrates the sample-specific weather priors extracted by the CLIP image encoder with the distribution-specific information (as learned by a set of parameters) and embeds these elements using a cross-attention mechanism. Extensive experiments demonstrate that our proposed method can achieve state-of-the-art performance under various and severe adverse weather conditions. The code will be made available.
Zhentao Tan, Qiankun Liu 0001, Qi Chu 0001, Le Lu 0001, Jieping Ye, Nenghai Yu
IEEE Trans. Image Process.3
2024 AnimeDiff: Customized Image Generation of Anime Characters Using Diffusion Model
abstract
Due to the unprecedented power of text-to-image diffusion models, customizing these models to generate new concepts has gained increasing attention. Existing works have achieved some success on real-world concepts, but fail on the concepts of anime characters. We empirically find that such low quality comes from the newly introduced identifier text tokens, which are optimized to identify different characters. In this paper, we proposeAnimeDiffwhich focuses on customized image generation of anime characters. Our AnimeDiff directly binds anime characters with their names and keeps the embeddings of text tokens unchanged. Furthermore, when composing multiple characters in a single image, the model tends to confuse the properties of those characters. To address this issue, our AnimeDiff incorporates aCut-and-Pastedata augmentation strategy that produces multi-character images for training by cutting and pasting multiple characters onto background images. Experiments are conducted to prove the superiority of AnimeDiff over other methods.
Qiankun Liu 0001, Dongdong Chen 0001, Lu Yuan 0001, Ying Fu 0001
IEEE Trans. Multim.2
2023 Dual-Feature Enhancement for Weakly Supervised Temporal Action Localization
abstract
Weakly-supervised Temporal Action Localization (WTAL) aims at localizing actions in untrimmed videos with only video-level labels. Most existing methods embrace a "localization by classification" paradigm and adopt a model that pre-trained with recognition task for feature extraction. The gap between recognition and localization tasks leads to inferior performance. Some recent works attempt to utilize feature enhancement to obtain better feature for localization and boost the performance to some extent. However, they are limited to intra-video information exploiting, while ignoring meaningful inter-video information in the dataset. In this paper, we propose a novel Dual-Feature Enhancement (DFE) method for WTAL, which can utilize both intra-and inter-video information. For intra-video, a local feature enhancement module is designed to promote the feature interaction along the temporal dimension within each video. For inter-video information, a global memory module is firstly designed to learn the representations for different categories across different videos. Then, a global feature enhancement module is used to enhance the video features with the help of those global representations in the memory. Besides, to reduce the extra computational cost caused by global enhancement module in the inference stage, a distillation loss is applied to enforce the local branch to learn the information from global branch, so the global enhancement module could be removed during inference. The proposed method achieves state-of-the-art performance on popular benchmarks.
Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu
ICASSP2
2023 MGT: Modality-Guided Transformer for Infrared and Visible Image Fusion
Taoying Zhang, Hesong Li, Qiankun Liu 0001, Xiaoyong Wang, Ying Fu 0001
PRCV (1)3
2023 Semantic Probability Distribution Modeling for Diverse Semantic Image Synthesis
abstract
Semantic image synthesis, translating semantic layouts to photo-realistic images, is a one-to-many mapping problem. Though impressive progress has been recently made, diverse semantic synthesis that can efficiently produce semantic-level or even instance-level multimodal results, still remains a challenge. In this article, we propose a novel diverse semantic image synthesis framework from the perspective of semantic class distributions, which naturally supports diverse generation at both semantics and instance level. We achieve this by modeling class-level conditional modulation parameters as continuous probability distributions instead of discrete values, and sampling per-instance modulation parameters through instance-adaptive stochastic sampling that is consistent across the network. Moreover, we propose prior noise remapping, through linear perturbation parameters encoded from paired references, to facilitate supervised training and exemplar-based instance style control at test time. To further extend the user interaction function of the proposed method, we also introduce sketches into the network. In addition, specially designed generator modules, Progressive Growing Module and Multi-Scale Refinement Module, can be used as a general module to improve the performance of complex scene generation. Extensive experiments on multiple datasets show that our method can achieve superior diversity and comparable quality compared to state-of-the-art methods. Codes are available at https://github.com/tzt101/INADE.git.
Zhentao Tan, Qi Chu 0001, Menglei Chai, Dongdong Chen 0001, Jing Liao 0001, Qiankun Liu 0001, Bin Liu 0016, Gang Hua 0001, Nenghai Yu
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Reduce Information Loss in Transformers for Pluralistic Image Inpainting
abstract
Transformers have achieved great success in pluralistic image inpainting recently. However, we find existing transformer based solutions regard each pixel as a token, thus suffer from information loss issue from two aspects: 1) They downsample the input image into much lower resolutions for efficiency consideration, incurring information loss and extra misalignment for the boundaries of masked regions. 2) They quantize 2563RGB pixels to a small number (such as 512) of quantized pixels. The indices of quantized pixels are used as tokens for the inputs and prediction targets of transformer. Although an extra CNN network is used to upsample and refine the low-resolution results, it is difficult to retrieve the lost information back. To keep input information as much as possible, we propose a new transformer based framework “PUT”. Specifically, to avoid input downsampling while maintaining the computation efficiency, we design a patch-based auto-encoder P-VQVAE, where the encoder converts the masked image into non-overlapped patch tokens and the decoder recovers the masked regions from the inpainted tokens while keeping the unmasked regions unchanged. To eliminate the information loss caused by quantization, an Un-Quantized Transformer (UQ-Transformer) is applied, which directly takes the features from P-VQVAE encoder as input without quantization and regards the quantized tokens only as prediction targets. Extensive experiments show that PUT greatly outperforms state-of-the-art methods on image fidelity, especially for large masked regions and complex large-scale datasets.
Qiankun Liu 0001, Zhentao Tan, Dongdong Chen 0001, Qi Chu 0001, Xiyang Dai, Yinpeng Chen, Mengchen Liu, Lu Yuan 0001, Nenghai Yu
CVPR1
2022 UIA-ViT: Unsupervised Inconsistency-Aware Method Based on Vision Transformer for Face Forgery Detection
Wanyi Zhuang, Qi Chu 0001, Zhentao Tan, Qiankun Liu 0001, Changtao Miao, Zixiang Luo, Nenghai Yu
ECCV (5)4
2022 Online multi-object tracking with unsupervised re-identification learning and occlusion estimation
Qiankun Liu 0001, Dongdong Chen 0001, Qi Chu 0001, Lu Yuan 0001, Bin Liu 0016, Lei Zhang 0001, Nenghai Yu
Neurocomputing1
2020 GSM: Graph Similarity Model for Multi-Object Tracking
abstract
The popular tracking-by-detection paradigm for multi-object tracking (MOT) focuses on solving data association problem, of which a robust similarity model lies in the heart. Most previous works make effort to improve feature representation for individual object while leaving the relations among objects less explored, which may be problematic in some complex scenarios. In this paper, we focus on leveraging the relations among objects to improve robustness of the similarity model. To this end, we propose a novel graph representation that takes both the feature of individual object and the relations among objects into consideration. Besides, a graph matching module is specially designed for the proposed graph representation to alleviate the impact of unreliable relations. With the help of the graph representation and the graph matching module, the proposed graph similarity model, named GSM, is more robust to the occlusion and the targets sharing similar appearance. We conduct extensive experiments on challenging MOT benchmarks and the experimental results demonstrate the effectiveness of the proposed method.
Qiankun Liu 0001, Qi Chu 0001, Bin Liu 0016, Nenghai Yu
IJCAI1
2019 Enhanced Video Segmentation with Object Tracking
Zheran Hong, Zhentao Tan, Qiankun Liu 0001, Bin Liu 0016, Nenghai Yu
ICIG (1)4
2018 Object-Oriented Anomaly Detection in Surveillance Videos
abstract
Detecting and localizing anomalies in surveillance videos is an ongoing challenge. Most existing methods are patch or trajectory-based, which lack semantic understanding of scenes and may split targets into pieces. To handle this problem, this paper proposes a novel and effective algorithm by incorporating deep object detection and tracking with full utilization of spatial and temporal information. We propose a new dynamic image by fusing both appearance and motion information and feed it into object detection network, which can detect and classify objects precisely even in dim and crowd scenes. Based on the detected objects, we develop an effective and scale-insensitive feature, named histogram variance of optical flow angle (HVOFA), together with motion energy to find abnormal motion patterns. In order to further discover missing anomalies and reduce false detected ones, we conduct a post-processing step with abnormal object tracking. The proposed algorithm outperforms state-of-the-art methods on standard benchmarks.
Weihai Li, Bin Liu 0016, Qiankun Liu 0001, Nenghai Yu
ICASSP4
2017 TCCF: Tracking Based on Convolutional Neural Network and Correlation Filters
Qiankun Liu 0001, Bin Liu 0016, Nenghai Yu
ICIG (1)1