VLDB 2026 Research / reviewers in the wild / expert
Yu Liu 0008
dblp:97/2274-8
· DBLP profile ↗
36ranked-venue papers
2as first author
12since 2021 · last 2025
0000-0002-3914-1252ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | NTR-Gaussian: Nighttime Dynamic Thermal Reconstruction with 4D Gaussian Splatting Based on ThermodynamicsabstractThermal infrared imaging enables a non-invasive measurement of the surface temperature of objects with all-weather applicability. Leveraging such techniques for 3D reconstruction can accurately reflect the temperature distribution of a scene, thereby supporting applications such as building monitoring and energy management. However, existing approaches predominantly focus on static 3D reconstruction for a single time period, overlooking the dynamic nature of thermal radiation phenomena, and failing to predict or analyze temperature variations over time. In this paper, we introduce a novel method, termed NTR-Gaussian, grounded in thermodynamics to address the challenge of nighttime dynamic thermal reconstruction using 4D Gaussian Splatting. Specifically, We utilize neural networks to predict thermodynamic parameters, such as emissivity, convective heat transfer coefficient, and heat capacity, etc. By means of integration, we numerically solve the infrared temperature of the scene at each moment during the night, so as to predict the temperature of the nighttime scene more accurately. To further advance research in this domain, we release a comprehensive dataset of dynamic thermal reconstruction spanning four distinct regions. Extensive experiments demonstrate that NTR-Gaussian significantly outperforms comparison methods in thermal reconstruction, achieving a predicted temperature error within 1 degree Celsius. The code is available at https://github.com/NPUCVPG/NTR-Gaussian. Zeyu Cui, Yu Liu 0008, Maojun Zhang, Shen Yan 0002 |
CVPR | 4 |
| 2025 | LoD-Loc v2: Aerial Visual Localization Over Low Level-of-Detail City Models using Explicit Silhouette AlignmentabstractWe propose a novel method for aerial visual localization over low Level-of-Detail (LoD) city models. Previous wireframe-alignment-based method LoD-Loc has shown promising localization results leveraging LoD models. However, LoD-Loc mainly relies on high-LoD (LoD3 or LoD2) city models, but the majority of available models and those many countries plan to construct nationwide are low-LoD (LoD1). Consequently, enabling localization on low-LoD city models could unlock drones' potential for global urban localization. To address these issues, we introduce LoD-Loc v2, which employs a coarse-to-fine strategy using explicit silhouette alignment to achieve accurate localization over low-LoD city models in the air. Specifically, given a query image, LoD-Loc v2 first applies a building segmentation network to shape building silhouettes. Then, in the coarse pose selection stage, we construct a pose cost volume by uniformly sampling pose hypotheses around a prior pose to represent the pose probability distribution. Each cost of the volume measures the degree of alignment between the projected and predicted silhouettes. We select the pose with maximum value as the coarse pose. In the fine pose estimation stage, a particle filtering method incorporating a multi-beam tracking approach is used to efficiently explore the hypothesis space and obtain the final pose estimation. To further facilitate research in this field, we release two datasets with LoD1 city models covering 10.7 km , along with real RGB queries and ground-truth pose annotations. Experimental results show that LoD-Loc v2 improves estimation accuracy with high-LoD models and enables localization with low-LoD models for the first time. Moreover, it outperforms state-of-the-art baselines by large margins, even surpassing texture-model-based methods, and broadens the convergence basin to accommodate larger prior errors. Juelin Zhu, Shuaibang Peng, Hanlin Tan, Yu Liu 0008, Maojun Zhang, Shen Yan 0002 |
ICCV | 5 |
| 2024 | LoD-Loc: Aerial Visual Localization using LoD 3D Map with Neural Wireframe AlignmentabstractWe propose a new method named LoD-Loc for visual localization in the air. Unlike existing localization algorithms, LoD-Loc does not rely on complex 3D representations and can estimate the pose of an Unmanned Aerial Vehicle (UAV) using a Level-of-Detail (LoD) 3D map. LoD-Loc mainly achieves this goal by aligning the wireframe derived from the LoD projected model with that predicted by the neural network. Specifically, given a coarse pose provided by the UAV sensor, LoD-Loc hierarchically builds a cost volume for uniformly sampled pose hypotheses to describe pose probability distribution and select a pose with maximum probability. Each cost within this volume measures the degree of line alignment between projected and predicted wireframes. LoD-Loc also devises a 6-DoF pose optimization algorithm to refine the previous result with a differentiable Gaussian-Newton method. As no public dataset exists for the studied problem, we collect two datasets with map levels of LoD3.0 and LoD2.0, along with real RGB queries and ground-truth pose annotations. We benchmark our method and demonstrate that LoD-Loc achieves excellent performance, even surpassing current state-of-the-art methods that use textured 3D models for localization. The code and dataset will be made available upon publication. Juelin Zhu, Shen Yan 0002, Shengyue Zhang, Yu Liu 0008, Maojun Zhang |
NeurIPS | 5 |
| 2024 | A survey on weakly supervised 3D point cloud semantic segmentationabstractAbstract With the popularity and advancement of 3D point cloud data acquisition technologies and sensors, research into 3D point clouds has made considerable strides based on deep learning. The semantic segmentation of point clouds, a crucial step in comprehending 3D scenes, has drawn much attention. The accuracy and effectiveness of fully supervised semantic segmentation tasks have greatly improved with the increase in the number of accessible datasets. However, these achievements rely on time‐consuming and expensive full labelling. In solve of these existential issues, research on weakly supervised learning has recently exploded. These methods train neural networks to tackle 3D semantic segmentation tasks with fewer point labels. In addition to providing a thorough overview of the history and current state of the art in weakly supervised semantic segmentation of 3D point clouds, a detailed description of the most widely used data acquisition sensors, a list of publicly accessible benchmark datasets, and a look ahead to potential future development directions is provided. Yu Liu 0008, Hanlin Tan, Maojun Zhang |
IET Comput. Vis. | 2 |
| 2024 | Noise2Variance: Dual networks with variance constraint for self-supervised real-world image denoisingabstractAbstract Image denoising aims to restore a clean image from a noisy image. Traditional methods utilizing convolutional neural networks (CNN) for denoising are trained using pairs of noisy and clean images to comprehend the transformation from a noisy image to a clean one. However, the acquisition of such image pairs in real‐world scenarios presents a challenge. Hence, numerous self‐supervised denoising techniques have been developed that do not require clean images for training. This study demonstrates that a straightforward loss design, concentrating on variance, can effectively train a standard CNN denoiser in a self‐supervised fashion. A novel theoretical framework is introduced for training a basic CNN denoising model using three constraints: mean, variance, and augmentation. The variance constraint is crucial as it prevents the trained model from converging to trivial solutions such as identity or zero mapping. This theory provides valuable insights for the development of new self‐supervised denoising methods. Furthermore, a method that applies this theory to proposed dual networks is developed, which consist of two standard CNN models predicting both the clean image and the noise. This approach enhances model capacity during training while minimizing computational costs during inference. This method exemplifies the implementation of the variance constraint and introduces a data constraint for dual networks. Notably, the proposed method only assumes the presence of additive white noise, irrespective of the noise distribution. This minimal assumption enhances the model's robustness against noise with complex or unknown distributions in real‐world distorted images. Experimental results indicate that the proposed Noise2Variance method exhibits commendable performance on peak signal noise ratio and structural similarity metrics compared to existing self‐supervised denoising techniques. Visual comparison of results further substantiates the efficacy of the proposed method. A comparison of model complexity reveals that the method is efficient among the compared CNN‐based techniques. Hanlin Tan, Yu Liu 0008, Maojun Zhang |
IET Image Process. | 2 |
| 2024 | Target Detection With Spectral Graph Contrast Clustering Assignment and Spectral Graph Transformer in Hyperspectral ImageryabstractHyperspectral target detection (HTD) is a method that recognizes objects of interest in a scene by a priori target spectrum. Local details and global information on the spectra are critical for accurate target identification. Detectors with excellent discrimination of spectral differences can better highlight targets while suppressing background. To this end, this article proposes an HTD method based on spectral graph contrast clustering assignment and the spectral graph transformer (SGT) to solve these problems. Specifically, for local-global feature extraction of spectra, the pixel spectra are first constructed as the spectral graph. Then, the representations of the first- or higher-order neighbors of the nodes in the spectral graph are aggregated using graph convolutional networks to extract the local detail information of the spectra. The self-attention in Transformer is utilized to learn the global information of the spectra. Second, a novel spectral graph contrast clustering assignment method is proposed to equip the model with excellent spectral discrimination ability. It maintains clustering consistency by swapping predictive clustering assignments while maximizing the similarity of semantically similar graph clusters and keeping other semantically different graph clusters away from them to better discriminate differences between spectra. Finally, comparisons with seven state-of-the-art HTD methods on four real hyperspectral datasets and ablation studies verify the effectiveness of the proposed method in HTD. Xi Chen 0077, Maojun Zhang, Yu Liu 0008 |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | ATLoc: Aerial Thermal Images Localization via View SynthesisabstractWhile visual localization has made significant advances in recent years, it still lacks robustness in low-light situations. Thermal camera images, which capture temperature data, provide a potential solution for these environments. However, the scarcity of well-annotated, publicly available datasets for thermal localization, particularly those focused on absolute pose estimation, impedes further advancement in this field. In this research, we introduce a novel dataset that includes six-degree-of-freedom (6-DoF) absolute poses of query images for large-scale, realistic aerial localization of thermal images. Besides, we introduce a render-to-localization pipeline tailored for thermal image localization. This pipeline predicts the 6-DoF pose of a query using a synthetic technique based on geometric refined thermal model. Experimental results demonstrate the effectiveness of our method on this newly proposed dataset. Notably, our method achieves a median position error of less than 1.5 m and a median angle error of less than 1.5° under diverse test conditions. A comprehensive analysis of factors influencing localization accuracy is also provided. Our code and dataset will be available athttps://github.com/RingoWRW/ATLoc. Rouwan Wu, Shen Yan 0002, Xiaoya Cheng, Juelin Zhu, Yu Liu 0008, Maojun Zhang |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | Long-Term Visual Localization with Mobile SensorsabstractDespite the remarkable advances in image matching and pose estimation, image-based localization of a camera in a temporally-varying outdoor environment is still a challenging problem due to huge appearance disparity between query and reference images caused by illumination, seasonal and structural changes. In this work, we propose to leverage additional sensors on a mobile phone, mainly GPS, compass, and gravity sensor, to solve this challenging problem. We show that these mobile sensors provide decent initial poses and effective constraints to reduce the searching space in image matching and final pose estimation. With the initial pose, we are also able to devise a direct 2D-3D matching network to efficiently establish 2D-3D correspondences instead of tedious 2D-2D matching in existing systems. As no public dataset exists for the studied problem, we collect a new dataset that provides a variety of mobile sensor data and significant scene appearance variations, and develop a system to acquire ground-truth poses for query images. We benchmark our method as well as several state-of-the-art baselines and demonstrate the effectiveness of the proposed approach. Our code and dataset are available on the project page: https://zju3dv.github.io/sensloc/ Shen Yan 0002, Yu Liu 0008, Zehong Shen, Haomin Liu, Maojun Zhang, Guofeng Zhang 0001, Xiaowei Zhou 0001 |
CVPR | 2 |
| 2023 | Deep Active Contours for Real-time 6-DoF Object TrackingabstractThis paper solves the problem of real-time 6-DoF object tracking from an RGB video. Prior optimization-based methods optimize the object pose by aligning the projected model to the image based on handcrafted features, which are prone to suboptimal solutions. Recent learning-based methods use neural networks to predict the pose, which suffer from limited generalizability or computational efficiency. We propose a learning-based active contour model to make the best use of both worlds. Specifically, given an initial pose, we project the object model to the image plane to obtain the initial contour and use a lightweight network to predict how the contour should move to match the true object boundary, which provides the gradients to optimize the object pose. We also devise an efficient optimization algorithm to train our model end-to-end with pose supervision. Experimental results on semi-synthetic and real-world 6-DoF object tracking datasets demonstrate that our model outperforms state-of-the-art methods by a substantial margin in pose accuracy, while achieving real-time performance on mobile devices. Code is available on our project page: https://zju3dv.github.io/deep_ac/. Shen Yan 0002, Jianan Zhen, Yu Liu 0008, Maojun Zhang, Guofeng Zhang 0001, Xiaowei Zhou 0001 |
ICCV | 4 |
| 2023 | Render-and-Compare: Cross-view 6-DoF Localization from Noisy PriorabstractDespite the significant progress in 6-DoF visual localization, researchers are mostly driven by ground-level benchmarks. Compared with aerial oblique photography capture, ground-level map collection lacks scalability and complete coverage. In this work, we propose to go beyond the traditional ground-level setting and exploit cross-view 6-DoF localization from aerial to ground. We address this problem by formulating camera pose estimation as an iterative render-and-compare pipeline and enhancing the algorithm robustness through augmenting seeds from noisy initial priors. As no public dataset exists for the studied problem, we have collected a new dataset that provides a variety of cross-view images from smartphones and low-altitude drones and developed a semi-automatic system to acquire ground-truth poses for query images. We benchmark our method as well as several state-of-the-art baselines and demonstrate that our method outperforms other approaches by a large margin. Code is available at https://github.com/Choyaa/Render2Loc. Shen Yan 0002, Xiaoya Cheng, Juelin Zhu, Rouwan Wu, Yu Liu 0008, Maojun Zhang |
ICME | 6 |
| 2022 | View graph construction for scenes with duplicate structures via graph convolutional networkabstractAbstract View graph construction aims to effectively organise disordered image dataset through image retrieval technique before structure from motion (SfM). Existing view graph construction methods usually fail to handle scenes with duplicate structure, because these methods solely treat the construction of view graph as a process of image‐pair‐wise matching and lack in exploiting images' topological details in dataset. In this paper, we handle this problem from a novel perspective to construct view graph in a global paradigm by introducing an end‐to‐end graph convolutional network (GCN). First, a location‐aware embedding module is introduced to encode images into a feature space that takes into account the feature's location by using Vision Transformer architecture, improving the distinction between features of duplicate structure. Second, graph convolutional network that consists of topological relationship preserving module and feature metric learning module is proposed. Topological relationship preserving network is proposed to help nodes maintain their connected neighbourhood features. By merging the topological connected information into images' embedding, our method can process image matching in a global mode, thus improving the disambiguation ability for images with duplicate scenes. Then a feature metric learning network is embedded into GCN to dynamically compute the linkage prediction among nodes based on their features. Finally, our method combines these three parts to jointly optimise nodes' features and linkage prediction in an end‐to‐end paradigm. We make qualitative and quantitative comparisons based on three public benchmark datasets and demonstrate that our proposed method performs favourably against other state‐of‐the‐art methods. Shen Yan 0002, Yu Liu 0008, Maojun Zhang |
IET Comput. Vis. | 4 |
| 2021 | Image retrieval for Structure-from-Motion via Graph Convolutional NetworkabstractConventional image retrieval techniques for Structure-from-Motion (SfM) are limited in their ability to effectively distinguish symmetric or repetitive textured patterns and cannot guarantee an accurate generation of pairwise matches without costly redundancy. In this paper, we formulate the image retrieval task as a node binary classification problem with graph data: if a candidate node is marked as positive, it is believed to share the same scene with the query image. The key idea of our approach is that the local context in the feature space around a query image contains abundant information about the matchable relation between the image and its neighbours. By constructing a subgraph surrounding the query image as input data, we adopt a learnable Graph Convolutional Network (GCN) to determine whether nodes in the subgraph have overlapping regions with the query photograph. Experiments demonstrate that our method performs remarkably well on a challenging dataset of highly ambiguous and duplicated scenes. Furthermore, compared with state-of-the-art matchable retrieval methods , the proposed approach significantly reduces unnecessary attempted matches without sacrificing the accuracy and completeness of reconstruction. Shen Yan 0002, Maojun Zhang, Shiming Lai, Yu Liu 0008 |
Inf. Sci. | 4 |
| 2020 | Denoising real bursts with squeeze-and-excitation residual networkabstractThe goal of image denoising is to recover a clean image from noisy input(s). For single image denoising, utilising similarities (or priors) within and across an image dataset helps recover clean images. As the noise level increases, using multiple frames become feasible, which is defined as burst denoising. In this study, the authors propose a deep residual model with squeeze‐and‐excitation (SE) modules for the burst denoising. Unlike previous methods, the authors' model does not need an explicit aligning procedure, which is light‐weighted and fast. The network contains a noise estimation convolutional neural network, which makes it capable of blind denoising. Besides, by inverting the image processing pipeline and simulating real noise in bursts, their model can suppress real noise blindly. Since denoising performance is closely related to the noise level, frame displacement, and the number of frames (burst length), intensive experiments including ablation study are performed. Quantitative results show that the proposed method performs significantly better than previous state‐of‐the‐art methods V‐BM4D and KPN in removing Gaussian noise. Qualitative results show that the proposed method is also effective in removing real noise using bursts and the SE module is key to reduce blur in results. Hanlin Tan, Huaxin Xiao, Shiming Lai, Yu Liu 0008, Maojun Zhang |
IET Image Process. | 4 |
| 2020 | Online Meta Adaptation for Fast Video Object SegmentationabstractConventional deep neural networks based video object segmentation (VOS) methods are dominated by heavily fine-tuning a segmentation model on the first frame of a given video, which is time-consuming and inefficient. In this paper, we propose a novel method which rapidly adapts a base segmentation model to new video sequences with only a couple of model-update iterations, without sacrificing performance. Such attractive efficiency benefits from the meta-learning paradigm which leads to a meta-segmentation model and a novel continuous learning approach which enables online adaptation of the segmentation model. Concretely, we train a meta-learner on multiple VOS tasks such that the meta model can capture their common knowledge and gains the ability to fast adapt the segmentation model to new video sequences. Furthermore, to deal with unique challenges of VOS tasks from temporal variations in the video, e.g., object motion and appearance changes, we propose a principled online adaptation approach that continuously adapts the segmentation model across video frames by exploiting temporal context effectively, providing robustness to annoying temporal variations. Integrating the meta-learner with the online adaptation approach, the proposed VOS model achieves competitive performance against the state-of-the-arts and moreover provides faster per-frame processing speed. Huaxin Xiao, Bingyi Kang, Yu Liu 0008, Maojun Zhang, Jiashi Feng |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Transferable Semi-Supervised Semantic SegmentationabstractThe performance of deep learning based semantic segmentation models heavily depends on sufficient data with careful annotations. However, even the largest public datasets only provide samples with pixel-level annotations for rather limited semantic categories. Such data scarcity critically limits scalability and applicability of semantic segmentation models in real applications. In this paper, we propose a novel transferable semi-supervised semantic segmentation model that can transfer the learned segmentation knowledge from a few strong categories with pixel-level annotations to unseen weak categories with only image-level annotations, significantly broadening the applicable territory of deep segmentation models. In particular, the proposed model consists of two complementary and learnable components: a Label transfer Network (L-Net) and a Prediction transfer Network (P-Net). The L-Net learns to transfer the segmentation knowledge from strong categories to the images in the weak categories and produces coarse pixel-level semantic maps, by effectively exploiting the similar appearance shared across categories. Meanwhile, the P-Net tailors the transferred knowledge through a carefully designed adversarial learning strategy and produces refined segmentation results with better details. Integrating the L-Net and P-Net achieves 96.5% and 89.4% performance of the fully-supervised baseline using 50% and 0% categories with pixel-level annotations respectively on PASCAL VOC 2012. With such a novel transfer mechanism, our proposed model is easily generalizable to a variety of new categories, only requiring image-level annotations, and offers appealing scalability in real applications. Huaxin Xiao, Yunchao Wei, Yu Liu 0008, Maojun Zhang, Jiashi Feng |
AAAI | 3 |
| 2018 | MoNet: Deep Motion Exploitation for Video Object SegmentationabstractIn this paper, we propose a novel MoNet model to deeply exploit motion cues for boosting video object segmentation performance from two aspects, i.e., frame representation learning and segmentation refinement. Concretely, MoNet exploits computed motion cue (i.e., optical flow) to reinforce the representation of the target frame by aligning and integrating representations from its neighbors. The new representation provides valuable temporal contexts for segmentation and improves robustness to various common contaminating factors, e.g., motion blur, appearance variation and deformation of video objects. Moreover, MoNet exploits motion inconsistency and transforms such motion cue into foreground/background prior to eliminate distraction from confusing instances and noisy regions. By introducing a distance transform layer, MoNet can effectively separate motion-inconstant instances/regions and thoroughly refine segmentation results. Integrating the proposed two motion exploitation components with a standard segmentation network, MoNet provides new state-of-the-art performance on three competitive benchmark datasets. Huaxin Xiao, Jiashi Feng, Guosheng Lin, Yu Liu 0008, Maojun Zhang |
CVPR | 4 |
| 2018 | No-Reference Image Sharpness Assessment Using Scale and Directional ModelsabstractWe propose a new method of no-reference (NR) image sharpness assessment. Our method is based on a multiscale decomposition of non-overlapping blocks, and on analyzing the statistics of Local-Mean Magnitude (LMM) maps. With the detected high-activity blocks, a set of inter-scale and inter-direction sharpness ratios are found. These sharpness ratios are closely correlated with the level of blurring, and is capable of measuring image sharpness effectively from a multiscale and a directional view. They are linearly combined using automatic estimated weighting parameters to induce an overall effective sharpness model. To enhance accuracy, a sharpness factor based on the blur effect on DCT edges and a log-energy sharpness model are also incorporated into the method. Experiments on a large number of public images have shown that our method consistently produces predictions that are highly correlated with human perceptual of image sharpness, and outperforms the current state-of-the-art algorithms. Zheng Zhang 0011, Yu Liu 0008, Hanlin Tan, Xiaoqing Yi, Maojun Zhang |
ICME | 2 |
| 2018 | Salient object detection via robust dictionary representation
Huaxin Xiao, Weiya Ren, Wei Wang 0108, Yu Liu 0008, Maojun Zhang |
Multim. Tools Appl. | 4 |
| 2018 | Focus and Blurriness Measure Using Reorganized DCT Coefficients for an Autofocus ApplicationabstractIn this paper, two metrics for measuring image sharpness are presented and used for an autofocus (AF) application. Both measures exploit reorganized discrete cosine transform (DCT) representation. The first metric is a focus measure, which involves optimal high- and middle-frequency coefficients to evaluate relative sharpness. It is robust to noise while remaining sensitive to the best focus position. A psychometric function-based metric is introduced to quantify the focus measure. The second metric is a no-reference blurriness metric, which is used to measure absolute blurriness. It first constructs multiscale DCT edge maps using directional energy information and then determines image blurriness by combining change information in edge structures with image contrast. This metric gives predictions that are closely correlated with subjective perceived scores and shows performance comparable with that of state-of-the-art methods, especially for noisy images. For noisy situations, the two metrics are adjusted adaptively according to the estimated noise level. To prevent the introduction of extra computational load, an efficient noise-level estimation algorithm based on median absolute deviation is presented. This algorithm exploits only the available reorganized DCT coefficients. With the focus and blurriness measures, an AF method for which the two metrics play an important role was developed. Because of their high-quality performance, the realized AF function is able to locate the best focus position swiftly and reliably. Zheng Zhang 0011, Yu Liu 0008, Zhihui Xiong, Jing Li 0014, Maojun Zhang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2017 | Joint demosaicing and denoising of noisy bayer images with ADMMabstractImage demosaicing and denoising are import steps of image signal processing. Sequential executions of demosaicing and denoising have essential drawbacks that they degrade the results of each other. Joint demosaicing and denoising overcomes the difficulties by solving the two problems in one model. This paper introduces a unified object function with hidden priors and a variant of ADMM to recover a full-resolution color image with a noisy Bayer input. Experimental results demonstrate that our method performs better than state-of-the-art methods in both PSNR comparison and human vision. In addition, our method is much more robust to variations of noise level. Hanlin Tan, Xiangrong Zeng, Shiming Lai, Yu Liu 0008, Maojun Zhang |
ICIP | 4 |
| 2017 | A Simplified Low Rank and Sparse Model for Visual Tracking
Mi Wang, Huaxin Xiao, Yu Liu 0008, Wei Xu 0019, Maojun Zhang |
ICPRAM | 3 |
| 2017 | Blur kernel estimation using normal sinh-arcsinh model based on simple lens systemabstractFocused images captured by the lens system suffer image degradation due to factors, such as aberration, caused by the optical structure. In the simple lens system, aberration is more severe because of the simplification of the imaging system. In the existing imaging model, the blur kernel of the image is usually described by the point spread function. A few studies have shown that the blur kernel of the simple lens is close to the spatially deformed wedge. In the present study, a more reasonable normal sinh-arcsinh model is used to fit the blur kernel, and the parametric blur kernel is obtained by Powell algorithm. Finally, the clear image is restored and compared with other models to prove the advantages of our method. Dazhi Zhan, Xiangrong Zeng, Yu Liu 0008, Zhihui Xiong |
MMSP | 4 |
| 2016 | Fast anomaly detection in traffic surveillance video based on robust sparse optical flowabstractFast abnormal events detection in video is important for intelligent analysis of video. This paper proposes a fast anomaly detection algorithm based on sparse optical flow. We improve the efficiency of optical flow computation with foreground mask and spacial sampling and increase the robustness of optical flow with good feature (TK) points selecting and forward-backward filtering. A foreground channel is also added to the feature vector to help detect static or low speed objects. The algorithm is validated on real-life traffic surveillance to prove its effectiveness. It is also evaluated on a benchmark dataset and achieve detection results comparable to state-of-art methods and outperforms them at pixel-level when the false alarm rate is low. The strength of our algorithm is that it runs real-time on the benchmark dataset which is hundreds of times faster than comparative methods. Hanlin Tan, Yongping Zhai, Yu Liu 0008, Maojun Zhang |
ICASSP | 3 |
| 2016 | Foreground Segmentation for Moving Cameras under Low Illumination ConditionsabstractA foreground segmentation method, including image enhancement, trajectory classification and object segmentation,
is proposed for moving cameras under low illumination conditions. Gradient-field-based image
enhancement is designed to enhance low-contrast images. On the basis of the dense point trajectories obtained
in long frames sequences, a simple and effective clustering algorithm is designed to classify foreground
and background trajectories. By combining trajectory points and a marker-controlled watershed algorithm,
a new type of foreground labeling algorithm is proposed to effectively reduce computing costs and improve
edge-preserving performance. Experimental results demonstrate the promising performance of the proposed
approach compared with other competing methods. Wei Wang 0068, Xiaoqing Yin, Yu Liu 0008, Maojun Zhang |
ICPRAM | 4 |
| 2016 | Data Based Color ConstancyabstractColor constancy is an important task in computer vision. By analyzing the image formation model, color gamut data under one light source can be mapped to a hyperplane whose normal vector is only determined by its light source. Thus, the canonical light source is represented through the kernel method, which trains the color data. When an image is captured under an unknown illuminant, the image-corrected matrix is obtained through optimization. After being mapped to the high-dimensional space, the corrected color data are best fit for the hyperplane of the canonical illuminant. The proposed unsupervised feature-mining kernel method only depends on the color data without any other information. The experiments on the standard test datasets show that the proposed method achieves comparable performance with other state-of-the-art methods. Wei Xu 0019, Huaxin Xiao, Yu Liu 0008, Maojun Zhang |
ICPRAM | 3 |
| 2016 | Image Denoising Using Quadtree-Based Nonlocal Means With Locally Adaptive Principal Component AnalysisabstractIn this letter, we present an efficient image denoising method combining quadtree-based nonlocal means (NLM) and locally adaptive principal component analysis. It exploits nonlocal multiscale self-similarity better, by creating sub-patches of different sizes using quadtree decomposition on each patch. To achieve spatially uniform denoising, we propose a local noise variance estimator combined with denoiser based on locally adaptive principal component analysis. Experimental results demonstrate that our proposed method achieves very competitive denoising performance compared with state-of-the-art denoising methods, even obtaining better visual perception at high noise levels. Chenglin Zuo, Ljubomir Jovanov, Bart Goossens, Hiêp Quang Luong, Wilfried Philips, Yu Liu 0008, Maojun Zhang |
IEEE Signal Process. Lett. | 6 |
| 2015 | Data Separation of L1-minimization for Real-time Motion Detection
Yu Liu 0008, Huaxin Xiao, Zheng Zhang 0011, Wei Xu 0019, Maojun Zhang, Jianguo Zhang 0001 |
BMVC | 1 |
| 2015 | A robust motion detection algorithm on noisy videosabstractThe applicability and performance of motion detection methods dramatically degrade with the increasing noise. In this paper, we propose a robust dictionary-based background subtraction approach, which formulates background modeling as a linear and sparse combination of atoms in a pre-learned dictionary. Motion detection is then implemented to compare the difference between sparse representations of the current frame and the background model. The projection of noise over the dictionary being irregular and random guarantees the adaptability of our approach. Experimental results on synthetic and real noisy videos demonstrate the robustness of the proposed approach compared to other methods. Yu Liu 0008, Huaxin Xiao, Wei Wang 0068, Maojun Zhang |
ICASSP | 1 |
| 2015 | Rotation invariant similarity measure for non-local self-similarity based image denoisingabstractNon-local self-similarity based image denoising depends strictly on similarity measure. The denoising performance is determined based on the ability to reliably find sufficient number of similar patches. In this paper, we propose a rotation invariant similarity measure to fully exploit the image non-local self-similarity. Instead of using image patches, we employ local frequency descriptors, that are rotation invariant and robust to noise, to measure the similarity. Thus, both translational and rotational similarity can be handled even at high noise level. The comparative experimental results show that the proposed method is effective as a rotation invariant similarity measure, and it can consistently improve the performance of non-local means algorithm to achieve better denoising results. Chenglin Zuo, Ljubomir Jovanov, Hiêp Quang Luong, Bart Goossens, Wilfried Philips, Yu Liu 0008, Maojun Zhang |
ICIP | 6 |
| 2014 | Robust Sharpness Metrics Using Reorganized DCT Coefficients for Auto-Focus Application
Zheng Zhang 0011, Yu Liu 0008, Maojun Zhang |
ACCV (4) | 2 |
| 2014 | A Low Illumination Environment Motion Detection Method based on Dictionary LearningabstractThis paper proposes a dictionary-based motion detection method on video images captured under low light with serious noise. The proposed approach trains a dictionary by background images without foreground. It then reconstructs the test image according to the theory of sparse coding, and introduces the Structural Similarity Index Measurement (SSIM) as the detection standard to identify the detection caused by the brightness and contrast ratio changes. Experimental results show that compared to the mixture of Gaussian model and frame difference method, the proposed method can reach a better result under extreme low illumination circumstance. Huaxin Xiao, Yu Liu 0008, Bin Wang 0043, Shuren Tan, Maojun Zhang |
ICPRAM | 2 |
| 2014 | Omni-gradient-based total variation minimisation for sparse reconstruction of omni-directional imageabstractTotal variation (TV) minimisation algorithms have been successfully applied in compressive sensing (CS) recovery for natural images owing to its advantage of preserving edges. However, traditional TV is no longer appropriate for omni‐directional image processing because of the distortions in catadioptric imaging systems. The omni‐gradient computing method combined with the characteristics of omni‐directional imaging is proposed in this study. To reconstruct the image from its compressive samples, the omni‐total variation (omni‐TV) regularisation based on omni‐gradient is utilised instead of traditional TV during the image restoration. The experimental results show that the omni‐directional images can be reconstructed effectively and accurately. Compared with the classical TV minimisation model, the images recovered based on omni‐TV model can provide higher quality both in subjective evaluation and objective evaluation. Jingtao Lou, Yongle Li, Yu Liu 0008, Shuren Tan, Maojun Zhang |
IET Image Process. | 3 |
| 2013 | Human action recognition with Optimized Video Densely SamplingabstractDense sample video patches have been used for video representation in action recognition and achieve better performance than sparse spatiotemporal local features. However, two problems of this method must be considered. First one, many video patches are from background other than human body. Second one, the descriptor is not reliable, since it is neither shift nor scale invariant. To solve these two problems, we proposed an Optimized Video Dense Sampling (OVDS) method combing with dense sampling and spatiotemporal interest points detector. OVDS densely sampled video patches with optimizing the position and scale parameters to guarantee the features are shift and scale invariant. To omit the action unrelated features, we extracted video patches only from human body regions instead of the whole videos. Experimental results on KTH, Weizmann, UCF, Hoollywood2 datasets showed that the features detected by OVDS are informative and reliable for action recognition, and achieve better performance over the existing spatiotemporal local features. Bin Wang 0043, Yu Liu 0008, Wenhua Xiao, Zhihui Xiong, Wei Wang 0068, Maojun Zhang |
ICME | 2 |
| 2013 | Action recognition using Feature Position Constrained Linear CodingabstractRecently space-time interest points (STIPs) using bag-of-feature (BOF) in action recognition has been highly successful. Despite its popularity, The quantization error and the lost of semantic meaning among STIPs are the main weaknesses that severely limit the effectiveness of this method. To overcome these limitations, this paper incorporated the feature position information into coding procedure and proposed a novel Feature Position Constrained Linear Coding (FPLC) method by extending the Locality Constrained Linear Coding (LLC) approach. It first project the features into the human ROI, then codes the features locally using FPLC with the consideration of feature position. Owning to that the local area of human ROI often aggregate features extracted from the same part of human body and those features should exhibit similar values, this local coding strategy helps to alleviate the quantization error and enhance correlation between features at the same time, which helps to improve the recognition accuracy. Compared with the state-of-the-art action recognition method, experiment results demonstrated the effectiveness of the proposed method. Wenhua Xiao, Bin Wang 0043, Yu Liu 0008, Wei Xu 0019, Wei Wang 0068, Weidong Bao 0001, Maojun Zhang |
ICME | 3 |
| 2012 | Fish-eye distortion correction based on midpoint circle algorithmabstractThis paper presents a novel embedded real-time fisheye image distortion correction algorithm with application in IP network camera. A fast and simple distortion correction method is introduced based on Midpoint Circle Algorithm (MCA) which aims to determine the pixel positions along a circle circumference based on incremental calculation of decision parameters. Although only the vertical is rectilinearised, experimental results show that our correction method based on MCA is efficient and effective. In particular, our method can be applied without considering planar checkerboard, iterative fitting of model parameters, complex computation, or traditional lookup tables. Therefore, our algorithm is suitable for embedded camera platform without any extra hardware resources. Yongle Li, Maojun Zhang, Yu Liu 0008, Zhihui Xiong |
SMC | 3 |
| 2012 | Coded aperture techniques for catadioptric omni-directional image defocus deblurringabstractThe defocus blur in catadioptric omni-directional imaging is caused by large apertures and mirror curvatures. This problem becomes more obviously when introducing high resolution sensors. In order to overcome this drawback, this work proposes a simple modification to a conventional catadioptric system that allows for the recovery of an all-focus omni-directional image. The modification is designed by inserting a patterned occluder within the aperture of the camera lens, creating a coded aperture. Then this work introduces a specific deconvolution method, which can recover an all-focus omni-directional image from photograph(s) taken by the camera with coded aperture. Comparing to the conventional aperture, the coded aperture techniques identify the blur scale easier and more accurate. The recovered sharp image eliminates the defocus blur and shows the efficiency of the algorithm. The obtained sharp image can be combined for various catadioptric applications, including omni-directional monitoring systems, intelligent omni-directional systems and robotics, etc. Yu Liu 0008, Yongle Li, Maojun Zhang |
SMC | 2 |