VLDB 2026 Research / reviewers in the wild / expert
Shaohui Liu
dblp:59/962
· DBLP profile ↗
133ranked-venue papers
12as first author
62since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 94 · 7 first-author · 43 since 2021Artificial intelligence and machine learning · 48 · 4 first-author · 27 since 2021Computer networks · 6 · 5 since 2021Systems, architecture and hardware · 5 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Security and privacy · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 1 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | FDM-Net: A frequency-decoupled network with adaptive masking for time series forecasting
Shaohui Liu, Chunzhi Yi, Baichun Wei, Haiqi Zhu |
Expert Syst. Appl. | 3 |
| 2026 | UGD-IML: A unified generative diffusion-based framework for constrained and unconstrained image manipulation localization
Yachun Mi, Xingyang He, Shixin Sun, Shaohui Liu |
Neurocomputing | 8 |
| 2026 | CLiF-VQA+: Enhancing video quality assessment by incorporating human objective and subjective feelings
Yachun Mi, Shaohui Liu |
Neurocomputing | 5 |
| 2026 | Reconstructing Temporal Heterogeneity: A Multidomain Collaborative Analysis Framework for Robust Time-Series Forecasting
Hengrui Li, Wenxue Cui, Yifeng Wang 0001, Chunshan Dong, Wenju Li, Jiangpeng Shi, Yongbing Zhang 0002, Shaohui Liu |
IEEE Internet Things J. | 10 |
| 2026 | GEOMR: Integrating image geographic features and human reasoning knowledge for image geolocalization
Siyi Qian, Shaohui Liu |
Knowl. Based Syst. | 3 |
| 2025 | Relative Pose Estimation through Affine Corrections of Monocular Depth PriorsabstractMonocular depth estimation (MDE) models have undergone significant advancements over recent years. Many MDE models aim to predict affine-invariant relative depth from monocular images, while recent developments in large-scale training and vision foundation models enable reasonable estimation of metric (absolute) depth. However, effectively leveraging these predictions for geometric vision tasks, in particular relative pose estimation, remains relatively under explored. While depths provide rich constraints for cross-view image alignment, the intrinsic noise and ambiguity from the monocular depth priors present practical challenges to improving upon classic keypoint-based solutions. In this paper, we develop three solvers for relative pose estimation that explicitly account for independent affine (scale and shift) ambiguities, covering both calibrated and uncalibrated conditions. We further propose a hybrid estimation pipeline that combines our proposed solvers with classic point-based solvers and epipolar constraints. We find that the affine correction modeling is beneficial to not only the relative depth priors but also, surprisingly, the "metric" ones. Results across multiple datasets demonstrate large improvements of our approach over classic keypoint-based baselines and PnP-based solutions, under both calibrated and uncalibrated setups. We also show that our method improves consistently with different feature matchers and MDE models, and can further benefit from very recent advances on both modules. Code is available at https://github.com/MarkYu98/madpose. Yifan Yu 0003, Shaohui Liu, Rémi Pautrat, Marc Pollefeys, Viktor Larsson |
CVPR | 2 |
| 2025 | Reloc3r: Large-Scale Training of Relative Camera Pose Regression for Generalizable, Fast, and Accurate Visual LocalizationabstractVisual localization aims to determine the camera pose of a query image relative to a database of posed images. In recent years, deep neural networks that directly regress camera poses have gained popularity due to their fast inference capabilities. However, existing methods struggle to either generalize well to new scenes or provide accurate camera pose estimates. To address these issues, we present Reloc3r, a simple yet effective visual localization framework. It consists of an elegantly designed relative pose regression network, and a minimalist motion averaging module for absolute pose estimation. Trained on approximately eight million posed image pairs, Reloc3r achieves surprisingly good performance and generalization ability. We conduct extensive experiments on six public datasets, consistently demonstrating the effectiveness and efficiency of the proposed method. It provides high-quality camera pose estimates in real time and generalizes to novel scenes. Code: https://github.com/ffrivera0/reloc3r. Siyan Dong, Shuzhe Wang, Shaohui Liu, Lulu Cai, Qingnan Fan, Juho Kannala, Yanchao Yang 0001 |
CVPR | 3 |
| 2025 | Image Compressive Sensing With Adaptive Sampling by Median FilteringabstractDeep unfolding compressive sensing (CS) has experienced remarkable advancements. However, there still exist two challenges: (1) Many algorithms either use uniform block-based sampling, which ignore the fact that the content of different blocks is different, or allocate the sampling rate referring to complete signal before CS sampling, which is not always feasible in real-world scenarios. (2) Traditional CNN is difficult to capture broader contextual priors during iterative recovery. In this paper, we propose a novel network ASMFNet to solve the above two issues. Specifically, to address the first issue, we introduce a dual-branch network featuring a basic sampling branch to acquire reference image and an adaptive sampling branch by median filtering for allocating remaining sampling rate adaptively. For the second problem, we use Swin Transformer and feature fusion block to increase the feature interactions. Experimental results demonstrate that our proposed method outperforms existing methods. Ronghua Liao, Shaohui Liu, Debin Zhao |
ICASSP | 4 |
| 2025 | Benchmarking Egocentric Visual-Inertial SLAM at City Scale
Anusha Krishnan, Shaohui Liu, Paul-Edouard Sarlin, Oscar Gentilhomme, David Caruso, Maurizio Monge, Richard Newcombe, Jakob J. Engel, Marc Pollefeys |
ICCV | 2 |
| 2025 | MVQA: Mamba with Unified Sampling for Efficient Video Quality AssessmentabstractThe rapid growth of long-duration, high-definition videos has made efficient video quality assessment (VQA) a critical challenge. Existing research typically tackles this problem through two main strategies: reducing model parameters and resampling inputs. However, light-weight Convolution Neural Networks (CNN) and Transformers often struggle to balance efficiency with high performance due to the requirement of long-range modeling capabilities. Recently, the state-space model, particularly Mamba, has emerged as a promising alternative, offering linear complexity with respect to sequence length. Meanwhile, efficient VQA heavily depends on resampling long sequences to minimize computational costs, yet current resampling methods are often weak in preserving essential semantic information. In this work, we present MVQA, a Mamba-based model designed for efficient VQA along with a novel Unified Semantic and Distortion Sampling (USDS) approach. USDS combines semantic patch sampling from low-resolution videos and distortion patch sampling from original-resolution videos. The former captures semantically dense regions, while the latter retains critical distortion details. To prevent computation increase from dual inputs, we propose a fusion mechanism using pre-defined masks, enabling a unified sampling strategy that captures both semantic and quality information without additional computational burden. Experiments show that the proposed MVQA, equipped with USDS, achieve comparable performance to state-of-the-art methods while being $2\times$ as fast and requiring only $1/5$ GPU memory. Yachun Mi, Weicheng Meng, Chaofeng Chen, Shaohui Liu |
ICCV | 6 |
| 2025 | MAMF-Net: Modality-Adaptive Masked Fusion Network for Speech Emotion RecognitionabstractThis paper introduces a novel multimodal emotion recognition model, the Modality-Adaptive Masked Fusion Network (MAMF-Net), designed to mitigate information loss and improve cross-modal alignment during the fusion of speech and text modalities. MAMF-Net employs an audio-guided text encoder to enhance the semantic representation of text by leveraging the temporal resolution and contextual information inherent in speech, thereby ensuring accurate alignment of modal features. Additionally, the model utilizes a modality transfer-based MAE masking strategy, which effectively captures complementary information between modalities by partially masking transferred information, thus improving fusion effectiveness and system stability. The experimental results show that MAMF-Net outperforms existing methods on datasets such as CMU-MOSI and CMU-MOSEI, highlighting its significant potential for multimodal emotion analysis. Hengrui Li, Xiaopei Chen, Shaohui Liu |
ICME | 6 |
| 2025 | BPCLIP: A Bottom-up Image Quality Assessment from Distortion to Semantics Based on CLIPabstractImage Quality Assessment (IQA) aims to evaluate the perceptual quality of images based on human subjective perception. Existing methods generally combine multiscale features to achieve high performance, but most rely on straightforward linear fusion of these features, which may not adequately capture the impact of distortions on semantic content. To address this, we propose a bottom-up image quality assessment approach based on the Contrastive Language-Image Pre-training (CLIP, a recently proposed model that aligns images and text in a shared feature space), named BPCLIP, which progressively extracts the impact of low-level distortions on high-level semantics. Specifically, we utilize an encoder to extract multiscale features from the input image and introduce a bottom-up multiscale cross attention module designed to capture the relationships between shallow and deep features. In addition, by incorporating 40 image quality adjectives across six distinct dimensions, we enable the pre-trained CLIP text encoder to generate representations of the intrinsic quality of the image, thereby strengthening the connection between image quality perception and human language. Our method achieves superior results on most public Full-Reference (FR) and No-Reference (NR) IQA benchmarks, while demonstrating greater robustness. Chenyue Song, Wei Zhang 0192, Haiqi Zhu, Shaohui Liu, Feng Jiang 0001 |
ICME | 5 |
| 2025 | ETRT-Net: Efficient Non-Local Transformer and Residual Triplet Attention for Breast Lesion SegmentationabstractBreast cancer stands as one of the leading causes of mortality among women. Accurate lesion segmentation is crucial for timely clinical intervention. However, the varying tumor morphologies and unclear boundaries present in breast ultrasound images pose significant challenges for existing methods, often hindering their ability to produce satisfactory results. This study proposes an Efficient Non-Local Transformer and Residual Triple Attention Network (ETRT-Net). The network combines three modules: the Efficient Non-Local Transformer Block (ENLTB), the Residual Triple Attention Block (RTAB), and the Residual Triple Conv Attention Block (RTCAB). The ENLTB module extends the receptive field of the model with a time complexity of O(N), which aids in accurately identifying breast lesion boundaries. Additionally, we innovatively propose the RTAB and RTCAB modules, which employ a triple attention mechanism to capture cross-dimensional feature dependencies across three parallel branches, enhancing the model’s sensitivity to feature correlations. Through experiments on two open-access breast ultrasound datasets, our method provides more accurate and reliable segmentation, highlighting its potential for improving diagnostic capabilities in clinical settings. Jingxing Cao, Shaohui Liu |
IJCNN | 4 |
| 2025 | UAV-MaLO: Mamba-Augmented YOLO Hybrid Architecture for UAV Micro-Object Detection in Autonomous RoboticsabstractThe rapid advancement of drone technology has led to the widespread application of micro-object detection in Unmanned Aerial Vehicle (UAV) systems. However, with the constraint of real-time computation, critical challenges remain in addressing extreme scale variations, low-resolution signatures and dense occlusions. For object detection task, although YOLO-based detectors outperform transformer models in efficiency-accuracy balance, their limited capacity for global context modeling and feature discriminability in complex aerial environments hinders optimal performance. To overcome these limitations, we introduce UAV-MaLO, a novel framework that incorporates state space modeling principles into YOLO’s architecture. By introducing the abilities of long-range dependency modeling and adaptive spatial-frequency fusion, the proposed approach dynamically optimizes receptive fields while suppressing background interference, achieving robust micro-object localization in cluttered scenarios. Furthermore, the parallelized attention mechanism and the hierarchical feature refinement further ensure real-time processing capabilities without compromising detection precision, establishing a new paradigm for UAV deployment. Our experimental results on the VisDrone-2019-DET dataset reveal a significant improvement in various variants of average precision (AP), indicating the extraordinary performance of our UAV-MaLO. Lennox Wei, Shixin Sun, Yachun Mi, Xiangyu Sui, Shaohui Liu |
IROS | 7 |
| 2025 | LVPNet: A Latent-Variable-Based Prediction-Driven End-to-End Framework for Lossless Compression of Medical Images
Chenyue Song, Wei Zhang 0192, Siqiao Li, Haiqi Zhu, Shengping Zhang, Shaohui Liu, Feng Jiang 0001 |
MICCAI (8) | 9 |
| 2025 | GAOT: Generating Articulated Objects Through Text-Guided Diffusion ModelsabstractArticulated object generation has seen increasing advancements, yet existing models often lack the ability to be conditioned on text prompts. To address the significant gap between textual descriptions and 3D articulated object representations, we propose GAOT, a three-phase framework that generates articulated objects from text prompts, leveraging diffusion models and hypergraph learning in a three-step process. Lei Fan 0007, Donglin Di, Shaohui Liu |
MMAsia | 4 |
| 2025 | Context-Aware Hierarchical Learning: A Two-Step Paradigm towards Safer LLMsabstractLarge Language Models (LLMs) have emerged as powerful tools for diverse applications. However, their uniform token processing paradigm introduces critical vulnerabilities in instruction handling, particularly when exposed to adversarial scenarios. In this work, we identify and propose a novel class of vulnerabilities, termed Tool-Completion Attack (TCA), which exploits function-calling mechanisms to subvert model behavior. To evaluate LLM robustness against such threats, we introduce the Tool-Completion benchmark, a comprehensive security assessment framework, which reveals that even state-of-the-art models remain susceptible to TCA, with surprisingly high attack success rates. To address these vulnerabilities, we introduce Context-Aware Hierarchical Learning (CAHL), a sophisticated mechanism that dynamically equilibrates semantic comprehension with role-specific instruction constraints. CAHL leverages the contextual correlations between different instruction segments to establish a robust, context-aware instruction hierarchy. Extensive experiments demonstrate that CAHL significantly enhances LLM robustness against both conventional attacks and the proposed TCA, exhibiting strong generalization capabilities in zero-shot evaluations while still preserving model performance on generic tasks. Our code is available at https://github.com/S2AILab/CAHL. Tengyun Ma, Daojing He, Shihao Peng, Yu Li 0007, Shaohui Liu, Zhuotao Tian |
NeurIPS | 6 |
| 2025 | RFformer: Rectified Flow Transformer for Time Series Anomaly Detection
Danni Hui, Haiqi Zhu, Shaohui Liu, Muyun Yang, Chunzhi Yi, Baichun Wei |
PRCV (1) | 3 |
| 2025 | AMH-Net: Adaptive Multi-Band Hybrid-Aware Network for Emotion Recognition in SpeechabstractSpeech emotion recognition (SER) technology analyzes speech signals to automatically identify the speaker's emotional state. However, existing methods overlook feature extraction based on human acoustic characteristics. In this paper, we propose AMH-Net, an Adaptive Multi-band Hybridaware Network designed for SER. The model leverages formant (F1, F2, F3) from speech science, which describe the human vocal tract, to partition speech signals into multiple frequency bands. A variable-depth residual network structure is employed for more precise extraction of emotional characteristics. In addition, a hybrid attention mechanism is integrated to combine information, resulting in a more comprehensive emotional representation. Experimental evaluations of six diverse datasets show that AMH-Net outperforms state-of-the-art methods, achieving improvements of 2.11% and 2.64% in average UAR and WAR, respectively, on each corpus. The code is publicly available at https://github.com/hengruili1997/AMH-net Hengrui Li, Yongbing Zhang 0002, Shaohui Liu |
IEEE Signal Process. Lett. | 3 |
| 2025 | Image Compressive Sensing With Scale-Variable Adaptive Sampling and Hybrid-Attention Transformer ReconstructionabstractRecently, a large number of image compressive sensing (CS) methods with deep unfolding networks (DUNs) have been proposed. However, existing methods either use fixed-scale blocks for sampling that leads to limited insights into the image content or employ a plain convolutional neural network (CNN) in each iteration that weakens the perception of broader contextual prior. In this paper, we propose a novel DUN (dubbed SVASNet) for image compressive sensing, which achieves scale-variable adaptive sampling and hybrid-attention Transformer reconstruction with a single model. Specifically, for scale-variable sampling, a sampling matrix-based calculator is first employed to evaluate the reconstruction distortion, which only requires measurements without access to the ground truth image. Then, a Block Scale Aggregation (BSA) strategy is presented to compute the reconstruction distortion under block divisions at different scales and select the optimal division scale for sampling. To realize hybrid-attention reconstruction, a dual Cross Attention (CA) submodule in the gradient descent step and a Spatial Attention (SA) submodule in the proximal mapping step are developed. The CA submodule introduces inter-phase inertial forces in the gradient descent, which improves the memory effect between adjacent iterations. The SA submodule integrates local and global prior representations of CNN and Transformer, and explores local and global affinities between dense feature representations. Extensive experimental results show that the proposed SVASNet achieves significant improvements over the state-of-the-art methods. Debin Zhao, Weisi Lin, Shaohui Liu, Feng Jiang 0001 |
IEEE Trans. Multim. | 4 |
| 2025 | Progressively Learning to Reach Remote Goals by Continuously Updating Boundary GoalsabstractTraining an effective policy on complex goal-reaching tasks with sparse rewards is an open challenge. It is more difficult for the task of reaching remote goals (RRG), as the unavailability of the original rewards and large Wasserstein distance between the distributions of desired goals and initial states make existing methods for common goal-reaching tasks inefficient or even completely ineffective. In this article, we propose progressively learning to reach remote goals by continuously updating boundary goals (PLUB), which solves RRG tasks by reducing the Wasserstein distance between the distributions of boundary goals and desired goals. Specifically, the concept of boundary goal is introduced, which is the set of the closest achieved goals for each desired goal. In addition, to reduce the computational complexity caused by the Wasserstein distance, the closest moving distance is introduced, which is its upper bound, and also the expectation of the distance between the desired goal and the closest boundary goal. By selecting the appropriate intermediate goal from all boundary goals and continuously updating boundary goals, both the closest moving distance and the Wasserstein distance can be reduced. As a result, RRG tasks degenerate into common goal-reaching tasks that can be efficiently solved by a combination of hindsight relabeling and the learning from demonstrations (LfD) method. Extensive experiments on several robotic manipulation tasks demonstrate that PLUB can bring substantial improvements over the existing methods. Mengxuan Shao, Haiqi Zhu, Debin Zhao, Feng Jiang 0001, Shaohui Liu, Wei Zhang 0192 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2024 | Handbook on Leveraging Lines for Two-View Relative Pose EstimationabstractWe propose an approach for estimating the relative pose between calibrated image pairs by jointly exploiting points, lines, and their coincidences in a hybrid manner. We investigate all possible configurations where these data modalities can be used together and review the minimal solvers available in the literature. Our hybrid framework combines the advantages of all configurations, enabling robust and accurate estimation in challenging environments. In addition, we design a method for jointly estimating multiple vanishing point correspondences in two images, and a bundle adjustment that considers all relevant data modalities. Experiments on various indoor and outdoor datasets show that our approach outperforms point-based methods, improving AUC@ 10° by 1-7 points while running at comparable speeds. The source code of the solvers and hybrid framework will be made public. Petr Hruby, Shaohui Liu, Rémi Pautrat, Marc Pollefeys, Daniel Barath |
3DV | 2 |
| 2024 | Data Augmentation Techniques for Chinese Disease Name NormalizationabstractDisease name normalization is an important task in the medical domain. It classifies disease names written in various formats into standardized names, serving as a fundamental component in smart healthcare systems for various disease-related functions. Nevertheless, the most significant obstacle to existing disease name normalization systems is the severe shortage of training data. Consequently, we present a novel data augmentation approach that includes a series of data augmentation techniques and some supporting modules to help mitigate the problem. Through extensive experimentation, we illustrate that our proposed approach exhibits significant performance improvements across various baseline models and training objectives, particularly in scenarios with limited training data1. Wenqian Cui, Xiangling Fu, Shaohui Liu, Mingjun Gu, Xien Liu, Ji Wu 0002, Irwin King |
BIBM | 3 |
| 2024 | 3D Neural Edge ReconstructionabstractReal-world objects and environments are predominantly composed of edge features, including straight lines and curves. Such edges are crucial elements for various applications, such as CAD modeling, surface meshing, lane mapping, etc. However, existing traditional methods only prioritize lines over curves for simplicity in geometric modeling. To this end, we introduce EMAP, a new method for learning 3D edge representations with a focus on both lines and curves. Our method implicitly encodes 3D edge distance and direction in Unsigned Distance Functions (UDF) from multi-view edge maps. On top of this neural representation, we propose an edge extraction algorithm that robustly abstracts parametric 3D edges from the inferred edge points and their directions. Comprehensive evaluations demonstrate that our method achieves better 3D edge reconstruction on multiple challenging datasets. We further show that our learned UDF field enhances neural surface reconstruction by capturing more details. Songyou Peng, Zehao Yu 0002, Shaohui Liu, Rémi Pautrat, Xiaochuan Yin, Marc Pollefeys |
CVPR | 4 |
| 2024 | Robust Incremental Structure-from-Motion with Hybrid Features
Shaohui Liu, Yidan Gao, Rémi Pautrat, Johannes L. Schönberger, Viktor Larsson, Marc Pollefeys |
ECCV (36) | 1 |
| 2024 | ZE-FESG: A Zero-Shot Feature Extraction Method Based on Semantic Guidance for No-Reference Video Quality AssessmentabstractAlthough the current deep neural network based no-reference video quality assessment (NR-VQA) methods can effectively simulate the human visual system (HVS), their interpretability is getting worse. The current methods only extract the low-level features of space and time of the video and do not consider the impact of high-level semantics. However, the high-level semantic information in the video related to human subjective perception and related to its own quality can be perceived by the HVS. In this work, we design the multidimensional feature extractor (MDFE), which takes the text descriptions related to video quality factors as semantic guidance, and uses the Contrastive Language-Image Pre-training (CLIP) model to perform zero-shot multidimensional feature extraction. Then, we further propose a zero-shot feature extraction method based on semantic guidance (ZE-FESG), which treats the MDFE as a feature extractor and acquires all the semantically corresponding features of the video by sliding over each frame of the video. Extensive experiments show that the proposed ZE-FESG has better interpretability and performance than the current mainstream 2D-CNN based feature extraction methods for NR-VQA. The code will be released on https://github.com/xiao-mi-d/ZE-FESG. Yachun Mi, Shaohui Liu |
ICASSP | 4 |
| 2024 | SC-HVPPNet: Spatial and Channel Hybrid-Attention Video Post-Processing Network with CNN and TransformerabstractConvolutional Neural Network (CNN) and Transformer have attracted much attention recently for video post-processing (VPP). However, the interaction between CNN and Transformer in existing VPP methods is not fully explored, leading to inefficient communication between the local and global extracted features. In this paper, we explore the interaction between CNN and Transformer in the task of VPP, and propose a novel Spatial and Channel Hybrid-Attention Video Post-Processing Network (SC-HVPPNet), which can cooperatively exploit the image priors in both spatial and channel domains. Specifically, in the spatial domain, a novel spatial attention fusion module is designed, in which two attention weights are generated to fuse the local and global representations collaboratively. In the channel domain, a novel channel attention fusion module is developed, which can blend the deep representations at the channel dimension dynamically. Extensive experiments show that SC-HVPPNet notably boosts video restoration quality, with average bitrate savings of 5.29%, 12.42%, and 13.09% for Y, U, and V components in the VTM-11.0-NNVC RA configuration. Wenxue Cui, Shaohui Liu, Feng Jiang 0001 |
ICME | 3 |
| 2024 | S2-CSNet: Scale-Aware Scalable Sampling Network for Image Compressive SensingabstractDeep network-based image Compressive Sensing (CS) has attracted much attention in recent years. However, there still exist the following two issues: 1) Existing methods typically use fixed-scale sampling, which leads to limited insights into the image content. 2) Most pre-trained models can only handle fixed sampling rates and fixed block scales, which restricts the scalability of the model. In this paper, we propose a novel scale-aware scalable CS network (dubbed S2-CSNet), which achieves scale-aware adaptive sampling, fine granular scalability and high-quality reconstruction with one single model. Specifically, to enhance the scalability of the model, a structural sampling matrix with a predefined order is first designed, which is a universal sampling matrix that can sample multi-scale image blocks with arbitrary sampling rates. Then, based on the universal sampling matrix, a distortion-guided scale-aware scheme is presented to achieve scale-variable adaptive sampling, which predicts the reconstruction distortion at different sampling scales from the measurements and select the optimal division scale for sampling. Furthermore, a multi-scale hierarchical sub-network under a well-defined compact framework is put forward to reconstruct the image. In the multi-scale feature domain of the sub-network, a dual spatial attention is developed to explore the local and global affinities between dense feature representations for deep fusion. Extensive experiments manifest that the proposed S2-CSNet outperforms existing state-of-the-art CS methods. Haiqi Zhu, Shuya Yan, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
ACM Multimedia | 4 |
| 2024 | CLiF-VQA: Enhancing Video Quality Assessment by Incorporating High-Level Semantic Information related to Human FeelingsabstractVideo Quality Assessment (VQA) aims to simulate the process of perceiving video quality by the Human Visual System (HVS). Although subjective studies have shown that the judgments of HVS are strongly influenced by human feelings, it remains unclear how video content relates to human feelings. The recent rapid development of Vision-Language pre-trained models (VLM) has established a solid link between language and vision. And human feelings can be accurately described by language, which means that VLM can extract information related to human feelings from visual content with linguistic prompts. In this paper, we propose CLiF-VQA, which innovatively utilizes the visual linguistic capabilities of VLM to introduce human feelings features based on traditional spatio-temporal features to more accurately simulate the perceptual process of HVS. In order to efficiently extract features related to human feelings from videos, we pioneer the exploration of the consistency between Contrastive Language-Image Pre-training (CLIP) and human feelings in video perception. In addition, we design effective prompts, i.e., a variety of objective and subjective descriptions closely related to human feelings, as prompts. Extensive experiments show that the proposed CLiF-VQA exhibits excellent performance on several VQA datasets. The results show that introducing human feelings features on top of spatio-temporal features is an effective way to obtain better performance. Yachun Mi, Puchao Zhou, Shaohui Liu |
ACM Multimedia | 6 |
| 2024 | AlphaTablets: A Generic Plane Representation for 3D Planar Reconstruction from Monocular VideosabstractWe introduce AlphaTablets, a novel and generic representation of 3D planes that features continuous 3D surface and precise boundary delineation. By representing 3D planes as rectangles with alpha channels, AlphaTablets combine the advantages of current 2D and 3D plane representations, enabling accurate, consistent and flexible modeling of 3D planes. We derive differentiable rasterization on top of AlphaTablets to efficiently render 3D planes into images, and propose a novel bottom-up pipeline for 3D planar reconstruction from monocular videos. Starting with 2D superpixels and geometric cues from pre-trained models, we initialize 3D planes as AlphaTablets and optimize them via differentiable rendering. An effective merging scheme is introduced to facilitate the growth and refinement of AlphaTablets. Through iterative optimization and merging, we reconstruct complete and accurate 3D planes with solid surfaces and clear boundaries. Extensive experiments on the ScanNet dataset demonstrate state-of-the-art performance in 3D planar reconstruction, underscoring the great potential of AlphaTablets as a generic 3D plane representation for various applications. Wang Zhao 0001, Shaohui Liu, Yubin Hu 0001, Yushi Bai, Yu-Hui Wen, Yong-Jin Liu 0001 |
NeurIPS | 3 |
| 2024 | De2Net: Under-display camera image restoration with feature deconvolution and kernel decomposition
Hangyan Zhu, Shaohui Liu, Ming Liu 0018, Zifei Yan, Wangmeng Zuo |
Comput. Vis. Image Underst. | 2 |
| 2024 | An Interpretable Multivariate Time-Series Anomaly Detection Method in Cyber-Physical Systems Based on Adaptive MaskabstractThe high complexity and wide applications of Cyber-Physical Systems (CPSs) pose a large requirement on both accuracy and interpretability of the time-series anomaly detection algorithms. While a large number of deep learning algorithms have achieved excellent accuracy, the interpretability is often limited, especially when considering retaining correlations in multivariate time-series. In this paper, we propose a novel multivariate time-series anomaly detection method based on adaptive masking mechanism to improve both accuracy and interpretability, which contains a specially designed series saliency module. For more intuitive and interpretable results, a learnable adaptive mask is introduced in the series saliency module, which can disclose the influence on anomalies in both feature and temporal dimensions. The original time-series and their versions with adaptive perturbations added are then mixed via the mask forming an adaptive data augmentation method to improve the accuracy of anomaly detection. Furthermore, the anomaly detection module is model-agnostic, whether based on forecasting or reconstruction. The optimization of the training objectives will lead to more accurate and interpretable detection results. With four real-world datasets, we demonstrate that the adaptive mask can provide more accurate anomaly detection results with meaningful interpretations in the form of a mask matrix. Haiqi Zhu, Chunzhi Yi, Seungmin Rho, Shaohui Liu, Feng Jiang 0001 |
IEEE Internet Things J. | 4 |
| 2024 | Rate-Adaptive Neural Network for Image Compressive SensingabstractDeep learning-based image compressive sensing (CS) methods have achieved great success in the past few years. However, most of them are content-independent, with a spatially uniform sampling rate allocation for the entire image. Such practises may potentially degrade the performance of image CS with block-based sampling, since the content of different blocks in an image is different. In this article, we propose a novel rate-adaptive image CS neural network (dubbed RACSNet) to achieve adaptive sampling rate allocation based on the content characteristics of the image with a single model. Specifically, a measurement domain-based reconstruction distortion is first used to guide the sampling rate allocation for different blocks in an image without access to the ground truth image. Then, a step-wise training strategy is designed to train a reusable sampling matrix, which is capable of sampling image blocks to generate the compressed measurements under arbitrary sampling rates. Subsequently, a pyramid-shaped initial reconstruction sub-network and a hierarchical deep reconstruction sub-network that fuse the measurement information of different scales are put forward to reconstruct image blocks from the compressed measurements. Finally, a reconstruction distortion map and an improved loss function are developed to eliminate the blocking artifacts and further enhance the CS reconstruction. Experimental results on both objective metrics and subjective visual qualities show that the proposed RACSNet achieves significant improvements over the state-of-the-art methods. Shengping Zhang, Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Multim. | 4 |
| 2024 | Deep Network for Image Compressed Sensing Coding Using Local Structural SamplingabstractExisting image compressed sensing (CS) coding frameworks usually solve an inverse problem based on measurement coding and optimization-based image reconstruction, which still exist the following two challenges: (1) the widely used random sampling matrix, such as the Gaussian Random Matrix (GRM), usually leads to low measurement coding efficiency, and (2) the optimization-based reconstruction methods generally maintain a much higher computational complexity. In this article, we propose a new convolutional neural network based image CS coding framework using local structural sampling (dubbed CSCNet) that includes three functional modules: local structural sampling, measurement coding, and Laplacian pyramid reconstruction. In the proposed framework, instead of GRM, a new local structural sampling matrix is first developed, which is able to enhance the correlation between the measurements through a local perceptual sampling strategy. Besides, the designed local structural sampling matrix can be jointly optimized with the other functional modules during the training process. After sampling, the measurements with high correlations are produced, which are then coded into final bitstreams by the third-party image codec. Last, a Laplacian pyramid reconstruction network is proposed to efficiently recover the target image from the measurement domain to the image domain. Extensive experimental results demonstrate that the proposed scheme outperforms the existing state-of-the-art CS coding methods while maintaining fast computational speed. Wenxue Cui, Xiaopeng Fan 0001, Shaohui Liu, Xinwei Gao, Debin Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 4 |
| 2023 | 3D Line Mapping RevisitedabstractIn contrast to sparse keypoints, a handful of line segments can concisely encode the high-level scene layout, as they often delineate the main structural elements. In addition to offering strong geometric cues, they are also omnipresent in urban landscapes and indoor scenes. Despite their apparent advantages, current line-based reconstruction methods are far behind their point-based counterparts. In this paper we aim to close the gap by introducing LIMAP, a library for 3D line mapping that robustly and efficiently creates 3D line maps from multi-view imagery. This is achieved through revisiting the degeneracy problem of line triangulation, carefully crafted scoring and track building, and exploiting structural priors such as line coincidence, parallelism, and orthogonality. Our code integrates seamlessly with existing point-based Structure-from-Motion methods and can leverage their 3D points to further improve the line reconstruction. Furthermore, as a byproduct, the method is able to recover 3D association graphs between lines and points / vanishing points (VPs). In thorough experiments, we show that LIMAP significantly outperforms existing approaches for 3D line mapping. Our robust 3D line maps also open up new research directions. We show two example applications: visual localization and bundle adjustment, where integrating lines alongside points yields the best results. Code is available at https://github.com/cvg/limap. Shaohui Liu, Yifan Yu 0003, Rémi Pautrat, Marc Pollefeys, Viktor Larsson |
CVPR | 1 |
| 2023 | EI2SR: Learning an Enhanced Intra-Instance Semantic Relationship for Arbitrary-Shaped Scene Text DetectionabstractText detection in natural scenarios, has made significant progress with the deep learning architecture. Towards arbitrary-shaped text detection, fracture detection is the major concern due to the lack of semantic relationship within an instance in existing methods. To circumvent this dilemma, we propose a novel network to learn an Enhanced Intra-Instance Semantic Relationship (EI2SR) which consists of Text-Specific Attention Mechanism (TAM) and Border Attraction Grouping (BAG). The former models the rich semantic information between different coarse-grained text regions to guide the fine-grained learning of corresponding text representations. The latter enhances the border-center semantic correlation by establishing high-dimension embedding space to attract and group the border at both ends to their corresponding center. Extensive experimental results show that the proposed EI2SR achieves state-of-the-art or competitive performance on existing benchmarks. Shaohui Liu, Yu Zhou 0015, Feng Jiang 0001 |
ICASSP | 2 |
| 2023 | Aprogressive Image Dehazing Framework with inter and Intra Contrastive LearningabstractImage dehazing, aims to estimate latent haze-free images from hazy images, suffering from a lot of lost information. Existing contrastive learning methods tend to utilize hazefree images as positive samples without consideration of negative samples. Even if negative samples are employed, the connection between patches within an image is always ignored. In addition, it is hard to train end-to-end dehazing networks due to the enormous gap between hazy images and corresponding clear images. In this paper, we propose a novel progressive image dehazing framework with inter and intra contrastive learning to solve the above problems. Specifically, the Inter and Intra Contrastive Learning (IICL) is proposed, in which the brightest and darkest patches within the same image are considered for contrastive learning. Furthermore, a progressive image dehazing framework consisting of an efficient Pre-restore Module (PRM) and an Alternative Restored Module (ARM) is proposed to facilitate the end-to-end model training. It is noted that our framework can be a complement to existing image dehazing methods. Extensive experiments on the dehazing benchmark demonstrate that our framework benefits various dehazing models which surpass previous state-of-the-art image dehazing methods. Shaohui Liu, Feng Jiang 0001 |
ICASSP | 2 |
| 2023 | Vanishing Point Estimation in Uncalibrated Images with Prior Gravity DirectionabstractWe tackle the problem of estimating a Manhattan frame, i.e. three orthogonal vanishing points, and the unknown focal length of the camera, leveraging a prior vertical direction. The direction can come from an Inertial Measurement Unit that is a standard component of recent consumer devices, e.g., smartphones. We provide an exhaustive analysis of minimal line configurations and derive two new 2-line solvers, one of which does not suffer from singularities affecting existing solvers. Additionally, we design a new non-minimal method, running on an arbitrary number of lines, to boost the performance in local optimization. Combining all solvers in a hybrid robust estimator, our method achieves increased accuracy even with a rough prior. Experiments on synthetic and real-world datasets demonstrate the superior accuracy of our method compared to the state of the art, while having comparable runtimes. We further demonstrate the applicability of our solvers for relative rotation estimation. The code is available at https://github.com/cvg/VP-Estimation-with-Prior-Gravity. Rémi Pautrat, Shaohui Liu, Petr Hruby, Marc Pollefeys, Daniel Barath |
ICCV | 2 |
| 2023 | LNPL-MIL: Learning from Noisy Pseudo Labels for Promoting Multiple Instance Learning in Whole Slide ImageabstractGigapixel Whole Slide Images (WSIs) aided patient diagnosis and prognosis analysis are promising directions in computational pathology. However, limited by expensive and time-consuming annotation costs, WSIs usually only have weak annotations, including 1) WSI-level Annotations (WA) and 2) Limited Patch-level Annotations (LPA). Currently, Multiple Instance Learning (MIL) often exploits WA, while LPA usually assign pseudo-labels for unlabeled data. Intuitively, pseudo-labels can serve as a practical guide for MIL, but the unreliable prediction caused by LPA inevitably introduce noise. Furthermore, WA-supervised MIL training inevitably suffers from the semantical unalignment between instances and bag-level labels. To address these problems, we design a framework called Learning from Noisy Pseudo Labels for promoting Multiple Instance Learning (LNPL-MIL), which considers both types of weak annotation. Specifically, for the LPA-trained weak classifier, we design a Super-Patch-based LNPL (SP-LNPL) method to reduce false positives in the noisy pseudo-labels and then select more accurate Top-K key instances. In MIL, we propose a Transformer aware of instance Order and Distribution (TOD-MIL) that strengthens instances correlation and weakens semantical unalignment in the bag. We validate our LNPL-MIL on Tumor Diagnosis and Survival Prediction, achieving state-of-the-art performance with at least 2.7%/2.9% AUC and 2.6%/2.3% C-Index improvement with the patches labeled for two scale. Ablation study and visualization analysis further verify the effectiveness. Zhuchen Shao, Yifeng Wang 0001, Yang Chen 0036, Hao Bian, Shaohui Liu, Haoqian Wang, Yongbing Zhang 0002 |
ICCV | 5 |
| 2023 | Uformer++: Light Uformer for Image Restoration
Shaohui Liu |
ICONIP (13) | 2 |
| 2023 | G2-DUN: Gradient Guided Deep Unfolding Network for Image Compressive SensingabstractInspired by certain optimization solvers, the deep unfolding network (DUN) usually inherits a multi-phase structure for image compressive sensing (CS). However, in existing DUNs, the message transmission within and between phases still faces two issues: 1) the roughness of transmitted information, e.g., the low-dimensional representations. 2) the inefficiency of transmitted policy, e.g., simply concatenating deep features. In this paper, by unfolding the Proximal Gradient Descent (PGD) algorithm, a novel gradient guided DUN (G2 -DUN) for image CS is proposed, in which a gradient map is delicately introduced within each phase for providing richer informational guidance at both intra-phase and inter-phase levels. Specifically, corresponding to the gradient descent (GD) of PGD, a gradient guided GD module is designed, in which the gradient map can adaptively guide step size allocation for different textures of input image, realizing a content-aware gradient updating. On the other hand, corresponding to the proximal mapping (PM) of PGD, a gradient guided PM module is developed, in which the gradient map can dynamically guide the exploring of deep textural priors in multi-scale space, achieving the dynamic perception of the proposed deep model. By introducing the gradient map, the proposed message transmission system not only facilitates the informational communication between different functional modules within each phase, but also strengthens the inferential cooperation among cascaded phases. Extensive experiments manifest that the proposed G2 -DUN outperforms existing state-of-the-art CS methods. Wenxue Cui, Xiaopeng Fan 0001, Shaohui Liu, Debin Zhao |
ACM Multimedia | 4 |
| 2023 | Perceiving Ambiguity and Semantics without Recognition: An Efficient and Effective Ambiguous Scene Text DetectorabstractAmbiguous scene text detection is an extremely challenging task. Existing text detectors that rely solely on visual cues often suffer from confusion due to being evenly distributed in rows/columns or incomplete detection owing to large character spacing. To overcome these challenges, the previous method recognizes a large number of proposals and utilizes semantic information predicted from recognition results to eliminate ambiguity. However, this method is inefficient, which limits their practical applications. In this paper, we propose a novel efficient and effective ambiguous text detector, which can Perceive Ambiguity and SEmantics without Recognition, termed PASER. On the one hand, PASER can perceive semantics without recognition with a light Perceiving Semantics (PerSem) module. In this way, proposals without reasonable semantics are filtered out, which largely speeds up the overall detection process. On the other hand, to detect both ambiguous and regular texts with a unified framework, PASER employs a Perceiving Ambiguity (PerAmb) module to distinguish ambiguous texts and regular texts, so that only the ambiguous proposals will be processed by PerSem while the regular texts are not, which further ensures the high efficiency. Extensive experiments show that our detector achieves state-of-the-art results on both ambiguous and regular scene text detection benchmarks. Notably, over 6 times faster speed and superior accuracy are achieved on TDA-ReCTS simultaneously. Wei Wang 0315, Yu Zhou 0015, Shaohui Liu, Aoting Zhang, Dongbao Yang, Weiping Wang 0005 |
ACM Multimedia | 4 |
| 2023 | VVA: Video Values Analysis
Yachun Mi, Shaohui Liu, Feng Jiang 0001 |
PRCV (7) | 4 |
| 2023 | Depth-Guided Optimization of Neural Radiance Fields for Indoor Multi-View StereoabstractIn this work, we present a new multi-view depth estimation method NerfingMVS that utilizes both conventional reconstruction and learning-based priors over the recently proposed neural radiance fields (NeRF). Unlike existing neural network based optimization method that relies on estimated correspondences, our method directly optimizes over implicit volumes, eliminating the challenging step of matching pixels in indoor scenes. The key to our approach is to utilize the learning-based priors to guide the optimization process of NeRF. Our system first adapts a monocular depth network over the target scene by finetuning on its MVS reconstruction from COLMAP. Then, we show that the shape-radiance ambiguity of NeRF still exists in indoor environments and propose to address the issue by employing the adapted depth priors to monitor the sampling process of volume rendering. Finally, a per-pixel confidence map acquired by error computation on the rendered image can be used to further improve the depth quality. We further present NerfingMVS++, where a coarse-to-fine depth priors training strategy is proposed to directly utilize sparse SfM points and the uniform sampling is replaced by Gaussian sampling to boost the performance. Experiments show that our NerfingMVS and its extension NerfingMVS++ achieve state-of-the-art performances on indoor datasets ScanNet and NYU Depth V2. In addition, we show that the guided optimization scheme does not sacrifice the original synthesis capability of neural radiance fields, improving the rendering quality on both seen and novel views. Code is available at https://github.com/weiyithu/NerfingMVS. Yi Wei 0003, Shaohui Liu, Jie Zhou 0001, Jiwen Lu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Image Compressed Sensing Using Non-Local Neural NetworkabstractDeep network-based image Compressed Sensing (CS) has attracted much attention in recent years. However, the existing deep network-based CS schemes either reconstruct the target image in a block-by-block manner that leads to serious block artifacts or train the deep network as a black box that brings about limited insights of image prior knowledge. In this paper, a novel image CS framework using non-local neural network (NL-CSNet) is proposed, which utilizes the non-local self-similarity priors with deep network to improve the reconstruction quality. In the proposed NL-CSNet, two non-local subnetworks are constructed for utilizing the non-local self-similarity priors in the measurement domain and the multi-scale feature domain respectively. Specifically, in the subnetwork of measurement domain, the long-distance dependencies between the measurements of different image blocks are established for better initial reconstruction. Analogically, in the subnetwork of multi-scale feature domain, the affinities between the dense feature representations are explored in the multi-scale space for deep reconstruction. Furthermore, a novel loss function is developed to enhance the coupling between the non-local representations, which also enables an end-to-end training of NL-CSNet. Extensive experiments manifest that NL-CSNet outperforms existing state-of-the-art CS methods, while maintaining fast computational speed. Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Multim. | 2 |
| 2022 | ParticleSfM: Exploiting Dense Point Trajectories for Localizing Moving Cameras in the Wild
Wang Zhao 0001, Shaohui Liu, Hengkai Guo, Wenping Wang 0001, Yong-Jin Liu 0001 |
ECCV (32) | 2 |
| 2022 | Source Camera Identification with Multi-Scale Feature Fusion NetworkabstractSource camera identification (SCI) technology has attracted increasing attentions over the past few years. However, the existing methods suppress image content with denoising filters that are largely agnostic to the specific sensor pattern noise (SPN) signal of interest. Such practices may potentially degrade the performance of SPN-based SCI due to un-reliable SPNs, especially when forensic images are transmitted through social networking platforms. In this paper, we address the problem of SPN-based device identification and propose a multi-scale feature fusion network (MSFFN) to boost the sensor-based source camera identification attribution. Specifically, several image patches of different scales are selected and input into the MSFFN to extract the SPN. The MSFFN is a multi-scale encoder-decoder structure, which is used to suppress image content and improve source attribution. Subsequently, the content-independent SPN features of different scales are fused. At last, the fused features are used for image source identification. Experimental results compared with the state-of-the-art demonstrate that the proposed scheme achieves significant improvements, especially in the accuracy of social networking image source identification. Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICME | 3 |
| 2022 | Multi-Channel Adaptive Partitioning Network for Block-Based Image Compressive SensingabstractImage compressive sensing (CS) technology has attracted increasing attentions in the past few years, and a great deal deep learning-based methods have been proposed. However, the existing methods use fixed-scale blocks for sampling and re-construction. Such practice will inevitably result in the in-ability to distinguish between significant regions and background regions, and even waste excessive sampling resources on the background ones to a large extent. In this paper, we propose a novel multi-channel adaptive partitioning network for block-based image CS, in which image blocks of different scales are utilized to distinguish regions of different saliency. Specifically, an adaptive block partitioning method based on image saliency is put forward, using which significant regions are divided into large blocks and background regions are divided into small blocks. Subsequently, blocks of different scales are fed to different-channel networks for sampling to yield the compressed measurements. To improve the re-construction quality of the image, a scalable multi-scale re-construction network is proposed to recover the compressed measurements into the reconstructed image. Experimental results compared with the state-of-the-art show that the proposed scheme achieves significant improvements in terms of objective metrics and subjective visual image quality. Shaohui Liu, Feng Jiang 0001 |
ICME | 2 |
| 2022 | Learning from Hindsight Demonstrations
Mengxuan Shao, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICONIP (5) | 3 |
| 2022 | Hindsight Balanced Reward Shaping
Mengxuan Shao, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICONIP (5) | 3 |
| 2022 | Fast Hierarchical Deep Unfolding Network for Image Compressed SensingabstractBy integrating certain optimization solvers with deep neural network, deep unfolding network (DUN) has attracted much attention in recent years for image compressed sensing (CS). However, there still exist several issues in existing DUNs: 1) For each iteration, a simple stacked convolutional network is usually adopted, which apparently limits the expressiveness of these models. 2) Once the training is completed, most hyperparameters of existing DUNs are fixed for any input content, which significantly weakens their adaptability. In this paper, by unfolding the Fast Iterative Shrinkage-Thresholding Algorithm (FISTA), a novel fast hierarchical DUN, dubbed FHDUN, is proposed for image compressed sensing, in which a well-designed hierarchical unfolding architecture is developed to cooperatively explore richer contextual prior information in multi-scale spaces. To further enhance the adaptability, series of hyperparametric generation networks are developed in our framework to dynamically produce the corresponding optimal hyperparameters according to the input content. Furthermore, due to the accelerated policy in FISTA, the newly embedded acceleration module makes the proposed FHDUN save more than 50% of the iterative loops against recent DUNs. Extensive CS experiments manifest that the proposed FHDUN outperforms existing state-of-the-art CS methods, while maintaining fewer iterations. Wenxue Cui, Shaohui Liu, Debin Zhao |
ACM Multimedia | 2 |
| 2022 | Adversarial training of LSTM-ED based anomaly detection for complex time-series in cyber-physical-social systems
Haiqi Zhu, Shaohui Liu, Feng Jiang 0001 |
Pattern Recognit. Lett. | 2 |
| 2022 | Spatio-Temporal Context Based Adaptive Camcorder Recording WatermarkingabstractVideo watermarking technology has attracted increasing attention in the past few years, and a great deal of traditional and deep learning-based methods have been proposed. However, these existing methods usually suffer from the following two challenges: First, most algorithms cannot resist camcorder recording attack, which limits their practical application. Second, watermark embedding may cause substantial degradation of video quality. Through analyzing the unique distortions presented in the camcorder recording process, including geometric distortion, temporal sampling distortion, sensor distortion and processing distortion, this paper proposes a novel spatio-temporal context based adaptive camcorder recording watermarking scheme STACR. In STACR, considering the geometric distortion and video visual quality, we embed the watermark by constructing a spatio-temporal histogram and incorporate a content features based adaptive locating algorithm to select embedding blocks and embedding strengths. As for the temporal sampling attack, we put forward a watermark correlation-based synchronization algorithm and combine it with cross-validation. Moreover, to resist the sensor distortion, we design a local matching-based algorithm to improve the extraction accuracy. In addition, grouped and repeated embedding strategies are combined to cope with the processing distortion. Experimental results compared with the state-of-the-art show that the proposed scheme achieves high video quality and is robust to geometric attacks, compression, scaling, transcoding, recoding, frame rate changes and especially for camcorder recording. Shaohui Liu, Wuzhen Shi, Feng Jiang 0001, Debin Zhao |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2022 | Hiding Message Using a Cycle Generative Adversarial NetworkabstractTraining an image steganography is an unsupervised problem, because it is impossible to obtain an ideal supervised steganographic image corresponding to the cover image and secret message. Inspired by the success of cycle generative adversarial networks in unsupervised tasks such as style transfer, this article proposes to use a cycle generative adversarial network to solve the problem of unsupervised image steganography. Specifically, this article jointly trains five networks, i.e., a steganographic network, an inverse steganographic network, a hidden message reconstruction network, and two discriminative networks, which together constitute a hidden message cycle generative adversarial network (HCGAN). Compared with the recent image steganography based on generative adversative network, HCGAN provides more accurate supervised information, which makes the training process of HCGAN converge faster and the performance of the trained image steganography network is better. In addition, this article introduces an image steganographic network based on residual learning and shows that residual learning can effectively improve the performance of steganography. Furthermore, to the best of our knowledge, we are the first to propose an inverse steganographic network for eliminating steganographic message from steganographic images, which can be used to avoid steganographic message being discovered or acquired by a third party. The experimental results show that compared with the steganography based on generative adversarial network, the proposed HCGAN has a higher correct decoding rate, better visual quality of steganographic image, and higher secrecy. Wuzhen Shi, Shaohui Liu |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2021 | NerfingMVS: Guided Optimization of Neural Radiance Fields for Indoor Multi-view StereoabstractIn this work, we present a new multi-view depth estimation method that utilizes both conventional SfM reconstruction and learning-based priors over the recently proposed neural radiance fields (NeRF). Unlike existing neural network based optimization method that relies on estimated correspondences, our method directly optimizes over implicit volumes, eliminating the challenging step of matching pixels in indoor scenes. The key to our approach is to utilize the learning-based priors to guide the optimization process of NeRF. Our system firstly adapts a monocular depth network over the target scene by finetuning on its sparse SfM reconstruction. Then, we show that the shape-radiance ambiguity of NeRF still exists in indoor environments and propose to address the issue by employing the adapted depth priors to monitor the sampling process of volume rendering. Finally, a per-pixel confidence map acquired by error computation on the rendered image can be used to fur ther improve the depth quality. Experiments show that our proposed framework significantly outperforms state-of-the-art methods on indoor scenes, with surprising findings presented on the effectiveness of correspondence-based opti-mization and NeRF-based optimization over the adapted depth priors. In addition, we show that the guided opti-mization scheme does not sacrifice the original synthesis capability of neural radiance fields, improving the rendering quality on both seen and novel views. Code is available at https://github.com/weiyithu/NerfingMVS. Yi Wei 0003, Shaohui Liu, Yongming Rao, Wang Zhao 0001, Jiwen Lu, Jie Zhou 0001 |
ICCV | 2 |
| 2021 | A Confidence-based Iterative Solver of Depths and Surface Normals for Deep Multi-view StereoabstractIn this paper, we introduce a deep multi-view stereo (MVS) system that jointly predicts depths, surface normals and per-view confidence maps. The key to our approach is a novel solver that iteratively solves for per-view depth map and normal map by optimizing an energy potential based on the locally planar assumption. Specifically, the algorithm updates depth map by propagating from neigh-boring pixels with slanted planes, and updates normal map with local probabilistic plane fitting. Both two steps are monitored by a customized confidence map. This solver is not only effective as a post-processing tool for plane-based depth refinement and completion, but also differentiable such that it can be efficiently integrated into deep learning pipelines. Our multi-view stereo system employs multiple optimization steps of the solver over the initial prediction of depths and surface normals. The whole system can be trained end-to-end, decoupling the challenging problem of matching pixels within poorly textured regions from the cost-volume based neural network. Experimental results on ScanNet and RGB-D Scenes V2 demonstrate state-of-the-art performance of the proposed deep MVS system on multi-view depth estimation, with our proposed solver consistently improving the depth quality over both conventional and deep learning based MVS pipelines. Code is available at https://github.com/thuzhaowang/idn-solver. Wang Zhao 0001, Shaohui Liu, Yi Wei 0003, Hengkai Guo, Yong-Jin Liu 0001 |
ICCV | 2 |
| 2021 | Adaptive Flexible 3D Histogram WatermarkingabstractWatermarking technology has attracted increasing attentions in the past few years, and a great deal traditional and deep learning-based methods have been proposed. However, these methods usually suffer from the following three challenges: First, the current algorithms are designed separately for images or videos, and there is no universal solution. Second, most algorithms cannot resist screen recording, which limits its application. Third, some algorithms can only embed fixed-length watermarks and cannot handle the embedding capacity flexibly. In this paper, a novel watermarking scheme is proposed based on spatial and temporal histograms, in which two types of histogram watermarking are designed: One is constructed in the spatial domain, using the low-frequency characteristics of the image to change the shape of the histogram and embed the watermark. The other is established in the time domain, which uses the similarity of adjacent frames and combines texture features to modify the shape of the temporal histogram to embed the watermark. Experimental results compared with the state-of-the-art demonstrate that the proposed scheme achieves superior performance. Shaohui Liu, Wenxue Cui, Jinghua Zeng, Feng Jiang 0001, Debin Zhao |
ICME | 2 |
| 2021 | Small Object Recognition Using a Spatio-Temporal Neural NetworkabstractObject recognition at different scales has been a fundamental problem in computer vision. In particular, small object recognition attracts increasing attention recently. However, because of working on a single frame only, many recognizers’ performances become unacceptable in many practical application scenarios: very low resolutions, invisible small targets, extremely similar appearances etc. Motivated by the way humans deal with these challenging scenarios of object recognition, this paper introduces frame sequence and attention mechanism to compensate for mutilated information. Specifically, this paper proposes a spatiotemporal neural network (dubbed STNet) for small object recognition. STNet fixes the regions of interest with a super-resolution module, and focuses on the discriminative region with a spatio-temporal attention module. In addition, STNet applies a double layer long short-term memory subnet to make full use of the inter-frame information. Furthermore, this paper presents a challenging air-target recognition dataset ATSETC4 for evaluating the performance of each method in identifying small targets. Our model outperforms many state-of-the-art models on ATSETC4, including MobileNetV2 and SENet. In particular, STNet surpasses VGG11 at an average of 3.67%, even reaches 87.50% and 82.50% on 28 scale and 14 scale on AT-SETC4 respectively. Zhibo Liang, Shaohui Liu, Wuzhen Shi, Feng Jiang 0001 |
ICME | 2 |
| 2021 | Learning Outfit Compatibility with Graph Attention Network and Visual-Semantic EmbeddingabstractFashion recommendation is an essential component of user shopping that it is capable of selecting and presenting fascinating items to customers. The fact that humans exhibit inconsistencies for fashion items in their choice is known to all due to the visual aesthetic features and fine-grained differences of fashion items. Previous research on fashion recommendations mainly focuses on sequential models, most of them only consider complex similarity relationships in fashion compatibility while neglecting the real-world compatible information often desired in practical applications. To learn the fashion compatibility and generate for the outfit, we propose an approach that jointly learns latent fashion concepts in visual-semantic space to measure compatibility between items. The fashion concepts are shaped by design elements such as color, material, and silhouette. Accordingly, we model a unified representation to learn different notions of similarity by mapping text descriptors and images into latent space to learn high-level representations. Experimental results reveal that our method effectively reaches the aimed results on the fill-in-the-blank and outfit compatibility tasks. Xiaochun Cheng, Ruomei Wang 0001, Shaohui Liu |
ICME | 4 |
| 2021 | DFD-Net: lung cancer detection from denoised CT scan image using deep learning
Worku Jifara Sori, Feng Jiang 0001, Arero W. Godana, Shaohui Liu |
Frontiers Comput. Sci. | 4 |
| 2021 | Combining Fields of Experts (FoE) and K-SVD methods in pursuing natural image priors
Feng Jiang 0001, Zhiyuan Chen 0007, Amril Nazir, Wuzhen Shi, Wei Xiang Lim, Shaohui Liu, Seungmin Rho |
J. Vis. Commun. Image Represent. | 6 |
| 2021 | Video Compressed Sensing Using a Convolutional Neural NetworkabstractRecently, a few image compressed sensing (CS) methods based on deep learning have been developed, which achieve remarkable reconstruction quality with low computational complexity. However, these existing deep learning-based image CS methods focus on exploring intraframe correlation while ignoring interframe cues, resulting in inefficiency when directly applied to video CS. In this paper, we propose a novel video CS framework based on a convolutional neural network (dubbed VCSNet) to explore both intraframe and interframe correlations. Specifically, VCSNet divides the video sequence into multiple groups of pictures (GOPs), of which the first frame is a keyframe that is sampled at a higher sampling ratio than the other nonkeyframes. In a GOP, the block-based framewise sampling by a convolution layer is proposed, which leads to the sampling matrix being automatically optimized. In the reconstruction process, the framewise initial reconstruction by using a linear convolutional neural network is first presented, which effectively utilizes the intraframe correlation. Then, the deep reconstruction with multilevel feature compensation is proposed, which compensates the nonkeyframes with the keyframe in a multilevel feature compensation manner. Such multilevel feature compensation allows the network to better explore both intraframe and interframe correlations. Extensive experiments on six benchmark videos show that VCSNet provides better performance over state-of-the-art video CS methods and deep learning-based image CS methods in both objective and subjective reconstruction quality. Wuzhen Shi, Shaohui Liu, Feng Jiang 0001, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | DIST: Rendering Deep Implicit Signed Distance Function With Differentiable Sphere TracingabstractWe propose a differentiable sphere tracing algorithm to bridge the gap between inverse graphics methods and the recently proposed deep learning based implicit signed distance function. Due to the nature of the implicit function, the rendering process requires tremendous function queries, which is particularly problematic when the function is represented as a neural network. We optimize both the forward and backward pass of our rendering layer to make it run efficiently with affordable memory consumption on a commodity graphics card. Our rendering method is fully differentiable such that losses can be directly computed on the rendered 2D observations, and the gradients can be propagated backward to optimize the 3D geometry. We show that our rendering method can effectively reconstruct accurate 3D shapes from various inputs, such as sparse depth and multi-view images, through inverse optimization. With the geometry based reasoning, our 3D shape prediction methods show excellent generalization capability and robustness against various noises. Shaohui Liu, Yinda Zhang 0001, Songyou Peng, Boxin Shi, Marc Pollefeys, Zhaopeng Cui |
CVPR | 1 |
| 2020 | Towards Better Generalization: Joint Depth-Pose Learning Without PoseNetabstractIn this work, we tackle the essential problem of scale inconsistency for self supervised joint depth-pose learning. Most existing methods assume that a consistent scale of depth and pose can be learned across all input samples, which makes the learning problem harder, resulting in degraded performance and limited generalization in indoor environments and long-sequence visual odometry application. To address this issue, we propose a novel system that explicitly disentangles scale from the network estimation. Instead of relying on PoseNet architecture, our method recovers relative pose by directly solving fundamental matrix from dense optical flow correspondence and makes use of a two-view triangulation module to recover an up-to-scale 3D structure. Then, we align the scale of the depth prediction with the triangulated point cloud and use the transformed depth map for depth error computation and dense reprojection check. Our whole system can be jointly trained end-to-end. Extensive experiments show that our system not only reaches state-of-the-art performance on KITTI depth and flow estimation, but also significantly improves the generalization ability of existing self-supervised depth-pose learning methods under a variety of challenging scenarios, and achieves state-of-the-art results among self-supervised learning-based methods on KITTI Odometry and NYUv2 dataset. Furthermore, we present some interesting findings on the limitation of PoseNet-based relative pose estimation methods in terms of generalization ability. Code is available at https://github.com/B1ueber2y/TrianFlow. Wang Zhao 0001, Shaohui Liu, Yezhi Shu, Yong-Jin Liu 0001 |
CVPR | 2 |
| 2020 | Multi-Stage Residual Hiding for Image-Into-Audio SteganographyabstractThe widespread application of audio communication technologies has speeded up audio data flowing across the Internet, which made it a popular carrier for covert communication. In this paper, we present a cross-modal steganography method for hiding image content into audio carriers while preserving the perceptual fidelity of the cover audio. In our framework, two multi-stage networks are designed: the first network encodes the decreasing multilevel residual errors inside different audio subsequences with the corresponding stage sub-networks, while the second network decodes the residual errors from the modified carrier with the corresponding stage sub-networks to produce the final revealed results. The multi-stage design of proposed framework not only make the controlling of payload capacity more flexible, but also make hiding easier because of the gradual sparse characteristic of residual errors. Qualitative experiments suggest that modifications to the carrier are unnoticeable by human listeners and that the decoded images are highly intelligible. Wenxue Cui, Shaohui Liu, Feng Jiang 0001, Yongliang Liu, Debin Zhao |
ICASSP | 2 |
| 2020 | Classify and Explain: An Interpretable Convolutional Neural Network For Lung Cancer DiagnosisabstractThe deep network-based computer-aided diagnosis systems have encountered many difficulties in practical applications because of its "black box" feature. The crux of the problem is that these models should be explainable - the model should provide doctors rationales that can explain the diagnosis. In this paper, we present a novel network structure for visually interpretable lung nodule diagnosis. Our proposed model works in an end-to-end manner, consisting of an importance estimation network and a classification network. The former produces a diagnostic visual interpretation for each case, and the latter diagnoses the case. Based on a computed tomography image dataset (LUNA16) on pulmonary nodule, extensive experiments have been conducted, demonstrating that the proposed model can produce state-of-the-art diagnostic visual interpretations compared with all baseline methods. Donghao Gu, Zhaojing Wen, Feng Jiang 0001, Shaohui Liu |
ICASSP | 5 |
| 2020 | An Effective Design to Improve the Efficiency of DPUs on FPGAabstractConvolutional neural networks (CNNs) have been widely used in various complicated problems, such as image classification, objection detection, semantic segmentation. To meet diversified CNN structures, the deep learning processing unit (DPU) is designed as a general accelerator on field programmable gate array (FPGA) to support various CNN layers, such as convolution, pooling, activation, etc. However, low DPU utilization and schedule efficiency appear when DPU used to multitask application completed by CNN models. In this paper, an effective design including multi-core with different size (MCDS) and DPU Plus is proposed to improve the efficiency of DPUs usage from the two dimensions of time and space. Through increasing the number of DPU cores on an FPGA and the utilization of single DPU core, the design of MCDS can effectively improve the overall throughput with restricted on-chip resources. Furthermore, the design of DPU Plus is proposed to improve the schedule efficiency of DPUs through simultaneously implementing DPU with other significant auxiliary modules of the application system on the same FPGA. Finally, a color space conversion module is implemented cooperate to the DPU cores to testify its performance, and the experimen shows that compared with running on the the CPU completely, it achieves16.2x acceleration, and increases the throughput of the entire system by 3.0x. Qingyong Deng, Saiqin Long, Shaohui Liu, Sangyoon Oh 0001 |
ICPADS | 4 |
| 2020 | Obstructive sleep apnea detection using ecg-sensor with convolutional neural networks
Maowei Cheng, Yefu Wang, Shaohui Liu, Zhihong Tian 0001, Feng Jiang 0001 |
Multim. Tools Appl. | 4 |
| 2020 | Multi-Scale Convolutional Neural Network-Based Intra Prediction for Video CodingabstractIn both H.264/AVC and HEVC, the angular prediction is adopted for intra coding, which only exploits the spatial correlation between the current block and its neighboring single line reference. This angular prediction can handle the main directional patterns well, however, lacks the ability to deal with other directions. In this paper, a multi-scale convolutional neural network based intra prediction is proposed to address this problem. Specifically, a predicted block is first generated by the angular prediction, then fed into the proposed network with its neighboring reconstructed L-shape to generate a more accurate predicted block. On one hand, the L-shape of multiple lines provides more reliable reconstructed pixels and more contextual information to get better prediction; on the other hand, the multi-scale feature extraction takes the advantage of the feature maps in different scales to further enhance the prediction. With this multi-scale structure, the L-shape can be used to refine both left-above and right-bottom pixels in the predicted block during the convolution operation. Experimental results demonstrate that compared with HEVC reference software HM 16.9, the proposed intra prediction can achieve an average of 3.4% (up to 5.6%) bitrate saving with all intra configuration. Yang Wang 0048, Xiaopeng Fan 0001, Shaohui Liu, Debin Zhao, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Image Compressed Sensing Using Convolutional Neural NetworkabstractIn the study of compressed sensing (CS), the two main challenges are the design of sampling matrix and the development of reconstruction method. On the one hand, the usually used random sampling matrices (e.g. GRM) are signal independent, which ignore the characteristics of the signal. On the other hand, the state-of-the-art image CS methods (e.g. GSR and MH) achieve quite good performance, however with much higher computational complexity. To deal with the two challenges, we propose an image CS framework using convolutional neural network (dubbed CSNet) that includes a sampling network and a reconstruction network, which are optimized jointly. The sampling network adaptively learns the sampling matrix from the training images, which makes the CS measurements retain more image structural information for better reconstruction. Specifically, three types of sampling matrices are learned, i.e. floating-point matrix, {0,1}-binary matrix, and {-1,+1}-bipolar matrix. The last two matrices are specially designed for easy storage and hardware implementation. The reconstruction network, which contains a linear initial reconstruction network and a non-linear deep reconstruction network, learns an end-to-end mapping between the CS measurements and the reconstructed images. Experimental results demonstrate that CSNet offers state-of-the-art reconstruction quality, while achieving fast running speed. In addition, CSNet with {0,1}-binary matrix, and {-1,+1}-bipolar matrix gets comparable performance with the existing deep learning based CS methods, and outperforms the traditional CS methods. What's more, the experimental results further suggest that the learned sampling matrices can improve the traditional image CS reconstruction methods significantly. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
IEEE Trans. Image Process. | 3 |
| 2020 | VINet: A Visually Interpretable Image Diagnosis NetworkabstractRecently, due to the black box characteristics of deep learning techniques, the deep network-based computer-aided diagnosis (CADx) systems have encountered many difficulties in practical applications. The crux of the problem is that these models should be explainable the model should give doctors rationales that can explain the diagnosis. In this paper, we propose a visually interpretable network (VINet) which can generate diagnostic visual interpretations while making accurate diagnoses. VINet is an end-to-end model consisting of an importance estimation network and a classification network. The former produces a diagnostic visual interpretation for each case, and the classifier diagnoses the case. In the classifier, by exploring the information in the diagnostic visual interpretation, the irrelevant information in the feature maps is eliminated by our proposed feature destruction process. This allows the classification network to concentrate on the important features and use them as the primary references for classification. Through a joint optimization of higher classification accuracy and eliminating as many irrelevant features as possible, a precise, fine-grained diagnostic visual interpretation, along with an accurate diagnosis, can be produced by our proposed network simultaneously. Based on a computed tomography image dataset (LUNA16) on pulmonary nodule, extensive experiments have been conducted, demonstrating that the proposed VINet can produce state-of-the-art diagnostic visual interpretations compared with all baseline methods. Donghao Gu, Feng Jiang 0001, Zhaojing Wen, Shaohui Liu, Wuzhen Shi, Guangming Lu 0001, Changsheng Zhou |
IEEE Trans. Multim. | 5 |
| 2020 | Delving Deeper in Drone-Based Person Re-Id by Employing Deep Decision Forest and Attributes FusionabstractDeep learning has revolutionized the field of computer vision and image processing. Its ability to extract the compact image representation has taken the person re-identification (re-id) problem to a new level. However, in most cases, researchers are focused on developing new approaches to extract more fruitful image representation and use it in the re-id task. The extra information about images is rarely taken into account because the traditional person re-id datasets usually do not have it. Nevertheless, the research in multimodal machine learning has demonstrated that the utilization of the information from different sources leads to better performance. In this work, we demonstrate how a person re-id problem can benefit from the utilization of multimodal data. We have used the UAV drone to collect and label the new person re-id dataset, which is composed of pedestrian images and its attributes. We have manually annotated this dataset with attributes, and in contrast to the recent research, we do not use the deep network to classify them. Instead, we employ the continuous bag-of-words model to extract the word embeddings from text descriptions and fuse it with features extracted from images. Then the deep neural decision forest is used for pedestrians classification. The extensive experiments on the collected dataset demonstrate the effectiveness of the proposed model. Aleksei Grigorev, Shaohui Liu, Zhihong Tian 0001, Jianxin Xiong, Seungmin Rho, Feng Jiang 0001 |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2019 | Normalized DiversificationabstractGenerating diverse yet specific data is the goal of the generative adversarial network (GAN), but it suffers from the problem of mode collapse. We introduce the concept of normalized diversity which force the model to preserve the normalized pairwise distance between the sparse samples from a latent parametric distribution and their corresponding high-dimensional outputs. The normalized diversification aims to unfold the manifold of unknown topology and non-uniform distribution, which leads to safe interpolation between valid latent variables. By alternating the maximization over the pairwise distance and updating the total distance (normalizer), we encourage the model to actively explore in the high-dimensional output space. We demonstrate that by combining the normalized diversity loss and the adversarial loss, we generate diverse data without suffering from mode collapsing. Experimental results show that our method achieves consistent improvement on unsupervised image generation, conditional image generation and hand pose estimation over strong baselines. Shaohui Liu, Jianqiao Wangni, Jianbo Shi |
CVPR | 1 |
| 2019 | Scalable Convolutional Neural Network for Image Compressed SensingabstractRecently, deep learning based image Compressed Sensing (CS) methods have been proposed and demonstrated superior reconstruction quality with low computational complexity. However, the existing deep learning based image CS methods need to train different models for different sampling ratios, which increases the complexity of the encoder and decoder. In this paper, we propose a scalable convolutional neural network (dubbed SCSNet) to achieve scalable sampling and scalable reconstruction with only one model. Specifically, SCSNet provides both coarse and fine granular scalability. For coarse granular scalability, SCSNet is designed as a single sampling matrix plus a hierarchical reconstruction network that contains a base layer plus multiple enhancement layers. The base layer provides the basic reconstruction quality, while the enhancement layers reference the lower reconstruction layers and gradually improve the reconstruction quality. For fine granular scalability, SCSNet achieves sampling and reconstruction at any sampling ratio by using a greedy method to select the measurement bases. Compared with the existing deep learning based image CS methods, SCSNet achieves scalable sampling and quality scalable reconstruction at any sampling ratio with only one model. Experimental results demonstrate that SCSNet has the state-of-the-art performance while maintaining a comparable running speed with the existing deep learning based image CS methods. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
CVPR | 3 |
| 2019 | Conditional Single-View Shape Generation for Multi-View Stereo ReconstructionabstractIn this paper, we present a new perspective towards image-based shape generation. Most existing deep learning based shape reconstruction methods employ a single-view deterministic model which is sometimes insufficient to determine a single groundtruth shape because the back part is occluded. In this work, we first introduce a conditional generative network to model the uncertainty for single-view reconstruction. Then, we formulate the task of multi-view reconstruction as taking the intersection of the predicted shape spaces on each single image. We design new differentiable guidance including the front constraint, the diversity constraint, and the consistency loss to enable effective single-view conditional generation and multi-view synthesis. Experimental results and ablation studies show that our proposed approach outperforms state-of-the-art methods on 3D reconstruction test error and demonstrates its generalization ability on real world data. Yi Wei 0003, Shaohui Liu, Wang Zhao 0001, Jiwen Lu |
CVPR | 2 |
| 2019 | Image Inpainting With Learnable Bidirectional Attention MapsabstractMost convolutional network (CNN)-based inpainting methods adopt standard convolution to indistinguishably treat valid pixels and holes, making them limited in handling irregular holes and more likely to generate inpainting results with color discrepancy and blurriness. Partial convolution has been suggested to address this issue, but it adopts handcrafted feature re-normalization, and only considers forward mask-updating. In this paper, we present a learnable attention map module for learning feature re-normalization and mask-updating in an end-to-end manner, which is effective in adapting to irregular holes and propagation of convolution layers. Furthermore, learnable reverse attention maps are introduced to allow the decoder of U-Net to concentrate on filling in irregular holes instead of reconstructing both holes and known regions, resulting in our learnable bidirectional attention maps. Qualitative and quantitative experiments show that our method performs favorably against state-of-the-arts in generating sharper, more coherent and visually plausible inpainting results. The source code and pre-trained models will be available at: https://github.com/Vious/LBAM_inpainting/. Chaohao Xie, Shaohui Liu, Chao Li 0034, Ming-Ming Cheng, Wangmeng Zuo, Xiao Liu 0022, Shilei Wen, Errui Ding |
ICCV | 2 |
| 2019 | RepPoints: Point Set Representation for Object DetectionabstractModern object detectors rely heavily on rectangular bounding boxes, such as anchors, proposals and the final predictions, to represent objects at various recognition stages. The bounding box is convenient to use but provides only a coarse localization of objects and leads to a correspondingly coarse extraction of object features. In this paper, we present RepPoints (representative points), a new finer representation of objects as a set of sample points useful for both localization and recognition. Given ground truth localization and recognition targets for training, RepPoints learn to automatically arrange themselves in a manner that bounds the spatial extent of an object and indicates semantically significant local areas. They furthermore do not require the use of anchors to sample a space of bounding boxes. We show that an anchor-free object detector based on RepPoints can be as effective as the state-of-the-art anchor-based detection methods, with 46.5 AP and 67.4 AP50 on the COCO test-dev detection benchmark, using ResNet-101 model. Code is available at https://github.com/microsoft/RepPoints. Ze Yang 0003, Shaohui Liu, Han Hu 0001, Liwei Wang 0001, Stephen Lin 0001 |
ICCV | 2 |
| 2019 | Continuous Bidirectional Optical Flow for Video Frame Sequence InterpolationabstractExisting optical flow-based frame interpolation frameworks usually suffer from two problems. First, it is difficult to accurately estimate both large motion and fine motion in the optical flow estimation stage. Second, the hole problem and occlusion problem cannot be efficiently solved in the pixel synthesis step. In this paper, we propose a novel optical flowbased frame interpolation framework, which consists of two submodules: optical flow network and pixel synthesis network. In the optical flow network, we estimate bidirectional optical flow sequences iteratively, which makes full use of the continuity of motion and therefore improves the accuracy of the optical flow estimation. Besides, a novel multi-scale architecture is developed to capture finer motions. In the pixel synthesis network, we fuse the statistical information generated during forward warping to solve the hole problem and the occlusion problem. Experimental results demonstrate that the proposed method achieves superior performance compared to state-of-the-art methods. Donghao Gu, Zhaojing Wen, Wenxue Cui, Rui Wang 0093, Feng Jiang 0001, Shaohui Liu |
ICME | 6 |
| 2019 | A Video Post-Filter Deblocking Method Based on Temporal Boosting Residual NetworksabstractBlock-based hybrid coding is widely used in video compression. As the bit rate decreases, the quantization becomes lossy resulting in unacceptable blocking artifacts. Although advanced video encoder with loop filter has achieved promising results, the problem still remains unsolved. Most of the previous solutions regard a video sequence as a group of independent frames, without considering their temporal relationships. To address the above issue, this paper designs a temporal boosting residual network aiming to integrate the temporal information and the structural information. And the information is represented as the residual values of previous frames which are extracted by the deep residual network. The entire framework is a cascading architecture that implies a coarse-to-fine processing. The proposed solution does not modify any module of the codec, it only takes the lossy frames as input, and outputs the enhanced frames, which is a typical end-to-end mapping. Experimental results validate that our framework has achieved 0.6-1.0 dB improvement on average based on HEVC software and outperformed the state-of-the-art methods in both objective and perceptual quality. Shaohui Liu, Feng Jiang 0001, Xiaoshuai Sun, Yongliang Liu |
ICME | 2 |
| 2019 | Medical image denoising using convolutional neural network: a residual learning approach
Worku Jifara Sori, Feng Jiang 0001, Seungmin Rho, Maowei Cheng, Shaohui Liu |
J. Supercomput. | 5 |
| 2018 | Multi-Scale Deep Networks for Image Compressed SensingabstractAs a successful deep model applied in image compressed sensing, the Compressed Sensing Network (CSNet) has demonstrated superior performance to the previous handcrafted models in both running speed and reconstruction quality. However, CSNet trains different models for different sampling rates that hinders it from practical usage since too many models need to store. In this paper, we propose multi-scale deep network for image compressed sensing. We still use a sampling network to learn the sampling operator and implement the compressed sampling process. Given the compressed measurements, the reconstruction network directly maps them to the desired reconstructed images. There are three main differences in comparison with CSNet. Firstly, this paper proposes to use an unified deep reconstruction network for all sampling rates that decreases large amount of storage requirements. Secondly, we redesign a better deep reconstruction network using the popular residual learning technology. Finally, we investigate an image local smooth prior based loss function to enhance image structural information. Extensive experimental results show that the proposed multi-scale deep network based image compressed sensing method outperforms many other state-of-the-art methods. Wuzhen Shi, Feng Jiang 0001, Shaohui Liu, Debin Zhao |
ICIP | 3 |
| 2018 | Classification Guided Deep Convolutional Network for Compressed SensingabstractCompressed Sensing (CS) has been successfully applied to image compression in the past few years. However, there are still several challenges that restrict its applications in practice including large memory requirement and unsatisfactory reconstruction performance. To address these challenges, in this paper, we propose a classification guided deep convolutional network for image compressed sensing (CCSNet), which includes a sampling sub-network and a reconstruction sub-network. In the sampling sub-network, multiple convolutional layers are used to sample the original image, which significantly reduces the parameters of the sampling matrix while causes performance degradation moderately compared against existing convolution based sampling methods. In the reconstruction sub-network, a novel two-branch architecture is proposed to improve the adaptability of the model to various textures in natural images. The first branch, named the classification branch, is to classify the sampled measurements of the original image to one of the predefined textural classes. The second branch, named the reconstruction branch, consists of multiple sub-branches, which are responsible for reconstructing the original images belonging to the corresponding textural classes. By jointly utilizing two sub-networks, the entire network can be trained in the form of end-to-end metric with a joint loss function. Experimental results demonstrate that the proposed method provides a significant quality improvement in terms of PSNR compared against state-of-the-art methods. Wenxue Cui, Shaohui Liu, Shengping Zhang, Yashu Liu 0003, Heyao Xu, Xinwei Gao, Feng Jiang 0001, Debin Zhao |
ICPR | 2 |
| 2018 | A hybrid framework of data hiding and encryption in H.264/SVC
Shaohui Liu, Seungmin Rho, Worku Jifara Sori, Feng Jiang 0001 |
Discret. Appl. Math. | 1 |
| 2018 | Hyperspectral classification based on spectral-spatial convolutional neural networks
Feng Jiang 0001, Chifu Yang, Seungmin Rho, Weizheng Shen, Shaohui Liu |
Eng. Appl. Artif. Intell. | 6 |
| 2018 | ACDIN: Bridging the gap between artificial and real bearing damages for bearing fault diagnosis
Yuanhang Chen, Gaoliang Peng, Chaohao Xie, Wei Zhang 0185, Chuanhao Li 0002, Shaohui Liu |
Neurocomputing | 6 |
| 2018 | Feature-preserving mesh denoising based on guided normal filtering
Shaohui Liu, Seungmin Rho, Feng Jiang 0001 |
Multim. Tools Appl. | 1 |
| 2018 | An End-to-End Compression Framework Based on Convolutional Neural NetworksabstractDeep learning, e.g., convolutional neural networks (CNNs), has achieved great success in image processing and computer vision especially in high-level vision applications, such as recognition and understanding. However, it is rarely used to solve low-level vision problems such as image compression studied in this paper. Here, we move forward a step and propose a novel compression framework based on CNNs. To achieve high-quality image compression at low bit rates, two CNNs are seamlessly integrated into an end-to-end compression framework. The first CNN, named compact convolutional neural network (ComCNN), learns an optimal compact representation from an input image, which preserves the structural information and is then encoded using an image codec (e.g., JPEG, JPEG2000, or BPG). The second CNN, named reconstruction convolutional neural network (RecCNN), is used to reconstruct the decoded image with high quality in the decoding end. To make two CNNs effectively collaborate, we develop a unified end-to-end learning algorithm to simultaneously learn ComCNN and RecCNN, which facilitates the accurate reconstruction of the decoded image using RecCNN. Such a design also makes the proposed compression framework compatible with existing image coding standards. Experimental results validate that the proposed compression framework greatly outperforms several compression frameworks that use existing image coding standards with the state-of-the-art deblocking or denoising post-processing methods. Feng Jiang 0001, Wen Tao, Shaohui Liu, Jie Ren 0016, Xun Guo 0002, Debin Zhao |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2017 | How many zero crossings? A method for structure-texture image decomposition
Xiaolei Jiang, Hongxun Yao, Shaohui Liu |
Comput. Graph. | 3 |
| 2017 | A Comprehensive Survey on Sampling-Based Image MattingabstractAbstract Sampling‐based image matting is currently playing a significant role and showing great further development potentials in image matting. However, the consequent survey articles and detailed classifications are still rare in the field of corresponding research. Furthermore, besides sampling strategies, most of the sampling‐based matting algorithms apply additional operations which actually conceal their real sampling performances. To inspire further improvements and new work, this paper makes a comprehensive survey on sampling‐based matting in the following five aspects: (i) Only the sampling step is initially preserved in the matting process to generate the final alpha results and make comparisons. (ii) Four basic categories including eight detailed classes for sampling‐based matting are presented, which are combined to generate the common sampling‐based matting algorithms. (iii) Each category including two classes is analysed and experimented independently on their advantages and disadvantages. (iv) Additional operations, including sampling weight, settling manner, complement and pre‐ and post‐processing, are sequentially analysed and added into sampling. Besides, the result and effect of each operation are also presented. (v) A pure sampling comparison framework is strongly recommended in future work. Guilin Yao, Zhijie Zhao, Shaohui Liu |
Comput. Graph. Forum | 3 |
| 2017 | Structured entropy of primitive: big data-based stereoscopic image quality assessmentabstractThe ultimate receiver of image and video is human visual system (HVS). It is an important problem in the domain of image and video processing that how to establish visual information representation model meeting the HVS perception property. In this study, authors give theory analysis and experiment results to prove that l_1 norm‐based entropy of primitive (EoP) is superior to the l_0 norm‐based EoP for the monocular cue in image quality assessment. By developing the concept of mutual information of primitive (MIP) as the binocular cue, an l_1 EoP‐based stereoscopic image quality assessment metric is proposed. With EoP as monocular cue and MIP as binocular cue, the relative entropy between the original stereoscopic image and the distorted one is explored to predict the quality score with support vector regression. To avoid destroying image's structured information, the structured EoP (SEoP) is further explored to measure the stereoscopic image information. Extensive experimental results demonstrate that the stereoscopic image quality assessment algorithm with SEoP as monocular cue and MIP as binocular cue outperforms many state‐of‐the‐art ones. Chifu Yang, Seungmin Rho, Shaohui Liu, Feng Jiang 0001 |
IET Image Process. | 4 |
| 2017 | Depth estimation from single monocular images using deep hybrid network
Aleksei Grigorev, Feng Jiang 0001, Seungmin Rho, Worku Jifara Sori, Shaohui Liu, Sergey V. Sai |
Multim. Tools Appl. | 5 |
| 2017 | 3D visual saliency detection model with generated disparity map
Debin Zhao, Shaohui Liu, Xiaopeng Fan 0001 |
Multim. Tools Appl. | 3 |
| 2017 | Hyperspectral image compression based on online learning spectral features dictionary
Worku Jifara Sori, Feng Jiang 0001, Huapeng Wang, Aleksei Grigorev, Shaohui Liu |
Multim. Tools Appl. | 7 |
| 2016 | Hierarchical frame based spatial-temporal recovery for video compressive sensing coding
Xinwei Gao, Feng Jiang 0001, Shaohui Liu, Wenbin Che, Xiaopeng Fan 0001, Debin Zhao |
Neurocomputing | 3 |
| 2016 | 3D object retrieval with multi-feature collaboration and bipartite graph matching
Yan Zhang 0109, Feng Jiang 0001, Seungmin Rho, Shaohui Liu, Debin Zhao, Rongrong Ji |
Neurocomputing | 4 |
| 2015 | View-based 3D object retrieval via multi-modal graph learning
Sicheng Zhao, Hongxun Yao, Yanhao Zhang 0001, Yasi Wang, Shaohui Liu |
Signal Process. | 5 |
| 2015 | A game theory-based block image compression method in encryption domain
Shaohui Liu, Anand Paul 0001, Guochao Zhang, Gwanggil Jeon |
J. Supercomput. | 1 |
| 2013 | Improved total variation based image compressive sensing recovery by nonlocal regularizationabstractRecently, total variation (TV) based minimization algorithms have achieved great success in compressive sensing (CS) recovery for natural images due to its virtue of preserving edges. However, the use of TV is not able to recover the fine details and textures, and often suffers from undesirable staircase artifact. To reduce these effects, this paper presents an improved TV based image CS recovery algorithm by introducing a new nonlocal regularization constraint into CS optimization problem. The nonlocal regularization is built on the well known nonlocal means (NLM) filtering and takes advantage of self-similarity in images, which helps to suppress the staircase effect and restore the fine details. Furthermore, an efficient augmented Lagrangian based algorithm is developed to solve the above combined TV and nonlocal regularization constrained problem. Experimental results demonstrate that the proposed algorithm achieves significant performance improvements over the state-of-the-art TV based algorithm in both PSNR and visual perception. Jian Zhang 0018, Shaohui Liu, Ruiqin Xiong, Siwei Ma 0001, Debin Zhao |
ISCAS | 2 |
| 2013 | Natural images scale invariance and high-fidelity image restorationabstractOne of the most striking properties of natural image statistics is their scale invariance. Intuitively, a natural image always contains the same contents of different scales and dually the same contents of same scale exist throughout scales of the image. Different from the previous scale invariance related work decomposing an image to its local band-pass filter components, this paper seeks a general model of the natural image paths distribution to describe the scale invariance in the visual world and then a novel strategy for high-fidelity image restoration is presented by characterizing nonlocal self-similarity of natural images throughout scales in a unified statistical manner, which offers a powerful mechanism of combining natural images scale invariance and nonlocal self-similarity simultaneously to ensure a more reliable and robust estimation. Extensive experiments on image restoration from partial random samples manifest that the proposed algorithm achieves significant performance improvements over the current state-of-the-art schemes. Feng Jiang 0001, Shaohui Liu, Debin Zhao |
VCIP | 3 |
| 2013 | An improved image compression scheme with an adaptive parameters set in encrypted domainabstractA growing societal awareness about privacy and security push the development of signal processing techniques in the encrypted domain. Data compression in encrypted domain attracts much attention recently years due to its avoiding the leakage of data source during compression. This paper proposes an improved block-by-block compression scheme of encrypted image with flexible compression ratio. The original image is encrypted by permuting the blocks of the image and then permuting the pixels in the blocks. In the compression, pixels chosen randomly used as reference information, and remaining pixels are compressed by coset code. At the decoder side, side information (SI) which is generated by combining correlation among blocks and image restoration from partial random samples (IRPRS) is utilized to assist the decompression. Moreover, an adaptive system parameters selection method is also given in this paper. The experimental results show that the proposed method can achieve a better reconstructed result compared with the earlier method. Guochao Zhang, Shaohui Liu, Feng Jiang 0001, Debin Zhao, Wen Gao 0001 |
VCIP | 2 |
| 2013 | Entropy of primitive: A top-down methodology for evaluating the perceptual visual informationabstractIn this paper, we aim at evaluating the perceptual visual information based on a novel top-down methodology: entropy of primitive (EoP). The EoP is determined by the distribution of the atoms in describing an image, and is demonstrated to exhibit closely correlation with the perceptual image quality. Based on the visual information evaluation, we further demonstrate that the EoP is effective in predicting the perceptual lossless of natural images. Inspired by this observation, in order to distinguish whether the loss of input signal is visual noticeable to human visual system (HVS), we introduce the EoP based perceptual lossless profile (PLP). Extensive experiments verify that, the proposed EoP based perceptual lossless profile can efficiently measure the minimum noticeable visual information distortion and achieve better performance compared to the-state-of-the-art just-noticeable difference (JND) profile. Xiang Zhang 0004, Shiqi Wang 0001, Siwei Ma 0001, Shaohui Liu, Wen Gao 0001 |
VCIP | 4 |
| 2013 | Robust visual tracking based on online learning sparse representation
Shengping Zhang, Hongxun Yao, Huiyu Zhou 0001, Xin Sun 0003, Shaohui Liu |
Neurocomputing | 5 |
| 2012 | A fast multiview video transcoder for bitrate reductionabstractVideo transcoding is an efficient way to reduce the bitrate or convert the format of the original video stream to meet the requirements of different applications and various channel capacity. In this paper, we propose a fast multiview video transcoder (MVT) for bitrate reduction. Different from the H.264 transcoder, the inter-view prediction information in the input video stream is utilized to reduce the complexity of transcoding. Besides, we also utilize the mode and selected reference frame information in original stream to accelerate RD optimization calculations. Experimental results show that the proposed transcoder can achieve significant computation reduction while maintaining close RD performance compared to the fully decode and re-encode transcoder (FDET). Xiaopeng Fan 0001, Shaohui Liu, Yan Liu 0014, Debin Zhao, Wen Gao 0001 |
VCIP | 3 |
| 2012 | Viewpoint-independent hand gesture recognition systemabstractIn this paper, we creatively present a viewpoint-free hand gesture recognition system based on Kinect sensor. Through depth image, we build Point Clouds of user. Then, we estimate the current optimal viewpoint, i.e., the front, and project Point Clouds to that direction. Through that process we in great extent overcome the viewpoint-dependency issue. To match hand types, we propose an improved shape context to describe each hand gesture and use the Hungarian algorithm to calculate match degree. Our method is quite straightforward, however the experimental results prove that by this means gestures can be recognized independent of viewpoints with great accuracy. Besides, it is fast and robust, thus can be applied under various realistic scenarios in realtime. Feng Jiang 0001, Debin Zhao, Shaohui Liu, Wen Gao 0001 |
VCIP | 4 |
| 2012 | Robust Visual Tracking Using an Effective Appearance Model Based on Sparse CodingabstractIntelligent video surveillance is currently one of the most active research topics in computer vision, especially when facing the explosion of video data captured by a large number of surveillance cameras. As a key step of an intelligent surveillance system, robust visual tracking is very challenging for computer vision. However, it is a basic functionality of the human visual system (HVS). Psychophysical findings have shown that the receptive fields of simple cells in the visual cortex can be characterized as being spatially localized, oriented, and bandpass, and it forms a sparse, distributed representation of natural images. In this article, motivated by these findings, we propose an effective appearance model based on sparse coding and apply it in visual tracking. Specifically, we consider the responses of general basis functions extracted by independent component analysis on a large set of natural image patches as features and model the appearance of the tracked target as the probability distribution of these features. In order to make the tracker more robust to partial occlusion, camouflage environments, pose changes, and illumination changes, we further select features that are related to the target based on an entropy-gain criterion and ignore those that are not. The target is finally represented by the probability distribution of those related features. The target search is performed by minimizing the Matusita distance between the distributions of the target model and a candidate using Newton-style iterations. The experimental results validate that the proposed method is more robust and effective than three state-of-the-art methods. Shengping Zhang, Hongxun Yao, Xin Sun 0003, Shaohui Liu |
ACM Trans. Intell. Syst. Technol. | 4 |
| 2011 | Saliency Detection: A Self-Adaption Sparse Representation ApproachabstractSaliency detection is essential to visual attention modelling and various computer vision tasks. Representation and measurement are two important issues for saliency models. Good representation and reasonable measurement are both critical issues in modelling visual saliency mechanism. For every input image, we obtain a self-adaptive dictionary that describes the image content effectively and image prior that forces sparsity in every location in the image using the K-SVD algorithm. For saliency measurement, background firing rate (BFR) is defined for each sparse features and it is followed by feature activation rate (FAR) computation to measure the bottom-up visual saliency. Gaoxiang Zhang, Feng Jiang 0001, Debin Zhao, Xiaoshuai Sun, Shaohui Liu |
ICIG | 5 |
| 2011 | Visual attention based image quality assessmentabstractInspired by the success of structural similarity index (SSIM), some image quality assessment (IQA) methods have been developed recently. To achieve better performance, this paper proposes a new visual attention (VA) model that combines saliency based VA and visual importance based VA, under the assumptions that humans often pay more attention to the regions with important content in the beginning of evaluating a given image and then the regions with poor quality. Then the proposed VA model is incorporated into SSIM. The experiments on LIVE database and TID2008 database demonstrate its improvements over the latest state-of-the-art IQA methods and the information content weighted SSIM measure (IW-SSIM). Anan Guo, Debin Zhao, Shaohui Liu, Xiaopeng Fan 0001, Wen Gao 0001 |
ICIP | 3 |
| 2010 | Visual tracking via weakly supervised learning from multiple imperfect oraclesabstractLong-term persistent tracking in ever-changing environments is a challenging task, which often requires addressing difficult object appearance update problems. To solve them, most top-performing methods rely on online learning-based algorithms. Unfortunately, one inherent problem of online learning-based trackers is drift, a gradual adaptation of the tracker to non-targets. To alleviate this problem, we consider visual tracking in a novel weakly supervised learning scenario where (possibly noisy) labels but no ground truth are provided by multiple imperfect oracles (i.e., trackers), some of which may be mediocre. A probabilistic approach is proposed to simultaneously infer the most likely object position and the accuracy of each tracker. Moreover, an online evaluation strategy of trackers and a heuristic training data selection scheme are adopted to make the inference more effective and fast. Consequently, the proposed method can avoid the pitfalls of purely single tracking approaches and get reliable labeled samples to incrementally update each tracker (if it is an appearance-adaptive tracker) to capture the appearance changes. Extensive comparing experiments on challenging video sequences demonstrate the robustness and effectiveness of the proposed method. Bineng Zhong 0001, Hongxun Yao, Sheng Chen 0007, Rongrong Ji, Xiao-Tong Yuan, Shaohui Liu, Wen Gao 0001 |
CVPR | 6 |
| 2010 | Robust visual tracking using feature-based visual attentionabstractPsychophysical findings have shown that human vision system has an ability to improve target search by enhancing the representation of image components that are related to the searched target, which is the so-called feature-based visual attention. In this paper, motivated by these psychophysical findings, we propose a robust visual tracking algorithm by simulating such feature-based visual attention. Specially, we consider the general sparse basis functions extracted on a large set of natural image patches as features. We define that a feature is related to the target when succeeding activations of that feature cannot increase system's entropy. The target is finally represented by the probability distribution of those related features. The target search is performed by minimizing the Matusita distance measure between the distributions of the target model and candidate using Newton-style iterations. The experimental results verify that the proposed method is more robust and effective than widely used mean shift based methods. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICASSP | 3 |
| 2010 | Robust background modeling via standard variance featureabstractIn this paper, a novel standard variance feature is proposed for background modeling in dynamic scenes involving waving trees and ripples in water. The standard variance feature is the standard variance of a set of pixels' feature values, which captures mainly co-occurrence statistics of neighboring pixels in an image patch. The background modeling method based on standard variance feature includes two main components. First, we divide image into patches and represent each image patch as a standard variance feature. Then, assuming that standard variance feature fits a mixture of Gaussians distribution, we use mixture of Gaussians models to model it. Experimental results on several challenging video sequences demonstrate the effectiveness of our method. Bineng Zhong 0001, Hongxun Yao, Shaohui Liu |
ICASSP | 3 |
| 2010 | An Image Data Hiding Method Using Pixel-Based JND Model
Shaohui Liu, Feng Jiang 0001, Hongxun Yao, Debin Zhao |
ICIC (3) | 1 |
| 2010 | Visual saliency as sequential eye fixation probabilityabstractHuman vision system acquires essential information from the environment by sequentially sampling visual contents at important locations under the control of selective attention mechanism. We propose that bottom-up saliency is not based on global statistics but on information sampled at prior eye fixations. Our model calculates visual saliency using sequential eye fixation probability. However, the proposed model needs fixation priors, which are hard to simulate given current fixation data and experimental conditions. An approximation is proposed to generate a single saliency map by fusing all possible conditions of fixation prior. Our method outperforms all state-of-the-art models in predicting eye fixations, and shows reasonable response to various psychological patterns. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xianming Liu 0005, Shaohui Liu |
ICIP | 6 |
| 2010 | Saliency detection based on short-term sparse representationabstractRepresentation and measurement are two important issues for saliency models. Different with previous works that learnt sparse features from large scale natural statistics, we propose to learn features from short-term statistics of single images. For saliency measurement, we define background firing rate (BFR) for each sparse feature, and then we propose to use feature activation rate (FAR) to measure the bottom-up visual saliency. The proposed FAR measure is biological plausible and easy to compute, also with satisfied performance. Experiments on human eye fixations and psychological patterns demonstrate the effectiveness and robustness of our proposed method. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Pengfei Xu 0001, Xianming Liu 0005, Shaohui Liu |
ICIP | 6 |
| 2010 | A steganography strategy based on equivalence partitions of hiding unitsabstractThis paper designs a novel hiding strategy based on an equivalence relation, which can remarkably enhance the quality of stego image without sacrificing the security and capacity of original steganography schemes. According to a constructed equivalence relation based on the capacity of hiding units, all hiding units can be partitioned into equivalence classes. Following that, the hiding procedure is performed in predefined order in equivalence classes as the traditional steganography scheme. Because of considering the relation between the length of message and capacity, the performance of the hiding method using proposed hiding strategy outperforms the original approaches when embedding the same length message. Experimental results indicate that the gain from the proposed strategy over existing hiding schemes can reach up to 4.0 dB. Shaohui Liu, Hongxun Yao, Shengping Zhang, Wen Gao 0001 |
ICME | 1 |
| 2010 | Fovea based image quality assessmentabstractHumans are the ultimate receivers of the visual information contained in an image, so the reasonable method of image quality assessment (IQA) should follow the properties of the human visual system (HVS). In recent years, IQA methods based on HVS-models are slowly replacing classical schemes, such as mean squared error (MSE) and Peak Signal-to-Noise Ratio (PSNR). IQA-structural similarity (SSIM) regarded as one of the most popular HVS-based methods of full reference IQA has apparent improvements in performance compared with traditional metrics in nature, however, it performs not very well when the images' structure is destroyed seriously or masked by noise. In this paper, a new efficient fovea based structure similarity image quality assessment (FSSIM) is proposed. It enlarges the distortions in the concerned positions adaptively and changes the importances of the three components in SSIM. FSSIM predicts the quality of an image through three steps. First, it computes the luminance, contrast and structure comparison terms; second, it computes the saliency map by extracting the fovea information from the reference image with the features of HVS; third, it pools the above three terms according to the processed saliency map. Finally, a commonly experimental database LIVE IQA is used for evaluating the performance of the FSSIM. Experimental results indicate that the consistency and relevance between FSSIM and mean opinion score (MOS) are both better than SSIM and PSNR clearly. Anan Guo, Debin Zhao, Shaohui Liu, Guangyao Cao |
VCIP | 3 |
| 2010 | Robust object tracking based on sparse representationabstractIn this paper, we propose a novel and robust object tracking algorithm based on sparse representation. Object tracking is formulated as a object recognition problem rather than a traditional search problem. All target candidates are considered as training samples and the target template is represented as a linear combination of all training samples. The combination coefficients are obtained by solving for the minimum l1-norm solution. The final tracking result is the target candidate associated with the non-zero coefficient. Experimental results on two challenging test sequences show that the proposed method is more effective than the widely used mean shift tracker. Shengping Zhang, Hongxun Yao, Xin Sun 0003, Shaohui Liu |
VCIP | 4 |
| 2010 | Partial occlusion robust object tracking using an effective appearance modelabstractPartial occlusion is one of the most challenging difficulties for object tracking. In this paper, we present an approach to address this problem by using an effective appearance model which has two innovations. First, in contrast to widely used color histogram that models the appearance of an object using only color information, we assert that both color and texture are important cues for tracking, especially in the presence of complex background. We thus propose a novel local descriptor, named local color texture pattern (LCTP), to model the appearance of the object with color and texture information simultaneously. Second, global color histogram completely ignores the spatial layout information of an object and are sensitive to partial occlusion. In this work, we overcome this limitation based on a block-dividing way: 1) divide target into multiple blocks and then represent each block with LCTP histogram, 2) with a selectivity strategy, we select blocks that are not occluded and then combine similarities of those selected blocks to obtain final similarity measure. Experimental results demonstrate that the proposed method is more robust to partial occlusion than two state-of-the-art algorithms. Shengping Zhang, Hongxun Yao, Shaohui Liu |
VCIP | 3 |
| 2009 | Universal Steganalysis Based on Statistical Models Using Reorganization of Block-based DCT CoefficientsabstractThe goal of steganography is to hide information into media without disclosing the fact of existing communication. Currently, stganography such as least significant bit (LSB), quantization index modulation (QIM) and spread spectrum (SS), has become increasingly widespread. Steganalysis as a counterpart of stganography is to detect the presence of it. In this paper, we present a new universal steganalysis method based on statistical models of the imagepsilas discrete cosine transform (DCT) coefficients. In fact, the block-based DCT by proper reorganization of its coefficients can have similar characteristics to wavelet transforms. The presented universal steganalysis method utilizes these characteristics to build statistical models of the image and its prediction-error image. Features extracted from the re-organization DCT blocks of host images and theirs prediction-error images and features extracted from steg images and theirs prediction-error images are used to train the SVM classifier. In the testing, features from those potential images are inputted the trained-well classifier to determine where the potential images are stego images or not. The experiments have shown that the proposed method outperforms in general prior-arts of steganalysis methods based on wavelet transform domain. Shaohui Liu, Hongxun Yao, Debin Zhao |
IAS | 1 |
| 2009 | Local Spatial Co-occurrence for Background Subtraction via Adaptive Binned Kernel Estimation
Bineng Zhong 0001, Shaohui Liu, Hongxun Yao |
ACCV (3) | 2 |
| 2009 | Multl-resolution background subtraction for dynamic scenesabstractDynamic scenes (e.g. waving trees, ripples in water, illumination changes, camera jitters etc.) challenge many traditional background subtraction methods. In this paper, we present a novel background subtraction approach for dynamic scenes, in which the background is modeled in a multi-resolution framework. First, for each level of the pyramid, we run an independent mixture of Gaussians Models (GMM) that outputs a background subtraction map. Second, these background subtraction maps are combined via AND operator to finally get a more robust and accurate background subtraction map. This is a natural fusion because the original resolution and low resolution images have complementary strengths, which original resolution image contains rich information and low resolution image is insensitive to the noises and the small movement of dynamic scene. Experimental result shows that this real-time algorithm is able to detect moving objects accurately even in dynamic scenes. Bineng Zhong 0001, Shaohui Liu, Hongxun Yao, Baochang Zhang 0001 |
ICIP | 2 |
| 2009 | Neighboring Image Patches Embedding for background modelingabstractWe present a novel feature extraction framework, Neighboring Image Patches Embedding (NIPE), for robust and efficient background modeling. We divide image into patches and represent each image patch as a NIPE vector. Then, the background model of each image patch is constructed as a group of weighted adaptive NIPE vectors. The NIPE feature vector, whose components are similarities between current image patch and its neighbors, describes mainly the mutual relationship between neighboring patches. Since neighboring image patches tend to be similarly affected by environmental effects (e.g., dynamic background), the NIPE vectors are more robust in these conditions comparing with the conventional method. Experimental results demonstrate the efficiency and effectiveness of our proposed NIPE method. Bineng Zhong 0001, Hongxun Yao, Shaohui Liu |
ICIP | 3 |
| 2009 | Spatial-temporal nonparametric background subtraction in dynamic scenesabstractTraditional background subtraction methods model only temporal variation of each pixel. However, there is also spatial variation in real word due to dynamic background such as waving trees, spouting fountain and camera jitters, which causes the significant performance degradation of traditional methods. In this paper, a novel spatial-temporal nonparametric background subtraction approach (STNBS) is proposed to effectively handle dynamic background by modeling the spatial and temporal variations simultaneously. Specially, for each pixel in an image, we adaptively maintain a sample consisting of pixels observed in previous frames. At current frame, for a particular pixel, the proposed method estimates the probabilities of observing this pixel based on samples of its neighboring pixels. The pixel is labeled as background if one of these estimated probabilities is larger than a fixed threshold. All samples are adaptively updated over time. Experimental results on several challenging sequences show that the proposed method achieves the best performance than two state-of-the-art algorithms. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICME | 3 |
| 2009 | Photo assessment based on computational visual attention modelabstractIt is difficult to be satisfied for automatic photo assessment using only low level visual features such as brightness, lighting, hue, contrast, color distribution and so on. Instead of using low level visual features, we present a novel computational visual attention model to assess photos. Firstly, a face-sensitive saliency map analysis is deployed to estimate attention distribution. Then, a Rate of Focused Attention (RFA) measurement is proposed to quantify photo quality. By integrating top-down supervision into the visual attention model, we further achieve personalized photo assessment to take user preference into quality evaluation, which can be extended into object or semantic oriented photo assessment scenarios. Experiments on personal photo albums with comparison to ground-truth user evaluations demonstrate the effeteness of the proposed method. Xiaoshuai Sun, Hongxun Yao, Rongrong Ji, Shaohui Liu |
ACM Multimedia | 4 |
| 2009 | Geometric and Algebraic Approaches of Planar Structure Recovery Based on Properties of Dual ConicabstractIn this article, approaches of planar Euclidean Structure recovery are proposed based on the geometric and algebraic characteristics of the dual conic. Algebraic relations between the matrix representation of the dual conic and the circular point enveloped are extensively analyzed, to attain a generalized framework for the problem. The dual conic has the decomposition with consistent geometric interpretation from which the circular point envelope can be solved. From the geometric viewpoint, based on the characteristics of the common elements of dual conics (their bitangents), a stratified approach is presented to establish the solution. Experiments on both synthetic and real images are conducted to demonstrate the robustness and accuracy of the presented approach. Liang Wang 0004, Hongxun Yao, Shaohui Liu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2009 | Dynamic Background Subtraction Based on Local Dependency HistogramabstractTraditional background subtraction methods perform poorly when scenes contain dynamic backgrounds such as waving tree branches, spouting fountain, illumination changes, camera jitters, etc. In this paper, from the view of spatial context, we present a novel and effective dynamic background method with three contributions. First, we present a novel local dependency descriptor, called local dependency histogram (LDH), to effectively model the spatial dependencies between a pixel and its neighboring pixels. The spatial dependencies contain substantial evidence for differentiating dynamic background regions from moving objects of interest. Second, based on the proposed LDH, an effective approach to dynamic background subtraction is proposed, in which each pixel is modeled as a group of weighted LDHs. Labeling a pixel as foreground or background is done by comparing the LDH computed in current frame against its model LDHs. The model LDHs are adaptively updated by the current LDH. Finally, unlike traditional approaches using a fixed threshold to judge whether a pixel matches to its model, an adaptive thresholding technique is also proposed. Experimental results on a diverse set of dynamic scenes validate that the proposed method significantly outperforms traditional methods for dynamic background subtraction. Shengping Zhang, Hongxun Yao, Shaohui Liu |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2008 | Dynamic background modeling and subtraction using spatio-temporal local binary patternsabstractTraditional background modeling and subtraction methods have a strong assumption that the scenes are of static structures with limited perturbation. These methods will perform poorly in dynamic scenes. In this paper, we present a solution to this problem. We first extend the local binary patterns from spatial domain to spatio-temporal domain, and present a new online dynamic texture extraction operator, named spatio-temporal local binary patterns (STLBP). Then we present a novel and effective method for dynamic background modeling and subtraction using STLBP. In the proposed method, each pixel is modeled as a group of STLBP dynamic texture histograms which combine spatial texture and temporal motion information together. Compared with traditional methods, experimental results show that the proposed method adapts quickly to the changes of the dynamic background. It achieves accurate detection of moving objects and suppresses most of the false detections for dynamic changes of nature scenes. Shengping Zhang, Hongxun Yao, Shaohui Liu |
ICIP | 3 |
| 2008 | A covariance-based method for dynamic background subtractionabstractBackground subtraction in dynamic scenes is an important and challenging task. In this paper, we present a novel and effective method for dynamic background subtraction based on covariance matrix descriptor. The algorithm integrates two distinct levels: pixel level and region level. At the pixel level, spatial properties that are obtained from pixel coordinate values, and appearance properties, i.e., intensity, texture, gradient, etc, are used as features of each pixel. In the region level, the correlation of features extracted at the pixel level is represented by a covariance matrix that is calculated over a rectangle region around the pixel. Each pixel is modeled as a group of weighted adaptive covariance matrices. Experimental results on a diverse set of dynamic scenes show that the proposed method dramatically out-performs traditional methods for dynamic background subtraction. Shengping Zhang, Hongxun Yao, Shaohui Liu, Xilin Chen 0001, Wen Gao 0001 |
ICPR | 3 |
| 2007 | Minimizing the Distortion Spatial Data Hiding Based on Equivalence Class
Shaohui Liu, Hongxun Yao, Wen Gao 0001, Dingguo Yang |
ICIC (1) | 1 |
| 2007 | MSMiner - a developing platform for OLAP
Zhongzhi Shi, Youping Huang, Qing He 0003, Shaohui Liu, Liangxi Qin, Ziyan Jia, Jiayou Li, Huijing Huang |
Decis. Support Syst. | 5 |
| 2003 | A Lattice Based General Blind Watermark Scheme
Yongliang Liu, Wen Gao 0001, Zhao Wang 0008, Shaohui Liu |
ICICS | 4 |
| 2003 | Neural network based steganalysis in still imagesabstractSeganalysis has recently attracted researchers' interests with the development of information hiding techniques. In this paper we propose a new method based neural network to get statistics features of images to identify the underlying hidden data. We first extract features of image embedded information, then input them into neural network to get output. And experiment results indicate this method is valid in steganalysis. This method will be used for Internet/network security, watermarking and so on. Shaohui Liu, Hongxun Yao, Wen Gao 0001 |
ICME | 1 |
| 2003 | A Teamwork Protocol for Multi-agent System
Qiujian Sheng, Zhi-Kun Zhao, Shaohui Liu, Zhongzhi Shi |
PRIMA | 3 |
| 2002 | CSIM: a document clustering algorithm based on swarm intelligenceabstractThis paper presents a document clustering algorithm based on swarm intelligence and k-means: CSIM. First, a document clustering algorithm based on swarm intelligence is employed. It is derived from a basic model interpreting ant colony organization of cemeteries. Swarm intelligence for flexibility, self-organization and robustness has been applied in a variety of areas. Taking advantage of these traits, good initial clusters are obtained in the first phase in CSIM. We then combine it with the classical k-means clustering method by using the clusters as initial centers. CSIM inherits the prominent properties of both swarm intelligence and k-means. It also offsets the weakness of those two techniques. Experimental results show the good performance of the hybrid document clustering algorithm. Bin Wu 0001, Shaohui Liu, Zhongzhi Shi |
IEEE Congress on Evolutionary Computation | 3 |