Jing Xiao 0004

dblp:207/2614 · DBLP profile ↗
← Back
54ranked-venue papers
6as first author
28since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 44 · 5 first-author · 21 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
YearPublicationVenuePosition
2026 CHARM: Collaborative Harmonization Across Arbitrary Modalities for Modality-Agnostic Semantic Segmentation
abstract
Modality-agnostic Semantic Segmentation (MaSS) aims to achieve robust scene understanding across arbitrary combinations of input modality. Existing methods typically rely on explicit feature alignment to achieve modal homogenization, which dilutes the distinctive strengths of each modality and destroys their inherent complementarity. To achieve cooperative harmonization rather than homogenization, we propose CHARM, a novel complementary learning framework designed to implicitly align content while preserving modality-specific advantages through two components: (1) Mutual Perception Unit (MPU), enabling implicit alignment through window-based cross-modal interaction, where modalities serve as both queries and contexts for each other to discover modality-interactive correspondences; (2) A dual-path optimization strategy that decouples training into Collaborative Learning Strategy (CoL) for complementary fusion learning and Individual Enhancement Strategy (InE) for protected modality-specific optimization. Experiments across multiple datasets and backbones indicate that CHARM consistently outperform the baselines, with significant increment on the fragile modalities. This work shifts the focus from model homogenization to harmonization, enabling cross-modal complementarity for true harmony in diversity.
Lekang Wen, Jing Xiao 0004, Mi Wang
AAAI2
2025 DAWA: Dynamic Ambiguity-Wise Adaptation for Real-Time Domain Adaptive Semantic Segmentation
Taorong Liu, Zhen Zhang 0046, Jing Xiao 0004, Chia-Wen Lin
PRCV (18)4
2025 Transref: Multi-scale reference embedding transformer for reference-guided image inpainting
Taorong Liu, Delin Chen, Jing Xiao 0004, Zheng Wang 0007, Chia-Wen Lin, Shin'ichi Satoh 0001
Neurocomputing4
2025 Multiscale Spatial-Spectral Attention Network for Hyperspectral Image Compressed Sensing
Shen Dong, Jing Xiao 0004, Zhen Zhang 0046
IEEE Geosci. Remote. Sens. Lett.2
2025 Luminance decomposition and reconstruction for high dynamic range Video Quality Assessment
Jifan Yang, Zhongyuan Wang 0001, Baojin Huang, Jiaxin Ai, Yuhong Yang 0001, Jing Xiao 0004, Zixiang Xiong
Pattern Recognit.6
2024 Boosting Image Quality Assessment Through Efficient Transformer Adaptation with Local Feature Enhancement
abstract
Image Quality Assessment (IQA) constitutes a funda-mental task within the field of computer vision, yet it re-mains an unresolved challenge, owing to the intricate dis-tortion conditions, diverse image contents, and limited availability of data. Recently, the community has wit-nessed the emergence of numerous large-scale pretrained foundation models. However, it remains an open problem whether the scaling law in high-level tasks is also appli-cable to IQA tasks which are closely related to low-level clues. In this paper, we demonstrate that with a proper in-jection of local distortion features, a larger pretrained vision transformer (ViT) foundation model performs better in IQA tasks. Specifically, for the lack of local distortion structure and inductive bias of the large-scale pretrained ViT, we use another pretrained convolution neural networks (CNNs), which is well known for capturing the local structure, to extract multi-scale image features. Further, we propose a local distortion extractor to obtain local distortion features from the pretrained CNNs and a local distortion in-jector to inject the local distortion features into ViT. By only training the extractor and injector, our method can benefit from the rich knowledge in the powerful foundation models and achieve state-of-the-art performance on popular IQA datasets, indicating that IQA is not only a low-level problem but also benefits from stronger high-level features drawn from large-scale pretrained models. Codes are publicly available at: https://github.com/NeosXu/LoDa.
Kangmin Xu, Jing Xiao 0004, Chaofeng Chen, Haoning Wu 0001, Qiong Yan, Weisi Lin
CVPR3
2024 RefScale: Multi-temporal Assisted Image Rescaling in Repetitive Observation Scenarios
abstract
With the continuous development of imaging technology and the gradual expansion of the amount of image data, how to achieve high compression efficiency of high-resolution images is a challenge problem for storage and transmission. Image rescaling aims to reduce the original data amount through downscaling to facilitate data transmission and storage before encoding, and reconstruct the quality through upscaling after decoding, which is a key technology to assist in high-ratio image compression. However, existing rescaling approaches are more focused on reconstruction quality rather than image compressibility. In repetitive observation scenarios, multi-temporal images brought by periodic observations provide an opportunity to alleviate the conflict between reconstruction quality and compressibility, that is, the historical images as reference indicates what information can be dropped at downscaling to reduce the information content in downscaled image and provides the dropped information to improve the image restoration quality at upscaling. Based on this consideration, we propose a novel multi-temporal assisted reference-based image rescaling framework (RefScale). Specifically, a referencing network is proposed to calculate the similarity map to provide the referencing condition, which is then injected into the conditional invertible neural network to guide the information drop at the downscaling stage and information fusion at the upscaling stage. Additionally, a low-resolution guidance loss is proposed to further constrain the data amount of the downscaled image. Experiments conducted on both satellite imaging and autonomous driving show the superior performance of our approach over the state-of-the-art methods.
Zhen Zhang 0046, Jing Xiao 0004, Mi Wang
ACM Multimedia2
2024 A Multi-scale Framework towards Human-Machine Friendly Remote Sensing Image Coding
Yingkai He, Zhen Zhang 0046, Jing Xiao 0004
MMAsia3
2024 Latent Variables Coding for Perceptual Image Compression
Yingkai He, Zhen Zhang 0046, Jing Xiao 0004
MMAsia4
2024 A Quantization Loss Compensation Network for Remote Sensing Image Compression
abstract
High-resolution remote sensing images (HRRSIs) contain abundant details and texture information. Existing lossy compression methods employ quantization to eliminate redundant information, but this leads to irreversible effects on the subtle details of HRRSIs. This paper proposes a quantization loss compensation network to address this issue. We use a trainable variational autoencoder to learn the details and texture information of HRRSIs from the error between pre-and post-quantized latent representations. During the encoding of HRRSIs, the quantization errors are inputted into the encoding module of the compensation network to generate bitstreams. When it comes to image decoding, using the decoding module of the compensation network to generate quantization loss compensation information, which, together with the quantization latent representations, contributes to the HRRSIs reconstruction. To validate the effectiveness of our approach in reconstructing details of HRRSIs, we conducted experiments on two remote sensing datasets. The experimental results also indicate that our method exhibits superior compression performance.
Shao Xiang, Jing Xiao 0004, Mi Wang
PCS2
2024 SDM-Car: A Dataset for Small and Dim Moving Vehicles Detection in Satellite Videos
abstract
Vehicle detection and tracking in satellite video is essential in remote sensing (RS) applications. However, upon the statistical analysis of existing datasets, we find that the dim vehicles with low radiation intensity and limited contrast against the background are rarely annotated, which leads to the poor effect of existing approaches in detecting moving vehicles under low radiation conditions. In this letter, we address the challenge by building a small and dim moving cars (SDM-Car) dataset with a multitude of annotations for dim vehicles in satellite videos, which is collected by the Luojia 3–01 satellite and comprises 99 high-quality videos. Furthermore, we propose a method based on image enhancement and attention mechanisms to improve the detection accuracy of dim vehicles, serving as a benchmark for evaluating the dataset. Finally, we assess the performance of several representative methods on SDM-Car and present insightful findings. The dataset is openly available athttps://github.com/TanedaM/SDM-Car.
Zhen Zhang 0046, Jing Xiao 0004, Mi Wang
IEEE Geosci. Remote. Sens. Lett.4
2023 Only a Few Classes Confusing: Pixel-Wise Candidate Labels Disambiguation for Foggy Scene Understanding
abstract
Not all semantics become confusing when deploying a semantic segmentation model for real-world scene understanding of adverse weather. The true semantics of most pixels have a high likelihood of appearing in the few top classes according to confidence ranking. In this paper, we replace the one-hot pseudo label with a candidate label set (CLS) that consists of only a few ambiguous classes and exploit its effects on self-training-based unsupervised domain adaptation. Specifically, we formulate the problem as a coarse-to-fine process. In the coarse-level process, adaptive CLS selection is proposed to pick a minimal set of confusing candidate labels based on the reliability of label predictions. Then, representation learning and label rectification are iteratively performed to facilitate feature clustering in an embedding space and to disambiguate the confusing semantics. Experimentally, our method outperforms the state-of-the-art methods on three realistic foggy benchmarks.
Wenyi Chen, Zhen Zhang 0046, Jing Xiao 0004, Chia-Wen Lin, Shin'ichi Satoh 0001
AAAI4
2023 Bat: Bi-Alignment Based On Transformation in Multi-Target Domain Adaptation for Semantic Segmentation
abstract
While enlightening progress has been made recently in single-target domain adaptive semantic segmentation (ST-DASS), the multi-peak distributed multi-target domain cannot be directly aligned well with the single-peak distributed source domain. As a result, it is impossible for existing methods to handle the more realistic multi-target domain adaptive semantic segmentation (MT-DASS) tasks. To solve this problem, we propose a Bi-Alignment framework based on Transformation (BAT). Specifically, we employ the Fourier style transform to convert the style of the source domain to that of the target domain without training any style transfer networks. In this way, we transform the single-peak distributed source domain into a multi-peak distribution that resembles the multi-target domain. Then, we perform fine-grained global and local dual distribution alignment between the same style of source-target domain pairs to achieve a multi-to-multi distribution alignment. Finally, self-training is utilized to further improve the network’s discriminability. Experimental results show that our approach achieves competitive results over state-of-the-art methods.
Xian Zhong, Jing Xiao 0004, Wenxuan Liu 0008, Wenxin Huang, Zheng Wang 0007
ICASSP4
2023 ASVFI: Audio-Driven Speaker Video Frame Interpolation
abstract
Due to limited data transmission, the video frame rate is low during the online conference, severely affecting user experience. Video frame interpolation can solve the problem by interpolating intermediate frames to increase the video frame rate. Generally, most existing video frame interpolation methods are based on the linear motion assumption. However, the mouth motion is nonlinear, and these methods can not generate superior intermediate frames in speaker video. Considering the strong correlation between mouth shape and vocalization, a new method is proposed, named Audio-driven Speaker Video Frame Interpolation(ASVFI). First, we extract the audio feature from Audio Net(ANet). Second, we use Video Net(VNet) encoder to extract the video feature. Finally, we fuse the audio and video features by AVFusion and decode out the intermediate frame in the VNet decoder. The experimental results show that the PSNR is nearly 0.13dB higher than the baseline of interpolating one frame. When interpolating seven frames, the PSNR is 0.33dB higher than the baseline.
Qianrui Wang, Dengshi Li, Jing Xiao 0004
ICIP6
2023 Noise adaptive speech intelligibility enhancement based on improved StarGAN*
abstract
When people communicate in noisy environments with the phone, it is difficult for listeners to obtain information even if the device outputs clear speech. Previous studies have focused on speech intelligibility enhancement (IENH) via normal speech and different levels of Lombard speech conversion. However, these methods often lead to speech distortion and impair the overall speech quality. We propose an IENH framework based on an improved Star Generative Adversarial Network (StarGAN) named D2StarGAN. It has two main advantages: 1) Inspired by the dual-discriminator idea, we add a speech metric discriminator based on StarGAN to optimize multiple intelligibility-related metrics simultaneously; 2) The framework can adapt to different far-and-near-end noise levels and different noise types. Experimental results using both objective measurements and subjective listening tests indicate that the proposed method outperforms the baseline method. It adaptively converts all the mobile communication scenarios with far-and-near-end noise, thus making IENH more widely used in practice.
Lanxin Zhao, Dengshi Li, Jing Xiao 0004, Chenyi Zhu
ICME3
2023 Cloud Coverage Estimation Network for Remote Sensing Images
abstract
The main purpose of cloud detection is to estimate cloud coverage and thus determine whether to transmit remote sensing images to earth or execute subsequent tasks based on cloud coverage. Fast and accurate cloud coverage estimation is a necessary preprocessing step on board. Therefore, we propose a new approach for cloud coverage estimation using a regression network to directly predict the coverage. A cloud coverage estimation network, which is termed$\text{C}^{2}\text{E}$-Net, is proposed in this work. The proposed network consists of three modules, including an encoder for representation feature extraction, a coverage estimation for predicting the cover rate of clouds, and an auxiliary supervision module for improving the performance of the model. To verify the effectiveness of our method, experiments are performed on two open-source datasets (Landset 8 Biome dataset and GaoFen-1 WFV dataset). Our method effectively improves the efficiency of cloud detection by at least doubling, while keeping the estimation error low.
Shao Xiang, Mi Wang, Jing Xiao 0004, Guangqi Xie, Peng Tang 0004
IEEE Geosci. Remote. Sens. Lett.3
2023 Uplink-Assist Downlink Remote-Sensing Image Compression via Historical Referencing
abstract
The traditional strategy of acquiring satellite images involves transmitting compressed satellite data to ground stations solely via the downlink, without utilizing the uplink. In this paper, we propose an enhanced remote sensing (RS) image compression approach that utilizes uplink assistance to improve compression efficiency. By leveraging the uplink, historical images from ground stations can serve as reference images for on-orbit compression, effectively eliminating spatio-temporal redundancy in RS images. However, due to radiation variations among RS images captured on different dates, pixel-wise referencing as employed in the prior codec paradigm is insufficient. To address this, we propose a novel dual-end referencing downsampling-based coding (RefDBC) framework. At the encoder, relevance embedding evaluates reconstructability and records information to restore texture details from the reference prior to downsampling. At the decoder, relevance-based super-resolution uses the identical reference and recorded relevance information to reconstruct the decoded low-resolution image. By incorporating relevance referencing, RefDBC effectively mitigates fake texture generation caused by downsampling and compression, achieving significant bitrate savings ranging from 35%-70% compared to standard, learning-based, and DBC compression baselines in experiments on Spot-5 and Luojia3 images. Code, data, and pretrained models are available online at https://github.com/WHW1233/RefDBC.
Jing Xiao 0004, Weisi Lin, Mi Wang
IEEE Trans. Geosci. Remote. Sens.3
2023 Progressive Motion Boosting for Video Frame Interpolation
abstract
Video frame interpolation has made great progress in estimating advanced optical flow and synthesizing in-between frames sequentially. However, frame interpolation involving various resolutions and motions remains challenging due to limited or fixed pre-trained networks. Inspired by the success of the coarse-to-fine scheme for video frame interpolation, i.e., gradually interpolating frames of different resolutions, we propose a progressive boosting network (ProBoost-Net) based on a multi-scale framework to achieve flexible recurrent scales and then gradually optimize optical flow estimation and frame interpolation. Specifically, we designed a dense motion boosting (DMB) module to transfer features close to real motion to the decoded features from the later scales, which provides complementary information to refine the motion further. Furthermore, to ensure the accuracy of the estimated motion features at each scale, we propose a motion adaptive fusion (MAF) module that adaptively deals with motions with different receptive fields according to the motion conditions. Thanks to the framework's flexible recurrent scales, we can customize the number of scales and make trade-offs between computation and quality depending on the application scenario. Extensive experiments with various datasets demonstrated the superiority of our proposed method over state-of-the-art approaches in various scenarios.
Jing Xiao 0004, Kangmin Xu, Mengshun Hu, Zheng Wang 0007, Chia-Wen Lin, Mi Wang, Shin'ichi Satoh 0001
IEEE Trans. Multim.1
2022 Spatial-Temporal Space Hand-in-Hand: Spatial-Temporal Video Super-Resolution via Cycle-Projected Mutual Learning
abstract
Spatial-Temporal Video Super-Resolution (ST-VSR) aims to generate super-resolved videos with higher resolution (HR) and higher frame rate (HFR). Quite intuitively, pioneering two-stage based methods complete ST-VSR by directly combining two sub-tasks: Spatial Video Super-Resolution (S-VSR) and Temporal Video Super-Resolution (T-VSR) but ignore the reciprocal relations among them. Specifically, 1) T-VSR to S-VSR: temporal correlations help accurate spatial detail representation with more clues; 2) S-VSR to T-VSR: abundant spatial information contributes to the refinement of temporal prediction. To this end, we propose a one-stage based Cycle-projected Mutual learning network (CycMu-Net) for ST-VSR, which makes full use of spatial-temporal correlations via the mutual learning between S-VSR and T-VSR. Specifically, we propose to exploit the mutual information among them via iterative up-and-down projections, where the spatial and temporal features are fully fused and distilled, helping the high-quality video reconstruction. Besides extensive experiments on benchmark datasets, we also compare our proposed CycMu-Net with S-VSR and T-VSR tasks, demonstrating that our method significantly outperforms state-of-the-art methods. Codes are publicly available at: https://github.com/hhhhhumengshun/CycMuNet.
Mengshun Hu, Kui Jiang, Jing Xiao 0004, Junjun Jiang, Zheng Wang 0007
CVPR4
2022 ECL: Exclusive Curriculum Learning for Video Super-Resolution
abstract
Video super-resolution (VSR) problem has gained a soaring development along with deep learning methods. However, the further progress requires the blessing of more complex architectures. Unlike them, this paper promotes VSR performance from a new perspective of sample difficulty. We propose an exclusive curriculum learning strategy for VSR, which can improve the representation power without noticeable computation increment. Specifically, this paper memorizes the performance track of every sample and calculate a customized weight for each sample according to it. In this way, the model can automatically concentrate on the easy samples first and gradually focus on the hard ones. Experimental analysis on training process and benchmark datasets demonstrate that our method can substantially boost the performance with a superior convergence speed and a limited number of parameters.
Sicheng Hu, Zhongyuan Wang 0001, Peng Yi 0002, Zheng He 0001, Jinsheng Xiao, Jing Xiao 0004
ICME6
2022 Progressive Spatial-temporal Collaborative Network for Video Frame Interpolation
abstract
Most video frame interpolation (VFI) algorithms infer the intermediate frame with the help of adjacent frames through the cascaded motion estimation and content refinement.However, the intrinsic correlations between motion and content are barely investigated, commonly producing interpolated results with inconsistency and blurry contents.Specifically, we first discover a simple yet essential domain knowledge that contents and motions characteristics should be homogeneous to a certain degree from the same objects, and formulate the consistency into the loss function for model optimization. Based on this, we propose to learn the collaborative representation between motions and contents, and construct a novel progressive spatial-temporal Collaborative network (Prost-Net) for video frame interpolation.Specifically, we develop a content-guided motion module (CGMM) and a motion-guided content module (MGCM) for individual content and motion representation. In particular, the predicted motion in CGMM is used to guide the fusion and distillation of contents for intermediate frame interpolation, and vice versa. Furthermore, by considering collaborative strategy in a multi-scale framework, our Prost-Net progressively optimizes motions and contents in a coarse-to-fine manner, making it robust to various challenging scenarios (occlusion and large motions) in VFI. Extensive experiments on the benchmark datasets demonstrate that our method significantly outperforms state-of-the-art methods.
Mengshun Hu, Kui Jiang, Zhixiang Nie, Jing Xiao 0004, Zheng Wang 0007
ACM Multimedia5
2022 UoLMM'22: 2nd International Workshop on Robust Understanding of Low-quality Multimedia Data: Unitive Enhancement, Analysis and Evaluation
abstract
Low-quality multimedia data (including low resolution, low illumination, defects, blurriness, etc.) often pose a challenge for content understanding, as algorithms are typically developed under ideal conditions (high resolution and good visibility). To alleviate this problem, data enhancement techniques (e.g., super-resolution, low-light enhancement, derain, and inpainting) have been proposed to restore low-quality multimedia data. Efforts are also being made to develop robust content understanding algorithms in adverse weather and lighting conditions. Some quality assessment techniques aiming at evaluating the analytical quality of data have also emerged. Even though these topics are mostly studied independently, they are tightly related in terms of ensuring a robust understanding of multimedia content. For example, enhancement should maintain the semantic consistency of the analysis, while quality assessment should consider the comprehensibility of the multimedia data. The purpose of this workshop is to bring together individuals in three areas: enhancement, analysis, and evaluation, for sharing ideas and discussion on current developments and future directions.
Yang Wu 0001, Xiao Wang 0029, Jing Xiao 0004
ACM Multimedia5
2022 Adaptive Speech Intelligibility Enhancement for Far-and-Near-end Noise Environments Based on Self-attention StarGAN
Dengshi Li, Lanxin Zhao, Jing Xiao 0004, Duanzheng Guan, Qianrui Wang
MMM (2)3
2022 Speech Intelligibility Enhancement By Non-Parallel Speech Style Conversion Using CWT and iMetricGAN Based CycleGAN
Jing Xiao 0004, Dengshi Li, Lanxin Zhao, Qianrui Wang
MMM (1)1
2022 Capturing Small, Fast-Moving Objects: Frame Interpolation via Recurrent Motion Enhancement
abstract
Interpolating video frames involving large motions remains an elusive challenge. In case that frames involve small and fast-moving objects, conventional feed-forward neural network-based approaches that estimate optical flow and synthesize in-between frames sequentially often result in loss of motion features and thus blurred boundaries. To address the problem, we propose a novel Recurrent Motion-Enhanced Interpolation Network (ReMEI-Net) by assigning attention to the motion features of small objects from both the intra-scale and inter-scale perspectives. Specifically, we add recurrent feedback blocks in the existing multi-scale autoencoder pipeline, aiming to iteratively enhance the motion information of small objects across different scales. Second, to further refine the motion features of the highly moving objects, we propose a Multi-Directional ConvLSTM (MD-ConvLSTM) block to capture the global spatial contextual information of motion from multiple directions. In this way, the coarse-scale features can be utilized to correct and enhance the fine-scale features through the feedback mechanism. Extensive experiments on various datasets demonstrate the superiority of our proposed method over state-of-the-art approaches in terms of clear locations and complete shape.
Mengshun Hu, Jing Xiao 0004, Zheng Wang 0007, Chia-Wen Lin, Mi Wang, Shin'ichi Satoh 0001
IEEE Trans. Circuits Syst. Video Technol.2
2022 Vehicle Counting in Very Low-Resolution Aerial Images via Cross-Resolution Spatial Consistency and Intraresolution Time Continuity
abstract
Vehicle counting is important for smart city applications such as logistics management, traffic estimation, and financial analysis. To perform vehicle counting using aerial images, researchers have proposed many algorithms, including detection-based, regression-based and density-based methods. However, most of these algorithms are only applicable to high-resolution images, which require clear vehicle outlines. For the reasons of acquisition difficulty, frequency and cost, it is necessary to explore methods for vehicle counting using low-resolution or even very low-resolution images. We build a cross-resolution vehicle counting (CRVC) dataset, including 192 very low-resolution images and 8 high-resolution images of a port from 2016 to 2019. For this task, we propose a novel vehicle counting via cross-resolution spatial consistency and intra-resolution time continuity constraints. The segmentation map is first obtained by semantic segmentation with the prior information above. The vehicle coverage rate relative to the located parking lot is calculated and then converted to vehicle area. Finally, the relationship between the area and the number of vehicles is established by regression. Experiments show that the vehicle counting results obtained by our method are highly consistent with the annotations and outperform other state-of-the-art methods. Our method is also applicable for images with a lower resolution of 10m and other locations. Code, data and pre-trained models are available online at https://github.com/hbsszq/Vehicle-Counting-in-Very-Low-Resolution-Aerial-Images.
Jing Xiao 0004, Zheng Wang 0007, Xujie Ma, Mi Wang, Shin'ichi Satoh 0001
IEEE Trans. Geosci. Remote. Sens.2
2022 Unsupervised Foggy Scene Understanding via Self Spatial-Temporal Label Diffusion
abstract
Understanding foggy image sequence in driving scene is critical for autonomous driving, but it remains a challenging task due to the difficulty in collecting and annotating real-world images of adverse weather. Recently, self-training strategy has been considered as a powerful solution for unsupervised domain adaptation, which iteratively adapts the model from the source domain to the target domain by generating target pseudo labels and re-training the model. However, the selection of confident pseudo labels inevitably suffers from the conflict between sparsity and accuracy, both of which will lead to suboptimal models. To tackle this problem, we exploit the characteristics of the foggy image sequence of driving scenes to densify the confident pseudo labels. Specifically, based on the two discoveries of local spatial similarity and adjacent temporal correspondence of the sequential image data, we propose a novel Target-Domain driven pseudo label Diffusion (TDo-Dif) scheme. It employs superpixels and optical flows to identify the spatial similarity and temporal correspondence, respectively, and then diffuses the confident but sparse pseudo labels within a superpixel or a temporal corresponding pair linked by the flow. Moreover, to ensure the feature similarity of the diffused pixels, we introduce local spatial similarity loss and temporal contrastive loss in the model re-training stage. Experimental results show that our TDo-Dif scheme helps the adaptive model achieve 51.92% and 53.84% mean intersection-over-union (mIoU) on two publicly available natural foggy datasets (Foggy Zurich and Foggy Driving), which exceeds the state-of-the-art unsupervised domain adaptive semantic segmentation methods. The proposed method can also be applied to non-sequential images in the target domain by considering only spatial similarity.
Wenyi Chen, Jing Xiao 0004, Zheng Wang 0007, Chia-Wen Lin, Shin'ichi Satoh 0001
IEEE Trans. Image Process.3
2021 Image Inpainting Guided by Coherence Priors of Semantics and Textures
abstract
Existing inpainting methods have achieved promising performance in recovering defective images of specific scenes. However, filling holes involving multiple semantic categories remains challenging due to the obscure se-mantic boundaries and the mixture of different semantic textures. In this paper, we introduce coherence priors between the semantics and textures which make it possible to concentrate on completing separate textures in a semantic-wise manner. Specifically, we adopt a multi-scale joint optimization framework to first model the coherence priors and then accordingly interleaving optimize image inpainting and semantic segmentation in a coarse-to-fine manner. A Semantic-Wise Attention Propagation (SWAP) module is devised to refine completed image textures across scales by exploring non-local semantic coherence, which effectively mitigates the mix-up of textures. We also propose two coherence losses to constrain the consistency between the semantics and the inpainted image in terms of the overall structure and detailed textures. Experimental results demonstrate the superiority of our proposed method for challenging cases with complex holes.
Jing Xiao 0004, Zheng Wang 0007, Chia-Wen Lin, Shin'ichi Satoh 0001
CVPR2
2020 Mining on Heterogeneous Manifolds for Zero-Shot Cross-Modal Image Retrieval
abstract
Most recent approaches for the zero-shot cross-modal image retrieval map images from different modalities into a uniform feature space to exploit their relevance by using a pre-trained model. Based on the observation that manifolds of zero-shot images are usually deformed and incomplete, we argue that the manifolds of unseen classes are inevitably distorted during the training of a two-stream model that simply maps images from different modalities into a uniform space. This issue directly leads to poor cross-modal retrieval performance. We propose a bi-directional random walk scheme to mining more reliable relationships between images by traversing heterogeneous manifolds in the feature space of each modality. Our proposed method benefits from intra-modal distributions to alleviate the interference caused by noisy similarities in the cross-modal feature space. As a result, we achieved great improvement in the performance of the thermal v.s. visible image retrieval task. The code of this paper: https://github.com/fyang93/cross-modal-retrieval
Fan Yang 0038, Zheng Wang 0007, Jing Xiao 0004, Shin'ichi Satoh 0001
AAAI3
2020 Guidance and Evaluation: Semantic-Aware Image Inpainting for Mixed Scenes
Jing Xiao 0004, Zheng Wang 0007, Chia-Wen Lin, Shin'ichi Satoh 0001
ECCV (27)2
2020 Motion Feedback Design for Video Frame Interpolation
abstract
This paper introduces a feedback-based approach to interpolate video frames involving small and fast-moving objects. Unlike the existing feedforward-based methods that estimate optical flow and synthesize in-between frames sequentially, we introduce a motion-oriented component that adds a feedback block to the existing multi-scale autoencoder pipeline, which feedbacks information of small objects shared between architectures of two different scales. We show that feeding this additional information enables more robust detection of optical flow caused by small objects in fast motion. Using experiments on various datasets, we show that the feedback mechanism allows our method to achieve state-of-the-art results, both qualitatively and quantitatively.
Mengshun Hu, Jing Xiao 0004, Lin Gu 0003, Shin'ichi Satoh 0001
ICASSP3
2020 Ts-Fen: Probing Feature Selection Strategy for Face Anti-Spoofing
abstract
Deep features extracted from different domains have shown great advantages in the face anti-spoofing task. Previous extraction strategies consider less of the extent variation in distinction among feature properties. Many of them straightly make classification using the extracted information but generalize weakly. In this paper, we propose a novel Two-Stream Feature Extraction Network (TS-FEN) based on depth and chrominance cues, guiding both sparsity and density of the feature distribution. We specifically design a Feature Enhancement (FE) Structure in the depth stream to strengthen the discrimination capacity, as well as a Feature Selection (FS) Module in the chroma stream to keep feature diversity and distinction. Besides, the specially designed bias arcface loss aims to enlarge the central distribution dispersion of opposite categories. Extensive experiments on three benchmark datasets validate that our proposed approach achieves explicit improvement on both intra-testing and cross-testing.
Dongmei Peng, Jing Xiao 0004
ICASSP2
2020 Learning latent geometric consistency for 6D object pose estimation in heavily cluttered scenes
Qingnan Li, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001, Yu Chen 0021
J. Vis. Commun. Image Represent.3
2020 Long-Term Background Redundancy Reduction for Earth Observatory Video Coding
abstract
Huge earth observatory video data (EOVD) and the limited transmission bandwidth from satellites to terrestrial devices pose serious challenges to compression efficiency for satellite video. In this work, we deeply explore long-term background redundancy caused by periodical satellite revisit, and make utilization of long-term background reference (LTBR) to design a high-efficiency coding method specific for EOVD. Firstly, we use data of Google Earth as prior background knowledge and construct LTBR from it. Then, two novel prediction methods are proposed to make full use of prior information of LTBR, including color and definition correction based inter prediction with LTBR (CDIP-LTBR), structure and texture constrained intra prediction with LTBR (STIP-LTBR). Lastly, the proposed two new prediction schemes are integrated into one unified coding framework along with HEVC to achieve improved coding performance, and an improved RDO method is designed for additional prediction modes selection problem in the framework to obtain a higher prediction efficiency. Extensive experiments on real-world EOVD show that the proposed coding scheme exhibit significant improvement over HEVC and H.264, in terms of BD-Rate, BD-PSNR and rate distortion comparison.
Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004, Shin'ichi Satoh 0001
IEEE Trans. Circuits Syst. Video Technol.4
2020 Multi-Channel and Multi-Model-Based Autoencoding Prior for Grayscale Image Restoration
abstract
Image restoration (IR) is a long-standing challenging problem in low-level image processing. It is of utmost importance to learn good image priors for pursuing visually pleasing results. In this paper, we develop a multi-channel and multi-model-based denoising autoencoder network as image prior for solving IR problem. Specifically, the network that trained on RGB-channel images is used to construct a prior at first, and then the learned prior is incorporated into single-channel grayscale IR tasks. To achieve the goal, we employ the auxiliary variable technique to integrate the higher-dimensional network-driven prior information into the iterative restoration procedure. In addition, according to the weighted aggregation idea, a multi-model strategy is put forward to enhance the network stability that favors to avoid getting trapped in local optima. Extensive experiments on image deblurring and deblocking tasks show that the proposed algorithm is efficient, robust, and yields state-of-the-art restoration quality on grayscale images.
Sanqian Li, Binjie Qin, Jing Xiao 0004, Qiegen Liu, Yuhao Wang 0001, Dong Liang 0001
IEEE Trans. Image Process.3
2019 CISI-net: Explicit Latent Content Inference and Imitated Style Rendering for Image Inpainting
abstract
Convolutional neural networks (CNNs) have presented their potential in filling large missing areas with plausible contents. To address the blurriness issue commonly existing in the CNN-based inpainting, a typical approach is to conduct texture refinement on the initially completed images by replacing the neural patch in the predicted region using the closest one in the known region. However, such a processing might introduce undesired content change in the predicted region, especially when the desired content does not exist in the known region. To avoid generating such incorrect content, in this paper, we propose a content inference and style imitation network (CISI-net), which explicitly separate the image data into content code and style code. The content inference is realized by performing inference in the latent space to infer the content code of the corrupted images similar to the one from the original images. It can produce more detailed content than a similar inference procedure in the pixel domain, due to the dimensional distribution of content being lower than that of the entire image. On the other hand, the style code is used to represent the rendering of content, which will be consistent over the entire image. The style code is then integrated with the inferred content code to generate the complete image. Experiments on multiple datasets including structural and natural images demonstrate that our proposed approach out-performs the existing ones in terms of content accuracy as well as texture details.
Jing Xiao 0004, Qiegen Liu, Ruimin Hu
AAAI1
2019 Long Term Background Reference Based Satellite Video Coding
abstract
Video transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellite calls for higher coding efficiency. In this paper, we propose a high efficiency satellite video coding method based on long term background reference (LTBR) to eliminate redundancy caused by periodical revisit. Firstly, data of Google Earth is used to provide prior information for establishing LTBR. Then a novel intra prediction method guided by pixels' cluster information from LTBR is introduced. Experiments demonstrate that our method outperforms HEVC and H.264 , in terms of rate-distortion, BD-PSNR and BD-Rate performance.
Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004
ICASSP4
2019 Multisource Surveillance Video Coding by Exploiting 3D and 2D Knolwedge
abstract
The rapidly increasing surveillance video data has challenged the existing video coding standards. Even though knowledge based video coding scheme proposed for moving objects so far has achieved high efficiency, it does not take full advantages of local information and highly relies on the accuracy of pose parameter of the objects, thus leading to large prediction residuals. In this paper, a novel surveillance video coding utilizing 3D and 2D knowledge is proposed. On the one hand, we generate a knowledge based reference frame from 3D models of the objects and incorporate it into the block based coding framework to remove global redundancy while improve the robustness to pose errors. On the other hand, 2D knowledge in the form of visual appearances of the objects in the previously encoded frames is employed to rectify the knowledge based reference frame for local redundancy removal. Experimental results demonstrate the effectiveness of our proposed method against HEVC and the knowledge based coding method.
Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001
ICASSP3
2019 Multisource surveillance video coding with synthetic reference frame
Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001
J. Vis. Commun. Image Represent.3
2019 Multisource surveillance video data coding with hierarchical knowledge library
Yu Chen 0021, Ruimin Hu, Jing Xiao 0004, Liang Xu 0010, Zhongyuan Wang 0001
Multim. Tools Appl.3
2019 Multi-filters guided low-rank tensor coding for image inpainting
Qiegen Liu, Sanqian Li, Jing Xiao 0004
Signal Process. Image Commun.3
2018 Edge-Aware Context Encoder for Image Inpainting
abstract
We present Edge-aware Context Encoder (E-CE): an image inpainting model which takes scene structure and context into account. Unlike previous CE which predicts the missing regions using context from entire image, E-CE learns to recover the texture according to edge structures, attempting to avoid context blending across boundaries. In our approach, edges are extracted from the masked image, and completed by a full-convolutional network. The completed edge map together with the original masked image are then input into the modified CE network to predict the missing region. The experiments demonstrate that E-CE can generate images with better shapes and structures than CE.
Ruimin Hu, Jing Xiao 0004, Zhongyuan Wang 0001
ICASSP3
2018 Two-Level Segment-Based Bitrate Control for Live ABR Streaming
Jing Xiao 0004, Gen Zhan, Xu Wang 0015, Zhongyuan Wang 0001
MMM (1)2
2018 Virtual Background Reference Frame Based Satellite Video Coding
abstract
Video transmission from satellites to terrestrial devices usually requires a large amount of channel resources due to the huge amount of satellite video data. Subject to limited transmission bandwidth in space environment, the video encoder for video satellites calls for higher coding efficiency. In this paper, we propose a high efficiency satellite video compression method to eliminate long-term redundancy among multiple periodically revisited videos, based on virtual background reference frame (VBRF) obtained from Google Earth data. First, we make full use of Google Earth data to create VBRF for representing the constant ground background. Then, we encode all I frames by referring to VBRF and performing interprediction. Experiments demonstrate that our method outperforms H.264 and HEVC, in terms of rate distortion, bjøntegaard delta peak signal-to-noise rate (BD-PSNR), and BD-Rate performance.
Xu Wang 0015, Ruimin Hu, Zhongyuan Wang 0001, Jing Xiao 0004
IEEE Signal Process. Lett.4
2017 Cruise UAV Video Compression Based on Long-Term Wide-Range Background
abstract
With the rapid development of Unmanned Aerial Vehicle (UAV), the compression of video data captured by UAV has become a growing critical issue. However, most advanced coding schemes, like H.264 and HEVC, are oriented for common videos and thus cannot afford ideal coding efficiency when applied to UAV platform. Considering the characteristics of UAV video, much more improvement could be imposed onto current coding schemes to make full use of UAV's sensor information. In this paper, we exploit long term redundancy existing in the video data captured by cruise UAV. Firstly, we establish a long-term wide-range background set for reference. Then we separate each frame into new-area part and overlapped part. Lastly, we use GPS information of each frame to get reference from background set and compress two parts individually. In the experiments, by comparing to standard HEVC, our method has given more than 20% reduction in bitrate and meanwhile more than 4% gain in PSNR.
Xu Wang 0015, Jing Xiao 0004, Ruimin Hu, Zhongyuan Wang 0001
DCC2
2017 A sensitive object-oriented approach to big surveillance data compression for social security applications in smart cities
abstract
Summary Surveillance has become a fairly common practice with the global boom in “smart cities”. How to efficiently store and manage the vast quantities of surveillance data is a persistent challenge in terms of analyzing social security problems. Developing data compression technology under the analytic requirements of surveillance data is the key to solving the storage problem. Criminal investigation demands the quality preservation of sensitive objects, typically pedestrians, human faces, vehicles, and license plates; however, the analytical value of surveillance data is rapidly lost as the compression ratio increases. In this paper, we propose a sensitive object‐oriented regions of interest‐based coding strategy for preserving the analytical value of surveillance data. In the proposed method, instead of generating a saliency map based on human visual perception, we consider saliency as a set of characteristics important for object detection and recognition. By making this modification, almost all sensitive objects necessary in a criminal investigation are assigned high saliency value rather than only one or two salient regions. Motions in the temporal domain are integrated to place emphasis on moving objects, namely moving sensitive objects, which then gain the highest saliency. Finally, a saliency‐based rate control algorithm embedded in High Efficiency Video Coding is used to maintain the quality of sensitive objects in the encoded video under a fixed bitrate. Experiments were conducted on two analytical indexes: Feature similarity and object detection accuracy. The results showed that by achieving the same feature similarity and object detection accuracy, our method can save 20% and 40% bitrate over High Efficiency Video Coding, respectively, for the storage of big surveillance data. Copyright © 2016 John Wiley & Sons, Ltd.
Jing Xiao 0004, Zhongyuan Wang 0001, Yu Chen 0021, Jun Xiao 0004, Gen Zhan, Ruimin Hu
Softw. Pract. Exp.1
2016 Spatiotemporal saliency based on location prior model
abstract
Saliency detection for images and videos becomes increasingly popular due to its wide applicability. Enormous research efforts have been focused on saliency detection, but it still has some issues in maintaining spatiotemporal consistency of videos and uniformly highlighting entire objects. To address these issues, this paper proposes a superpixel-level spatiotemporal saliency model for saliency detection in videos. To detect salient object, we extract multiple spatiotemporal features combined with intra-consistency motion information preliminarily. Meanwhile, considering inter-consistency of foreground in videos, a set of foreground locations are obtained from previous frames. Then, we introduce foreground-background and local foreground contrast saliency cues of those features using the location prior information of foreground. These two improved contrast saliency cues uniformly highlight the entire object and suppress the background effectively. Finally, we use an interactively dynamic fusion method to integrate the output spatial and temporal saliency maps. The proposed approach is validated on challenging sets of video sequences. Subjective observations and objective evaluations demonstrate that the proposed model achieves a better performance on saliency detection compared with the state-of-the-art spatiotemporal saliency methods.
Liuyi Hu, Zhongyuan Wang 0001, Mang Ye, Jing Xiao 0004, Ruimin Hu
IJCNN4
2016 Filling Kinect depth holes via position-guided matrix completion
Zhongyuan Wang 0001, Shizheng Wang, Jing Xiao 0004, Ruimin Hu
Neurocomputing4
2016 Knowledge-Based Coding of Objects for Multisource Surveillance Video Data
abstract
Global object redundancy (GOR), as opposed to local spatial/temporal redundancies in a single video clip, is a new form of redundancy common in multisource surveillance video data (MSVD). GOR is induced by the repetition of foreground objects across multiple cameras, and becomes influential as the number of objects increases. Eliminating GOR considerably improves MSVD coding efficiency. In an effort to accomplish this, this study first proposes a knowledge-based representation of objects based on careful analysis of GOR composition. The representation contains a constant part and a variational part: the former is used to represent the common knowledge shared by an object across multiple cameras, while the latter is used to represent local variations on the object's surfaces. Based on the proposed representation, a knowledge-based coding (KBC) method is then proposed in which each foreground object is encoded with a hybrid prediction scheme, where the constant part of the object is generated via global prediction from a model library and the variational part is predicted via local reference frames with pose-based, short-term prediction. Experimental results showed that the KBC method saves more than 39% bits on average for encoding foreground objects in high-resolution video clips (compared to 16% for the entire videos). Applying the proposed coding method to surveillance videos in large spatial and temporal scale allows storage savings at the PB level.
Jing Xiao 0004, Ruimin Hu, Yu Chen 0021, Zhongyuan Wang 0001, Zixiang Xiong
IEEE Trans. Multim.1
2015 Joint Weighted Sparse Representation Based Median Filter for Depth Video Coding
abstract
In order to promote the development of auto-stereoscopic display, MPEG has proposed multi-view plus depth (MVD) format. The depth video is encoded and transmitted with color video to synthesize virtual views at the receiver side. The existing video coding standards such as H.264/AVC introduces coding artifacts along the depth boundaries, which may seriously affects the synthesized view quality and coding efficiency. Many in-loop depth filters such as joint depth filter have been proposed to remove the artifacts in compressed depth video. However, their performance is unstable and affected by the outliers due to the weighted summation. In this paper, based on the sparse prior characteristic in local region of depth map, we propose a joint weighted sparse representation based median filter to select the most relevant neighboring depth pixel as the output during the filter process. Experimental results show the proposed method is more effective in improving the depth video coding efficiency.
Ruimin Hu, Yu Chen 0021, Jing Xiao 0004, Ruolin Ruan
DCC5
2015 Global Coding of Multi-source Surveillance Video Data
abstract
In this paper, we exploit a new type of data redundancy in the multisource surveillance video to reduce the huge gap between the growth rate of the data and the video compression rate. Global redundancy caused by correlated appearances of moving objects in multiple videos consists of model similarity, spatial correlation and temporal consistency. Therefore, we propose a global coding scheme of moving objects to eliminate the global redundancy: a model based object reconstruction is initially employed to reconstruct the objects in the video, then a pose-based residual error prediction is developed to compensate the difference between the real video appearance and the initial reconstruction from model. The experiment with two simulated surveillance videos has proved that the proposed coding scheme can achieve better coding performance than the main profile of HEVC and surveillance profile of IEEE 1857-2013.
Jing Xiao 0004, Yu Chen 0021, Ruimin Hu
DCC1
2015 A Block-Based Background Model for Surveillance Video Coding
abstract
Background model can help to improve the compression efficiency for surveillance video coding, but the existing frame-based background model is inefficient in some situations, for example, when a region of background changes frequently or periodically. In this paper, a block-based background model is proposed to solve this problem. We save the background blocks recognized from each reconstructed frame into a buffer, thus the background blocks are collected gradually. At the same time, we compose a new background frame for each frame to be encoded based on the background blocks currently available in the buffer. Compared with the pre-built background frame, the instantly composed background frame often predicts more accurately because of the accumulated information about background. Experimental results show that the proposed model achieves better rate-distortion performance over the existing frame-based model in most cases, while keeping almost the same computation complexity.
Liming Yin, Ruimin Hu, Jing Xiao 0004
DCC4
2015 Exploiting effects of parts in fine-grained categorization of vehicles
abstract
Fine-grained categorization has become a hot topic in computer vision. Based on the theory that part information is crucial for fine-grained categorization, we proposed a part-based categorization method for vehicles, consisting vehicle parts localization, part-based vehicle representation and classification. There were three contributions we made in this work: 1) we analyzed discriminative powers of parts for fine-grained categorization; 2) we proposed a frame of how to integrate discriminative powers of parts into categorization, and proved that it can achieve better performance than treating every part equally; 3) we provided an annotated dataset with parts for vehicle categorization.
Ruimin Hu, Jun Xiao 0004, Jing Xiao 0004, Jun Chen 0001
ICIP5
2015 A Unsupervised Person Re-identification Method Using Model Based Representation and Ranking
abstract
As a core technique supporting the multi-camera tracking task, person re-identification attracts increasing research interests in both academic and industrial communities. Its aim is to match individuals across a group of spatially non-overlapping surveillance cameras, which are usually interfered by various imaging conditions and object motions. Current methods mainly focus on robust feature representation and accurate distance measure, where intensive computations and expensive training samples prohibit their practical applications. To address the above problems, this paper proposes a new unsupervised person re-identification method featured by its competitive accuracy and high efficiency. Both merits stem from model based person image representation and ranking, with which, merely 4-dimension pixel-level features can achieve over 20% matching rate at Rank 1 on the challenging VIPeR dataset.
Chao Liang 0001, Bingyue Huang, Ruimin Hu, Chunjie Zhang 0001, Xiaoyuan Jing, Jing Xiao 0004
ACM Multimedia6