EDBT 2026 Demo / reviewers in the wild / expert
Xingsong Hou
dblp:88/4367
· DBLP profile ↗
56ranked-venue papers
12as first author
23since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 38 · 6 first-author · 15 since 2021Artificial intelligence and machine learning · 9 · 2 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 4 first-author · 4 since 2021Computer networks · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Enhancing single image compressive sensing via pyramid-based multi-scale sampling
Huake Wang, Xingsong Hou, Xiaoyang Yan, Jutao Li |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | Learning-based haze visibility ranking score for real-world traffic surveillance images
Yu Cao 0016, Yuanjing Feng, Xingsong Hou, Xueming Qian |
J. Vis. Commun. Image Represent. | 4 |
| 2026 | From structure to detail: A conditional diffusion framework for extremely low-bitrate image compression
Yiyang Zou, Xingsong Hou, Zhixuan Guo |
Signal Process. | 3 |
| 2026 | Towards Better Distortion Feature Learning for Object Detection in Top-View Fisheye CamerasabstractWith the development of deep learning in recent years, the performance of object detection under conventional cameras has been significantly improved. Nevertheless, due to the distortion caused by the fisheye cameras, detecting objects in this scenario remains a significant challenge. The dominant approaches focus on modifying the shape of the bounding box to better align the boundaries of the distorted object. However, these methods neglect the learning of spatial distortion information, which prevents them from satisfactory results. In this paper, we propose a novel fisheye camera detection network to learn distortion features better, dubbed SDANet. SDANet is composed of a series of SDABlocks, which are designed to learn spatial distortion features. Each SDABlock consists of multiple convolution kernels of different sizes, and it can generate the most suitable kernel based on the current input's distortion characteristics. Moreover, to address the limitations of the scarcity and uneven spatial distribution of fisheye image datasets on performance improvement, we propose a dedicated data augmentation strategy called Prominent Fisheye Distortion Augmentation (PFDAug). PFDAug can further introduce distortions to fisheye images, effectively alleviating these problems. Experimental results on the CEPDOF, MW-R, HABBOF, LOAF, and FishEye8k fisheye image datasets demonstrate that our method achieves state-ofthe-art performance. Pengbo Guo, Chengxu Liu 0001, Xingsong Hou, Xueming Qian |
IEEE Trans. Multim. | 3 |
| 2025 | Learning scalable Omni-scale distribution for crowd counting
Huake Wang, Xingsong Hou, Kaibing Zhang, Minqi Li, Wenke Sun, Xueming Qian |
J. Vis. Commun. Image Represent. | 2 |
| 2025 | Extracting Noise and Darkness: Low-Light Image Enhancement via Dual Prior GuidanceabstractThe complex entanglement between darkness and noise hinders the advance of low-light image enhancement. Most existing methods adopted lightening-then-denoising or embedded a special denoising module into enhancement network without specific noise knowledge as supervision to restore low-light images. However, they either fail to remove the amplified noise or blur the detail information. Against above drawbacks, we propose a novel dual prior guidance method for low-light image enhancement that relights darkness and suppresses noise simultaneously. Concretely, the main novelties of our proposed method are three-fold. Firstly, our formulation originates from a statistic observation that darkness can be disentangled into luminance channel, yet noise still exists each channel when low-light images are transformed from RGB space to YCbCr space. It inspires us to design an ingenious method, extracting noise and darkness, termed END, to enhance low-light images. Secondly, we propose a prior extraction network with prior composition module to extract luminance and noise priors from different channels. Thirdly, an image enhancement network deployed with prior guidance module is proposed to progressively lighten the darkness and remove noise. Extensive experiments on multiple benchmarks demonstrate that our proposed method achieves remarkable performance compared to other state-of-the-art low-light image enhancement methods. The source code and trained model can be found inhttps://github.com/WHK-Huake/END. Huake Wang, Xiaoyang Yan, Xingsong Hou, Kaibing Zhang, Yujie Dun |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Exploring Distortion Prior With Latent Diffusion Models for Remote Sensing Image CompressionabstractLearning-based image compression algorithms typically focus on designing encoding and decoding networks and improving the accuracy of entropy model estimation to enhance the rate-distortion (RD) performance. However, few algorithms leverage the compression distortion prior from existing compression algorithms to improve RD performance. In this paper, we propose a latent diffusion model-based remote sensing image compression (LDM-RSIC) method, which aims to enhance the final decoding quality of RS images by utilizing the generated distortion prior from a LDM. Our approach consists of two stages. In Stage I, a self-encoder learns prior from the high-quality input image. In Stage II, the prior is generated through a LDM, conditioned on the decoded image of an existing learning-based image compression algorithm, to be used as auxiliary information for generating the texture-rich enhanced images. To better utilize the prior, a channel attention and gate-based dynamic feature attention module (DFAM) is embedded into a Transformer-based multi-scale enhancement network (MEN) for image enhancement. Extensive experimental results demonstrate the proposed LDM-RSIC outperforms existing state-of-the-art traditional and learning-based image compression algorithms in terms of both subjective perception and objective metrics. The code will be available at https://github.com/mlkk518/LDM-RSIC. Jutao Li, Xingsong Hou, Huake Wang |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2025 | Multi-Scale Retinex Unfolding Network for Low-Light Image EnhancementabstractRetinex theory-based low-light image enhancement methods have received increasing attention and achieved tremendous advancements. However, there still exist two seldom-explored issues: 1) The above methods only formally simulate the Retinex decomposition, resulting in lacking explicit interpretability. 2) They usually are performed in single-scale space, leading to suboptimal enhancement results. In this paper, we propose an interpretable Multi-scale Retinex Unfolding Network (MRUNet) for low-light image enhancement, which can tackle both of the aforementioned issues simultaneously. Specifically, we formulate low-light image enhancement as a multi-scale Retinex optimization problem and design an iteration minimization solution to solve it. The optimization solution is further unfolded to fabricate MRUNet, which is empowered with clear physical significance and multi-scale prior knowledge in favor of image enhancement. However, it will aggravate model size and efficiency when exploiting multiple proximal mapping networks to extract multi-scale prior from multi-scale inputs. To surmount the issue, we propose a Scale-Aware Proximal mapping Module (SAPM), which efficiently collect multi-scale prior knowledge via the weight sharing strategy. In SAPM, we tailor a scale-aware transformer to model the specific scale-similarity among different scales. Extensive experiments manifest that MRUNet surpasses other Retinex-based low-light image enhancement methods on multiple benchmarks. Huake Wang, Xingsong Hou, Jutao Li, Yadi Yan, Wenke Sun, Kaibing Zhang, Xiangyong Cao |
IEEE Trans. Multim. | 2 |
| 2025 | Discard Significant Bits of Compressed Sensing: A Robust Image Coding for Resource-Limited ContextsabstractCompressed sensing (CS) provides a robust and simple framework for compressing images in resource-constrained environments. However, CS-based image coding schemes often have poor rate-distortion (R-D) performance, particularly due to the quantization process. Our research indicates that leveraging the image prior enables the estimation of most significant bits (MSBs) from least significant bits (LSBs), which provides a quantization strategy to improve R-D performance without increasing coding complexity. That is discarding MSBs of measurements, and only transmitting LSBs to the decoder side. At the decoder side, we reconstruct images by solving an inverse-quantization set-constrained CS optimization problem. Our approach further employs a tailored designed deep denoiser as the proximal operator to enhance the reconstructed image quality. Extensive experimental results demonstrate that the proposed scheme achieves satisfactory performance, with promising R-D results (PSNR gains over 1.71 dB than JPEG at 0.50 bpp compression ratio), and robust bit error and loss resilience (reconstructed 29.98 dB even with 50% bit loss at 0.50 bpp compression ratio), meanwhile having lower encoding complexity (less than half encoding time of CCSDS-IDC). Tao Wang 0147, Jun Li 0123, Yuanjing Feng, Xueming Qian, Xingsong Hou |
ACM Trans. Multim. Comput. Commun. Appl. | 7 |
| 2024 | QueryCDR: Query-Based Controllable Distortion Rectification Network for Fisheye Images
Pengbo Guo, Chengxu Liu 0001, Xingsong Hou, Xueming Qian |
ECCV (15) | 3 |
| 2024 | AMP-BCS: AMP-based image block compressed sensing with permutation of sparsified DCT coefficients
Xingsong Hou, Huake Wang, Shuhao Bi, Xueming Qian |
J. Vis. Commun. Image Represent. | 2 |
| 2024 | Division gets better: Learning brightness-aware and detail-sensitive representations for low-light image enhancement
Huake Wang, Xiaoyang Yan, Xingsong Hou, Yujie Dun, Kaibing Zhang |
Knowl. Based Syst. | 3 |
| 2024 | The Design of an Adaptive Enhanced AMP-Based Image Block Compressed Sensing and Its Application to Image EncryptionabstractCompressed sensing (CS) has become a widely employed technique in the field of image encryption. Despite achieving a high level of data encryption security, the resulting decrypted images often lack satisfactory quality. In this paper, we develop an adaptive enhanced approximate message passing (AMP) block CS (AE-AMP-BCS) algorithm for image BCS and its application to image encryption. Initially, we present an adaptive energy-based anti-aliasing filtering strategy (AAFS) for preprocessing the input image, mitigating the noise effect during reconstruction. The filtered image is then transformed into the Haar domain for enhanced security, followed by a twofold perturbation operation and a piece-wise linear chaotic system-based measurement generation approach for ratio allocation and security consideration. Finally, the cipher measurements are decrypted using an existing AMP-based algorithm. It is noteworthy that the keys for perturbation operations and the chaotic system are derived using the SHA512 hash. Comprehensive experimental results demonstrate that the proposed AE-AMP-BCS significantly outperforms state-of-the-art BCS methods in terms of image reconstruction quality. Its application in image encryption showcases competitive encryption capabilities, strong robustness, and outstanding image decryption quality compared to other CS-based image encryption algorithms. Xingsong Hou |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Object-Fidelity Remote Sensing Image Compression With Content-Weighted Bitrate Allocation and Patch-Based Local AttentionabstractAs the contradiction between high-resolution remote sensing (RS) image acquisition and limited storage space or bandwidth becomes increasingly prominent, the importance of prioritizing the compression of object regions has become evident. However, how to achieve object-fidelity RS image compression at low bitrates remains a challenging task. In this paper, we use a deep neural network to develop a novel object-fidelity RS image compression (OF-RSIC) method. Firstly, an object detection algorithm is employed to split the image into object and background regions, followed by background region smoothing through bilinear downsampling to reduce the bitrate required for background region encoding. Subsequently, a Transformer-based content-weighted attention module (CWAM) is developed for adaptive bits allocation. This module allows the model to capture global correlations among pixels and generates a more compact representation for the latent features. Additionally, a patch-based local attention module (PLAM) is proposed to reweight the local information of the entropy model, thereby improving the rate-distortion (RD) performance. Finally, to constrain the bitrate allocation between the background and object regions, a region-differentiated loss is introduced for the model training. To assess the efficacy of the proposed method, we employ an object detection algorithm to select four classes of object images from the DIOR dataset for experimentation. Comprehensive experimental results demonstrate that OF-RSIC outperforms state-of-the-art image compression algorithms in terms of object fidelity at lower bitrates. Xingsong Hou |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | CRS-Diff: Controllable Remote Sensing Image Generation With Diffusion ModelabstractThe emergence of generative models has revolutionized the field of remote sensing (RS) image generation. Despite generating high-quality images, existing methods are limited in relying mainly on text control conditions, and thus do not always generate images accurately and stably. In this article, we propose CRS-Diff, a new RS generative framework specifically tailored for RS image generation, leveraging the inherent advantages of diffusion models while integrating more advanced control mechanisms. Specifically, CRS-Diff can simultaneously support text-condition, metadata-condition, and image-condition control inputs, thus enabling more precise control to refine the generation process. To effectively integrate multiple condition control information, we introduce a new conditional control mechanism to achieve multiscale feature fusion (FF), thus enhancing the guiding effect of control conditions. To the best of our knowledge, CRS-Diff is the first multiple-condition controllable RS generative model. Experimental results in single-condition and multiple-condition cases have demonstrated the superior ability of our CRS-Diff to generate RS images both quantitatively and qualitatively compared with previous methods. Additionally, our CRS-Diff can serve as a data engine that generates high-quality training data for downstream tasks, e.g., road extraction. The code is available athttps://github.com/Sonettoo/CRS-Diff. Datao Tang, Xiangyong Cao, Xingsong Hou, Zhongyuan Jiang, Junmin Liu, Deyu Meng |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Hierarchical Kernel Interaction Network for Remote Sensing Object CountingabstractDifferent from object counting in surveillance scenes, remote sensing object counting encounters knotty challenges due to its tiny scale and cluttered background. However, existing remote sensing counting methods pursue favorable performance by sacrificing resolution to obtain semantic information, resulting in the loss of significant features of tiny-scale objects. To surmount the above issue, we propose a novel hierarchical kernel interaction network, dubbed HKINet, for remote sensing object counting. HKINet is comprised of several hierarchical kernel interaction modules (HKIMs) to simultaneously preserve high-resolution features and extract deep-layer semantic information. Specifically speaking, HKIM hierarchically performs multiresolution convolutions to avoid information loss in low resolution. Moreover, a scale interaction block (SIB) is used to combine multiresolution features for semantic information interaction. Finally, hierarchical resolutions are fused to output the prediction density map. To validate the superiority of our proposed HKINet, we conduct extensive experiments on four remote sensing object counting datasets, e.g., RSOC, CARPK, PUCPR+, and DroneCrowd datasets, and experimental results demonstrate HKINet outperforms other state-of-the-art remote sensing counting methods in terms of mean absolute error (MAE) and root mean squared error (RMSE). Huake Wang, Jinjiang Wei, Xingsong Hou, Hengfeng Wu, Kaibing Zhang |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2023 | Privacy-Enhanced Zero-Shot Learning via Data-Free Knowledge TransferabstractConsidering the increasing concerns about data copyright and sensitivity issues, we present a novel Privacy-Enhanced Zero-Shot Learning (PE-ZSL) paradigm. The key innovation is to involve a teacher model as the data safeguard to guide the PE-ZSL model training without data sharing. The PE-ZSL model consists of a generator and student network, which can achieve data-free knowledge transfer while maintaining the performance of teacher model. We investigate ‘black-’ and ‘white-box’ scenarios in PE-ZSL task as different levels of framework privacy. Besides, we provide the discussion of teacher model in both omniscient and quasi-omniscient settings according to the knowledge space. Despite simple implementations and data-missing disadvantages, our PE-ZSL framework can retain state-of-the-art ZSL and GZSL performance under the ‘white-box’ scenario. Extensive qualitative and quantitative analysis also demonstrates promising results when deploying the model under ‘black-box’ scenario. Fan Wan, Daniel Organisciak, Jiyao Pu, Haoran Duan 0001, Peng Zhang 0058, Xingsong Hou, Yang Long 0001 |
ICME | 7 |
| 2023 | Robust image compression-encryption via scrambled block bernoulli sampling with diffusion noiseabstractAbstract This paper proposed an image compression‐encryption scheme based on compressive sensing theory, which achieves high security, strong robustness, and high rate‐distortion performance. First, the denoising preprocessing strategy is applied at the encoder side, which can enhance the rate‐distortion performance without sacrificing security and robustness. Second, the preprocessed image is randomly down‐sampled using scrambled block Bernoulli sampling with diffusion noise (SBBS‐DN), which is generated by combining a hyper‐chaotic system and SHA256 hash of the plain image. Third, a deep‐learned plug‐and‐play is embedded prior for plain image reconstruction at the decoder side. Simulation results show that the proposed scheme has desirable security performance (being resistant to different attacks), high R‐D performance (PSNR gains over 1.3 dB than JPEG at 0.50 bpp compression ratio), and high error resilience (reconstructed 29.92 dB at 0.50 bpp compression ratio even with 50% bit loss). Chaocheng Ma, Tao Wang 0147, Yuanjing Feng, Xingsong Hou, Xueming Qian |
IET Image Process. | 5 |
| 2023 | Directional lifting wavelet transform domain image steganography with deep-based compressive sensing
Chaocheng Ma, Yuanjing Feng, Xingsong Hou, Xueming Qian |
Multim. Tools Appl. | 4 |
| 2023 | Versatile Denoising-Based Approximate Message Passing for Compressive SensingabstractApproximate message passing-based compressive sensing reconstruction has received increasing attention, the performance of which depends heavily on the ability of the denoising operator. However, most methods only employ an off-the-shelf denoising model as the denoising operator of the iteration solver, which imposes an unfavorable limit on reconstruction performance of compressive sensing. To solve the aforementioned issue, we propose a novel versatile denoising-based approximate message passing model, abbreviated as VD-AMP, for compressive sensing (CS) recovery. To be specific, we meticulously design a double encoder-decoder denoising network (DEDNet), which manifests the impressive performance in Gaussian denoising. Moreover, a fine-grained noise level division (FNLD) solution is proposed to release the potential of the well-designed DEDNet so as to improve the reconstruction performance. However, strengthening the denoiser alone fails to remove the distortion artifact of reconstruction images at low sampling rates. To alleviate the defect, we propose an anti-aliasing sampling (AS), which firstly maps the input image to a smoothing sub-space using the proposed DEDNet before vanilla sampling, reducing aliasing between high-frequency and low-frequency information on measurement. Extensive experiments on benchmark datasets demonstrate that the proposed VD-AMP significantly outperforms state-of-the-art CS reconstruction models by a large margin, e.g., up to 2 dB gains on PSNR. Huake Wang, Xingsong Hou |
IEEE Trans. Image Process. | 3 |
| 2023 | Visual-Semantic Aligned Bidirectional Network for Zero-Shot LearningabstractZero-shot learning (ZSL) aims to recognize unknown categories that are unavailable during training. Recently, generative models have shown the potential to address this challenging problem by synthesizing unseen features conditioned on semantic embeddings such as attributes. However, unidirectional generative models cannot guarantee the effective coupling between visual and semantic spaces. To this end, we propose a visual-semantic aligned bidirectional network with cycle consistency to alleviate the gap between these two spaces, generating unseen features of high quality. More importantly, we incorporate two carefully designed strategies into our bidirectional framework to improve the overall ZSL performance. Specifically, we enhance the intra-domain class divergence in both visual and semantic spaces, and in the meantime, mitigate the inter-domain shift to preserve seen-unseen domain discrimination. Experimental results on four standard benchmarks show the superiority of our framework over existing state-of-the-art methods under both conventional and generalized ZSL settings. Xingsong Hou, Jie Qin 0004, Yuming Shen, Yang Long 0001, Li Liu 0004, Zhao Zhang 0001, Ling Shao 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | Compressive Sensing Image Steganography via Directional Lifting Wavelet Transform
Chaocheng Ma, Yuanjing Feng, Xingsong Hou |
ICICS (2) | 4 |
| 2021 | Learning Deformable and Attentive Network for image restoration
Xingsong Hou, Yujie Dun, Jie Qin 0004, Li Liu 0004, Xueming Qian, Ling Shao 0001 |
Knowl. Based Syst. | 2 |
| 2020 | Personalized location recommendation by fusing sentimental and spatial context
Guoshuai Zhao 0001, Peiliang Lou, Xueming Qian, Xingsong Hou |
Knowl. Based Syst. | 4 |
| 2020 | Compressive Sensing Multi-Layer Residual Coefficients for Image CodingabstractCompressive sensing (CS)-based image coding scheme has been enthusiastically studied, but it still has a poor rate-distortion performance compared with the traditional image coding techniques. In this paper, we propose a CS multi-layer residual coding scheme to rectify this problem to a certain extent. By dividing CS measurements into multi-layers and predicting a particular layer's measurements with all its preceding layers' measurements, we can transform CS measurements into multi-layer residual coefficients, which are easier to compress. By calculating the residual between the quantized ground-truth CS measurements and their corresponding quantized inference measurements and using Huffman coding to associate each residual quantization index with a binary code, we can reduce the redundancies among CS measurements efficiently. Besides, the prediction and quantization process is designed to be layer-independent, which can save much of the encoding time. The proposed approach introduces a novel framework for using CS in the compression domain. The experimental results show that the proposed scheme can significantly outperform JPEG2000 and approach or reach the performance of HEVC-Intra on some test images. Xingsong Hou, Ling Shao 0001, Chen Gong 0001, Xueming Qian |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2020 | Sketch-Based Image Retrieval With Multi-Clustering Re-RankingabstractTo improve the performance of sketch-based image retrieval (SBIR) methods, most existing SBIR methods develop brand new SBIR methods. In fact, a re-ranking approach, which can refine the retrieval results of SBIR methods, is also beneficial. Inspired by this, in this paper, an SBIR re-ranking approach based on multi-clustering is proposed. In order to make the re-ranking approach invisible to users and adaptive to different types of image datasets, we made it an unsupervised method using blind feedback. Distinguished from the existing methods, this re-ranking approach uses the semantic information of three types of images: edge maps, object images (images with black background and natural images' foreground objects) and natural images themselves. With the initial retrieval results of an SBIR method, our approach first does the clustering operation for three types of images. Then, we utilize the clustering results to generate a cluster score for each initial retrieval result. Finally, the cluster score is used to calculate the final retrieval scores for the initial retrieval results. The experiments on different SBIR datasets are conducted. Experimental results demonstrate that, by implementing our re-ranking approach, the retrieval accuracy of a variety of SBIR methods is increased. Furthermore, the comparisons between our re-ranking method and the existing re-ranking methods are given. Luo Wang, Xueming Qian, Xingjun Zhang, Xingsong Hou |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | Zero-VAE-GAN: Generating Unseen Features for Generalized and Transductive Zero-Shot LearningabstractZero-shot learning (ZSL) is a challenging task due to the lack of unseen class data during training. Existing works attempt to establish a mapping between the visual and class spaces through a common intermediate semantic space. The main limitation of existing methods is the strong bias towards seen class, known as the domain shift problem, which leads to unsatisfactory performance in both conventional and generalized ZSL tasks. To tackle this challenge, we propose to convert ZSL to the conventional supervised learning by generating features for unseen classes. To this end, a joint generative model that couples variational autoencoder (VAE) and generative adversarial network (GAN), called Zero-VAE-GAN, is proposed to generate high-quality unseen features. To enhance the class-level discriminability, an adversarial categorization network is incorporated into the joint framework. Besides, we propose two self-training strategies to augment unlabeled unseen features for the transductive extension of our model, addressing the domain shift problem to a large extent. Experimental results on five standard benchmarks and a large-scale dataset demonstrate the superiority of our generative model over the state-of-the-art methods for conventional, especially generalized ZSL tasks. Moreover, the further improvement of the transductive setting demonstrates the effectiveness of the proposed self-training strategies. Xingsong Hou, Jie Qin 0004, Jiaxin Chen 0002, Li Liu 0004, Fan Zhu 0001, Zhao Zhang 0001, Ling Shao 0001 |
IEEE Trans. Image Process. | 2 |
| 2020 | Location Recommendation for Enterprises by Multi-Source Urban Big Data AnalysisabstractEffective location recommendation is an important problem in both research and industry. Much research has focused on personalized recommendation for users. However, there are more uses such as site selection for firms and factories. In this study, we try to solve site selection problem by recommending some locations satisfying special requirements. There are many factors affecting it, including functions of architecture, building cost, pollution discharge etc. We focus on the specific site selection of meteorological observation stations in this paper with leveraging the factors of functions of architecture and building cost from multi-source urban big data. We consider not only recommending the locations that can provide more accurate prediction and cover more areas, but also minimizing the cost of building new stations. We design an extensible two-stage framework for the station placing including prediction model and recommendation model. It is very convenient for executives to add more real-life factors into our approach. We have some empirical findings and evaluate the proposed approach using the real meteorological data of Shaanxi province, China. Experiment results show the better performance of our approach than existing commonly used methods. Guoshuai Zhao 0001, Tianlei Liu, Xueming Qian, Huan Wang 0002, Xingsong Hou, Zhetao Li |
IEEE Trans. Serv. Comput. | 6 |
| 2019 | Compressive-Sensed Image Coding via Multi-layer Closed-Loop PredictionabstractThese years have seen the advance of compressive sensing (CS), but the CS-based image coding scheme still has a poor rate-distortion (R-D) performance compared with the traditional image coding techniques. In this paper, we propose an image coding scheme based on the CS paradigm via multi-layer closed-loop prediction. In the scheme, we divide CS measurements into multi-layers and predict a particular layer's measurements with all its preceding layers' measurements, which can reduce the redundancies between CS measurements efficiently. The produced measurement residuals are then quantized into binary codes, which are tremendously reduced compared to quantizing the CS measurements directly. Furthermore, We provide a non-local low-rank CS reconstruction algorithm corresponding to our multi-layer closed-loop prediction scheme. Experimental results verify that the proposed scheme can significantly outperform JPEG2000, and the reconstruction quality of our scheme is no worse or even better than that of HEVC-Intra. Xingsong Hou, Ling Shao 0001 |
DCC | 2 |
| 2018 | Recurrent Transformer Network for Remote Sensing Scene Categorisation
Xingsong Hou, Ling Shao 0001 |
BMVC | 3 |
| 2018 | POI Summarization by Aesthetics Evaluation From Crowd Source Social MediaabstractPlace-of-Interest (POI) summarization by aesthetics evaluation can recommend a set of POI images to the user and it is significant in image retrieval. In this paper, we propose a system that summarizes a collection of POI images regarding both aesthetics and diversity of the distribution of cameras. First, we generate visual albums by a coarse-to-fine POI clustering approach and then generate 3D models for each album by the collected images from social media. Second, based on the 3D to 2D projection relationship, we select candidate photos in terms of the proposed crowd source saliency model. Third, in order to improve the performance of aesthetic measurement model, we propose a crowd-sourced saliency detection approach by exploring the distribution of salient regions in the 3D model. Then, we measure the composition aesthetics of each image and we explore crowd source salient feature to yield saliency map, based on which, we propose an adaptive image adoption approach. Finally, we combine the diversity and the aesthetics to recommend aesthetic pictures. Experimental results show that the proposed POI summarization approach can return images with diverse camera distributions and aesthetics. Xueming Qian, Ke Lan, Xingsong Hou, Zhetao Li, Junwei Han 0001 |
IEEE Trans. Image Process. | 4 |
| 2018 | Efficient and Robust Image Coding and Transmission Based on Scrambled Block Compressive SensingabstractImage transmission in a wireless visual sensor network (WVSN) with limited resources over an unreliable and bandwidth-limited wireless channel is challenging. This paper presents a highly efficient and robust image coding and transmission scheme with a simple encoder based on compressive sensing (CS) for WVSNs. First, an image measurement based on scrambled block compressive sampling with a separable sensing operator is proposed to simplify the encoder. Second, a progressive nonuniform quantization, which exploits the measurement distribution at the encoder side and the measurement dependencies at the decoder side, is designed to improve the rate-distortion (R-D) performance while maintaining low complexity at the encoder. Third, to further improve the R-D performance, a progressive non-local low-rank reconstruction is designed at the decoder. The experimental results show that the proposed scheme can achieve higher R-D performance compared with the benchmark CS-based image coding and transmission schemes. Higher robustness can be achieved compared with the traditional source-channel coding, such as Consultative Committee for Space Data Systems$-$Image Data Compression (CCSDS-IDC) with Raptor codes under a time-varying packet loss channel, and the encoding time can be significantly reduced compared with the traditional image coding schemes. The experimental results also show that the proposed scheme achieves state-of-the-art coding efficiency with lower computational complexity at the encoder while still supporting error resilience. Xingsong Hou, Xueming Qian, Chen Gong 0001 |
IEEE Trans. Multim. | 2 |
| 2017 | Image Location Inference by Multisaliency EnhancementabstractLocations of images have been widely used in many application scenarios for large geotagged image corpora. As to images that are not geographically tagged, we estimate their locations with the help of the large geotagged image set by content-based image retrieval. Bag-of-words image representation has been utilized widely. However, the individual visual word-based image retrieval approach is not effective in expressing the salient relationships of image region. In this paper, we present an image location estimation approach by multisaliency enhancement. We first extract region-of-interests (ROIs) by mean-shift clustering on the visual words and salient map of the image based on which we further determine the importance of the ROI. Then, we describe each ROI by the spatial descriptors of visual words. Finally, region-based visual phrases are generated to further enhance the saliency in image location estimation. Experiments show the effectiveness of our proposed approach. Xueming Qian, Huan Wang 0002, Yisi Zhao, Xingsong Hou, Richang Hong, Meng Wang 0001, Yuan Yan Tang |
IEEE Trans. Multim. | 4 |
| 2016 | Compressive sensing reconstruction for compressible signal based on projection replacement
Xingsong Hou, Chen Gong 0001, Xueming Qian |
Multim. Tools Appl. | 2 |
| 2015 | SAR complex image data compression based on quadtree and zerotree Coding in Discrete Wavelet Transform Domain: A Comparative Study
Xingsong Hou, Chen Gong 0001, Xueming Qian |
Neurocomputing | 1 |
| 2015 | Landmark Summarization With Diverse ViewpointsabstractLandmark summarization with diverse viewpoints is very important in landmark retrieval, as it can create a comprehensive description of a landmark for users. In this paper we present an approach for summarizing a collection of landmark images from diverse viewpoints. First, we group landmark images with content overlap by viewpoint album (VA) generation. Second, we model the relative viewpoint of each image within the VA based on the spatial layout of distinctive descriptors of a landmark. Third, we express the relative viewpoint of an image with a 4-D viewpoint vector, including horizontal, vertical, scale, and rotation. Finally, we summarize the landmarks in terms of viewpoints. Experimental results show the effectiveness of the proposed landmark summarization approach. Xueming Qian, Xiyu Yang, Yuan Yan Tang, Xingsong Hou, Tao Mei 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2014 | Robust and efficient SAR image coding transmission based on compressive sensingabstractIn this work, a new robust and efficient airborne synthetic aperture radar (SAR) image coding transmission scheme based on compressive sensing (CS) against lossy channels is proposed. The robustness is achieved using the democracy of CS. Considering the poor R-D performance of the traditional CS due to SAR image's weak sparsity, we use directional lifting wavelet transform (DLWT) as sparse representation and sparse-filtering to eliminate the interference of small coefficients. By exploiting the inter-scale dependency of DLWT coefficients, an efficient Bayesian reconstruction algorithm is adopted. Furthermore, optimal tradeoff between bit-depth and measurement rate is used. Experimental results show that the proposed scheme is more robust against packet loss compared with the traditional joint source-channel coding (JSCC) scheme. When the packet loss rate (PLR) is excessive, the JSCC scheme easily leads to cliff effect, however, the R-D performance of the proposed scheme decreases more gracefully while achieving a comparative R-D performance. Xingsong Hou, Wenwen Tian |
ICIP | 1 |
| 2014 | Personalized tag recommendation for Flickr usersabstractSocial image share websites such as Flickr allow users to manually annotate their images with their own words, which can be used to facilitating image retrieval and other image applications. For the vast number of online images contributed by social users, existing methods on tag recommendation haven't taken users' characteristics and tagging habits into consideration. In this paper, we propose a personalized tag recommendation system for Flickr users. It can recommend users personalized tags for their newly uploaded photos based on the history information in their social communities. We carry out the personalized tag recommendation from three aspects. First, the tags we recommend to users are users' own vocabularies. Second, different recommendation methods are implemented to different users. Third, different users are recommended with different number of tags based on their tagging habits. The experimental results indicate that our personalized tag recommendation is effective. Xueming Qian, Dan Lu 0003, Xingsong Hou |
ICME | 4 |
| 2014 | SAR image Bayesian compressive sensing exploiting the interscale and intrascale dependencies in directional lifting wavelet transform domain
Xingsong Hou, Chen Gong 0001, Jinqiang Sun, Xueming Qian |
Neurocomputing | 1 |
| 2014 | HWVP: hierarchical wavelet packet descriptors and their applications in scene categorization and semantic concept retrieval
Xueming Qian, Danping Guo, Xingsong Hou, Zhi Li 0003, Huan Wang 0002, Guizhong Liu |
Multim. Tools Appl. | 3 |
| 2014 | Video text detection and localization in intra-frames of H.264/AVC compressed video
Xueming Qian, Huan Wang 0002, Xingsong Hou |
Multim. Tools Appl. | 3 |
| 2013 | Tagging photos using users' vocabularies
Xueming Qian, Youtian Du, Xingsong Hou |
Neurocomputing | 5 |
| 2013 | Complex SAR Image Compression Based on Directional Lifting Wavelet Transform With High Clustering CapabilityabstractWe propose two synthetic aperture radar (SAR) complex image compression schemes based on DLWT_IQ and DLWT_FFT. DLWT_IQ encodes the real parts and imaginary parts of the images using directional lifting wavelet transform (DLWT) and bit plane encoder (BPE), while DLWT_FFT encodes the real images converted by fast Fourier transform (FFT). Compared with discrete wavelet transform-IQ (DWT_IQ), DLWT_IQ improves the peak signal-to-noise ratio (PSNR) up to 1.28 dB and reduces the mean phase error (MPE) up to 21.74%; and compared with DWT_FFT, DLWT_FFT improves the PSNR up to 1.22 dB and reduces the MPE up to 20.32%. Moreover, the proposed schemes increase the PSNR up to 3.34 dB and decrease the MPE up to 50.43% as compared with the set partitioning in hierarchical trees (SPIHT) algorithm. In addition to this, we observe a novel phenomenon, that is, DLWT with direction prediction achieves a higher clustering capability for complex SAR images than DWT. Then, coding algorithm based on DLWT requires fewer coding bits than DWT for the same number of coding coefficients, and DLWT outperforms DWT in terms of rate-distortion performance even if the K-term nonlinear approximation of DWT is better than that of DLWT. Xingsong Hou, Guifeng Jiang, Xueming Qian |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2012 | Tag filtering based on similar compatible principleabstractIn social image sharing websites, users provide several descriptive tags to annotate their shared images. Usually, the raw tags are noisy, biased and incomplete. How to filter the tags is important for tag based applications. In this paper, a similar compatible principle based tag filtering approach is proposed. We classify tags into two sets. One is relevant to image content. The other is irrelevant to image content. We filter the tags by ranking high relevant tags ahead of the tags with low relevance by the similar compatible principle. This approach determines the ranks of user annotated tags by maximizing the compatible value of changing the labels of the tags from irrelevant to relevant at each step. Experiments on crawled Flickr dataset demonstrate the effectiveness of the proposed approach. Xueming Qian, Xian-Sheng Hua 0001, Xingsong Hou |
ICIP | 3 |
| 2012 | HMM based soccer video event detection using enhanced mid-level semantic
Xueming Qian, Huan Wang 0002, Guizhong Liu, Xingsong Hou |
Multim. Tools Appl. | 4 |
| 2011 | Joint optimization coding for level and map information in H.264/AVC
Xingsong Hou, Duan Xue, Baiping Jin, Lijuan Cao |
Signal Process. Image Commun. | 1 |
| 2010 | Directional lifting wavelet and universal trellis coded quantization based image coding algorithm and objective quality evaluationabstractIn this paper, an image coding algorithm based on directional lifting wavelet transform (DLWT) and universal trellis coded quantization (UTCQ) is presented, and the coding performance is evaluated with multi-scale structural similarity index (MSSIM) and peak signal-to-noise ratio (PSNR). Compared with discrete wavelet transform (DWT), DLWT can provide an efficient representation of edges, but shows a similar ability in representing the smooth area. In order to alleviate blurring artifacts in the smooth area, UTCQ is adopted to quantizing the wavelet coefficients. Experimental results show that the proposed algorithm has the best MSSIM performance among the compared algorithms (including JPEG2000), and its decoded images at low bit-rate are visually more appealing in both edges and smooth areas. The experimental results also show that UTCQ does perform better than scalar quantization (SQ) in MSSIM and improves the subjective visual quality, although it is not necessary better than SQ in PSNR. Xingsong Hou, Guifeng Jiang, Rongjing Ji, Chenglong Shi |
ICIP | 1 |
| 2007 | Investigation of H.264 intra coding for SAR imageabstractIn this paper we investigate the performance of H.264 Intra Coding for Synthetic Aperture Radar(SAR) image. The results show that H.264 Intra Coding is a high performance coder for the SAR image. However, when the SAR image is despeckled, H.264 Intra coding is not so efficient than the wavelet based image coder, such as JPEG2000 and SPIHT. Then more efficient representation is needed when H.264 Intra Coding is used to code the despeckled SAR image. Xingsong Hou, Yujie Dun, Rongjing Ji |
IGARSS | 1 |
| 2004 | MRF based construction of statistical operator and its application
Hongliang Li 0001, Guizhong Liu, Yongli Li 0001, Xingsong Hou |
Sci. China Ser. F Inf. Sci. | 4 |
| 2004 | SAR image data compression using wavelet packet transform and universal-trellis coded quantizationabstractA wavelet packet image coding algorithm for synthetic aperture radar (SAR) image data compression is proposed in this paper. High rate-distortion (R-D) performance of this algorithm [wavelet packet quantization universal trellis-coded quantization (WPQUTCQ)] is achieved by incorporating the wavelet packet transform for representing the rich texture information of SAR image data, the quadtree technique for classifying the wavelet packet coefficients, and the universal trellis-coded quantization (UTCQ). For a typical SAR image, WPQUTCQ outperforms set partitioning in hiearchical trees and JPEG2000 by about 1.21 and 0.81 dB in peak signal-to-noise ratio (PSNR) at 2 bpp, respectively. The effect of best basis selection on the R-D performance of coding algorithm for SAR image data compression is also evaluated through comparing the effects of different cost functions (e.g., entropy, the texture energy, which is designed for SAR image data compression particularly, and the cost function, based on the actual coding strategy proposed by us) on the R-D performance of WPQUTCQ. The experimental results show the suitability of the cost function based on the actual coding strategy for SAR image data compression when compared with other cost functions. For a typical SAR image, the coding results by WPQUTCQ corresponding to the cost function based on the actual coding strategy achieve gains 0.59 and 0.49 dB on average in PSNR when compared with the coding results corresponding to entropy and the texture energy, respectively. Xingsong Hou, Guizhong Liu, Yiyang Zou |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2003 | Embedded quadtree-based image compression in DCT domainabstractThe success in discrete cosine transform(DCT) image coding is mainly attributed to recognition of the importance of data organization and representation. In this paper, we proposed an embedded image coder based on quadtree set partition in DCT domain (EZDCT) which is suitable for many kinds of DCT coefficients reorganization schemes. The experimental results show that it is among the state-of-the-art DCT-based image coders when compared with the famous DCT-based image coders, such as EZDCT and MRDCT. For example, for the Barbara image, EQDCT outperforms JPEG EZDCT and MRDCT by 3.3,1.71,1.70 dB in peak-signal-to-noise ratio at 0.25 bpp, respectively. Xingsong Hou, Guizhong Liu, Yiyang Zou |
ICASSP (3) | 1 |
| 2003 | A wavelet packet image coding algorithm based on quadtree classification and UTCQabstractIn this paper, we present a wavelet packet image coding algorithm based on quadtree classification and UTCQ. It is composed of four parts: (1) wavelet packet decomposition and best basis selection based on a new cost function, (2) a quadtree classification procedure, used to classify the wavelet packet coefficients into two sets: a significant one and an insignificant one, (3) the universal trellis coded quantization, used to code the significant coefficients sets, (4) the entropy coder, used to code the indices of the universal trellis coded quantizer output. The image coding results, calculated from actual file sizes and images reconstructed by the decoding algorithm are either comparable to or surpass previous results for texture-rich images. Xingsong Hou, Guizhong Liu |
ICME | 1 |
| 2003 | Wavelet packet remote-sensing images coding algorithm based on quadtree classification and UTCQabstractAbstract – In this paper, we propose a new wavelet packet image coding technique for synthetic aperture radar(SAR) remote-sensing image data compression. The image compressibility of this algorithm(WPQTCQ)is improved by incorporating wavelet packet transform for rich texture component of SAR image data, the quadtree set partition for sorting the wavelet packet transform coefficients and universal trellis coded quantization. The experimental results show that WPQTCQ is among the state-of-theart image coders for SAR image data. To a typical SAR image at 2bpp, WPQTCQ out performs SPIHT and JPEG2000 by about 2.45dB and 1.34dB in peak signal-to-noise ratio, respectively. Xingsong Hou, Guizhong Liu, Yiyang Zou |
IGARSS | 1 |
| 2002 | An embedded wavelet packet image coding algorithmabstractIn this paper, we presented a novel wavelet packet image coding approach which provides the functionality of fine granular bitstream scalability. The proposed progressive wavelet packet image coding scheme consists of three parts, wavelet packet decomposition, a quadtree sorting procedure for classifying wavelet coefficients and universal trellis-coded quantization for quantizing the sorted coefficients. The image coding results, calculated in PSNR and images reconstructed by the decoding algorithm, are either comparable to or surpass previous results, due to the flexible representation ability of wavelet packet, the effective quadtree classifier and the improved granular fidelity of the UTCQ over scalar quantization. Xingsong Hou, Guizhong Liu, Hongliang Li 0001, Yongli Li 0001 |
ICASSP | 1 |
| 2002 | Wavelet-based analysis of hurst parameter estimation for self-similar trafficabstractIn order to guarantee quality of service (QoS) over Internet, traffic analysis, traffic management have been active research areas. A lot of facts show the Internet traffic and variable bit rate videos streaming all are characterized by self-similar property. Hurst parameter as an important factor that reflects the self-similar property is a key to traffic management and QoS. In this paper existing wavelet methods for the estimation of the Hurst parameter of self-similar traffic is systematically analyzed and examined. The effects of wavelet functions, vanishing moments and wavelet decomposition levels to the results of wavelet methods for acquiring the Hurst parameter are investigated via numerical experiments. Some useful conclusions are drawn on the relationship between the accuracy of the methods and the selection of the order of vanishing moments and the selection of wavelet functions. Yongli Li 0001, Guizhong Liu, Hongliang Li 0001, Xingsong Hou |
ICASSP | 4 |
| 2002 | The construction of a statistical prediction lifting operator and its applicationabstractA new method of nonseparable nonlinear wavelet decomposition is proposed, which is suited for the task of image compression, especially for lossless coding applications. It is based on a certain statistical operator that is defined here according to the Markov random field theory. In contrast to the previous nonlinear predictors such as the median or morphological operators, this statistical operator can sufficiently take advantage of the statistical correlation between neighboring pixels. It can be used to realize integer-valued wavelet transforms, which can avoid quantization with the image detail signals being zero (or almost zero) in the smooth gray-level variation areas at a big probability. Numerical results show that the entropy of the coefficients in the transform domain obtained with this new method is smaller than that obtained with the other nonlinear transform methods. Hongliang Li 0001, Guizhong Liu, Yongli Li 0001, Xingsong Hou |
ICIP (1) | 4 |