EDBT 2026 Demo / reviewers in the wild / expert
Haiyong Xu
dblp:13/10180
· DBLP profile ↗
48ranked-venue papers
2as first author
36since 2021 · last 2026
0000-0003-1590-6799ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 32 · 1 first-author · 23 since 2021Artificial intelligence and machine learning · 12 · 10 since 2021Computer networks · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ARMLF: Anomalous region representation learning for multi-exposure fused light field image quality assessment
Guanglong Liao, Gangyi Jiang, Linwei Zhu, Yeyao Chen, Yueli Cui, Ting Luo 0001, Haiyong Xu |
Expert Syst. Appl. | 7 |
| 2026 | LatentDark: Reflectance guided latent diffusion model for low-light image enhancement
Renzhi Hu, Ting Luo 0001, Gangyi Jiang, Leiming Liu, Yeyao Chen, Haiyong Xu, Zhouyan He |
Signal Process. | 6 |
| 2026 | Water-KAN: Efficient Underwater Image Enhancement via Kolmogorov-Arnold Networks
Peiyuan Jin, Haiyong Xu, Yeyao Chen, Gangyi Jiang |
IEEE Signal Process. Lett. | 2 |
| 2026 | DiffW: Multi-Encoder Based on Conditional Diffusion Model for Robust Image WatermarkingabstractThe existing deep-learning based robust watermarking model generally applies a discriminator to form generative adversarial network (GAN) for increasing the quality of encoded images, and adopts a single encoder to embed watermark. However, GAN training is unstable, and the single encoder cannot fully adjust the watermarking distribution, thus affecting the watermarking performance. To address those limitations, this paper presents the multi-encoder based on conditional diffusion model (CDM) for robust image watermarking, namely, DiffW. To enhance the stability, the multi-encoder structure based on CDM replaces GAN for optimizing the watermarking distribution iteratively. Specifically, the operation of each timestep in the forward and reverse diffusion processes of the CDM is regarded as an encoder to overcome the shortcomings of the single encoder structure. At the training stage, under the guidance of the conditional noisy image, the forward process trains each encoder to fuse the image and watermark to generate high-quality encoded images. During the testing stage, only a small number of trained encoders of the forward process are used, so as to reduce the time complexity. Furthermore, to improve watermarking robustness, the channel attention module (CAM) is designed to extract main watermark features by mining channel correlations for multi-layer fusion, so that watermark can be embedded into imperceptible and texture areas. The experimental results reveal that compared with the existing watermarking model, the proposed DiffW can achieve better results in terms of watermarking invisibility and robustness. Ting Luo 0001, Renzhi Hu, Zhouyan He, Gangyi Jiang, Haiyong Xu, Yang Song 0015, Chin-Chen Chang 0001 |
IEEE Trans. Multim. | 5 |
| 2025 | Uncovering the Impact of Chain-of-Thought Reasoning for Direct Preference Optimization: Lessons from Text-to-SQLabstractDirect Preference Optimization (DPO) has proven effective in complex reasoning tasks like math word problems and code generation. However, when applied to Text-to-SQL datasets, it often fails to improve performance and can even degrade it. Our investigation reveals the root cause: unlike math and code tasks, which naturally integrate Chain-of-Thought (CoT) reasoning with DPO, Text-to-SQL datasets typically include only final answers (gold SQL queries) without detailed CoT solutions. By augmenting Text-to-SQL datasets with synthetic CoT solutions, we achieve, for the first time, consistent and significant performance improvements using DPO.Our analysis shows that CoT reasoning is crucial for unlocking DPO’s potential, as it mitigates reward hacking, strengthens discriminative capabilities, and improves scalability. These findings offer valuable insights for building more robust Text-to-SQL models. To support further research, we publicly release the code and CoT-enhanced datasets: https://github.com/RUCKBReasoning/DPO_Text2SQL. Ruotong Chen, Haiyong Xu |
ACL (1) | 5 |
| 2025 | Mamba-Based Blind Stitched Wide Field of View Light Field Image Quality Assessment via Dual-Viewport SamplingabstractDue to the limitations of commercial light field camera hardware, the field of view (FOV) of light field images (LFIs) is relatively narrow. To expand the FOV, various LFI stitching algorithms have been developed. However, these algorithms inevitably introduce localized distortions and angular consistency disruptions, which conventional LFI quality assessment metrics struggle to evaluate effectively. To address this issue, a novel Mamba-based blind quality assessment metric for stitched wide field of view light field images (WLFIs) using dual-viewport sampling is proposed. Firstly, sub-aperture images from horizontal and vertical directions are stacked to characterize angular information, and a dual-viewport sampling pattern is designed to enhance data augmentation and capture spatial details. After that, a multi-scale state space block is proposed to improve distortion feature extraction, complemented by an auxiliary distortion discrimination task. Finally, experimental results demonstrate that the proposed metric outperforms state-of-the-art metrics on the benchmark WLFI dataset. Gangyi Jiang, Linwei Zhu, Yeyao Chen, Yueli Cui, Ting Luo 0001, Haiyong Xu |
ICME | 7 |
| 2025 | Combining independent and joint spatial-angular information learning for light field image super-resolution
Dezhang Ke, Yeyao Chen, Chongchong Jin, Haiyong Xu, Zhidi Jiang, Ting Luo 0001, Gangyi Jiang |
Knowl. Based Syst. | 4 |
| 2025 | DiffOSR: Latitude-aware conditional diffusion probabilistic model for omnidirectional image super-resolution
Leiming Liu, Ting Luo 0001, Gangyi Jiang, Yeyao Chen, Haiyong Xu, Renzhi Hu, Zhouyan He |
Knowl. Based Syst. | 5 |
| 2025 | DiffDark: Multi-prior integration driven diffusion model for low-light image enhancement
Renzhi Hu, Ting Luo 0001, Gangyi Jiang, Yeyao Chen, Haiyong Xu, Leiming Liu, Zhouyan He |
Pattern Recognit. | 5 |
| 2025 | Frequency domain-based latent diffusion model for underwater image enhancement
Jingyu Song, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Ting Luo 0001, Yang Song 0015 |
Pattern Recognit. | 2 |
| 2025 | Geometry-Aware RWKV for Heterogeneous Light Field Spatial Super-ResolutionabstractHeterogeneous Light Field (LF) spatial Super-Resolution (SR) aims to significantly enhance the spatial resolution of LF imaging by integrating an extra 2D digital camera. Inspired by the Receptance Weighted Key Value (RWKV), a simple yet effective heterogeneous LF spatial SR method is proposed. Specifically, a texture transfer module with channel correlation is designed, which leverages a feature distillation strategy to transfer texture information from the high-resolution 2D image to the low-resolution LF image. Meanwhile, a spatial-angular rectification module is constructed to restore the spatial-angular coherence damaged in texture transfer. It employs geometry-aware RWKV to capture the intrinsic geometric structure of LFs. Experimental results show that the proposed method outperforms the state-of-the-art methods in both quantitative and qualitative comparisons, while achieving higher efficiency in terms of inference time and memory usage. Zean Chen, Yeyao Chen, Linwei Zhu, Haiyong Xu, Gangyi Jiang |
IEEE Signal Process. Lett. | 4 |
| 2025 | StegMamba: Distortion-Free Immune-Cover for Multi-Image Steganography With State Space ModelabstractMulti-image steganography ensures privacy protection while avoiding suspicion from third parties by embedding multiple secret images within a cover image. However, existing multi-image steganographic methods fail to model global spatial correlations to reduce image damage at the low computation cost. Moreover, they do not account for the anti-distortion capability of the cover image, which is crucial for achieving imperceptible and ensuring security. To overcome these limitations, we propose StegMamba, a distortion-free immune-cover for multi-image steganography architecture with a state space model. Specifically, we first explore the potential of the linear computational cost model Mamba for data hiding tasks through a steganography Mamba block (SMB), whose efficiency makes it suitable for real-time applications. Subsequently, considering that images with distortion resistance reduce embedding damage, the original cover image is reconstructed through immune-cover construction module (ICCM) and associated with the steganography task. Moreover, well-coupled features facilitate fusion, and thus a wavelet-based interaction module (WIM) is designed for effective communication between the immune-cover and the secret images. Compared with the state-of-the-art global attention-based methods, the proposed StegMamba obtains PSNR gains of 3.30 dB, 1.37 dB, and 1.92 dB for the stego image, and two secret recovery images, respectively, and the reduction of 2.87% in detection accuracy for anti-steganalysis. This code is available athttps://github.com/YuhangZhouCJY/StegMamba. Ting Luo 0001, Zhouyan He, Gangyi Jiang, Haiyong Xu, Yushu Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Multi-Attention Learning and Exposure Guidance Toward Ghost-Free High Dynamic Range Light Field ImagingabstractDue to sensor limitations, the light field (LF) images captured by the LF camera suffer from low dynamic range and are prone to poor exposure. To solve this problem, combining multi-exposure technology with LF camera imaging can achieve high dynamic range (HDR) LF imaging. However, for dynamic scenes, this approach tends to produce disturbing ghosting artifacts and destroy the parallax structure of the generated results. To this end, this paper proposes a novel ghost-free HDR LF imaging method using multi-attention learning and exposure guidance. Specifically, the proposed method first designs a multi-scale cross-attention module to achieve efficient multi-exposure LF feature alignment. After that, a dual self-attention-driven Transformer block is constructed to excavate the geometric information of LF and fuse the aligned LF features. In particular, exposure masks derived from middle-exposure are introduced in the feature fusion to guide the network to focus on information recovery in low- and high-brightness regions. Besides, a local compensation module is integrated to cope with local alignment errors and refine details. Finally, a multi-objective reconstruction strategy combined with exposure masks is employed to restore high-quality HDR LF images. Extensive experimental results on the benchmark dataset show that the proposed method generates HDR LF results with high spatial-angular quality consistency and outperforms the state-of-the-art methods in quantitative and qualitative comparisons. Furthermore, the proposed method can enhance the performance of existing LF applications, such as depth estimation. Yeyao Chen, Gangyi Jiang, Chongchong Jin, Ting Luo 0001, Haiyong Xu, Mei Yu 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2024 | Hybrid Domain Learning towards Light Field Spatial Super-Resolution using Heterogeneous ImagingabstractLight field (LF) cameras usually capture dense angular samples, but suffer from low spatial resolution. Existing single-LF super-resolution methods struggle with textures at larger scales (e.g., 8×). To address this issue, this paper proposes a novel hybrid domain learning-based method to enhance LF spatial resolution from heterogeneous imaging (integrating an LF camera and a 2D digital camera). The proposed method consists of two core modules, namely LF feature alignment module and cross-domain multi-scale fusion module. The former combines optical flow and deformable convolution to gradually align the 2D high-resolution features with the low-resolution LF features. The latter progressively fuses the aligned multi-resolution LF features to enable high-quality reconstruction. Experimental results show the proposed method recovers fine textures and preserves accurate angular consistency, and outperforms the state-of-the-art methods in both quantitative and qualitative comparisons. Zean Chen, Yeyao Chen, Mei Yu 0001, Haiyong Xu, Gangyi Jiang |
ICASSP | 4 |
| 2024 | mQUIET: Maximum Quantum Communication Network Transmission Under Reliability ConstraintabstractQuantum communications using entangled photons as qubits offer a promising approach to quantum key distribution, enhancing network security. However, establishing successful entanglement links between distant nodes is challenging due to factors like distance and rapid entanglement decay, resulting in low network transmission efficiency for quantum key distribution. This paper studies the problem of finding the maximum quantum communication network transmission scheme for multiple source-destination pairs with considering the success rate of quantum entangle establishments. We propose the Feasible Quantum Communication Network Transmission (fQUIET) algorithm, which leverages network transformation and auxiliary network construction. The Maximum Quantum commUnIcation nEtwork Transmission (mQUIET) algorithm is further proposed by iteratively searching for augmenting paths. Through theoretical analysis, we demonstrate that the mQUIET algorithm is capable of achieving the maximum QCN transmission scheme. Extensive experiments are carried out and the results reveal that our suggested algorithms have superior performance compared to existing methods. Wei An 0002, Yanni Han, Haiyong Xu, Bo An 0010 |
WCNC | 4 |
| 2024 | Multi-exposure fused light field image quality assessment for dynamic scenes: Benchmark dataset and objective metric
Yun Liu 0048, Guanglong Liao, Gangyi Jiang, Yeyao Chen, Yueli Cui, Haiyong Xu, Mei Yu 0001 |
Expert Syst. Appl. | 6 |
| 2024 | CAISFormer: Channel-wise attention transformer for image steganographyabstractCurrent Transformer-based image steganography cannot embed data properly without considering the correlation of the cover image and the secret image . In addition, to save computational complexity, spatial-wise Transformer is often used to apply in small spatial windows, which limits the extraction of the global feature. To solve those limitations, we present a channel-wise attention Transformer model for image steganography (CAISFormer), which aims to construct long-range dependencies for identifying inconspicuous positions to embed data. A channel self-attention module (CSAM) is deployed to focus the feature channels suitable for data hiding by establishing channel relationships. Meanwhile, a non-linear enhancement (NLE) layer is employed to enhance the beneficial features while weaken the irrelevant ones. For building feature coupling between the cover image and the secret image, a channel-wise cross attention module (CCAM) is designed to fine-tune cover image features by capturing their cross-dependencies. In addition, for concealing data properly, a global–local aggregation module (GLAM) is deployed to adjust fused features by combining global and local attention, which can focus on inconspicuous and texture regions, respectively. The experimental results demonstrate that CAISFormer obtains PSNR gains of more than 0.36 dB and 0.90 dB for the cover/stego image pair and the secret/recovery image pair, respectively, and the detection ratio is decreased by 3.43%, in single image hiding compared to the state-of-the-art. Moreover, the generalization ability is also proved across a variety of datasets. The code will be made publicly available at https://github.com/YuhangZhouCJY/CAISFormer . Ting Luo 0001, Zhouyan He, Gangyi Jiang, Haiyong Xu, Chin-Chen Chang 0001 |
Neurocomputing | 5 |
| 2024 | A Transformer-based invertible neural network for robust image watermarking
Zhouyan He, Renzhi Hu, Ting Luo 0001, Haiyong Xu |
J. Vis. Commun. Image Represent. | 5 |
| 2024 | COC-UFGAN: Underwater image enhancement based on color opponent compensation and dual-subnet underwater fusion generative adversarial network
Zhenkai Liu, Xinxiao Fu, Chi Lin 0003, Haiyong Xu |
J. Vis. Commun. Image Represent. | 4 |
| 2024 | Vision graph convolutional network for underwater image enhancement
Zexuan Xing, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Yeyao Chen |
Knowl. Based Syst. | 2 |
| 2024 | HDR light field imaging of dynamic scenes: A learning-based method and a benchmark dataset
Yeyao Chen, Gangyi Jiang, Mei Yu 0001, Chongchong Jin, Haiyong Xu, Yo-Sung Ho |
Pattern Recognit. | 5 |
| 2024 | Underwater Monocular Depth Estimation Based on Physical-Guided TransformerabstractOwing to the light absorption and wavelength scattering in underwater environments, underwater images are severely degraded, which directly affects the depth estimation of underwater scenes. Accurate underwater depth estimation is essential for representing and understanding underwater scenes. However, the existing underwater depth estimation methods have not fully taken into account the distinctive physical properties of underwater environments, which has resulted in increased bias and feature distortion in the depth estimation results. In this paper, an underwater monocular depth estimation method based on physical-guided Transformer (UPGformer) is proposed, considering the characteristics of underwater imaging, including shallow feature extraction, encoding, decoding, and regression stages. Specifically, in the shallow feature extraction stage, considering the color deviation of underwater images and extracting richer primary features, an enrichment and extraction depth Transformer (EEDT) module is proposed, by interacting physically inverted transmission maps of the underwater dark channel prior (UDCP) with physical color-compensated underwater images through self-attention. In the encoding stage, considering the nonuniform degradation of underwater images (nonuniform local distortion and inconsistent channel degradation), the underwater physical Transformer interaction encoder (UPTE) module, which fuses the Transformer and physically inverted transmission maps, is proposed. Furthermore, in the decoding stage, to better recover features and reduce information loss, the underwater physical embedded decoding (UPED) module is proposed, which embeds the physically inverted transmission maps with the upsampling process. Finally, the depth map is constructed during the regression stage. The experimental results demonstrate that the proposed UPGformer outperforms existing methods, both qualitatively and quantitatively. Chen Wang 0141, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Yeyao Chen |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Underwater Image Quality Assessment from Synthetic to Real-world: Dataset and Objective MethodabstractThe complicated underwater environment and lighting conditions lead to severe influence on the quality of underwater imaging, which tends to impair underwater exploration and research. To effectively evaluate the quality of underwater images, an underwater image quality assessment dataset is constructed from synthetic to real-world, and then a new objective underwater image assessment method based on the characteristics of the underwater imaging is proposed (UICQA). Specifically, to address the lack of a publicly available datasets and more accurately quantify the quality of underwater images, a subjective underwater image quality assessment dataset from synthetic to real-world underwater images, named USRD, is constructed. Considering that the transmission map can effectively reflect the characteristics of the underwater imaging, statistical features are effectively extracted from the transmission map for distinguishing underwater images of different quality. Further, considering that the transmission map negatively correlates with scene depth, a local-to-global transmission map weighted contrast feature is constructed. Additionally, the color features of human perception and texture features based on fractal dimensions are proposed. Finally, the experimental results show that the proposed UICQA method exhibits the highest correlation with ground truth scores compared to state-of-the-art UIQA methods. Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Xuebo Zhang 0002, Hongwei Ying |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2023 | UDAformer: Underwater image enhancement based on dual attention transformer
Haiyong Xu, Ting Luo 0001, Yang Song 0015, Zhouyan He |
Comput. Graph. | 2 |
| 2023 | A bilateral attention based generative adversarial network for DIBR 3D image watermarking
Zhouyan He, Lingqiang He, Haiyong Xu, Tong-Yuen Chai, Ting Luo 0001 |
J. Vis. Commun. Image Represent. | 3 |
| 2023 | RDD-net: Robust duplicated-diffusion watermarking based on deep network
Guowei Jiang, Zhouyan He, Jiangtao Huang, Ting Luo 0001, Haiyong Xu, Chongchong Jin |
J. Vis. Commun. Image Represent. | 5 |
| 2023 | Blind light field image quality assessment with tensor color domain and 3D shearlet transform
Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu |
Signal Process. | 5 |
| 2023 | Deep Light Field Spatial Super-Resolution Using Heterogeneous ImagingabstractLight field (LF) imaging expands traditional imaging techniques by simultaneously capturing the intensity and direction information of light rays, and promotes many visual applications. However, owing to the inherent trade-off between the spatial and angular dimensions, LF images acquired by LF cameras usually suffer from low spatial resolution. Many current approaches increase the spatial resolution by exploring the four-dimensional (4D) structure of the LF images, but they have difficulties in recovering fine textures at a large upscaling factor. To address this challenge, this paper proposes a new deep learning-based LF spatial super-resolution method using heterogeneous imaging (LFSSR-HI). The designed heterogeneous imaging system uses an extra high-resolution (HR) traditional camera to capture the abundant spatial information in addition to the LF camera imaging, where the auxiliary information from the HR camera is utilized to super-resolve the LF image. Specifically, an LF feature alignment module is constructed to learn the correspondence between the 4D LF image and the 2D HR image to realize information alignment. Subsequently, a multi-level spatial-angular feature enhancement module is designed to gradually embed the aligned HR information into the rough LF features. Finally, the enhanced LF features are reconstructed into a super-resolved LF image using a simple feature decoder. To improve the flexibility of the proposed method, a pyramid reconstruction strategy is leveraged to generate multi-scale super-resolution results in one forward inference. The experimental results show that the proposed LFSSR-HI method achieves significant advantages over the state-of-the-art methods in both qualitative and quantitative comparisons. Furthermore, the proposed method preserves more accurate angular consistency. Yeyao Chen, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Yo-Sung Ho |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | A two-stage and two-branch generative adversarial network-based underwater image enhancement
Haiyong Xu, Chi Lin 0003, Ting Luo 0001 |
Vis. Comput. | 2 |
| 2022 | Robust HDR video watermarking method based on the HVS model and T-QR
Ting Luo 0001, Haiyong Xu, Yang Song 0015, Chunpeng Wang 0001, Li Li 0014 |
Multim. Tools Appl. | 3 |
| 2022 | Multi-Angle Projection Based Blind Omnidirectional Image Quality AssessmentabstractMost of the existing blind omnidirectional image quality assessment (BOIQA) methods are based on data-driven approach where the end-to-end neural network or deep learning tools are mainly used for feature extraction. However, it usually lacks interpretability and is difficult to discover the perceptual mechanism behind. In this paper, from the perspective of perception modeling, we propose a novel multi-angle projection based BOIQA (MP-BOIQA) method. Considering the omnibearing and near eye display characteristics with head mounted display, multiple color cubemap projection images with respect to different viewpoints are grouped as the color omnidirectional distortion (COD) units so as to simulate the user’s viewing behavior in subjective quality assessment. In the designed multi-angle projection based feature extractor, tensor decomposition is implemented on each COD unit for dimensionality reduction, and piecewise exponential fitting is used to get the distribution of mean subtracted contrast normalized coefficients of the unit’s feature matrices in tensor domain. Finally, the extracted features are pooled with random forest. The experimental results on three omnidirectional image quality datasets show that the MP-BOIQA method can deliver highly competitive performance compared with some representative full-reference quality assessment methods, as well as some state-of-the-art BOIQA methods. Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Haiyong Xu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2022 | Reinforced Swin-Convs Transformer for Simultaneous Underwater Sensing Scene Image Enhancement and Super-resolutionabstractUnderwater image enhancement (UIE) technology aims to tackle the challenge of restoring the degraded underwater images due to light absorption and scattering. Meanwhile, the ever-increasing requirement for higher resolution images from a lower resolution in the underwater domain cannot be overlooked. To address these problems, a novel U-Net-based reinforced Swin-Convs Transformer for simultaneous enhancement and superresolution (URSCT-SESR) method is proposed. Specifically, with the deficiency of U-Net based on pure convolutions, the Swin Transformer is embedded into U-Net for improving the ability to capture the global dependence. Then, given the inadequacy of the Swin Transformer capturing the local attention, the reintroduction of convolutions may capture more local attention. Thus, an ingenious manner is presented for the fusion of convolutions and the core attention mechanism to build a reinforced Swin-Convs Transformer block (RSCTB) for capturing more local attention, which is reinforced in the channel and the spatial attention of the Swin Transformer. Finally, experimental results on available datasets demonstrate that the proposed URSCT-SESR achieves the state-of-the-art performance compared with other methods in terms of both subjective and objective evaluations. The code is publicly available athttps://github.com/TingdiRen/URSCT-SESR. Tingdi Ren, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | Tensor Product and Tensor-Singular Value Decomposition Based Multi-Exposure Fusion of ImagesabstractConsidering multidimensional structure of the multi-exposure images, a new Tensor product and Tensor-singular value decomposition based Multi-Exposure image Fusion (TT-MEF) method is proposed. The main innovation of this work is to explore a new feature representation of multi-exposure images in the new tensor domain and design the fusion strategy on this basis. Specifically, the luminance and the chrominance channels are fused separately to maintain color consistency. For the luminance fusion, the luminance channel of multi-exposure images is divided into two parts, that is, de-mean term and mean term. The de-mean term is represented as a tensor to extract the feature. Then, the tensor product and tensor-singular value decomposition (T-SVD) are used to design a tensor feature extractor. Furthermore, a fusion strategy of the de-mean term is presented according to the visual saliency model, and a fusion strategy of the mean term is defined by the local and the global visual weights to control counterpoise between the local and global luminance. For the chrominance fusion, a new fusion strategy is also designed by the tensor product and T-SVD, similar to the luminance fusion. Finally, the fused image is obtained by combining the luminance and chrominance fusion. Experimental results show that the proposed TT-MEF method generally outperforms the existing state-of-the-art in terms of subjective visual quality and objective evaluation. Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Zhongjie Zhu, Yongqiang Bai, Yang Song 0015, Huifang Sun |
IEEE Trans. Multim. | 1 |
| 2022 | Robust HDR video watermarking method based on saliency extraction and T-SVD
Ting Luo 0001, Haiyong Xu, Yang Song 0015, Chunpeng Wang 0001 |
Vis. Comput. | 3 |
| 2021 | Reversible data hiding scheme for high dynamic range images based on multiple prediction error expansion
Yongqiang Bai, Gangyi Jiang, Zhongjie Zhu, Haiyong Xu, Yang Song 0015 |
Signal Process. Image Commun. | 4 |
| 2021 | Pseudo Video and Refocused Images-Based Blind Light Field Image Quality AssessmentabstractThe commercial light field camera is able to capture four-dimensional Light Field Image (LFI), which can be visualized to LFI contents on 2D displays by means of the Pseudo Video (PV) or the Refocused Images (RIs) generated with the refocusing function of LFI. However, the quality degradation of LFI will affect user’s visual experience of LFI contents. Hence, it is crucial to develop an effective LFI quality assessment method to monitor the LFI quality. Most existing subjective databases of LFI use PV and RIs visualization techniques to assess the quality of LFI. Therefore, as the way of presenting LFI on 2D display, PV and RIs are closely related to the subjective perception of LFI by human eyes. Based on these two visualization techniques, this article proposes a novel PV and RIs based blind LFI quality assessment method, in which the feature extraction is divided into two parts. In the first part, the PV’s structure, motion and disparity information are extracted with multi-scale and multi-directional Shearlet transform. In the other part, the spatial structure, depth and semantic information of the RIs are obtained. Finally, support vector regression is used to nonlinear map the perceptual features to quality score of LFI. The experimental results on four LFI databases show that the proposed method has better correlation with human visual perception, compared with the classical 2D image quality assessment methods as well as the state-of-the-art LFI quality assessment methods. Jianjun Xiang, Mei Yu 0001, Gangyi Jiang, Haiyong Xu, Yang Song 0015, Yo-Sung Ho |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2020 | VBLFI: Visualization-Based Blind Light Field Image Quality AssessmentabstractLight field image (LFI) contains the intensity and direction information of the scene. The huge amount of data and different visualization methods of LFI brings great challenges to LFI processing and its blind LFI quality assessment. This paper analyzes the human visual perception from the LFI's visualization, and proposes a novel Visualization-based Blind Light Field Image quality assessment (VBLFI) model. With LFI's visualization and its depth cues, we compute mean difference image from LFI to reduce redundant information of LFI and to describe depth and structural information of LFI. LFI's multi-scale expression with curvelet transform is used to reflect the multi-channel characteristics of human visual system. So, the corresponding natural scene statistical features and energy features are extracted from the mean difference image and sub-aperture images of LFI in curvelet transform domain to form the feature vector, further used to predict the LFI quality. Compared to the representative 2D image quality assessment models and the state-of-the-art LFIQA models, the proposed VBLFI model has better prediction accuracy and stability in the public LFI databases. Jianjun Xiang, Mei Yu 0001, Hua Chen 0004, Haiyong Xu, Yang Song 0015, Gangyi Jiang |
ICME | 4 |
| 2020 | Multi-exposure high dynamic range imaging with informative content enhanced network
Zhiyong Pan, Mei Yu 0001, Gangyi Jiang, Haiyong Xu, Zongju Peng |
Neurocomputing | 4 |
| 2020 | Multi-exposure image fusion based on tensor decomposition
Shengcong Wu, Ting Luo 0001, Yang Song 0015, Haiyong Xu |
Multim. Tools Appl. | 4 |
| 2019 | Convolutional neural networks-based stereo image reversible data hiding method
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Caiming Zhong, Haiyong Xu, Zhiyong Pan |
J. Vis. Commun. Image Represent. | 5 |
| 2019 | A novel robust color image watermarking method using RGB correlations
Fangyan Zhang, Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wujie Zhou |
Multim. Tools Appl. | 5 |
| 2019 | Ensemble clustering based on evidence extracted from the co-association matrix
Caiming Zhong, Lianyu Hu 0001, Ting Luo 0001, Haiyong Xu |
Pattern Recognit. | 6 |
| 2019 | Robust high dynamic range color image watermarking method based on feature map extraction
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wei Gao 0013 |
Signal Process. | 4 |
| 2018 | 3D visual discomfort predictor based on subjective perceived-constraint sparse representation in 3D display system
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Zongju Peng, Feng Shao 0001, Hao Jiang 0014 |
Future Gener. Comput. Syst. | 1 |
| 2018 | Sparse recovery based reversible data hiding method using the human visual system
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wei Gao 0013 |
Multim. Tools Appl. | 4 |
| 2017 | Stereoscopic image quality assessment by learning non-negative matrix factorization-based color visual characteristics and considering binocular interactions
Gangyi Jiang, Haiyong Xu, Mei Yu 0001, Ting Luo 0001, Yun Zhang 0002 |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Asymmetric self-recovery oriented stereo image watermarking method for three dimensional video system
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu |
Multim. Syst. | 4 |
| 2015 | Difference of Gaussian statistical features based blind image quality assessment: A deep learning approachabstractNowadays, natural scene statistics (NSS) based blind image quality assessment (BIQA) models trained by machine learning, tend to achieve excellent performance. However, BIQA is still a very challenging research topic due to the lack of reference images. The key of further improvement lies in feature mining and pooling strategy decision. In this work, a new BIQA model is proposed to utilize local normalized multi-scale difference of Gaussian (DoG) response in distorted images as features which show a high correlation with perceptual quality. Then, a three-step-framework based deep neural network (DNN) is designed and employed as the pooling strategy. Compared with the support vector machine (SVM), the proposed three-step-framework DNN can excavate better feature representation, leading to more accurate predictions and stronger generalization ability. The proposed model achieves state-of-the-art performance on two authoritative databases and excellent generalization ability in cross database experiments. Yaqi Lv, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Feng Shao 0001 |
ICIP | 4 |