Ting Luo 0001

dblp:20/7700-1 · DBLP profile ↗
← Back
63ranked-venue papers
9as first author
46since 2021 · last 2027
0000-0003-1762-148XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 40 · 8 first-author · 27 since 2021Artificial intelligence and machine learning · 15 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Systems, architecture and hardware · 2 · 1 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2027 CDV-PCQA: Content-distortion-guided dynamic viewpoint quality assessment for 3D point clouds
Qihao Liang, Li Li 0014, Ting Luo 0001, Gangyi Jiang, Wujie Zhou, Linwei Zhu, Zhouyan He
Expert Syst. Appl.3
2026 ARMLF: Anomalous region representation learning for multi-exposure fused light field image quality assessment
Guanglong Liao, Gangyi Jiang, Linwei Zhu, Yeyao Chen, Yueli Cui, Ting Luo 0001, Haiyong Xu
Expert Syst. Appl.6
2026 Artifact-suppressed 3D retinal microvascular segmentation via multi-scale topology regulation
Ting Luo 0001, Jinxian Zhang, Tao Chen 0003, Zhouyan He, Yanda Meng, Jiong Zhang 0004, Dan Zhang 0026
Medical Image Anal.1
2026 LatentDark: Reflectance guided latent diffusion model for low-light image enhancement
Renzhi Hu, Ting Luo 0001, Gangyi Jiang, Leiming Liu, Yeyao Chen, Haiyong Xu, Zhouyan He
Signal Process.2
2026 MLANet: Multilevel aggregation network for binocular eye-fixation prediction
Wujie Zhou, Jiabao Ma, Yulai Zhang, Lu Yu 0003, Weijia Gao, Ting Luo 0001
Signal Process. Image Commun.6
2026 DiffW: Multi-Encoder Based on Conditional Diffusion Model for Robust Image Watermarking
abstract
The existing deep-learning based robust watermarking model generally applies a discriminator to form generative adversarial network (GAN) for increasing the quality of encoded images, and adopts a single encoder to embed watermark. However, GAN training is unstable, and the single encoder cannot fully adjust the watermarking distribution, thus affecting the watermarking performance. To address those limitations, this paper presents the multi-encoder based on conditional diffusion model (CDM) for robust image watermarking, namely, DiffW. To enhance the stability, the multi-encoder structure based on CDM replaces GAN for optimizing the watermarking distribution iteratively. Specifically, the operation of each timestep in the forward and reverse diffusion processes of the CDM is regarded as an encoder to overcome the shortcomings of the single encoder structure. At the training stage, under the guidance of the conditional noisy image, the forward process trains each encoder to fuse the image and watermark to generate high-quality encoded images. During the testing stage, only a small number of trained encoders of the forward process are used, so as to reduce the time complexity. Furthermore, to improve watermarking robustness, the channel attention module (CAM) is designed to extract main watermark features by mining channel correlations for multi-layer fusion, so that watermark can be embedded into imperceptible and texture areas. The experimental results reveal that compared with the existing watermarking model, the proposed DiffW can achieve better results in terms of watermarking invisibility and robustness.
Ting Luo 0001, Renzhi Hu, Zhouyan He, Gangyi Jiang, Haiyong Xu, Yang Song 0015, Chin-Chen Chang 0001
IEEE Trans. Multim.1
2025 Mamba-Based Blind Stitched Wide Field of View Light Field Image Quality Assessment via Dual-Viewport Sampling
abstract
Due to the limitations of commercial light field camera hardware, the field of view (FOV) of light field images (LFIs) is relatively narrow. To expand the FOV, various LFI stitching algorithms have been developed. However, these algorithms inevitably introduce localized distortions and angular consistency disruptions, which conventional LFI quality assessment metrics struggle to evaluate effectively. To address this issue, a novel Mamba-based blind quality assessment metric for stitched wide field of view light field images (WLFIs) using dual-viewport sampling is proposed. Firstly, sub-aperture images from horizontal and vertical directions are stacked to characterize angular information, and a dual-viewport sampling pattern is designed to enhance data augmentation and capture spatial details. After that, a multi-scale state space block is proposed to improve distortion feature extraction, complemented by an auxiliary distortion discrimination task. Finally, experimental results demonstrate that the proposed metric outperforms state-of-the-art metrics on the benchmark WLFI dataset.
Gangyi Jiang, Linwei Zhu, Yeyao Chen, Yueli Cui, Ting Luo 0001, Haiyong Xu
ICME6
2025 TransMeter: Robust Water-Meter Reading under Half-Character Transitions via YOLO + CogVLM2 Fusion
abstract
Accurate reading of mechanical water meters is crucial for automated water billing and resource management. When a digit wheel advances between two positions, it often displays overlapping parts of two digits, forming a halfcharacter transition. Conventional vision models struggle to interpret these cases, leading to sequence-level misreads. To address this challenge, we present TransMeter, a robust reading framework that combines object detection with multimodal semantic reasoning. Specifically, YOLOv11n is employed for precise digit-wheel detection, while the vision-language large model CogVLM2 performs fine-grained reasoning to identify half-character states. A position-aware confidence fusion module then integrates visual and semantic cues to produce coherent readings. Experiments on a self-built dataset demonstrate that TransMeter corrects 40 misread cases (26 detection errors and 14 reasoning self-corrections) and improves overall accuracy from$\mathbf{9 3. 5 \%}$to$\mathbf{9 7. 1 \%}$, validating the effectiveness of vision-language fusion for transition-digit recognition.
Ting Luo 0001, Ximing Li 0001
ICPADS4
2025 Multi-level cross-modal attention guided DIBR 3D image watermarking
Qingmo Chen, Zhouyan He, Ting Luo 0001, Jiangtao Huang
J. Vis. Commun. Image Represent.4
2025 Combining independent and joint spatial-angular information learning for light field image super-resolution
Dezhang Ke, Yeyao Chen, Chongchong Jin, Haiyong Xu, Zhidi Jiang, Ting Luo 0001, Gangyi Jiang
Knowl. Based Syst.6
2025 DiffOSR: Latitude-aware conditional diffusion probabilistic model for omnidirectional image super-resolution
Leiming Liu, Ting Luo 0001, Gangyi Jiang, Yeyao Chen, Haiyong Xu, Renzhi Hu, Zhouyan He
Knowl. Based Syst.2
2025 DiffDark: Multi-prior integration driven diffusion model for low-light image enhancement
Renzhi Hu, Ting Luo 0001, Gangyi Jiang, Yeyao Chen, Haiyong Xu, Leiming Liu, Zhouyan He
Pattern Recognit.2
2025 Frequency domain-based latent diffusion model for underwater image enhancement
Jingyu Song, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Ting Luo 0001, Yang Song 0015
Pattern Recognit.6
2025 Highly applicable and imperceptible watermark attack network
Chunpeng Wang 0001, Qi Li 0029, Jian Li 0034, Ziqi Wei 0001, Ting Luo 0001, Bin Ma 0003
Signal Process.7
2025 Blind Light Field Image Quality Assessment via Frequency Domain Analysis and Auxiliary Learning
abstract
Due to the distortions occurring at various stages from acquisition to visualization, light field image quality assessment (LFIQA) is crucial for guiding the processing of light field images (LFIs). In this letter, we propose a new blind LFIQA metric via frequency domain analysis and auxiliary learning, termed as FABLFQA. First, spatial-angular patches are extracted from LFIs and further processed through discrete cosine transform to obtain light field frequency maps. Subsequently, a concise and efficient frequency-aware deep learning network is designed to extract frequency features, including the frequency descriptor, 3D ConvBlock, and frequency transformer. Finally, a distortion type discrimination auxiliary task is employed to facilitate the learning of the main quality assessment task. Experimental results on three representative LFI datasets show that the proposed metric outperforms the state-of-the-art metrics.
Gangyi Jiang, Linwei Zhu, Yueli Cui, Ting Luo 0001
IEEE Signal Process. Lett.5
2025 Enhancing RGB-D Mirror Segmentation With a Neighborhood-Matching and Demand-Modal Adaptive Network Using Knowledge Distillation
abstract
Recent breakthroughs in computer vision have led to remarkable progress in the areas of autonomous vehicles and robotics. However, ordinary objects such as mirrors pose unique challenges to computer vision systems owing to occlusion, reflection, and distortion. Moreover, existing deep learning models suffer from issues such as excessive parameters and high computational complexity, making it challenging to implement numerous studies offline. To address these issues, we propose an innovative solution: a neighborhood-matching and demand-modal adaptive network using knowledge distillation (KD), called NDANet-S$^{\ast }$, specifically designed for red-green-blue depth mirror segmentation. NDANet-S$^{\ast }$operates by iteratively matching detailed and semantic difference between neighborhood features during the encoding phase. It then complements information across different modalities through demand-modal adaptation, enhancing heteromodal cross-complementation during the KD stage. In the decoding phase, semantic enhancement features and iterative encoding features are deeply integrated, forming a strong foundation for multistage progressive knowledge transfer in the KD process. Furthermore, we introduce a multistage teacher-assisted KD scheme, guided by sample complexity, to work synergistically with the mirror segmentation model. This innovative scheme includes a sample complexity rater, heterogeneous cross-complementarity, and hierarchical progressive knowledge transfer. Experimental evaluations on publicly available datasets indicate that NDANet-S$^{\ast }$significantly enhances segmentation accuracy while preserving a consistent number of parameters. Additionally, it achieves state-of-the-art performance in mirror segmentation. The source code for our model is publicly available and can be accessed at:https://github.com/2021nihao/NMDANet. Note to Practitioners—This study presents a neighborhood-matching and demand-modal adaptive network for RGB-D mirror segmentation, incorporating knowledge distillation (KD). The combination of the KD framework and segmentation network enables sample complexity discrimination, cross-modal distillation, and multilevel distillation (guided by the former) to achieve targeted distillation. The newly proposed NDANet-S$^{\ast }$surpasses current state-of-the-art methods, with a reduction of 93.2% in parameters and 87.65% in floating-point operations (FLOPs) compared to NDANet-T.
Wujie Zhou, Yuanyuan Liu 0004, Ting Luo 0001
IEEE Trans Autom. Sci. Eng.4
2025 StegMamba: Distortion-Free Immune-Cover for Multi-Image Steganography With State Space Model
abstract
Multi-image steganography ensures privacy protection while avoiding suspicion from third parties by embedding multiple secret images within a cover image. However, existing multi-image steganographic methods fail to model global spatial correlations to reduce image damage at the low computation cost. Moreover, they do not account for the anti-distortion capability of the cover image, which is crucial for achieving imperceptible and ensuring security. To overcome these limitations, we propose StegMamba, a distortion-free immune-cover for multi-image steganography architecture with a state space model. Specifically, we first explore the potential of the linear computational cost model Mamba for data hiding tasks through a steganography Mamba block (SMB), whose efficiency makes it suitable for real-time applications. Subsequently, considering that images with distortion resistance reduce embedding damage, the original cover image is reconstructed through immune-cover construction module (ICCM) and associated with the steganography task. Moreover, well-coupled features facilitate fusion, and thus a wavelet-based interaction module (WIM) is designed for effective communication between the immune-cover and the secret images. Compared with the state-of-the-art global attention-based methods, the proposed StegMamba obtains PSNR gains of 3.30 dB, 1.37 dB, and 1.92 dB for the stego image, and two secret recovery images, respectively, and the reduction of 2.87% in detection accuracy for anti-steganalysis. This code is available athttps://github.com/YuhangZhouCJY/StegMamba.
Ting Luo 0001, Zhouyan He, Gangyi Jiang, Haiyong Xu, Yushu Zhang 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Toward Better Than Pseudo-Reference in Underwater Image Enhancement
abstract
Since degraded underwater images are not always accompanied with distortion-free counterparts in real-world situations, existing underwater image enhancement (UIE) methods are mostly learned on a paired set consisting of raw underwater images and their corresponding pseudo-reference labels. Although the existing UIE datasets manually select the best model-generated results as pseudo-References, such pseudo-reference labels do not always exhibit perfect visual quality. Therefore, it would be interesting to investigate whether it is possible to break through the performance bottleneck of UIE networks trained with imperfect pseudo-references. Motivated by these facts, this paper focuses on innovating more advanced loss functions rather than designing more complex network architectures. Specifically, a plug-and-play hybrid Performance SurPassing Loss (PSPL), consisting of a Quality Score Comparison Loss (QSCL) and a scene Depth-aware Unpaired Contrastive Loss (DUCL), is formulated to guide the training of UIE network. Functionally, QSCL aims to guide the UIE network to generate enhanced results with better visual quality than pseudo-references by constructing image quality score comparison losses from both image-level and region-level. Nevertheless, only using QSCL cannot guarantee obtaining desired results for those severely degraded distant regions. Therefore, we also design a tailored DUCL to handle this challenging issue from the scene depth perspective, i.e., DUCL encourages the distant regions of the enhanced results to be closer to the high-quality nearby regions (pull) and far away from the low-quality distant regions (push) of the pseudo-references. Extensive experimental results demonstrate the advantage of using PSPL over the state-of-the-arts even with an extremely simple and lightweight UIE network. The source code will be released at https://github.com/lewis081/PSPL.
Yi Liu 0085, Qiuping Jiang, Xingbo Li, Ting Luo 0001, Wenqi Ren
IEEE Trans. Image Process.4
2025 The Safety Illusion? Testing the Boundaries of Concept Removal in Diffusion Models
abstract
Text-to-image diffusion models are capable of producing high-quality images from textual descriptions; however, they present notable security concerns. These include the potential for generating Not-Safe-For-Work (NSFW) content, replicating artists' styles without authorization, or creating deepfakes. Recent advancements have proposed concept erasure techniques to eliminate sensitive concepts from these models, aiming to mitigate the generation of undesirable content. Nevertheless, the robustness of these techniques against a wide range of adversarial inputs has not been comprehensively investigated. To address this challenge, a novel two-stage optimization attack framework based on adversarial perturbations, referred to as Concept Embedding Adversary (CEA), was proposed in the present study. By leveraging the cross-modal alignment priors of the CLIP model, CEA iteratively adjusts adversarial embedding vectors to approximate the semantic expression of specific target concepts. This process enables the construction of deceptive adversarial prompts that exploit diffusion models, compelling them to regenerate previously erased concepts. The performance of concept erasure methods was evaluated, specifically when dealing with diversified adversarial prompts targeting erased concepts, such as NSFW content, artistic styles, and objects. Extensive experimental results demonstrate that existing concept erasure methods are unable to completely eliminate target concepts. In contrast, the proposed CEA framework exploits residual vulnerabilities within the generative latent space through a two-stage optimization process. By achieving precise cross-modal alignment, CEA attains significantly higher ASR in regenerating erased concepts.
Yixiang Pan, Ting Luo 0001, Wenpeng Xing
IEEE Trans. Image Process.2
2025 DA-Net: A Double Alignment Multimodal Learning Network for Point Cloud Quality Assessment
abstract
Existing multimodal point cloud quality assessment (PCQA) methods usually integrate 3D and 2D information to simulate human visual perception of distortions. However, due to the lack of consideration of spatial correspondence, they have difficulty to learn consistent distortion representations from different modalities in the same region of the PC. In addition, they also ignore the heterogeneity of modalities and rely on complex fusion mechanisms (e.g., attention) to integrate multimodal features. Both lead to limited performance and increased computational complexity. To address these limitations, we propose a novel double alignment multimodal learning network (DA-Net), which introduces two key alignment strategies. Specifically, the first is spatial pre-alignment strategy, which generates informative 2D patch for each 3D patch via an adaptive patch projection module (APPM), ensuring accurate spatial correspondence of different modalities prior to feature extraction. The second is a uniform feature alignment strategy, which includes feature disentanglement module (FDM) and feature mapping module (FMM) to relieve heterogeneity of modalities and guide the optimization of 2D and 3D encoder. Finally, multimodal features are simply integrated and regressed to obtain the quality score. Experimental results demonstrate that the DA-Net exhibits outstanding performance and generalization ability. It also achieves lower computational complexity compared with other multimodal PCQA methods. The source codes of DA-Net will be available at https://github.com/Rphone/DA-Net.
Xinqiang Wu, Zhouyan He, Ting Luo 0001, Gangyi Jiang, Wujie Zhou, Linwei Zhu, Weisi Lin
IEEE Trans. Image Process.3
2025 Local and Global Structure-Guided No-Reference Point Cloud Quality Assessment
abstract
As a crucial representation of 3D data, a point cloud (PC) can accurately capture the geometry, structure, and color information of objects. However, various quality problems arise owing to device noise, data acquisition errors, and compression algorithms, limiting the application of PCs. Therefore, assessing PC quality to determine its suitability for applications is a challenging task. In this work, a local and global structure-guided feature extraction and attention network (LGS-Net) is introduced for no-reference PC quality assessment (PCQA). This approach incorporates cluster construction (CC), local structure-guided cluster feature extraction (LSFE), and global structure-guided attention (GSA) modules. First, owing to the heightened sensitivity of the human visual system (HVS) to structural information, a graph filter is employed to identify high-frequency clusters. Within the LSFE module, a multiscale strategy is employed to ensure that structural information effectively influences both the geometry and color information. Simultaneously, the multiscale features within the cluster are dynamically fine-tuned using feature channel weight reassignment. To account for the impact of interclusters on overall quality, a GSA module is introduced to establish global dependencies between local clusters. This approach enables the extraction of final geometry, color, and structure information, which are ultimately used for accurate quality assessment. Extensive experimental results show that the proposed method outperforms the existing state-of-the-art PCQA methods using two publicly available subjective datasets.
Zhouyan He, Qihao Liang, Gangyi Jiang, Mei Yu 0001, Yeyao Chen, Ting Luo 0001, Wujie Zhou
IEEE Trans. Multim.6
2025 Underwater Image Enhancement With Cascaded Contrastive Learning
abstract
Underwater image enhancement (UIE) is a highly challenging task due to the complexity of underwater environment and the diversity of underwater image degradation. Due to the application of deep learning, current UIE methods have made significant progress. Most of the existing deep learning-based UIE methods follow a single-stage network which cannot effectively address the diverse degradations simultaneously. In this paper, we propose to address this issue by designing a two-stage deep learning framework and taking advantage of cascaded contrastive learning to guide the network training of each stage. The proposed method is called CCL-Net in short. Specifically, the proposed CCL-Net involves two cascaded stages, i.e., a color correction stage tailored to the color deviation issue and a haze removal stage tailored to improve the visibility and contrast of underwater images. To guarantee the underwater image can be progressively enhanced, we also apply contrastive loss as an additional constraint to guide the training of each stage. In the first stage, the raw underwater images are used as negative samples for building the first contrastive loss, ensuring the enhanced results of the first color correction stage are better than the original inputs. While in the second stage, the enhanced results rather than the raw underwater images of the first color correction stage are used as the negative samples for building the second contrastive loss, thus ensuring the final enhanced results of the second haze removal stage are better than the intermediate color corrected results. Extensive experiments on multiple benchmark datasets demonstrate that our CCL-Net can achieve superior performance compared to many state-of-the-art methods. In addition, a series of ablation studies also verify the effectiveness of each key component involved in the proposed CCL-Net.
Yi Liu 0085, Qiuping Jiang, Ting Luo 0001, Jingchun Zhou
IEEE Trans. Multim.4
2025 Multi-Attention Learning and Exposure Guidance Toward Ghost-Free High Dynamic Range Light Field Imaging
abstract
Due to sensor limitations, the light field (LF) images captured by the LF camera suffer from low dynamic range and are prone to poor exposure. To solve this problem, combining multi-exposure technology with LF camera imaging can achieve high dynamic range (HDR) LF imaging. However, for dynamic scenes, this approach tends to produce disturbing ghosting artifacts and destroy the parallax structure of the generated results. To this end, this paper proposes a novel ghost-free HDR LF imaging method using multi-attention learning and exposure guidance. Specifically, the proposed method first designs a multi-scale cross-attention module to achieve efficient multi-exposure LF feature alignment. After that, a dual self-attention-driven Transformer block is constructed to excavate the geometric information of LF and fuse the aligned LF features. In particular, exposure masks derived from middle-exposure are introduced in the feature fusion to guide the network to focus on information recovery in low- and high-brightness regions. Besides, a local compensation module is integrated to cope with local alignment errors and refine details. Finally, a multi-objective reconstruction strategy combined with exposure masks is employed to restore high-quality HDR LF images. Extensive experimental results on the benchmark dataset show that the proposed method generates HDR LF results with high spatial-angular quality consistency and outperforms the state-of-the-art methods in quantitative and qualitative comparisons. Furthermore, the proposed method can enhance the performance of existing LF applications, such as depth estimation.
Yeyao Chen, Gangyi Jiang, Chongchong Jin, Ting Luo 0001, Haiyong Xu, Mei Yu 0001
IEEE Trans. Vis. Comput. Graph.4
2024 CAISFormer: Channel-wise attention transformer for image steganography
abstract
Current Transformer-based image steganography cannot embed data properly without considering the correlation of the cover image and the secret image . In addition, to save computational complexity, spatial-wise Transformer is often used to apply in small spatial windows, which limits the extraction of the global feature. To solve those limitations, we present a channel-wise attention Transformer model for image steganography (CAISFormer), which aims to construct long-range dependencies for identifying inconspicuous positions to embed data. A channel self-attention module (CSAM) is deployed to focus the feature channels suitable for data hiding by establishing channel relationships. Meanwhile, a non-linear enhancement (NLE) layer is employed to enhance the beneficial features while weaken the irrelevant ones. For building feature coupling between the cover image and the secret image, a channel-wise cross attention module (CCAM) is designed to fine-tune cover image features by capturing their cross-dependencies. In addition, for concealing data properly, a global–local aggregation module (GLAM) is deployed to adjust fused features by combining global and local attention, which can focus on inconspicuous and texture regions, respectively. The experimental results demonstrate that CAISFormer obtains PSNR gains of more than 0.36 dB and 0.90 dB for the cover/stego image pair and the secret/recovery image pair, respectively, and the detection ratio is decreased by 3.43%, in single image hiding compared to the state-of-the-art. Moreover, the generalization ability is also proved across a variety of datasets. The code will be made publicly available at https://github.com/YuhangZhouCJY/CAISFormer .
Ting Luo 0001, Zhouyan He, Gangyi Jiang, Haiyong Xu, Chin-Chen Chang 0001
Neurocomputing2
2024 MJPNet-S*: Multistyle Joint-Perception Network With Knowledge Distillation for Drone RGB-Thermal Crowd Density Estimation in Smart Cities
abstract
Crowd density estimation has gained significant research interest owing to its potential in various industries and social applications. Therefore, this paper proposes a multistyle joint-perception network based on a knowledge distillation-trained student network (MJPNet-S*) for drone-based red–green–blue, thermal/depth (RGB-T/D) crowd density estimation tasks. To provide superior accuracy and efficiency, a novel trimodal working module effectively combines the modalities to facilitate comprehensive extraction and utilization. A two-step strategy comprising high-and low-level fusion is employed in which the high-level features capture relational reasoning and a one-dimensional projection relationship module captures multisensory field information with high-quality semantics. A shallow injection fusion module leverages the multiscale and channel relationships at the low level to combine full-text information interactively. Finally, to reduce resource consumption, a neighboring collaborative distillation method enables the lightweight student network to achieve superior performance by increasing the speed by 92 reducing the number of parameters by 83 of the teacher. Extensive experiments demonstrate that the proposed MJPNet-S* performs remarkably well on two RGB-T datasets. The code will be made public at https://github.com/WBangG/MJPNet.
Wujie Zhou, Xiena Dong, Meixin Fang, Weiqing Yan, Ting Luo 0001
IEEE Internet Things J.6
2024 A Transformer-based invertible neural network for robust image watermarking
Zhouyan He, Renzhi Hu, Ting Luo 0001, Haiyong Xu
J. Vis. Commun. Image Represent.4
2024 Vision graph convolutional network for underwater image enhancement
Zexuan Xing, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Yeyao Chen
Knowl. Based Syst.5
2024 DMFTNet: dense multimodal fusion transfer network for free-space detection
Jiabao Ma, Wujie Zhou, Meixin Fang, Ting Luo 0001
Multim. Syst.4
2024 Underwater Monocular Depth Estimation Based on Physical-Guided Transformer
abstract
Owing to the light absorption and wavelength scattering in underwater environments, underwater images are severely degraded, which directly affects the depth estimation of underwater scenes. Accurate underwater depth estimation is essential for representing and understanding underwater scenes. However, the existing underwater depth estimation methods have not fully taken into account the distinctive physical properties of underwater environments, which has resulted in increased bias and feature distortion in the depth estimation results. In this paper, an underwater monocular depth estimation method based on physical-guided Transformer (UPGformer) is proposed, considering the characteristics of underwater imaging, including shallow feature extraction, encoding, decoding, and regression stages. Specifically, in the shallow feature extraction stage, considering the color deviation of underwater images and extracting richer primary features, an enrichment and extraction depth Transformer (EEDT) module is proposed, by interacting physically inverted transmission maps of the underwater dark channel prior (UDCP) with physical color-compensated underwater images through self-attention. In the encoding stage, considering the nonuniform degradation of underwater images (nonuniform local distortion and inconsistent channel degradation), the underwater physical Transformer interaction encoder (UPTE) module, which fuses the Transformer and physically inverted transmission maps, is proposed. Furthermore, in the decoding stage, to better recover features and reduce information loss, the underwater physical embedded decoding (UPED) module is proposed, which embeds the physically inverted transmission maps with the upsampling process. Finally, the depth map is constructed during the regression stage. The experimental results demonstrate that the proposed UPGformer outperforms existing methods, both qualitatively and quantitatively.
Chen Wang 0141, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Yeyao Chen
IEEE Trans. Geosci. Remote. Sens.5
2024 Underwater Image Quality Assessment from Synthetic to Real-world: Dataset and Objective Method
abstract
The complicated underwater environment and lighting conditions lead to severe influence on the quality of underwater imaging, which tends to impair underwater exploration and research. To effectively evaluate the quality of underwater images, an underwater image quality assessment dataset is constructed from synthetic to real-world, and then a new objective underwater image assessment method based on the characteristics of the underwater imaging is proposed (UICQA). Specifically, to address the lack of a publicly available datasets and more accurately quantify the quality of underwater images, a subjective underwater image quality assessment dataset from synthetic to real-world underwater images, named USRD, is constructed. Considering that the transmission map can effectively reflect the characteristics of the underwater imaging, statistical features are effectively extracted from the transmission map for distinguishing underwater images of different quality. Further, considering that the transmission map negatively correlates with scene depth, a local-to-global transmission map weighted contrast feature is constructed. Additionally, the color features of human perception and texture features based on fractal dimensions are proposed. Finally, the experimental results show that the proposed UICQA method exhibits the highest correlation with ground truth scores compared to state-of-the-art UIQA methods.
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Xuebo Zhang 0002, Hongwei Ying
ACM Trans. Multim. Comput. Commun. Appl.5
2024 DHFNet: dual-decoding hierarchical fusion network for RGB-thermal semantic segmentation
Yuqi Cai, Wujie Zhou, Lu Yu 0003, Ting Luo 0001
Vis. Comput.5
2023 UDAformer: Underwater image enhancement based on dual attention transformer
Haiyong Xu, Ting Luo 0001, Yang Song 0015, Zhouyan He
Comput. Graph.3
2023 MENet: Lightweight multimodality enhancement network for detecting salient objects in RGB-thermal images
Wujie Zhou, Xiaohong Qian, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001
Neurocomputing6
2023 A bilateral attention based generative adversarial network for DIBR 3D image watermarking
Zhouyan He, Lingqiang He, Haiyong Xu, Tong-Yuen Chai, Ting Luo 0001
J. Vis. Commun. Image Represent.5
2023 RDD-net: Robust duplicated-diffusion watermarking based on deep network
Guowei Jiang, Zhouyan He, Jiangtao Huang, Ting Luo 0001, Haiyong Xu, Chongchong Jin
J. Vis. Commun. Image Represent.4
2023 A two-stage and two-branch generative adversarial network-based underwater image enhancement
Haiyong Xu, Chi Lin 0003, Ting Luo 0001
Vis. Comput.4
2022 Light-field image watermarking based on geranion polar harmonic Fourier moments
Chunpeng Wang 0001, Bin Ma 0003, Jian Li 0034, Ting Luo 0001, Qi Li 0029
Eng. Appl. Artif. Intell.6
2022 GCNet: Grid-like context-aware network for RGB-thermal semantic segmentation
Wujie Zhou, Yueli Cui, Lu Yu 0003, Ting Luo 0001
Neurocomputing5
2022 HFNet: Hierarchical feedback network with multilevel atrous spatial pyramid pooling for RGB-D saliency detection
Wujie Zhou, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001
Neurocomputing5
2022 Robust HDR video watermarking method based on the HVS model and T-QR
Ting Luo 0001, Haiyong Xu, Yang Song 0015, Chunpeng Wang 0001, Li Li 0014
Multim. Tools Appl.2
2022 Multi-Angle Projection Based Blind Omnidirectional Image Quality Assessment
abstract
Most of the existing blind omnidirectional image quality assessment (BOIQA) methods are based on data-driven approach where the end-to-end neural network or deep learning tools are mainly used for feature extraction. However, it usually lacks interpretability and is difficult to discover the perceptual mechanism behind. In this paper, from the perspective of perception modeling, we propose a novel multi-angle projection based BOIQA (MP-BOIQA) method. Considering the omnibearing and near eye display characteristics with head mounted display, multiple color cubemap projection images with respect to different viewpoints are grouped as the color omnidirectional distortion (COD) units so as to simulate the user’s viewing behavior in subjective quality assessment. In the designed multi-angle projection based feature extractor, tensor decomposition is implemented on each COD unit for dimensionality reduction, and piecewise exponential fitting is used to get the distribution of mean subtracted contrast normalized coefficients of the unit’s feature matrices in tensor domain. Finally, the extracted features are pooled with random forest. The experimental results on three omnidirectional image quality datasets show that the MP-BOIQA method can deliver highly competitive performance compared with some representative full-reference quality assessment methods, as well as some state-of-the-art BOIQA methods.
Hao Jiang 0014, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Haiyong Xu
IEEE Trans. Circuits Syst. Video Technol.4
2022 Reinforced Swin-Convs Transformer for Simultaneous Underwater Sensing Scene Image Enhancement and Super-resolution
abstract
Underwater image enhancement (UIE) technology aims to tackle the challenge of restoring the degraded underwater images due to light absorption and scattering. Meanwhile, the ever-increasing requirement for higher resolution images from a lower resolution in the underwater domain cannot be overlooked. To address these problems, a novel U-Net-based reinforced Swin-Convs Transformer for simultaneous enhancement and superresolution (URSCT-SESR) method is proposed. Specifically, with the deficiency of U-Net based on pure convolutions, the Swin Transformer is embedded into U-Net for improving the ability to capture the global dependence. Then, given the inadequacy of the Swin Transformer capturing the local attention, the reintroduction of convolutions may capture more local attention. Thus, an ingenious manner is presented for the fusion of convolutions and the core attention mechanism to build a reinforced Swin-Convs Transformer block (RSCTB) for capturing more local attention, which is reinforced in the channel and the spatial attention of the Swin Transformer. Finally, experimental results on available datasets demonstrate that the proposed URSCT-SESR achieves the state-of-the-art performance compared with other methods in terms of both subjective and objective evaluations. The code is publicly available athttps://github.com/TingdiRen/URSCT-SESR.
Tingdi Ren, Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001
IEEE Trans. Geosci. Remote. Sens.7
2022 Robust HDR video watermarking method based on saliency extraction and T-SVD
Ting Luo 0001, Haiyong Xu, Yang Song 0015, Chunpeng Wang 0001
Vis. Comput.2
2021 A Novel HDR Image Zero-Watermarking Based on Shift-Invariant Shearlet Transform
abstract
In this paper, a novel high dynamic range (HDR) image zero-watermarking algorithm against the tone mapping attack is proposed. In order to extract stable and invariant features for robust zero-watermarking, the shift-invariant shearlet transform (SIST) is used to transform the HDR image. Firstly, the HDR image is converted to CIELAB color space, and the L component is selected to perform SIST for obtaining the low-frequency subband containing the robust structure information of the image. Secondly, the low-frequency subband is divided into nonoverlapping blocks, which are transformed by using discrete cosine transform (DCT) and singular value decomposition (SVD) to obtain the maximum singular values for constructing a binary feature image. To increase the watermarking security, a hybrid chaotic mapping (HCM) is employed to get the scrambled watermark. Finally, an exclusive-or operation is performed between the binary feature image and the scrambled watermark to compute robust zero-watermark. Experimental results show that the proposed algorithm has a good capability of resisting tone mapping and other image processing attacks.
Shanshan Shi, Ting Luo 0001, Jiangtao Huang
Secur. Commun. Networks2
2021 Multiscale multilevel context and multimodal fusion for RGB-D salient object detection
Junwei Wu 0001, Wujie Zhou, Ting Luo 0001, Lu Yu 0003, Jingsheng Lei
Signal Process.3
2021 Multi-layer fusion network for blind stereoscopic 3D visual quality prediction
Wujie Zhou, Xinyang Lin, Jingsheng Lei, Lu Yu 0003, Ting Luo 0001
Signal Process. Image Commun.6
2020 Multi-exposure image fusion based on tensor decomposition
Shengcong Wu, Ting Luo 0001, Yang Song 0015, Haiyong Xu
Multim. Tools Appl.2
2019 Convolutional neural networks-based stereo image reversible data hiding method
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Caiming Zhong, Haiyong Xu, Zhiyong Pan
J. Vis. Commun. Image Represent.1
2019 Deep blind quality evaluator for multiply distorted images based on monogenic binary coding
Wujie Zhou, Lu Yu 0003, Yaguan Qian, Weiwei Qiu, Yang Zhou 0011, Ting Luo 0001
J. Vis. Commun. Image Represent.6
2019 Lossless fragile watermarking algorithm in compressed domain for multiview video coding
Wei Gao 0013, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001
Multim. Tools Appl.4
2019 A novel robust color image watermarking method using RGB correlations
Fangyan Zhang, Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wujie Zhou
Multim. Tools Appl.2
2019 Ensemble clustering based on evidence extracted from the co-association matrix
Caiming Zhong, Lianyu Hu 0001, Ting Luo 0001, Haiyong Xu
Pattern Recognit.4
2019 Robust high dynamic range color image watermarking method based on feature map extraction
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wei Gao 0013
Signal Process.1
2018 3D visual discomfort predictor based on subjective perceived-constraint sparse representation in 3D display system
Haiyong Xu, Gangyi Jiang, Mei Yu 0001, Ting Luo 0001, Zongju Peng, Feng Shao 0001, Hao Jiang 0014
Future Gener. Comput. Syst.4
2018 Sparse recovery based reversible data hiding method using the human visual system
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu, Wei Gao 0013
Multim. Tools Appl.1
2018 Local and Global Feature Learning for Blind Quality Evaluation of Screen Content and Natural Scene Images
abstract
The blind quality evaluation of screen content images (SCIs) and natural scene images (NSIs) has become an important, yet very challenging issue. In this paper, we present an effective blind quality evaluation technique for SCIs and NSIs based on a dictionary of learned local and global quality features. First, a local dictionary is constructed using local normalized image patches and conventional -means clustering. With this local dictionary, the learned local quality features can be obtained using a locality-constrained linear coding with max pooling. To extract the learned global quality features, the histogram representations of binary patterns are concatenated to form a global dictionary. The collaborative representation algorithm is used to efficiently code the learned global quality features of the distorted images using this dictionary. Finally, kernel-based support vector regression is used to integrate these features into an overall quality score. Extensive experiments involving the proposed evaluation technique demonstrate that in comparison with most relevant metrics, the proposed blind metric yields significantly higher consistency in line with subjective fidelity ratings.
Wujie Zhou, Lu Yu 0003, Yang Zhou 0011, Weiwei Qiu, Mingwei Wu 0001, Ting Luo 0001
IEEE Trans. Image Process.6
2017 Blind 3D image quality assessment based on self-similarity of binocular features
Wujie Zhou, Shuangshuang Zhang, Lu Yu 0003, Weiwei Qiu, Yang Zhou 0011, Ting Luo 0001
Neurocomputing7
2017 Stereoscopic image quality assessment by learning non-negative matrix factorization-based color visual characteristics and considering binocular interactions
Gangyi Jiang, Haiyong Xu, Mei Yu 0001, Ting Luo 0001, Yun Zhang 0002
J. Vis. Commun. Image Represent.4
2017 Blind quality estimator for 3D images based on binocular combination and extreme learning machine
Wujie Zhou, Lu Yu 0003, Yang Zhou 0011, Weiwei Qiu, Mingwei Wu 0001, Ting Luo 0001
Pattern Recognit.6
2016 Asymmetric self-recovery oriented stereo image watermarking method for three dimensional video system
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Haiyong Xu
Multim. Syst.1
2016 Utilizing binocular vision to facilitate completely blind 3D image quality measurement
Wujie Zhou, Lu Yu 0003, Weiwei Qiu, Ting Luo 0001, Zhongpeng Wang, Mingwei Wu 0001
Signal Process.4
2014 Disparity based stereo image reversible data hiding
abstract
As the popularity of three dimensional video, security of stereo image has become an evident issue to be solved. This paper presents a disparity based stereo image reversible data hiding by using histogram shifting, which can recover the original stereo image from marked stereo image without any distortion. Inter-correlations between left and right views of stereo image are utilized to predict pixels accurately. Then prediction error bins are constructed, and many points are around zero-valued bin for embedding data with low distortion of stereo images. The zero-valued bin is used twice to embed data, so that embedding capacity can reach more than 1 bit per pixel. Experimental results demonstrate that the proposed method outperforms the extended stereo image data hiding methods.
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng
ICIP1
2014 Stereo image watermarking scheme for authentication with self-recovery capability using inter-view reference sharing
Ting Luo 0001, Gangyi Jiang, Mei Yu 0001, Feng Shao 0001, Zongju Peng
Multim. Tools Appl.1