VLDB 2026 Research / reviewers in the wild / expert
Fayaz Ali Dharejo
dblp:226/3346
· DBLP profile ↗
24ranked-venue papers
5as first author
21since 2021 · last 2026
0000-0001-7685-3913ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 9 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Illuminating Darkness: Learning to Enhance Low-light Images In-the-WildabstractSingle-shot low-light image enhancement (SLLIE) remains challenging due to the limited availability of diverse, realworld paired datasets. To bridge this gap, we introduce the Low-Light Smartphone Dataset (LSD), a large-scale, high-resolution (4K+) dataset collected in the wild across a wide range of challenging lighting conditions (0.1–200 lux). LSD contains 6,425 precisely aligned low and normallight image pairs, selected from over 8,000 dynamic indoor and outdoor scenes through multi-frame acquisition and expert evaluation. To evaluate generalization and aesthetic quality, we collect 2,117 unpaired low-light images from previously unseen devices. To fully exploit LSD, we propose TFFormer, a hybrid model that encodes luminance and chrominance (LC) separately to reduce color-structure entanglement. We further propose a cross-attention-driven joint decoder for context-aware fusion of LC representations, along with LC refinement and LC-guided supervision to significantly enhance perceptual fidelity and structural consistency. TFFormer achieves state-of-the-art results on LSD (+2.45 dB PSNR) and substantially improves downstream vision tasks, such as low-light object detection (+6.80 mAP on ExDark). S. M. A. Sharif, Fayaz Ali Dharejo, Radu Timofte, Rizwan Ali Naqvi |
WACV | 4 |
| 2026 | Underwater visual tracking with a large scale dataset and image enhancementabstractThis paper presents a new dataset and general tracker enhancement method for Underwater Visual Object Tracking (UVOT). Despite its significance, underwater tracking has remained unexplored due to data inaccessibility. It poses distinct challenges; the underwater environment exhibits non-uniform lighting conditions, low visibility, lack of sharpness, low contrast, camouflage, and reflections from suspended particles. Performance of traditional tracking methods designed primarily for terrestrial or open-air scenarios drops in such conditions. We address the problem by proposing a novel underwater image enhancement algorithm designed specifically to boost tracking quality. The method has resulted in a significant performance improvement, of up to 5.0% AUC, of state-of-the-art (SOTA) visual trackers. To develop robust and accurate UVOT methods, large-scale datasets are required. To this end, we introduce a large-scale UVOT benchmark dataset consisting of 400 video segments and 275,000 manually annotated frames enabling underwater training and evaluation of deep trackers. The videos are labelled with several underwater-specific tracking attributes including watercolor variation, target distractors, camouflage, target relative size, and low visibility conditions. The UVOT400 dataset, tracking results, and the code are publicly available on: https://github.com/BasitAlawode/UVOT400 . Basit Alawode, Sajid Javed, Fayaz Ali Dharejo, Mehnaz Ummar, Arif Mahmoud, Fahad Shahbaz Khan, Jiri Matas |
Neurocomputing | 3 |
| 2026 | Spec-ViT: A Vision Transformer With Wavelet for Anti-Aliasing and Denoising in Medical Image ClassificationabstractMedical image analysis remains challenging due to inherent limitations in imaging modalities, where structural aliasing and noise artifacts persistently compromise diagnostic accuracy. While convolutional neural networks (CNNs) and vision transformers (ViTs) have achieved remarkable progress in feature extraction, their inherent sampling mechanisms and spectral biases often exacerbate these high-frequency distortions, leading to suboptimal lesion characterization. To address this critical limitation, we propose Spec-ViT, a novel wavelet-based anti-aliasing Transformer architecture that synergistically integrates adaptive spectral purification with hierarchical attentive learning. The Wavelet Antialiasing Module (WAM) first implements learnable smoothing factor in the wavelet domain to suppress highfrequency artifacts, while preserving clinically relevant lowfrequency structures and fine diagnostic details. Building upon this spectral foundation, the Lightweight Enhanced Attention (LEA) refines feature representations through a dual-path mechanism, coupling channel-spatial attention with global multi-head self-attention to enhance lesion context modeling. Finally, the Smoothed Convolutional Gate (SCG) further sharpens local discriminability through depthwise convolution and adaptive Swish gating, completing a coherent pipeline from frequency-aware purification to global-local attentive analysis. Extensive experiments on five benchmark medical image classification datasets demonstrate that Spec-ViT consistently outperforms both baseline and state-of-the-art methods, achieving up to 84.04% accuracy on the Pediatric Pneumonia Chest X-rays dataset in particular. Yanying Rao, Yuzheng Su, Fayaz Ali Dharejo, Radu Timofte, Guo-jun Mao, Moath Alathbah |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | Multi-Agent DRL for QKD-Enabled Resource Allocation in 6G TN-NTN Metaverse ServiceabstractThe integration of terrestrial and non-terrestrial networks (TN-NTN) in 6 G is essential to support real-time applications like the Metaverse and intelligent edge services, which demand ultra-reliable low-latency communications (xURLLC). Managing these networks and maintaining robust security presents significant challenges due to their complexity and high-dimensional environments. Quantum communication, particularly quantum key distribution (QKD), offers a promising solution by providing unbreakable encryption and enhancing security across TN-NTN architectures. In this paper, we propose a novel deep reinforcement learning approach for QKD-enabled resource allocation in 6 G TN-NTN Metaverse service and transform the joint resource allocation and QKD deployment cost optimization problem into a stochastic game model to ensure secure and efficient resource distribution across TN-NTN environment. We introduce a novel hierarchical multi-agent proximal policy optimization (MAPPO) framework to address the formulated optimization problem. This framework enables dynamic and secure allocation of Metaverse resources and services from multiple providers to users while minimizing QKD deployment costs. Our simulations demonstrate that the proposed framework significantly enhances network performance, reduces key generation costs, and optimizes resource utilization and service quality. Hayla Nahom Abishu, Fayaz Ali Dharejo, Aiman Erbad, Mounir Hamdi, Mohsen Guizani |
ICC | 3 |
| 2025 | SF-YOLO: A Novel YOLO Framework for Small Object Detection in Aerial ScenesabstractABSTRACT Object detection models are widely applied in the fields such as video surveillance and unmanned aerial vehicles to enable the identification and monitoring of various objects on a diversity of backgrounds. The general CNN‐based object detectors primarily rely on downsampling and pooling operations, often struggling with small objects that have low resolution and failing to fully leverage contextual information that can differentiate objects from complex background. To address the problems, we propose a novel YOLO framework called SF‐YOLO for small object detection. Firstly, we present a spatial information perception (SIP) module to extract contextual features for different objects through the integration of space to depth operation and large selective kernel module, which dynamically adjusts receptive field of the backbone and obtains the enhanced features for richer understanding of differentiation between objects and background. Furthermore, we design a novel multi‐scale feature weighted fusion strategy, which performs weighted fusion on feature maps by combining fast normalized fusion method and CARAFE operation, accurately assessing the importance of each feature and enhancing the representation of small objects. The extensive experiments conducted on VisDrone2019, Tiny‐Person and PESMOD datasets demonstrate that our proposed method enables comparable detection performance to state‐of‐the‐art detectors. Wangyu Jiang, Fayaz Ali Dharejo, Guojun Mao, Radu Timofte |
IET Image Process. | 4 |
| 2025 | Multi-Modal LLMs in Agriculture: A Comprehensive ReviewabstractGiven the rapid emergence and applications of Multi-Modal Large Language Models (MM-LLMs) across various scientific fields, insights regarding their applicability in agriculture are still only partially explored. This paper conducts an in-depth review of MM-LLMs in agriculture, focusing on understanding how MM-LLMs can be developed and implemented to optimize agricultural processes, increase efficiency, and reduce costs. Recent studies have explored the capabilities of MM-LLMs in agricultural information processing and decision-making. Despite these advancements, significant gaps persist, particularly in addressing domain-specific challenges such as variable data quality and availability, integration with existing agricultural systems, and the creation of robust training datasets that accurately represent complex agricultural environments. Moreover, a comprehensive understanding of the capabilities, challenges, and limitations of MM-LLMs in agricultural information processing and application is still missing. Exploring these areas is crucial to providing the community with a broader perspective and a clearer understanding of MM-LLMs’ applications, establishing a benchmark for the current state and emerging trends in this field. To bridge this gap, this survey reviews the progress of MM-LLMs and their utilization in agriculture, with an additional focus on 11 key research questions (RQs), where 4 RQs are general and 7 RQs are agriculture focused. By addressing these RQs, this review outlines the current opportunities and challenges, limitations, and future roadmap for MM-LLMs in agriculture. The findings indicate that multi-modal MM-LLMs not only simplify complex agricultural challenges but also significantly enhance decision-making and improve the efficiency of agricultural image processing. These advancements position MM-LLMs as an essential tool for the future of farming. For continued research and understanding, an organized and regularly updated list of papers on MM-LLMs is available at https://github.com/JiajiaLi04/Multi-Modal-LLMs-in-Agriculture. Ranjan Sapkota, Rizwan Qureshi, Muhammad Usman Hadi, Syed Zohaib Hassan, Ferhat Sadak, Maged Shoman, Fayaz Ali Dharejo, Paudel Achyut, John M. Shutske, Manoj Karkee |
IEEE Trans Autom. Sci. Eng. | 8 |
| 2025 | Unsupervised Dual Transformer Learning for 3-D Textured Surface SegmentationabstractAnalysis of the 3-D texture is indispensable for various tasks, such as retrieval, segmentation, classification, and inspection of sculptures, knit fabrics, and biological tissues. A 3-D texture represents a locally repeated surface variation (SV) that is independent of the overall shape of the surface and can be determined using the local neighborhood and its characteristics. Existing methods mostly employ computer vision techniques that analyze a 3-D mesh globally, derive features, and then utilize them for classification or retrieval tasks. While several traditional and learning-based methods have been proposed in the literature, only a few have addressed 3-D texture analysis, and none have considered unsupervised schemes so far. This article proposes an original framework for the unsupervised segmentation of 3-D texture on the mesh manifold. The problem is approached as a binary surface segmentation task, where the mesh surface is partitioned into textured and nontextured regions without prior annotation. The proposed method comprises a mutual transformer-based system consisting of a label generator (LG) and a label cleaner (LC). Both models take geometric image representations of the surface mesh facets and label them as texture or nontexture using an iterative mutual learning scheme. Extensive experiments on three publicly available datasets with diverse texture patterns demonstrate that the proposed framework outperforms standard and state-of-the-art unsupervised techniques and performs reasonably well compared to supervised methods. Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Sajid Javed, Syed Sadaf Ali, Naoufel Werghi |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2024 | 3D-TexSeg: Unsupervised Segmentation of 3D Texture Using Mutual Transformer LearningabstractAnalysis of the 3D Texture is indispensable for various tasks, such as retrieval, segmentation, classification, and inspection of sculptures, knitted fabrics, and biological tissues. A 3D texture is a locally repeated surface variation independent of the surface’s overall shape and can be determined using the local neighborhood and its characteristics. Existing techniques typically employ computer vision techniques that analyze a 3D mesh globally, derive features, and then utilize the obtained features for retrieval or classification. Several traditional and learning-based methods exist in the literature; however, only a few are on 3D texture, and nothing yet, to the best of our knowledge, on the unsupervised schemes. This paper presents an original framework for the unsupervised segmentation of the 3D texture on the mesh manifold. We approach this problem as binary surface segmentation, partitioning the mesh surface into textured and non-textured regions without prior annotation. We devise a mutual transformer-based system comprising a label generator and a cleaner. The two models take geometric image representations of the surface mesh facets and label them as texture or non-texture across an iterative mutual learning scheme. Extensive experiments on three publicly available datasets with diverse texture patterns demonstrate that the proposed framework outperforms standard and SOTA unsupervised techniques and competes reasonably with supervised methods. Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Sajid Javed, Syed Sadaf Ali, Naoufel Werghi |
3DV | 2 |
| 2024 | CPLIP: Zero-Shot Learning for Histopathology with Comprehensive Vision-Language AlignmentabstractThis paper proposes Comprehensive Pathology Language Image Pretraining (CPLIP), a new unsupervised technique designed to enhance the alignment of images and text in histopathology for tasks such as classification and segmentation. This methodology enriches vision-language models by leveraging extensive data without needing ground truth annotations. CPLIP involves constructing a pathology-specific dictionary, generating textual descriptions for images using language models, and retrieving relevant images for each text snippet via a pretrained model. The model is then fine-tuned using a many-to-many contrastive learning method to align complex interrelated concepts across both modalities. Evaluated across multiple histopathology tasks, CPLIP shows notable improvements in zero-shot learning scenarios, outperforming existing methods in both interpretability and robustness and setting a higher benchmark for the application of vision-language models in the field. To encourage further research and replication, the code for CPLIP is available on GitHub at https://cplip.github.io/ Sajid Javed, Arif Mahmood, Iyyakutti Iyappan Ganapathi, Fayaz Ali Dharejo, Naoufel Werghi, Mohammed Bennamoun |
CVPR | 4 |
| 2024 | Resource management in multi-heterogeneous cluster networks using intelligent intra-clustered federated learning
Fahad Razaque Mughal, Jingsha He, Nafei Zhu, Saqib Hussain, Zulfiqar Ali Zardari, Gulam Ali Mallah, Mohammad Jalil Piran, Fayaz Ali Dharejo |
Comput. Commun. | 8 |
| 2024 | FIND: Privacy-Enhanced Federated Learning for Intelligent Fake News DetectionabstractThe development and popularity of social networks have made information dissemination unprecedentedly convenient and speedy. However, the spread of fake news can often cause serious harm to society and individuals. Therefore, machine learning-based fake news detection methods have become increasingly important. The existing work often needs to collect sufficient user-side data for training, which also boosts the privacy leakage risk to the users. Therefore, this article proposes an intelligent fake news detection system based on federated learning (FL) called FIND, which can train a global model while keeping user data locally. At the same time, we also designed a sparsified update perturbation method to enhance the system security further. Finally, we conduct simulation experiments to study and discuss multiple acoustic factors and prove the feasibility of our system in terms of accuracy, security, and efficiency. Zhuotao Lian, Chen Zhang 0033, Chunhua Su, Fayaz Ali Dharejo, Mutiq Almutiq, Muhammad Hammad Memon |
IEEE Trans. Comput. Soc. Syst. | 4 |
| 2023 | Window-based transformer generative adversarial network for autonomous underwater image enhancement
Mehnaz Ummar, Fayaz Ali Dharejo, Basit Alawode, Taslim Mahbub, Mohammad Jalil Piran, Sajid Javed |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | View-aware attribute-guided network for vehicle re-identification
Saifullah Tumrani, Wazir Ali, Rajesh Kumar 0014, Abdullah Aman Khan, Fayaz Ali Dharejo |
Multim. Syst. | 5 |
| 2023 | Transfer learning-based quantized deep learning models for nail melanoma classification
Mujahid Hussain, Makhmoor Fiza, Aiman Khalil, Asad Ali Siyal, Fayaz Ali Dharejo, Waheeduddin Hyder, Antonella Guzzo, Moez Krichen, Giancarlo Fortino |
Neural Comput. Appl. | 5 |
| 2023 | Multimodal-Boost: Multimodal Medical Image Super-Resolution Using Multi-Attention Network With Wavelet TransformabstractMultimodal medical images are widely used by clinicians and physicians to analyze and retrieve complementary information from high-resolution images in a non-invasive manner. Loss of corresponding image resolution adversely affects the overall performance of medical image interpretation. Deep learning-based single image super resolution (SISR) algorithms have revolutionized the overall diagnosis framework by continually improving the architectural components and training strategies associated with convolutional neural networks (CNN) on low-resolution images. However, existing work lacks in two ways: i) the SR output produced exhibits poor texture details, and often produce blurred edges, ii) most of the models have been developed for a single modality, hence, require modification to adapt to a new one. This work addresses (i) by proposing generative adversarial network (GAN) with deep multi-attention modules to learn high-frequency information from low-frequency data. Existing approaches based on the GAN have yielded good SR results; however, the texture details of their SR output have been experimentally confirmed to be deficient for medical images particularly. The integration of wavelet transform (WT) and GANs in our proposed SR model addresses the aforementioned limitation concerning textons. While the WT divides the LR image into multiple frequency bands, the transferred GAN uses multi-attention and upsample blocks to predict high-frequency components. Additionally, we present a learning method for training domain-specific classifiers as perceptual loss functions. Using a combination of multi-attention GAN loss and a perceptual loss function results in an efficient and reliable performance. Applying the same model for medical images from diverse modalities is challenging, our work addresses (ii) by training and performing on several modalities via transfer learning. Using two medical datasets, we validate our proposed SR network against existing state-of-the-art approaches and achieve promising results in terms of structural similarity index (SSIM) and peak signal-to-noise ratio (PSNR). Fayaz Ali Dharejo, Muhammad Zawish, Farah Deeba, Yuanchun Zhou, Kapal Dev, Sunder Ali Khowaja, Nawab Muhammad Faseeh Qureshi |
IEEE ACM Trans. Comput. Biol. Bioinform. | 1 |
| 2022 | A novel image dehazing framework for robust vision-based intelligent systemsabstractApart from high-level computer vision tasks, deep learning has also made significant progress in low-level tasks, including single image dehazing. A well-detailed image looks realistic and natural with its clear edges and balanced colour. To achieve a clearer and vivid view, we exploit the role of edges and colours as a significant part of our proposed work. A progressive two-stage image dehazing network is presented to overcome the challenges of current image dehazing algorithms. The proposed image dehazing framework is divided into two steps; in the first stage, the multiscale image features of the encoder and decoder structure can be extracted. The second stage consists of the Color Correction Model (CCM), which retrieves balanced colour close to the ground truth. The encode-decoder network consists of a dense residual attention unit (DRAU) that comprises channel attention with pixel attention mechanisms. We have seen that weighted information and the haze difference is inconsistent across pixels without DRAU at the various channel-specific features. DRAU deals with different features and pixels unequally, which offers more versatility in handling knowledge of various types of detailed information. Our proposed two-stage network exceeds state-of-the-art algorithms in both visual and quantitative aspects. The findings are tested with the best-published peak signal-to-noise ratio metrics of 33.55–33.44 dB and SSIM 0.9619–0.9714 on SOTS indoor and outdoor test data sets. Farah Deeba, Fayaz Ali Dharejo, Muhammad Zawish, Fida Hussain Memon, Kapal Dev, Rizwan Ali Naqvi, Yuanchun Zhou, Yi Du 0010 |
Int. J. Intell. Syst. | 2 |
| 2022 | FuzzyAct: A Fuzzy-Based Framework for Temporal Activity Recognition in IoT Applications Using RNN and 3D-DWTabstractDespite massive research in deep learning, the human activity recognition (HAR) domain still suffers from key challenges in terms of accurate classification and detection. The core idea behind recognizing activities accurately is to assist Internet-of-things (IoT) enabled smart surveillance systems. Thereby, this work is based on the joint use of discrete wavelet transform (DWT) and recurrent neural network (RNN) to classify and detect human activities accurately. Recent approaches on HAR exploit the three-dimensional (3-D) convolutional neural networks (CNNs) to extract spatial information, which adds a computational burden. In our case, features are extracted using 3D-DWT instead of 3-D CNNs, performed in three steps of 1D-DWT to reflect the spatio-temporal features of human action. Given the features, the RNN produces an output label for each video clip taking care of the long-term temporal consistency among close predictions in the output sequence. It is noticed that feature extraction through 3D-DWT essentially recovers the multiple angles of an activity. Many HAR techniques distinguish an activity based on the posture of an image frame rather than learning the transitional relationship between postures in the temporal sequence, resulting in degraded accuracy. To address this problem, in this article, we designed a novel rank-based fuzzy approach that segregates activities precisely by ranking the probabilities of activities based on confidence scores. FuzzyAct achieved an average mean average precision (mAP) of 0.8012 mAP on the ActivityNet dataset, and outperformed the baseline counterparts and other state-of-the-art approaches on benchmark datasets. Finally, we present a mechanism to compress the proposed RNN for edge-enabled IoT applications. Fayaz Ali Dharejo, Muhammad Zawish, Yuanchun Zhou, Steven Davy, Kapal Dev, Sunder Ali Khowaja, Yanjie Fu, Nawab Muhammad Faseeh Qureshi |
IEEE Trans. Fuzzy Syst. | 1 |
| 2022 | Reinforcement learning based on routing with infrastructure nodes for data dissemination in vehicular networks (RRIN)
Arbelo Lolai, Xingfu Wang, Ammar Hawbani, Fayaz Ali Dharejo, Taiyaba Qureshi, Muhammad Umar Farooq 0002, Muhammad Mujahid, Abdul Hafeez Babar |
Wirel. Networks | 4 |
| 2021 | A plexus-convolutional neural network framework for fast remote sensing image super-resolution in wavelet domainabstractAbstract Satellite image processing has been widely used in recent years in a number of applications such as land classification, Identification transfer, resource exploration, super‐resolution image, etc. Due to the orbital location, revision time, quick view angle limitations, and weather impact, the satellite images are challenging to manage. There are many types of resolution, such as spatial, spectral, and temporal. Still, in our case, we concentrated on spatial image resolution to super resolve the images from low‐resolution images. For remote sensing image super‐resolution fast wavelet‐based super‐resolution (FWSR), we propose a novel, fast wavelet‐based plexus framework that performs super‐resolution convolutional neural network (SRCNN)‐like extraction of features based on three hidden layers. First, wavelet sub‐band images are combined into a pre‐defined full‐scale data training factor, including approximation and interchangeable stand‐alone units (frequency sub‐bands). Second, to speed up image recovery, mapping the sub‐band image of the wavelet is then measured using its approximate image. Third, the added sub‐pixel layer at the end of the network model is intended to reproduce image quality using a plexus framework. The approximation sub‐band images obtained after discrete wavelet transform wavelet decomposition are used as input rather than the original image because of their high‐frequency data and preserved characteristics. Five current super‐resolution neural network approaches are compared with the proposed technique and tested on three pubic satellite image datasets and two benchmark datasets. The experimental findings are well compared qualitatively and quantitatively. Farah Deeba, Yuanchun Zhou, Fayaz Ali Dharejo, Muhammad Ashfaq Khan, Bhagwan Das, Xuezhi Wang 0004, Yi Du 0010 |
IET Image Process. | 3 |
| 2021 | A remote-sensing image enhancement algorithm based on patch-wise dark channel prior and histogram equalisation with colour correctionabstractAbstract The object identification within an image captured during rough weather conditions (such as haze, fog) poses difficulty due to the reduction of an image. The rough weather conditions lead not only to the variation of the image's visual effect but also to the disadvantage of post‐processing of an image. Furthermore, it causes inconvenience of all types of instruments that rely on optical imaging, such as satellite remote‐sensing systems, aerial photo systems, outdoor monitoring systems, and object identification systems, respectively. Hence, the improvement and restorement of the visual effects and enhanced post‐processing are needed. This research introduces a new image enhancement approach for image dehazing based on dark channel prior and piecewise linear transformation; also, the histogram equalisation technique, i.e. contrast limited adaptive histogram equalisation is applied. A dark channel prior is well known for its simplicity and productivity. In this work, the dark channel prior to a new angle is analysed in the first step, where average patch sizes are estimated for the computation of haze densities. Furthermore, the sky is approximated up to 5–10% of the hazy images, which has a good effect in removing the haze from the image. Using the dark channel, the proposed algorithm significantly boosted the effects of the dark images as well as reduced the influence of haze and noise. Eventually, for colour correction, the piecewise linear transformation technique is applied, which enhances the colour close to the original image. Experimental results demonstrate that the proposed method significantly improves the visibility of the algorithm on dark remote‐sensing images as well as on hazy natural images. Fayaz Ali Dharejo, Yuanchun Zhou, Farah Deeba, Munsif Ali Jatoi, Yi Du 0010, Xuezhi Wang 0004 |
IET Image Process. | 1 |
| 2021 | TWIST-GAN: Towards Wavelet Transform and Transferred GAN for Spatio-Temporal Single Image Super ResolutionabstractSingle Image Super-resolution (SISR) produces high-resolution images with fine spatial resolutions from a remotely sensed image with low spatial resolution. Recently, deep learning and generative adversarial networks (GANs) have made breakthroughs for the challenging task of single image super-resolution (SISR) . However, the generated image still suffers from undesirable artifacts such as the absence of texture-feature representation and high-frequency information. We propose a frequency domain-based spatio-temporal remote sensing single image super-resolution technique to reconstruct the HR image combined with generative adversarial networks (GANs) on various frequency bands (TWIST-GAN). We have introduced a new method incorporating Wavelet Transform (WT) characteristics and transferred generative adversarial network. The LR image has been split into various frequency bands by using the WT, whereas the transfer generative adversarial network predicts high-frequency components via a proposed architecture. Finally, the inverse transfer of wavelets produces a reconstructed image with super-resolution. The model is first trained on an external DIV2 K dataset and validated with the UC Merced Landsat remote sensing dataset and Set14 with each image size of 256 × 256. Following that, transferred GANs are used to process spatio-temporal remote sensing images in order to minimize computation cost differences and improve texture information. The findings are compared qualitatively and qualitatively with the current state-of-art approaches. In addition, we saved about 43% of the GPU memory during training and accelerated the execution of our simplified version by eliminating batch normalization layers. Fayaz Ali Dharejo, Farah Deeba, Yuanchun Zhou, Bhagwan Das, Munsif Ali Jatoi, Muhammad Zawish, Yi Du 0010, Xuezhi Wang 0004 |
ACM Trans. Intell. Syst. Technol. | 1 |
| 2020 | Lossless digital image watermarking in sparse domain by using K-singular value decomposition algorithmabstractThe crucial hurdle faced by the watermarking technique is to maintain the steadiness corresponding to several attacks while assisting a sufficient level of security. In this study, a robust lossless sparse domain‐based watermarking approach combined with discrete cosine transform (DCT) is introduced to hide the secret message in the selected significant sparse elements of the host image. The proposed method takes advantage of a sparse representation‐based dictionary learning process. To enhance the security of the original image, the authors first apply the DCT on a secret message. These DCT coefficients with some regularised parameters will be inserted into the selected significant sparse coefficients. At the extraction stage, the secret message is extracted from those significant sparse coefficients by employing the sparse domain orthogonal matching pursuit algorithm. Finally, the inverse DCT is applied to extract the secret message without any information loss. To show the effectiveness of the proposed method, different commonly used attacks are simulated. Simulation results in terms of peak signal‐to‐noise ratio, structural similarity, normal correlation, and feature similarity indicate that the proposed method can recover the hidden secret message accurately against seven different types of attacks including speckle, Gaussian, salt and pepper, rotate, crop, fold, and blur attack. Farah Deeba, Kun She 0001, Fayaz Ali Dharejo, Yuanchun Zhou |
IET Image Process. | 3 |
| 2020 | Sparse representation based computed tomography images reconstruction by coupled dictionary learning algorithmabstractIt is very interesting to reconstruct high‐resolution computed tomography (CT) medical images that are very useful for clinicians to analyse the diseases. This study proposes an improved super‐resolution method for CT medical images in the sparse representation domain with dictionary learning. The sparse coupled K‐singular value decomposition (KSVD) algorithm is employed for dictionary learning purposes. Images are divided into two sets of low resolution (LR) and high resolution (HR), to improve the quality of low‐resolution images, the authors prepare dictionaries over LR and HR image patches using the KSVD algorithm. The main idea behind the proposed method is that sparse coupled dictionaries learn about each patch and establish the relationship between sparse coefficients of LR and HR image patches to recover the HR image patch for LR image. The proposed method is compared to conventional algorithms in terms of mean peak signal‐to‐noise ratio and structural similarity index measurements by using three different data set images, including CT chest, CT dental and CT brain images. The authors also analysed the proposed improved method for different dictionary sizes and patch size to obtain a similar high‐resolution image. These parameters play an essential role in the reconstruction of the HR images. Farah Deeba, Kun She 0001, Fayaz Ali Dharejo, Yuanchun Zhou |
IET Image Process. | 3 |
| 2020 | A Color Enhancement Scene Estimation Approach for Single Image Haze RemovalabstractThe presence of suspended particles, such as fog, smoke, and dust, in the atmosphere reduces the quality of the captured image. It is essential to overcome these particles, because these particles have a very dire effect on different applications of image processing. We propose a new image dehazing method for remote sensing (RS) applications. Since the hazy RS image is affected by multiple colors and contrast reduction, our goal is to focus on degraded objects, including color correction and color-contrast enhancement. A “Piecewise Linear Transformation (PWLT)” is used to correct the color distortion, and then the color contrast is improved by applying the proposed method. The intensity distribution manipulation is one of the most common methods used to strengthen image contrast in the past. Compared with advanced techniques, the proposed method is easy to implement and suitable for real-time applications. Moreover, it is not necessary to require prior imaging conditions. The experimental results show that, in terms of subjective and visual quality, the color, contrast, naturalness, and high brightness of the object increase in the image to be improved. Fayaz Ali Dharejo, Yuanchun Zhou, Farah Deeba, Yi Du 0010 |
IEEE Geosci. Remote. Sens. Lett. | 1 |