EDBT 2026 Demo / reviewers in the wild / expert
Guoping Qiu
dblp:43/835
· DBLP profile ↗
183ranked-venue papers
28as first author
47since 2021 · last 2026
0000-0002-5877-5648ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 127 · 20 first-author · 34 since 2021Artificial intelligence and machine learning · 58 · 10 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 1 since 2021Databases, data management, data science and information retrieval · 6 · 3 since 2021Security and privacy · 4Human-computer interaction and ubiquitous computing · 3Systems, architecture and hardware · 2Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CFFormer: Cross CNN-Transformer channel attention and spatial feature fusion for improved segmentation of heterogeneous medical images
Qing Xu 0014, Xiangjian He, Daokun Zhang, Ruili Wang 0001, Rong Qu, Guoping Qiu |
Expert Syst. Appl. | 8 |
| 2026 | Purifying medical images with generative adversarial networksabstractArtificial intelligence (AI) applications in medical imaging have the potential to revolutionize healthcare by improving diagnostic accuracy, reducing errors, and increasing efficiency. Deep neural networks (DNNs) have achieved a great success in diagnosing various diseases, often reaching or even surpassing the performance levels of human experts. However, DNNs remain vulnerable to adversarial attacks — carefully crafted inputs that can manipulate the neural networks into incorrect predictions. Current adversarial defense techniques, such as adversarial training, are limited in their effectiveness against specific attacks and necessitate model re-training, thereby hindering their widespread application. Adversarial purification is an emerging defense method that leverages generative models to refine training data and eliminate adversarial examples before classification. This paper explores various generative models and proposes AdvPurify-GAN, a novel Generative Adversarial Network (GAN) for purifying medical images against adversarial attacks. Our work stands out as the first to utilize the adversarial purification techniques for enhancing robustness in medical image analysis. We evaluate AdvPurify-GAN on multiple datasets and show favorable performance under standard attacks compared with existing adversarial training and purification baselines. Juexin Zhang 0001, Ying Weng, Boding Wang, Guoping Qiu |
Neurocomputing | 5 |
| 2026 | UDG-Prom: A unified dense-guided semantic prompting for cross-domain few-shot image segmentationabstract• MAF preserves low-level feature representations, while fusing global and local information to generate robust class-agnostic features. • TA2MP, as a unified feature transformation mechanism equipped with an automatic learnable prompt branch, reduces human reliance and disentangles domain- and class-specific information through contrastive learning. • UDG-Prom integrates the MAF and TA2MP modules to address the CD-FSS task with SAM. • Our model achieves competitive or superior performance compared to state-of-the-art methods on four CD-FSS benchmarks, and its strong generalization ability is comprehensively validated through evaluations on more difficult cross-domain datasets including CT-Lung (medical) and SUIM (underwater). Large Vision Models (LVMs), exemplified by SAM, contain powerful general knowledge from extensive pre-training, yet they often underperform in highly specialized domains. Building large models tailored for each domain is usually impractical due to the substantial cost of data collection and training. Therefore, a key challenge is how to tap into SAM’s strong knowledge base and transfer it effectively to new, domain-specific tasks, especially under Cross-Domain or Few-Shot constraints. Previous efforts have leveraged prior knowledge from foundation models for transfer learning; however, they typically target specific tasks and exhibit limited robustness in broader applications. To tackle this issue, we propose a Unified Dense-Guided Semantic Prompting framework (UDG-Prom), a new paradigm for Cross-Domain Few-Shot Segmentation (CD-FSS). First, a Multi-level Adaptation Framework (MAF) is used for integrated feature extraction as prior knowledge. Then, we incorporate a Task-Adaptive Auto Meta Prompt (TA 2 MP) module to enable the extraction of class-domain-agnostic features and generate high-quality, learnable visual prompts. By combining learnable prompts with a structured model and prototype disentanglement, this method retains SAM’s prior knowledge and effectively adapts to CD-FSS through category and domain cues. Extensive experiments on four benchmarks show that our model not only surpasses state-of-the-art CD-FSS approaches but also achieves a remarkable improvement in average accuracy. Xiangjian He, Xin Chen 0003, Jingxi Hu, LinLin Shen, Guoping Qiu |
Knowl. Based Syst. | 7 |
| 2026 | Ultra-High-Definition Video Quality Assessment Based on Adaptively Spatial Fusion and Temporal SelectionabstractUltra high definition video quality assessment (UHD VQA) is challenging due to the high spatio-temporal dimensions of UHD content. In this work, we propose a spatio-temporal framework based on feature fusion and selection to facilitate UHD VOA. First, we introduce a spatially adaptive fusion net (SAFNet) that integrates all patch-wise features into a compact frame-wise representation via local aggregation and global interaction. During the aggregation and the interaction, the contribution of each patch-wise feature is initially estimated via visual saliency-derived weights and then adaptively refined by contextual adjustments. Second, we design a temporally adaptive selection net (TASNet) to reduces temporal redundancy. It retains a small number of temporal positions through thresholding a dynamically estimated importance vector, which captures both representativeness and inter-frame variation. Finally, extensive experiments conducted on 4 datasets with various 4K videos demonstrate the effectiveness of the proposed method in comparison with some traditional and state-of-the-art VQA methods. Zhijie Liang, Wei Sheng, Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | Self-Iteration Image Haze Removal Using a Deep Curve-Dehazing ModelabstractThis paper proposes a novel dehazing method termed Haze-Restoration Curve Model (HRCM), which transforms the single-image dehazing task into a specific curve estimation problem, achieving haze removal through an intuitive and simple nonlinear curve mapping. Unlike methods based on Atmospheric Scattering Model (ASM), HRCM does not require the computation of complex physical parameters. Instead, it estimates two intuitive curvature adjustment coefficients. Moreover, compared to recent end-to-end dehazing methods, HRCM circumvents the challenging modeling of static mapping functions, thereby improving the generalization ability and dehazing performance of the model. All of these are attributed to a meticulously designed dehazing curve, which first reversing the hazy image to highlight obscured regions, and then specifies a set of high-order functions to remap hazy pixels for image restoration. Moreover, to estimate the curve parameters, we designed a dual-branch Deep Dehaze Curve Estimation Network(DDCEN), which consists of the Residual Swin Transformer Block(RTSB) and the Large kernel convolutional Attention Block(LAB). Specifically, RTSB captures the global fog density distribution features of foggy images by introducing window self-attention and shifted window mechanisms, providing support for global semantic information for subsequent parameter estimation. LAB captures local multi-scale features by constructing a large receptive field, and uses the attention mechanism of feature pooling in horizontal and vertical directions to focus on detail regions, refining the local details of the parameter map. Extensive experiments on synthetic and real-world hazy image datasets demonstrate that the proposed approach achieves superior performance in terms of quantitative accuracy and subjective visual quality compared to the current state-of-the-art methods. The source code of our HRCM is available at https://github.com/larrylanrui/HRCM. Wei Liu 0123, Rui Nan, Jiayi Ma 0001, Xin Chen 0003, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2026 | PSAvatar: A Point-Based Shape Model for Real-Time Head Avatar Animation With 3D Gaussian SplattingabstractDespite much progress, achieving real-time highfidelity head avatar animation is still difficult and existing methods have to trade-off between speed and quality. 3DMM based methods often fail to model non-facial structures such as eyeglasses and hairstyles, while neural implicit models suffer from deformation inflexibility and rendering inefficiency. Although 3D Gaussian has been demonstrated to possess promising capability for geometry representation and radiance field reconstruction, applying 3D Gaussian in head avatar creation remains a major challenge since it is difficult for 3D Gaussian to model the head shape variations caused by changing poses and expressions. In this paper, we introduce PSAvatar, a novel framework for animatable head avatar creation that utilizes discrete geometric primitive to create a parametric shape model and employs 3D Gaussian for fine detail representation and high fidelity rendering. The parametric shape model is a Point-based Shape Model (PSM) which uses points instead of meshes for 3D representation to achieve enhanced representation flexibility. Specifically, PSM first converts the FLAME mesh to points by sampling on the surfaces as well as off the meshes to enable the reconstruction of not only surface-like structures but also complex geometries such as eyeglasses and hairstyles. By aligning these points with the head shape in an analysis-by-synthesis manner, the PSM makes it possible to utilize 3D Gaussian for fine detail representation and appearance modeling, thus enabling the creation of high-fidelity avatars. We show that PSAvatar can reconstruct high-fidelity head avatars of varieties of subjects and the avatars can be animated in real-time. Zhenyu Bao, Qing Li 0029, Guoping Qiu, Kanglin Liu |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2025 | Towards Smart Point-and-Shoot PhotographyabstractHundreds of millions of people routinely take photos using their smartphones as point and shoot (PAS) cameras, yet very few would have the photography skills to compose a good shot of a scene. While traditional PAS cameras have built-in functions to ensure a photo is well focused and has the right brightness, they cannot tell the users how to compose the best shot of a scene. In this paper, we present a first of its kind smart point and shoot (SPAS) system to help users to take good photos. Our SPAS proposes to help users to compose a good shot of a scene by automatically guiding the users to adjust the camera pose live on the scene. We first constructed a large dataset containing 320K images with camera pose information from 4000 scenes. We then developed an innovative CLIP-based Composition Quality Assessment (CCQA) model to assign pseudo labels to these images. The CCQA introduces a unique learnable text embedding technique to learn continuous word embeddings capable of discerning subtle visual quality differences in the range covered by five levels of quality description words {bad, poor, fair, good, perfect}. And finally we have developed a camera pose adjustment model (CPAM) which first determines if the current view can be further improved and if so it outputs the adjust suggestion in the form of two camera pose adjustment angles. The two tasks of CPAM make decisions in a sequential manner and each involves different sets of training samples, we have developed a mixture-of-experts model with a gated loss function to train the CPAM in an end-to-end manner. We will present extensive results to demonstrate the performances of our SPAS system using publicly available image composition datasets. Jiawan Li, Fei Zhou 0001, Zhipeng Zhong, Jiongzhi Lin, Guoping Qiu |
CVPR | 5 |
| 2025 | SPC-GS: Gaussian Splatting with Semantic-Prompt Consistency for Indoor Open-World Free-view Synthesis from Sparse Inputsabstract3D Gaussian Splatting-based indoor open-world free-view synthesis approaches have shown significant performance with dense input images. However, they exhibit poor performance when confronted with sparse inputs, primarily due to the sparse distribution of Gaussian points and insufficient view supervision. To relieve these challenges, we propose SPC-GS, leveraging Scene-layout-based Gaussian Initialization (SGI) and Semantic-Prompt Consistency (SPC) Regularization for open-world free view synthesis with sparse inputs. Specifically, SGI provides a dense, scene-layout-based Gaussian distribution by utilizing view-changed images generated from the video generation model and view-constraint Gaussian points densification. Additionally, SPC mitigates limited view supervision by employing semantic-prompt-based consistency constraints developed by SAM2. This approach leverages available semantics from training views, serving as instructive prompts, to optimize visually overlapping regions in novel views with 2D and 3D consistency constraints. Extensive experiments demonstrate the superior performance of SPC-GS across Replica and ScanNet benchmarks. Notably, our SPC-GS achieves a 3.06 dB gain in PSNR for reconstruction quality and a 7.3% improvement in mIoU for open-world semantic segmentation. Project website at: https://gbliao.github.io/SPC-GS.github.io. Guibiao Liao, Qing Li 0029, Zhenyu Bao, Guoping Qiu, Kanglin Liu |
CVPR | 4 |
| 2025 | Low-Light Image Enhancement Through Learning a Simplified Inverse Rendering ModelabstractIt remains to be extremely difficult to capture high quality photographs of low-light scenes. Low light causes the low signal-to-noise ratio (SNR) problem which makes the image noisy. Such scenes almost always have the high dynamic range (HDR) problem caused by uneven lighting where a small area surrounding the light source is very bright while the rest of the scene is very dark, making it very difficult to simultaneously obtain high quality signals in both the dark and bright areas. This paper presents a new image restoration method for tackling the problems in low-light scenes. Fundamentally differing from existing approaches, the new method borrows ideas from inverse graphics rendering and re-renders the image with a canonical light source thus correcting the image from first principle. A deep learning based simplified inverse rendering model (SIRM) featuring implicit regularization is first developed for correcting uneven lighting and then an end-to-end convolutional neural network is constructed for reducing noise. Extensive experimental results are presented to demonstrate that the new method outperforms state-of-the-art methods, and is capable of effectively brightening up dark image regions while at the same time preserving details and color consistency. Our code is available at: https://github.com/pj0927/SIRNet. Wenhui Wu 0001, Jia Pang, Shuaibo Gao, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Geometric Distortion Guided Transformer for Omnidirectional Image Super-ResolutionabstractAs virtual and augmented reality applications gain popularity, omnidirectional image (ODI) super-resolution has become increasingly important. Unlike 2D plain images that are formed on a plane, ODIs are projected onto spherical surfaces. Applying established image super-resolution methods to ODIs, therefore, requires performing equirectangular projection (ERP) to map the ODIs onto a plane. ODI super-resolution needs to take into account geometric distortion resulting from ERP. However, without considering such geometric distortion of ERP images, previous methods only utilize a limited range of pixels and may easily miss self-similar textures for reconstruction. In this paper, we introduce a novel Geometric Distortion Guided Transformer for Omnidirectional image Super-Resolution (GDGT-OSR). Specifically, a distortion modulated rectangle-window selfattention mechanism, integrated with deformable self-attention, is proposed to better perceive the distortion and thus involve more self-similar textures. Distortion modulation is achieved through a newly devised distortion guidance generator that produces guidance for the rectangular windows by exploiting the variability of distortion across latitudes. Furthermore, we propose a dynamic feature aggregation scheme to adaptively fuse the features from different self-attention modules. We present extensive experimental results on public datasets and show that the new GDGT-OSR outperforms methods in existing literature. Cuixin Yang, Rongkang Dong, Jun Xiao 0010, Kin-Man Lam 0001, Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | LoopSparseGS: Loop-Based Sparse-View Friendly Gaussian SplattingabstractDespite the photorealistic novel view synthesis (NVS) performance achieved by the original 3D Gaussian splatting (3DGS), its rendering quality significantly degrades with sparse input views. This performance drop is mainly caused by the limited number of initial points generated from the sparse input, lacking reliable geometric supervision during the training process, and inadequate regularization of the oversized Gaussian ellipsoids. To handle these issues, we propose the LoopSparseGS, a loop-based 3DGS framework for the sparse novel view synthesis task. In specific, we propose a loop-based Progressive Gaussian Initialization (PGI) strategy that could iteratively densify the initialized point cloud using the rendered pseudo images during the training process. Then, the sparse and reliable depth from the Structure from Motion, and the window-based dense monocular depth are leveraged to provide precise geometric supervision via the proposed Depth-alignment Regularization (DAR). Additionally, we introduce a novel Sparse-friendly Sampling (SFS) strategy to handle oversized Gaussian ellipsoids leading to large pixel errors. Comprehensive experiments on four datasets demonstrate that LoopSparseGS outperforms existing state-of-the-art methods for sparse-input novel view synthesis, across indoor, outdoor, and object-level scenes with various image resolutions. Code is available at: https://github.com/pcl3dv/LoopSparseGS. Zhenyu Bao, Guibiao Liao, Kaichen Zhou, Kanglin Liu, Qing Li 0029, Guoping Qiu |
IEEE Trans. Image Process. | 6 |
| 2024 | 3D Reconstruction and Novel View Synthesis of Indoor Environments Based on a Dual Neural Radiance Field
Zhenyu Bao, Guibiao Liao, Kanglin Liu, Qing Li 0029, Guoping Qiu |
ACM Multimedia | 6 |
| 2024 | Vision-Language Knowledge Exploration for Video Saliency Prediction
Fei Zhou 0001, Baitao Huang, Guoping Qiu |
PRCV (9) | 3 |
| 2024 | Image Intrinsic Components Guided Conditional Diffusion Model for Low-Light Image EnhancementabstractThrough formulating the image restoration as a generation problem, the conditional diffusion model has been applied to low-light image enhancement (LIE) to restore the details in dark regions. However, in the previous diffusion model based LIE methods, the conditions used for guiding generation are degraded images, such as low-light image, signal-to-noise ratio map and color map, which suffer from severe degradation and are simply fed into diffusion model by rigidly concatenating with the noise. To avoid using degraded conditions resulting in sub-optimal performance in recovering details and enhancing brightness, we use the image intrinsic components originating from the Retinex model as guidance, whose multi-scale features are flexibly integrated into the diffusion model, and propose a novel conditional diffusion model for LIE. Specifically, the input low-light image is decomposed into reflectance and illumination by a Retinex decomposition module, where two components contain abundant physical property and lighting conditions of the scene. Then, we extract the latent features from two conditions through a component-dependent feature extraction module, which is designed according to the physical property of components. Finally, instead of previous rigid concatenation manner, a well-designed feature fusion mechanism is equipped to adaptively embed generative conditions into diffusion model. Extensive experimental results demonstrate that our method outperforms the state-of-the-art methods, and is capable of effectively restoring the local details while brightening the dark regions. Our codes are available athttps://github.com/Knossosc/ICCDiff. Sicong Kang, Shuaibo Gao, Wenhui Wu 0001, Xu Wang 0006, Shuoyao Wang, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Blind Image Quality Assessment Based on Separate Representations and Adaptive Interaction of Content and DistortionabstractThe visual quality of an image mainly relies on its content and its distortions. However, the adaptability between their contributions to the image quality has not be well investigated yet. Besides, albeit of many promising efforts, lacking sufficient labeled data still hinders the robust representation of quality-related information. In this work, we first design a self-supervised architecture, named collaborative autoencoder (COAE), to separately represent the content and the distortion information, and then develop a Self-Adaptive Weighting based quAlity predictoR (SAWAR) to balance the individual representations of the content and the distortions in the prediction of image quality. Specifically, the COAE is trained with large-scale unlabeled data, consisting of a content autoencoder (CAE) and a distortion autoencoder (DAE) that work collaboratively and individually. While the CAE is a standard autoencoder for the content representation, the design of the DAE is unique. We introduce the CAE-encoded content representation as an extra input to the decoder of the DAE to learn to reconstruct distorted images, thus effectively forcing it to extract the distortion representation. The SAWAR, whose parameter number is much smaller than that of the COAE, is trained with labeled data in existing IQA datasets. It takes advantage of the interaction between the image content and the distortions to adaptively balance their contributions. Extensive experiments show that the COAE effectively extracts quality-related representations and the SAWAR achieves the state-of-the-art performance. Zehong Zhou, Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2024 | A Dataset and Model for the Visual Quality Assessment of Inversely Tone-Mapped HDR VideosabstractTo enhance the viewer experience of standard dynamic range (SDR) video content on high dynamic range (HDR) displays, inverse tone mapping (ITM) is employed. Objective visual quality assessment (VQA) models are needed for effective evaluation of ITM algorithms. However, there is a lack of specialized VQA models for assessing the visual quality of inversely tone-mapped HDR videos (ITM-HDR-Videos). This paper addresses both an algorithmic and a dataset gap by introducing a novel SDR referenced HDR (SD-R-HD) VQA model tailored for ITM-HDR-Videos, along with the first public dataset specifically constructed for this purpose. The innovations of the SD-R-HD VQA model include 1) utilizing available SDR video as a reference signal, 2) extracting features that characterize standard ITM operations such as global mapping and local compensation, and 3) directly modeling interframe inconsistencies introduced by ITM operations. The newly created ITM-HDR-VQA dataset comprises 200 ITM-HDR-Videos annotated with mean opinion scores, gathered over 320 man-hours of psychovisual experiments. Experimental results demonstrate that the SD-R-HD VQA model significantly outperforms existing state-of-the-art VQA models. Fei Zhou 0001, Shuhong Yuan, Zhijie Liang, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 5 |
| 2024 | Generating Counterfactual Instances for Explainable Class-Imbalance LearningabstractExisting class imbalance learning paradigms focus on lifting the importance of minority instance, aiming to improve the model in terms of certain evaluation metrics (e.g., AUC and$F_{1}$-measure). One drawback of these methods is that they lack enough transparency, hence, cannot be fully trusted in vital domains. To this end, this paper deal with the class imbalance learning task with counterfactual instances. Given an instance and a classifier, a counterfactual is a fake instance which, while having smallest distance to the original instance, is classified as a different class by the classifier. Therefore, the most important features for a classifier can be identified by inspecting the difference between an instance and its counterfactual. To utilize counterfactuals, a novel Explainable Generative Adversarial Network (EXGAN) is proposed. EXGAN has a unique “two generatorsversusmultiple discriminators” architecture where the generators are used to generate effective counterfactuals and discriminators are trained for the class imbalance learning task. In addition to the architecture, an innovative ensemble loss function ensuring each discriminator complementing each other is designed to overcome the class imbalance issue. Extensive experiments prove that the counterfactuals generated by EXGAN can be used to produce effective local explanation and provide significant better class imbalance learning ability than existing competitors. Zhi Chen 0017, Jiang Duan, Rui Chen 0003, Guoping Qiu |
IEEE Trans. Knowl. Data Eng. | 6 |
| 2023 | Aesthetically Relevant Image CaptioningabstractImage aesthetic quality assessment (AQA) aims to assign numerical aesthetic ratings to images whilst image aesthetic captioning (IAC) aims to generate textual descriptions of the aesthetic aspects of images. In this paper, we study image AQA and IAC together and present a new IAC method termed Aesthetically Relevant Image Captioning (ARIC). Based on the observation that most textual comments of an image are about objects and their interactions rather than aspects of aesthetics, we first introduce the concept of Aesthetic Relevance Score (ARS) of a sentence and have developed a model to automatically label a sentence with its ARS. We then use the ARS to design the ARIC model which includes an ARS weighted IAC loss function and an ARS based diverse aesthetic caption selector (DACS). We present extensive experimental results to show the soundness of the ARS concept and the effectiveness of the ARIC model by demonstrating that texts with higher ARS’s can predict the aesthetic ratings more accurately and that the new ARIC model can generate more accurate, aesthetically more relevant and more diverse image captions. Furthermore, a large new research database containing 510K images with over 5 million comments and 350K aesthetic scores, and code for implementing ARIC, are available at https://github.com/PengZai/ARIC Zhipeng Zhong, Fei Zhou 0001, Guoping Qiu |
AAAI | 3 |
| 2023 | Improving Robustness of Single Image Super-Resolution Models with Monte Carlo MethodabstractDeep learning-based methods have achieved promising results in single image super-resolution (SISR). However, the performance of existing deep SISR methods is very sensitive to image degradation. In addition, these methods are deterministic and do not introduce any uncertainty to the generated images, so we have no way of knowing the reliability of these generated images. To address these two challenging issues, we propose a model-agnostic approach for existing deep SISR networks to improve their robustness under various degradations. Our proposed method follows a probabilistic framework and applies Monte Carlo dropout to existing deep SISR methods. Instead of performing point estimation, the proposed method predicts the posterior distribution of super-resolved images. Based on this, we can determine the uncertainty of the generated images. Experiment results show that the proposed method can effectively improve the robustness of existing deep SISR methods, leading to state-of-the-art performance when applied to images having different degradations. The code is available at https://github.com/YangTracy/MCD-SR. Cuixin Yang, Jun Xiao 0010, Yakun Ju, Guoping Qiu, Kin-Man Lam 0001 |
ICIP | 4 |
| 2023 | Collaborative Auto-encoding for Blind Image Quality AssessmentabstractBlind image quality assessment (BIQA) is a challenging problem with important real-world applications. Recent efforts attempting to exploit powerful representations by deep neural networks (DNN) are hindered by the lack of subjectively annotated data. This paper presents a novel BIQA method which overcomes this fundamental obstacle. Specifically, we design a pair of collaborative autoencoders (COAE) consisting of a content autoencoder (CAE) and a distortion autoencoder (DAE) that work together to extract content and distortion representations, which are shown to be highly descriptive of image quality. While the CAE follows a standard codec procedure, we introduce the CAE-encoded feature as an extra input to the DAE's decoder for reconstructing distorted images, thus effectively forcing DAE's encoder to extract distortion representations. The self-supervised learning framework allows the COAE including two feature extractors to be trained by almost unlimited amount of data, thus leaving limited samples with annotations to finetune a BIQA model. We will show that the proposed BIQA method achieves state-of-the-art performance and has superior generalization capability over other learning based models. All the related codes, trained models, and the supplementary material are available at: https://github.com/Macro-Zhou/NRIQA-VISOR/. Zehong Zhou, Fei Zhou 0001, Guoping Qiu |
ICME | 3 |
| 2023 | Video Inverse Tone Mapping Network with Luma and Chroma Mapping
Peihuan Huang, Gaofeng Cao, Fei Zhou 0001, Guoping Qiu |
ACM Multimedia | 4 |
| 2023 | P2I-NET: Mapping Camera Pose to Image via Adversarial Learning for New View Synthesis in Real Indoor EnvironmentsabstractGiven a new 6DoF camera pose in an indoor environment, we study the challenging problem of predicting the view from that pose based on a set of reference RGBD views. Existing explicit or implicit 3D geometry construction methods are computationally expensive while those based on learning have predominantly focused on isolated views of object categories with regular geometric structure. Differing from the traditional render-inpaint approach to new view synthesis in the real indoor environment, we propose a conditional generative adversarial neural network (P2I-NET) to directly predict the new view from the given pose. P2I-NET learns the conditional distribution of the images of the environment for establishing the correspondence between the camera pose and its view of the environment, and achieves this through a number of innovative designs in its architecture and training lost function. Two auxiliary discriminator constraints are introduced for enforcing the consistency between the pose of the generated image and that of the corresponding real world image in both the latent feature space and the real world pose space. Additionally a deep convolutional neural network (CNN) is introduced to further reinforce this consistency in the pixel space. We have performed extensive new view synthesis experiments on real indoor datasets. Results show that P2I-NET has superior performance against a number of NeRF based strong baseline models. In particular, we show that P2I-NET is 40 to 100 times faster than these competitor techniques while synthesising similar quality images. Furthermore, we contribute a new publicly available indoor environment dataset containing 22 high resolution RGBD videos where each frame also has accurate camera pose parameters. Xujie Kang, Kanglin Liu, Jiang Duan, Yuanhao Gong, Guoping Qiu |
ACM Multimedia | 5 |
| 2023 | Restoration of Multiple Image Distortions using a Semi-dynamic Deep Neural NetworkabstractRestoring multiple image distortions with a single model is difficult because different distortions require fundamentally different processing mechanisms, e.g., deblurring requires high-pass filtering, while denoising requires low-pass filtering operations. This paper presents a dynamic universal image restoration (DUIR) system capable of simultaneously processing multiple distortions. The new model features several innovative designs: (i) a distortion embedding module (DEM) to automatically encode the distortion information of an input, (ii) a distortion attention module (DAM) that uses a bi-directional long short-term memory (LSTM) to encode the distortion into a sequence of forward and backward interdependent modulating signals, and (iii) a dynamically adaptive image restoration deep convolutional neural network (DAIR-DCNN) featuring unique semi-dynamic layers (SDLs) in which part of their parameters are dynamically modulated by the distortion signals. DEM, DAM, and SDLs together make DAIR-DCNN adaptive to the distortions of the current input, which in turn equips the DUIR system with the capability of simultaneously processing multiple image distortions with a single trained model. We present extensive experimental results to show that the new technique achieves superior performance to state-of-the-art models on both synthetic and real data. We further demonstrate that a trained DUIR system can simultaneously handle different distortions, including those with conflicting demands, such as denoising, deblurring, and compression artifact removal. Hongming Luo, Fei Zhou 0001, Zehong Zhou, Kin-Man Lam 0001, Guoping Qiu |
ACM Multimedia | 5 |
| 2023 | Feature preserving 3D mesh denoising with a Dense Local Graph Neural Network
Wenming Tang, Yuanhao Gong, Guoping Qiu |
Comput. Vis. Image Underst. | 3 |
| 2023 | Supervised Anomaly Detection via Conditional Generative Adversarial Network and Ensemble Active LearningabstractAnomaly detection has wide applications in machine intelligence but is still a difficult unsolved problem. Major challenges include the rarity of labeled anomalies and it is a class highly imbalanced problem. Traditional unsupervised anomaly detectors are suboptimal while supervised models can easily make biased predictions towards normal data. In this paper, we present a new supervised anomaly detector through introducing the novel Ensemble Active Learning Generative Adversarial Network (EAL-GAN). EAL-GAN is a conditional GAN having a unique one generator versus multiple discriminators architecture where anomaly detection is implemented by an auxiliary classifier of the discriminator. In addition to using the conditional GAN to generate class balanced supplementary training data, an innovative ensemble learning loss function ensuring each discriminator makes up for the deficiencies of the others is designed to overcome the class imbalanced problem, and an active learning algorithm is introduced to significantly reduce the cost of labeling real-world data. We present extensive experimental results to demonstrate that the new anomaly detector consistently outperforms a variety of SOTA methods by significant margins. Zhi Chen 0017, Jiang Duan, Guoping Qiu |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Deep generative image priors for semantic face manipulationabstractPrevious works on generative adversarial networks (GANs) mainly focus on how to synthesize high-fidelity images. In this paper, we present a framework to leverage the knowledge learned by GANs for semantic face manipulation. In particular, we propose to control the semantics of synthesized faces by adapting the latent codes with an attribute prediction model. Moreover, in order to achieve a more accurate estimation of different facial attributes, we propose to pretrain the attribute prediction model by inverting the synthesized face images back to the GAN latent space. As a result, our method explicitly considers the semantics encoded in the latent space of a pretrained GAN and is able to faithfully edit various attributes like eyeglasses, smiling, bald, age, mustache and gender for high-resolution face images. Extensive experiments show that our method has superior performance compared to state of the art for both face attribute prediction and semantic face manipulation. Xianxu Hou, LinLin Shen, Zhong Ming 0001, Guoping Qiu |
Pattern Recognit. | 4 |
| 2023 | Super-resolving compressed images via parallel and series integration of artefacts removal and resolution enhancement
Hongming Luo, Fei Zhou 0001, Guangsen Liao, Guoping Qiu |
Signal Process. | 4 |
| 2023 | Super-resolution image visual quality assessment based on structure-texture features
Fei Zhou 0001, Wei Sheng, Zitao Lu, Bo Kang, Mianyi Chen, Guoping Qiu |
Signal Process. Image Commun. | 6 |
| 2023 | Learning Deep Co-Occurrence FeaturesabstractWe exploit the computational capability of deep convolutional neural network (CNN) architecture and the natural interpretability of the co-occurrence matrix (CM) to learn deep co-occurrence features (DCOFs). The DCOFs represent the statistics of the co-occurrences of pixels thus overcoming the black box nature of traditional deep representation learning while at the same time solving the inherent computational difficulty of CM. We propose a parametric co-occurrence matrix (PCM) model to approximate the CM with multivariate Gaussian functions, and have developed three approaches to decomposing the PCM model into linear and nonlinear operations such that the model can be easily implemented using standard CNN operations and to learn the DCOFs of arbitrary shapes. The CNN implementation of the PCM model, termed PCMCNN, can be used as a standard plugin module of a deep learning system and adaptively learns the DCOFs for downstream applications. We demonstrate the broad applicability of the DCOFs and their effectiveness in fine-grained image classification tasks such as texture classification and GAN (generative adversarial network) image detection. The introduction of the PCMCNN module makes it much more compact and efficient than conventional implementations of deep learning models, achieving comparable classification performances to state of the art methods on a variety of benchmarking datasets with models that are more than 30 folds smaller and 11 times less complex. The small model size for learning the DCOFs makes the new method particularly effective for few shot classification of large number of texture categories where the small number of training samples can easily cause traditional deep learning models to overfit. This work shows the potential benefits of combining the principles of traditional handcrafted features and deep representation learning to take advantage of both for advancing state of the art. Guanglin Li 0007, Bin Li 0011, Shunquan Tan, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Learning General Gaussian Mixture Model with Integral Cosine SimilarityabstractGaussian mixture model (GMM) is a powerful statistical tool in data modeling, especially for unsupervised learning tasks. Traditional learning methods for GMM such as expectation maximization (EM) require the covariance of the Gaussian components to be non-singular, a condition that is often not satisfied in real-world applications. This paper presents a new learning method called G$^2$M$^2$ (General Gaussian Mixture Model) by fitting an unnormalized Gaussian mixture function (UGMF) to a data distribution. At the core of G$^2$M$^2$ is the introduction of an integral cosine similarity (ICS) function for comparing the UGMF and the unknown data density distribution without having to explicitly estimate it. By maximizing the ICS through Monte Carlo sampling, the UGMF can be made to overlap with the unknown data density distribution such that the two only differ by a constant scalar, and the UGMF can be normalized to obtain the data density distribution. A Siamese convolutional neural network is also designed for optimizing the ICS function. Experimental results show that our method is more competitive in modeling data having correlations that may lead to singular covariance matrices in GMM, and it outperforms state-of-the-art methods in unsupervised anomaly detection. Guanglin Li 0001, Bin Li 0011, Changsheng Chen 0001, Shunquan Tan, Guoping Qiu |
IJCAI | 5 |
| 2022 | Restoration of User Videos Shared on Social MediaabstractUser videos shared on social media platforms usually suffer from degradations caused by unknown proprietary processing procedures, which means that their visual quality is poorer than that of the originals. This paper presents a new general video restoration framework for the restoration of user videos shared on social media platforms. In contrast to most deep learning-based video restoration methods that perform end-to-end mapping, where feature extraction is mostly treated as a black box, in the sense that what role a feature plays is often unknown, our new method, termed Video restOration through adapTive dEgradation Sensing (VOTES), introduces the concept of a degradation feature map (DFM) to explicitly guide the video restoration process. Specifically, for each video frame, we first adaptively estimate its DFM to extract features representing the difficulty of restoring its different regions. We then feed the DFM to a convolutional neural network (CNN) to compute hierarchical degradation features to modulate an end-to-end video restoration backbone network, such that more attention is paid explicitly to potentially more difficult to restore areas, which in turn leads to enhanced restoration performance. We will explain the design rationale of the VOTES framework and present extensive experimental results to show that the new VOTES method outperforms various state-of-the-art techniques both quantitatively and qualitatively. In addition, we contribute a large scale real-world database of user videos shared on different social media platforms. Codes and datasets are available at https://github.com/luohongming/VOTES.git Hongming Luo, Fei Zhou 0001, Kin-Man Lam 0001, Guoping Qiu |
ACM Multimedia | 4 |
| 2022 | A Novel Structure Adaptive Algorithm for Feature-preserving 3D Mesh DenoisingabstractIn this paper, we propose a novel algorithm for 3D mesh filtering (Structural Adaptive Filtering, SAF) based on mesh structural adaptation. As we all know, 3D meshes mainly have three types of geometric features: corners, edges, and planes. Therefore, we designed a protection mechanism for these three types of features to achieve the feature-preserving denoising. In the first step, for the faces normals to be processed, we build a variable set of similarity between the face normal and the neighborhood faces normal, calculate their coefficient of variation, variance, and quartile difference, and then select the neighborhood face normals with high similarity to update the current normals through self adaptation of these variables. In the second step, all vertices complete the iterative update of vertex coordinates according to the filtered face normals. Unlike existing 3D mesh denoising algorithms, which have too many parameters to manually set thresholds and are sensitive to parameters, SAF is based on the geometry of local faces (no need to manually set denoising thresholds). SAF only needs to set the iterative parameters to complete high-performance feature-preserving filtering, which has high practical value. We demonstrate through extensive experimental data that SAF outperforms or is comparable to state-of-the-art methods in feature-preserving denoising at different noise levels. Wenming Tang, Yuanhao Gong, Guoping Qiu |
MMSP | 3 |
| 2022 | Finding Beautiful and Happy Images for Mental Health and Well-Being Applications
Ruitao Xie, Connor S. Qiu, Guoping Qiu |
PRCV (3) | 3 |
| 2022 | Tone mapping high dynamic range images based on region-adaptive self-supervised deep learning
Fei Zhou 0001, Guangsen Liao, Jiang Duan, Guoping Qiu |
Signal Process. Image Commun. | 5 |
| 2022 | Deep Cross-Modal Representation Learning and Distillation for Illumination-Invariant Pedestrian DetectionabstractIntegrating multispectral data has been demonstrated to be an effective solution for illumination-invariant pedestrian detection, in particular, RGB and thermal images can provide complementary information to handle light variations. However, most of the current multispectral detectors fuse the multimodal features by simple concatenation, without discovering their latent relationships. In this paper, we propose a cross-modal feature learning (CFL) module, based on a split-and-aggregation strategy, to explicitly explore both the shared and modality-specific representations between paired RGB and thermal images. We insert the proposed CFL module into multiple layers of a two-branch-based pedestrian detection network, to learn the cross-modal representations in diverse semantic levels. By introducing a segmentation-based auxiliary task, the multimodal network is trained end-to-end by jointly optimizing a multi-task loss. On the other hand, to alleviate the reliance of existing multispectral pedestrian detectors on thermal images, we propose a knowledge distillation framework to train a student detector, which only receives RGB images as input and distills the cross-modal representations guided by a well-trained multimodal teacher detector. In order to facilitate the cross-modal knowledge distillation, we design different distillation loss functions for the feature, detection and segmentation levels. Experimental results on the public KAIST multispectral pedestrian benchmark validate that the proposed cross-modal representation learning and distillation method achieves robust performance. Tianshan Liu, Kin-Man Lam 0001, Rui Zhao 0012, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Towards Disentangling Latent Space for Unsupervised Semantic Face EditingabstractFacial attributes in StyleGAN generated images are entangled in the latent space which makes it very difficult to independently control a specific attribute without affecting the others. Supervised attribute editing requires annotated training data which is difficult to obtain and limits the editable attributes to those with labels. Therefore, unsupervised attribute editing in an disentangled latent space is key to performing neat and versatile semantic face editing. In this paper, we present a new technique termed Structure-Texture Independent Architecture with Weight Decomposition and Orthogonal Regularization (STIA-WO) to disentangle the latent space for unsupervised semantic face editing. By applying STIA-WO to GAN, we have developed a StyleGAN termed STGAN-WO which performs weight decomposition through utilizing the style vector to construct a fully controllable weight matrix to regulate image synthesis, and employs orthogonal regularization to ensure each entry of the style vector only controls one independent feature matrix. To further disentangle the facial attributes, STGAN-WO introduces a structure-texture independent architecture which utilizes two independently and identically distributed (i.i.d.) latent vectors to control the synthesis of the texture and structure components in a disentangled way. Unsupervised semantic editing is achieved by moving the latent code in the coarse layers along its orthogonal directions to change texture related attributes or changing the latent code in the fine layers to manipulate structure related ones. We present experimental results which show that our new STGAN-WO can achieve better attribute editing than state of the art methods. Kanglin Liu, Gaofeng Cao, Fei Zhou 0001, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 6 |
| 2022 | Class-Imbalanced Deep Learning via a Class-Balanced EnsembleabstractClass imbalance is a prevalent phenomenon in various real-world applications and it presents significant challenges to model learning, including deep learning. In this work, we embed ensemble learning into the deep convolutional neural networks (CNNs) to tackle the class-imbalanced learning problem. An ensemble of auxiliary classifiers branching out from various hidden layers of a CNN is trained together with the CNN in an end-to-end manner. To that end, we designed a new loss function that can rectify the bias toward the majority classes by forcing the CNN's hidden layers and its associated auxiliary classifiers to focus on the samples that have been misclassified by previous layers, thus enabling subsequent layers to develop diverse behavior and fix the errors of previous layers in a batch-wise manner. A unique feature of the new method is that the ensemble of auxiliary classifiers can work together with the main CNN to form a more powerful combined classifier, or can be removed after finished training the CNN and thus only acting the role of assisting class imbalance learning of the CNN to enhance the neural network's capability in dealing with class-imbalanced data. Comprehensive experiments are conducted on four benchmark data sets of increasing complexity (CIFAR-10, CIFAR-100, iNaturalist, and CelebA) and the results demonstrate significant performance improvements over the state-of-the-art deep imbalance learning methods. Zhi Chen 0017, Jiang Duan, Guoping Qiu |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2021 | Structure Adaptive Filtering for Edge-Preserving Image Smoothing
Wenming Tang, Yuanhao Gong, Linyu Su, Wenhui Wu 0001, Guoping Qiu |
ICIG (3) | 5 |
| 2021 | A Discrete Scheme for Computing Image's Weighted Gaussian CurvatureabstractWeighted Gaussian curvature is an important smoothness measurement for images. However, its conventional computation scheme has low performance, low accuracy and requires that the input image must be second order differentiable. To tackle these three issues, we propose a novel discrete computation scheme for the weighted Gaussian curvature. Our scheme does not require the second order differentiability. Moreover, our scheme is more accurate, has smaller support region and computationally more efficient than the conventional schemes. Therefore, our scheme holds promise for a large range of applications where the weighted Gaussian curvature is needed, for example, image smoothing, cartoon texture decomposition, optical flow estimation, etc. Yuanhao Gong, Wenming Tang, Lebin Zhou, Lantao Yu, Guoping Qiu |
ICIP | 5 |
| 2021 | Quarter Laplacian Filter For Edge Aware Image ProcessingabstractThis paper presents a quarter Laplacian filter that can preserve corners and edges during image smoothing. Its support region is $2\times 2$, which is smaller than the $3\times 3$ support region of the classical Laplacian filter. Thus, it is more local. Moreover, this filter can be implemented via the classical box filter, leading to high performance for real time applications. Finally, we show its edge preserving property in several image processing tasks, including image smoothing, texture enhancement, and low-light image enhancement. The proposed filter can be adopted in a wide range of image processing applications. Yuanhao Gong, Wenming Tang, Lebin Zhou, Lantao Yu, Guoping Qiu |
ICIP | 5 |
| 2021 | Self-Supervised Video Super-Resolution by Spatial Constraint and Temporal Fusion
Cuixin Yang, Hongming Luo, Guangsen Liao, Zitao Lu, Fei Zhou 0001, Guoping Qiu |
PRCV (3) | 6 |
| 2021 | Geographical and temporal huff model calibration using taxi trajectory data
Shuhui Gong, John Cartlidge, Ruibin Bai, Yang Yue 0001, Qingquan Li 0001, Guoping Qiu |
GeoInformatica | 6 |
| 2021 | Relative geometry-aware siamese neural network for 6DOF camera relocalization
Qing Li 0029, Jiasong Zhu, Rui Cao 0001, Ke Sun 0006, Jonathan M. Garibaldi, Qingquan Li 0001, Guoping Qiu |
Neurocomputing | 8 |
| 2021 | A hybrid data-level ensemble to enable learning from highly imbalanced dataset
Zhi Chen 0017, Jiang Duan, Guoping Qiu |
Inf. Sci. | 4 |
| 2021 | Dense graph convolutional neural networks on 3D meshes for 3D object segmentation and classification
Wenming Tang, Guoping Qiu |
Image Vis. Comput. | 2 |
| 2021 | Image Defogging Quality Assessment: Real-World Database and MethodabstractFog removal from an image is an active research topic in computer vision. However, current literature is weak in the following two areas which in many ways are hindering progress for developing defogging algorithms. First, there is no true real-world and naturally occurring foggy image datasets suitable for developing defogging models. Second, there is no suitable mathematically simple and easy to use image quality assessment (IQA) methods for evaluating the visual quality of defogged images. We address these two aspects in this paper. We first introduce a new foggy image dataset called multiple real-world foggy image dataset (MRFID). MRFID contains foggy and clear images of 200 outdoor scenes. For each scene, one clear image and 4 foggy images of different densities defined as slightly foggy, moderately foggy, highly foggy, and extremely foggy, are manually selected from images taken from these scenes over the course of one calendar year. We then process the foggy images of MRFID using 16 defogging methods to obtain 12,800 defogged images (DFIs) and perform a comprehensive subjective evaluation of the visual quality of the DFIs. Through collecting the mean opinion score (MOS) of 120 subjects and evaluating a variety of fog-relevant image features, we have developed a new Fog-relevant Feature based SIMilarity index (FRFSIM) for assessing the visual quality of DFIs. We present extensive experimental results to show that our new visual quality assessment measure, the FRFSIM, is more consistent with the MOS than other IQA methods and is therefore more suitable for evaluating defogged images than other state-of-the-art IQA methods. Our dataset and relevant code are available at http://www.vistalab.ac.cn/MRFID-for-defogging/. Wei Liu 0123, Fei Zhou 0001, Tao Lu 0001, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 5 |
| 2021 | End-to-End Fovea Localisation in Colour Fundus Images With a Hierarchical Deep Regression NetworkabstractAccurately locating the fovea is a prerequisite for developing computer aided diagnosis (CAD) of retinal diseases. In colour fundus images of the retina, the fovea is a fuzzy region lacking prominent visual features and this makes it difficult to directly locate the fovea. While traditional methods rely on explicitly extracting image features from the surrounding structures such as the optic disc and various vessels to infer the position of the fovea, deep learning based regression technique can implicitly model the relation between the fovea and other nearby anatomical structures to determine the location of the fovea in an end-to-end fashion. Although promising, using deep learning for fovea localisation also has many unsolved challenges. In this paper, we present a new end-to-end fovea localisation method based on a hierarchical coarse-to-fine deep regression neural network. The innovative features of the new method include a multi-scale feature fusion technique and a self-attention technique to exploit location, semantic, and contextual information in an integrated framework, a multi-field-of-view (multi-FOV) feature fusion technique for context-aware feature learning and a Gaussian-shift-cropping method for augmenting effective training data. We present extensive experimental results on two public databases and show that our new method achieved state-of-the-art performances. We also present a comprehensive ablation study and analysis to demonstrate the technical soundness and effectiveness of the overall framework and its various constituent components. Ruitao Xie, Jingxin Liu 0005, Rui Cao 0001, Connor S. Qiu, Jiang Duan, Jonathan M. Garibaldi, Guoping Qiu |
IEEE Trans. Medical Imaging | 7 |
| 2020 | End-to-End Illuminant Estimation Based on Deep Metric LearningabstractPrevious deep learning approaches to color constancy usually directly estimate illuminant value from input image. Such approaches might suffer heavily from being sensitive to the variation of image content. To overcome this problem, we introduce a deep metric learning approach named Illuminant-Guided Triplet Network (IGTN) to color constancy. IGTN generates an Illuminant Consistent and Discriminative Feature (ICDF) for achieving robust and accurate illuminant color estimation. ICDF is composed of semantic and color features based on a learnable color histogram scheme. In the ICDF space, regardless of the similarities of their contents, images taken under the same or similar illuminants are placed close to each other and at the same time images taken under different illuminants are placed far apart. We also adopt an end-to-end training strategy to simultaneously group image features and estimate illuminant value, and thus our approach does not have to classify illuminant in a separate module. We evaluate our method on two public datasets and demonstrate our method outperforms state-of-the-art approaches. Furthermore, we demonstrate that our method is less sensitive to image appearances, and can achieve more robust and consistent results than other methods on a High Dynamic Range dataset. Bolei Xu, Jingxin Liu 0005, Xianxu Hou, Guoping Qiu |
CVPR | 5 |
| 2020 | VHS to HDTV Video Translation Using Multi-task Adversarial Learning
Hongming Luo, Guangsen Liao, Xianxu Hou, Fei Zhou 0001, Guoping Qiu |
MMM (1) | 6 |
| 2020 | Marker controlled superpixel nuclei segmentation and automatic counting on immunohistochemistry staining imagesabstractMOTIVATION: For the diagnosis of cancer, manually counting nuclei on massive histopathological images is tedious and the counting results might vary due to the subjective nature of the operation. RESULTS: This paper presents a new segmentation and counting method for nuclei, which can automatically provide nucleus counting results. This method segments nuclei with detected nuclei seed markers through a modified simple one-pass superpixel segmentation method. Rather than using a single pixel as a seed, we created a superseed for each nucleus to involve more information for improved segmentation results. Nucleus pixels are extracted by a newly proposed fusing method to reduce stain variations and preserve nucleus contour information. By evaluating segmentation results, the proposed method was compared to five existing methods on a dataset with 52 immunohistochemically (IHC) stained images. Our proposed method produced the highest mean F1-score of 0.668. By evaluating the counting results, another dataset with more than 30 000 IHC stained nuclei in 88 images were prepared. The correlation between automatically generated nucleus counting results and manual nucleus counting results was up to R2 = 0.901 (P < 0.001). By evaluating segmentation results of proposed method-based tool, we tested on a 2018 Data Science Bowl (DSB) competition dataset, three users obtained DSB score of 0.331 ± 0.006. AVAILABILITY AND IMPLEMENTATION: The proposed method has been implemented as a plugin tool in ImageJ and the source code can be freely downloaded. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Jie Shu, Jingxin Liu 0005, Yongmei Zhang, Hao Fu 0001, Mohammad Ilyas, Giuseppe Faraci, Vincenzo Della Mea, Guoping Qiu |
Bioinform. | 9 |
| 2020 | HLO: Half-kernel Laplacian operator for surface smoothing
Wei Pan 0010, Xuequan Lu, Yuanhao Gong, Wenming Tang, Ying He 0001, Guoping Qiu |
Comput. Aided Des. | 7 |
| 2020 | Extracting activity patterns from taxi trajectory data: a two-layer framework using spatio-temporal clustering, Bayesian probability and Monte Carlo simulationabstractGlobal positioning system (GPS) data generated from taxi trips is a valuable source of information that offers an insight into travel behaviours of urban populations with high spatio-temporal resolution. However, in its raw form, GPS taxi data does not offer information on the purpose (or intended activity) of travel. In this context, to enhance the utility of taxi GPS data sets, we propose a two-layer framework to identify the related activities of each taxi trip automatically and estimate the return trips and successive activities after the trip, by using geographic point-of-interest (POI) data and a combination of spatio-temporal clustering, Bayesian inference and Monte Carlo simulation. Two million taxi trips in New York, the United States of America, and ten million taxi trips in Shenzhen, China, are used as inputs for the two-layer framework. To validate each layer of the framework, we collect 6,003 trip diaries in New York and 712 questionnaire surveys in Shenzhen. The results show that the first layer of the framework performs better than comparable methods published in the literature, while the second layer has high accuracy when inferring return trips. Shuhui Gong, John Cartlidge, Ruibin Bai, Yang Yue 0001, Qingquan Li 0001, Guoping Qiu |
Int. J. Geogr. Inf. Sci. | 6 |
| 2020 | Class-aware domain adaptation for improving adversarial robustness
Xianxu Hou, Jingxin Liu 0005, Bolei Xu, Guoping Qiu |
Image Vis. Comput. | 6 |
| 2020 | A contextual conditional random field network for monocular depth estimation
Qing Li 0029, Rui Cao 0001, Wenming Tang, Guoping Qiu |
Image Vis. Comput. | 5 |
| 2020 | Spectral regularization for combating mode collapse in GANs
Kanglin Liu, Guoping Qiu, Wenming Tang, Fei Zhou 0001 |
Image Vis. Comput. | 2 |
| 2020 | SMLBoost-adopting a soft-margin like strategy in boosting
Zhi Chen 0017, Jiang Duan, Guoping Qiu |
Knowl. Based Syst. | 5 |
| 2020 | Lipschitz constrained GANs via boundedness and continuityabstractAbstract One of the challenges in the study of generative adversarial networks (GANs) is the difficulty of its performance control. Lipschitz constraint is essential in guaranteeing training stability for GANs. Although heuristic methods such as weight clipping, gradient penalty and spectral normalization have been proposed to enforce Lipschitz constraint, it is still difficult to achieve a solution that is both practically effective and theoretically provably satisfying a Lipschitz constraint. In this paper, we introduce the boundedness and continuity (BC) conditions to enforce the Lipschitz constraint on the discriminator functions of GANs. We prove theoretically that GANs with discriminators meeting the BC conditions satisfy the Lipschitz constraint. We present a practically very effective implementation of a GAN based on a convolutional neural network (CNN) by forcing the CNN to satisfy the BC conditions (BC–GAN). We show that as compared to recent techniques including gradient penalty and spectral normalization, BC–GANs have not only better performances but also lower computational complexity. Kanglin Liu, Guoping Qiu |
Neural Comput. Appl. | 2 |
| 2020 | Fast and efficient implementation of image filtering using a side window convolutional neural network
Yuanhao Gong, Guoping Qiu |
Signal Process. | 3 |
| 2020 | End-to-End Single Image Fog Removal Using Enhanced Cycle Consistent Adversarial NetworksabstractSingle image defogging is a classical and challenging problem in computer vision. Existing methods towards this problem mainly include handcrafted priors based methods that rely on the use of the atmospheric degradation model and learning-based approaches that require paired fog-fogfree training example images. In practice, however, prior-based methods are prone to failure due to their own limitations and paired training data are extremely difficult to acquire. Moreover, there are few studies on the unpaired trainable defogging network in this field. Thus, inspired by the principle of CycleGAN network, we have developed an end-to-end learning system that uses unpaired fog and fogfree training images, adversarial discriminators and cycle consistency losses to automatically construct a fog removal system. Similar to CycleGAN, our system has two transformation paths; one maps fog images to a fogfree image domain and the other maps fogfree images to a fog image domain. Instead of one stage mapping, our system uses a two stage mapping strategy in each transformation path to enhance the effectiveness of fog removal. Furthermore, we make explicit use of prior knowledge in the networks by embedding the atmospheric degradation principle and a sky prior for mapping fogfree images to the fog images domain. In addition, we also contribute the first real world nature fog-fogfree image dataset for defogging research. Our multiple real fog images dataset (MRFID) contains images of 200 natural outdoor scenes. For each scene, there is one clear image and corresponding four foggy images of different fog densities manually selected from a sequence of images taken by a fixed camera over the course of one year. Qualitative and quantitative comparison against several state-of-the-art methods on both synthetic and real world images demonstrate that our approach is effective and performs favorably for recovering a clear image from a foggy image. Wei Liu 0123, Xianxu Hou, Jiang Duan, Guoping Qiu |
IEEE Trans. Image Process. | 4 |
| 2020 | Structure and Texture-Aware Image Decomposition via Training a Neural NetworkabstractStructure-texture image decomposition is a funda-mental but challenging topic in computational graphics and image processing. In this paper, we introduce a structure-aware and a texture-aware measures to facilitate the structure-texture de-composition (STD) of images. Edge strengths and spatial scales that have been widely-used in previous STD researches cannot describe the structures and textures of images well. The proposed two measures differentiate image textures from image structures based on their distinctive characteristics. Specifically, the first one aims to measure the anisotropy of local gradients, and the second one is designed to measure the repeatability degree of signal pat-terns in a neighboring region. Since these two measures describe different properties of image structures and textures, they are complementary to each other. The STD is achieved by optimizing an objective function based on the two new measures. As using traditional optimization methods to solve the optimization prob-lem will require designing different optimizers for different func-tional spaces, we employ an architecture of deep neural network to optimize the STD cost function in a unified manner. The ex-perimental results demonstrate that, as compared with some state-of-the-art methods, our method can better separate image structure and texture and result in shaper edges in the structural component. Furthermore, to demonstrate the usefulness of the proposed STD method, we have successfully applied it to several applications including detail enhancement, edge detection, and visual quality assessment of super-resolved images. Fei Zhou 0001, Guoping Qiu |
IEEE Trans. Image Process. | 4 |
| 2020 | Visual Saliency via Embedding Hierarchical Knowledge in a Deep Neural NetworkabstractDeep neural networks (DNNs) have been extensively applied in image processing, including visual saliency map pre-diction of images. A major difficulty in using a DNN for visual saliency prediction is the lack of labeled ground truth of visual saliency. A powerful DNN usually contains a large number of trainable parameters. This condition can easily lead to model over-fitting. In this study, we develop a novel method that over-comes such difficulty by embedding hierarchical knowledge of existing visual saliency models in a DNN. We achieve the objective of exploiting the knowledge contained in the existing visual sali-ency models by using saliency maps generated by local, global, and semantic models to tune and fix about 92.5% of the parame-ters in our network in a hierarchical manner. As a result, the number of trainable parameters that need to be tuned by the ground truth is considerably reduced. This reduction enables us to fully utilize the power of a large DNN and overcome the issue of over-fitting at the same time. Furthermore, we introduce a simple but very effective center prior in designing the learning cost function of the DNN by attaching high importance to the errors around the image center. We also present extensive experimental results on four commonly used public databases to demonstrate the superiority of the proposed method over classical and state-of-the-art methods on various evaluation metrics. Fei Zhou 0001, Rongguo Yao, Guangsen Liao, Guoping Qiu |
IEEE Trans. Image Process. | 5 |
| 2020 | Attention by Selection: A Deep Selective Attention Approach to Breast Cancer ClassificationabstractDeep learning approaches are widely applied to histopathological image analysis due to the impressive levels of performance achieved. However, when dealing with high-resolution histopathological images, utilizing the original image as input to the deep learning model is computationally expensive, while resizing the original image to achieve low resolution incurs information loss. Some hard-attention based approaches have emerged to select possible lesion regions from images to avoid processing the original image. However, these hard-attention based approaches usually take a long time to converge with weak guidance, and valueless patches may be trained by the classifier. To overcome this problem, we propose a deep selective attention approach that aims to select valuable regions in the original images for classification. In our approach, a decision network is developed to decide where to crop and whether the cropped patch is necessary for classification. These selected patches are then trained by the classification network, which then provides feedback to the decision network to update its selection policy. With such a co-evolution training strategy, we show that our approach can achieve a fast convergence rate and high classification accuracy. Our approach is evaluated on a public breast cancer histopathological image database, where it demonstrates superior performance compared to state-of-the-art deep learning approaches, achieving approximately 98% classification accuracy while only taking 50% of the training time of the previous hard-attention approach. Bolei Xu, Jingxin Liu 0005, Xianxu Hou, Jonathan M. Garibaldi, Ian O. Ellis, Andrew R. Green, LinLin Shen, Guoping Qiu |
IEEE Trans. Medical Imaging | 9 |
| 2019 | Side Window FilteringabstractLocal windows are routinely used in computer vision and almost without exception the center of the window is aligned with the pixels being processed. We show that this conventional wisdom is not universally applicable. When a pixel is on an edge, placing the center of the window on the pixel is one of the fundamental reasons that cause many filtering algorithms to blur the edges. Based on this insight, we propose a new Side Window Filtering (SWF) technique which aligns the window's side or corner with the pixel being processed. The SWF technique is surprisingly simple yet theoretically rooted and very effective in practice. We show that many traditional linear and nonlinear filters can be easily implemented under the SWF framework. Extensive analysis and experiments show that implementing the SWF principle can significantly improve their edge preserving capabilities and achieve state of the art performances in applications such as image smoothing, denoising, enhancement, structure-preserving texture-removing, mutual-structure extraction, and HDR tone mapping. In addition to image filtering, we further show that the SWF principle can be extended to other applications involving the use of a local window. Using colorization by optimization as an example, we demonstrate that implementing the SWF principle can effectively prevent artifacts such as color leakage associated with the conventional implementation. Given the ubiquity of window based operations in computer vision, the new SWF technique is likely to benefit many more applications. Yuanhao Gong, Guoping Qiu |
CVPR | 3 |
| 2019 | Spectral Regularization for Combating Mode Collapse in GANsabstractDespite excellent progress in recent years, mode collapse remains a major unsolved problem in generative adversarial networks (GANs). In this paper, we present spectral regularization for GANs (SR-GANs), a new and robust method for combating the mode collapse problem in GANs. Theoretical analysis shows that the optimal solution to the discriminator has a strong relationship to the spectral distributions of the weight matrix. Therefore, we monitor the spectral distribution in the discriminator of spectral normalized GANs (SN-GANs), and discover a phenomenon which we refer to as spectral collapse, where a large number of singular values of the weight matrices drop dramatically when mode collapse occurs. We show that there are strong evidence linking mode collapse to spectral collapse; and based on this link, we set out to tackle spectral collapse as a surrogate of mode collapse. We have developed a spectral regularization method where we compensate the spectral distributions of the weight matrices to prevent them from collapsing, which in turn successfully prevents mode collapse in GANs. We provide theoretical explanations for why SR-GANs are more stable and can provide better performances than SN-GANs. We also present extensive experimental results and analysis to show that SR-GANs not only always outperform SN-GANs but also always succeed in combating mode collapse where SN-GANs fail. Kanglin Liu, Guoping Qiu, Wenming Tang, Fei Zhou 0001 |
ICCV | 2 |
| 2019 | Learning Spatial-Aware Cross-View Embeddings for Ground-to-Aerial Geolocalization
Rui Cao 0001, Jiasong Zhu, Qing Li 0029, Qian Zhang 0018, Qingquan Li 0001, Guoping Qiu |
ICIG (1) | 7 |
| 2019 | Soft Tissue Removal in X-Ray Images by Half Window Dark Channel PriorabstractSoft tissue in X-ray images obscures the bone structure such that the details on bones are not clear. Conventional methods simultaneously enhance the image contrast for the soft tissue and the bones. Here we propose to remove all soft tissue in X-ray images, making the bone structure clear. For this purpose, we first propose a half window dark channel prior and a half window guided filter. Then, we apply this prior and filter on X-ray images. After processing, the bone details in X-ray images become clear and sharp. Several experiments confirm the effectiveness and efficiency of our method. Our method can be used for bone segmentation, classification, recognition, and clinical diagnosis. Yuanhao Gong, Jingxin Liu 0005, Guoping Qiu |
ICIP | 5 |
| 2019 | Salient Object Detection With Capsule-Based Conditional Generative Adversarial NetworkabstractSalient Object Detection (SOD) is one significant research area which is closely correlated to the attention of human beings. Most of the nowadays CNN-based approaches for SOD are based on an U-Net architecture. In this paper, we propose a novel capsule-based salient object detection framework by integrating the novel capsule blocks into both the generator and discriminator of GAN architecture. The experimental result showed that our approach is able to generate accurate saliency maps, which also highlighted the effectiveness of the capsule blocks. We also provide a challenging dataset that contains 3,299 images for SOD with difficult foreground objects and complex background contents. Chao Zhang 0020, Guoping Qiu, Qian Zhang 0018 |
ICIP | 3 |
| 2019 | FrameRank: A Text Processing Approach to Video SummarizationabstractVideo summarization has been extensively studied in the past decades. However, user-generated video summarization is much less explored since there lack large-scale video datasets within which human-generated video summaries are unambiguously defined and annotated. Toward this end, we propose a user-generated video summarization dataset - UGSum52 - that consists of 52 videos (207 minutes). In constructing the dataset, because of the subjectivity of user-generated video summarization, we manually annotate 25 summaries for each video, which are in total 1300 summaries. To the best of our knowledge, it is currently the largest dataset for user-generated video summarization. Based on this dataset, we present FrameRank, an unsupervised video summarization method that employs a frame-to-frame level affinity graph to identify coherent and informative frames to summarize a video. We use the Kullback-Leibler(KL)-divergence-based graph to rank temporal segments according to the amount of semantic information contained in their frames. We illustrate the effectiveness of our method by applying it to three datasets SumMe, TVSum and UGSum52 and show it achieves state-of-the-art results. Zhuo Lei, Chao Zhang 0020, Qian Zhang 0018, Guoping Qiu |
ICME | 4 |
| 2019 | Dual Adaptive Pyramid Network for Cross-Stain Histopathology Image Segmentation
Xianxu Hou, Jingxin Liu 0005, Bolei Xu, Xin Chen 0003, Mohammad Ilyas, Ian O. Ellis, Jonathan M. Garibaldi, Guoping Qiu |
MICCAI (2) | 9 |
| 2019 | Character Prediction in TV Series via a Semantic Projection Network
Ke Sun 0006, Zhuo Lei, Jiasong Zhu, Xianxu Hou, Guoping Qiu |
MMM (1) | 6 |
| 2019 | Improving variational autoencoder with deep feature consistent and generative adversarial training
Xianxu Hou, Ke Sun 0006, LinLin Shen, Guoping Qiu |
Neurocomputing | 4 |
| 2019 | A physics based generative adversarial network for single image defogging
Wei Liu 0123, Rongguo Yao, Guoping Qiu |
Image Vis. Comput. | 3 |
| 2019 | Deep reinforcement learning-based patch selection for illuminant estimation
Bolei Xu, Jingxin Liu 0005, Xianxu Hou, Guoping Qiu |
Image Vis. Comput. | 5 |
| 2019 | Side window guided filtering
Yuanhao Gong, Guoping Qiu |
Signal Process. | 3 |
| 2019 | Visual Quality Assessment for Super-Resolved Images: Database and MethodabstractImage super-resolution (SR) has been an active research problem which has recently received renewed interest due to the introduction of new technologies such as deep learning. However, the lack of suitable criteria to evaluate the SR performance has hindered technology development. In this paper, we fill a gap in the literature by providing the first publicly available database as well as a new image quality assessment (IQA) method specifically designed for assessing the visual quality of super-resolved images (SRIs). In constructing the quality assessment database for SRIs (QADS), we carefully selected 20 reference images and created 980 SRIs using 21 image SR methods. Mean opinion score (MOS) for these SRIs is collected through 100 individuals participating in a suitably designed psychovisual experiment. Extensive numerical and statistical analysis is performed to show that the MOS of QADS has excellent suitability and reliability. The psychovisual experiment has led to the discovery that, unlike distortions encountered in other IQA databases, artifacts of the SRIs degenerate the image structure as well as the image texture. Moreover, the structural and textural degenerations have distinctive perceptual properties. Based on these insights, we propose a novel method to assess the visual quality of SRIs by separately considering the structural and textural components of images. Observing that textural degenerations are mainly attributed to dissimilar texture or checkerboard artifacts, we propose to measure the changes of textural distributions. We also observe that structural degenerations appear as blurring and jaggies artifacts in SRIs and develop separate similarity measures for different types of structural degenerations. A new pooling mechanism is then used to fuse the different similarities together to give the final quality score for an SRI. The experiments conducted on the QADS demonstrate that our method significantly outperforms the classical as well as current state-of-the-art IQA methods. Fei Zhou 0001, Rongguo Yao, Guoping Qiu |
IEEE Trans. Image Process. | 4 |
| 2019 | An End-to-End Deep Learning Histochemical Scoring System for Breast Cancer TMAabstractOne of the methods for stratifying different molecular classes of breast cancer is the Nottingham prognostic index plus, which uses breast cancer relevant biomarkers to stain tumor tissues prepared on tissue microarray (TMA). To determine the molecular class of the tumor, pathologists will have to manually mark the nuclei activity biomarkers through a microscope and use a semi-quantitative assessment method to assign a histochemical score (H-Score) to each TMA core. Manually marking positively stained nuclei is a time-consuming, imprecise, and subjective process, which will lead to inter-observer and intra-observer discrepancies. In this paper, we present an end-to-end deep learning system, which directly predicts the H-Score automatically. Our system imitates the pathologists' decision process and uses one fully convolutional network (FCN) to extract all nuclei region (tumor and non-tumor), a second FCN to extract tumor nuclei region, and a multi-column convolutional neural network, which takes the outputs of the first two FCNs and the stain intensity description image as an input and acts as the high-level decision making mechanism to directly output the H-Score of the input TMA image. To the best of our knowledge, this is the first end-to-end system that takes a TMA image as the input and directly outputs a clinical score. We will present experimental results, which demonstrate that the H-Scores predicted by our model have very high and statistically significant correlation with experienced pathologists' scores and that the H-Score discrepancy between our algorithm and the pathologists is on par with the inter-subject discrepancy between the pathologists. Jingxin Liu 0005, Bolei Xu, Chi Zheng, Yuanhao Gong, Jonathan M. Garibaldi, Daniele Soria, Andrew R. Green, Ian O. Ellis, Wenbin Zou, Guoping Qiu |
IEEE Trans. Medical Imaging | 10 |
| 2018 | Capsule Based Image Synthesis for Interior Design Effect Rendering
Zheng Lu 0002, Guoping Qiu, Qian Zhang 0018 |
ACCV (5) | 3 |
| 2018 | Urban Land Use Classification Based on Aerial and Ground ImagesabstractUrban land use is key to rational urban planning and management. Traditional land use classification methods rely heavily on domain experts, which is both expensive and inefficient. In this paper, we explore to utilise deep neural network based approaches to label urban land use at pixel level using high-resolution aerial images and ground-level street images. We use a deep neural network to extract semantic features from sparsely distributed street images and interpolate them in the spatial domain to match the spatial resolution of the aerial images, which are then fused together through a deep neural network for classifying land use categories. We test our methods on a large publicly available aerial and street images dataset of New York City, and the results show that using aerial images alone can achieve relatively high classification accuracy and the ground-level street views contain useful information for urban land use classification. Fusing street image features with aerial images can improve classification accuracy to some extent but the improvement is somewhat limited. Rui Cao 0001, Guoping Qiu |
CBMI | 2 |
| 2018 | Annotating Images with Drawings Through Gameplay for Drawing Based Image RetrievalabstractDrawing based image retrieval offers an intuitive and engaging way to retrieve images from large repositories, and can be easily performed on touchscreen devices. However, there is usually a large visual gap between a drawing query and a desired image since most users are not drawing experts. Based on the hypothesis that color drawings aiming at the same target image by different users would be more similar to each other than between a drawing and the target image itself, we propose the novel concept of image annotation with drawings for drawing based image retrieval. Just as text based image retrieval would retrieve images through retrieving their associated annotation texts, drawing based image retrieval is performed through the retrieval of the annotation drawings of the images. To address the issue of creating drawing annotations for images, we design and implement an entertaining drawing game on mobile devices based on the concept of game with a purpose (GWAP). We have so far collected over 1200 drawings of 60 target images from 30 users who play our game. We will present experimental results to demonstrate the correctness of our hypothesis. Comparing with the traditional approaches that directly comparing drawings with target images, we show that drawing based image retrieval through retrieving drawing annotations of the target images significantly improves the retrieval accuracy. Guoping Qiu |
CBMI | 2 |
| 2018 | Video Salient Object Detection via Multiple Time-scale AnalysisabstractThis paper focuses on salient object detection in video by multiple time-scale analysis, which exploits the temporally consistent information under three different scales. In the first time-scale, we define an effective measure called motion contrast from both low-level cues and the optical flow fields. In the second time-scale, we propose a novel approach to repair the inaccurate motion contrast due to the mistake of optical flow. In the third time-scale, considering the low-contrast objects that stop moving for a certain amount of time and cannot remain prominent, we present a robust motion detection method based on point-tracking and trajectories clustering. Finally, the outcomes from the three time-scales jointly formulate the saliency detection by Bayesian inference. The proposed model is evaluated on the widely-used DAVIS and FBMS benchmark. Experiments demonstrate that our proposed model substantially outperforms the state-of-the-art saliency detection models. Yuhuan Chen, Limin Huang, Wenbin Zou, Xia Li 0006, Guoping Qiu |
ICPR | 5 |
| 2018 | Quality Classified Image Analysis with Application to Face Detection and RecognitionabstractMotion blur, out of focus, insufficient spatial resolution, lossy compression and many other factors can all cause an image to have poor quality. However, image quality is a largely ignored issue in traditional pattern recognition literature. In this paper, we use face detection and recognition as case studies to show that image quality is an essential factor which will affect the performances of traditional algorithms. We demonstrated that it is not the image quality itself that is the most important, but rather the quality of the images in the training set should have similar quality as those in the testing set. To handle real-world application scenarios where images with different kinds and severities of degradation can be presented to the system, we have developed a quality classified image analysis framework to deal with images of mixed qualities adaptively. We use deep neural networks first to classify images based on their quality classes and then design a separate face detector and recognizer for images in each quality class. We will present experimental results to show that our quality classified framework can accurately classify images based on the type and severity of image degradations and can significantly boost the performances of state-of-the-art face detector and recognizer in dealing with image datasets containing mixed quality images. Qian Zhang 0018, Miaohui Wang, Guoping Qiu |
ICPR | 4 |
| 2018 | Sub-window Box FilterabstractBox filter is a fundamental filter in image processing. However, it can not preserve edges or corners. In this paper, we present a simple but novel method that can make box filter both edge and corner preserving. More specifically, we combine the box filter with sub-window regression to achieve this task. This filter inherits some properties from box filter, such as O(1) running time with respect to the window radius. After analyzing its parameters, we show its corner and edge preserving property on real images and compare it with Guided filter. Yuanhao Gong, Xianxu Hou, Guoping Qiu |
VCIP | 4 |
| 2018 | Direct Application of Convolutional Neural Network Features to Image Quality AssessmentabstractWe take advantage of the popularity of deep convolutional neural networks (CNNs) and have developed a very simple image quality assessment method that rivals state of the art. We show that convolutional layer outputs (deep features) of a CNN compute the local structural information of spatial regions of different sizes in the input image. The learned convolutional kernels contain a much richer set of weights thus capturing much more local structural information than hand crafted ones. As the deep features learned from large datasets already contain very rich multi-resolutional structural image information, they can be directly used to calculate visual distortion of an image and it is not necessary to introduce further complicated computational process. We will present experimental results to demonstrate that this is indeed the case, and that simple cosine distance of the deep features is as good as state the art methods for full reference image quality assessment. Xianxu Hou, Ke Sun 0006, Yuanhao Gong, Jonathan M. Garibaldi, Guoping Qiu |
VCIP | 6 |
| 2018 | Random forest for label ranking
Yangming Zhou, Guoping Qiu |
Expert Syst. Appl. | 2 |
| 2018 | Riemannian competitive learning for symmetric positive definite matrices clustering
Ligang Zheng, Guoping Qiu, Jiwu Huang |
Neurocomputing | 2 |
| 2017 | Learning deep semantic attributes for user video summarizationabstractThis paper presents a Semantic Attribute assisted video SUMmarization framework (SASUM). Compared with traditional methods, SASUM has several innovative features. Firstly, we use a natural language processing tool to discover a set of keywords from an image and text corpora to form the semantic attributes of visual contents. Secondly, we train a deep convolution neural network to extract visual features as well as predict the semantic attributes of video segments which enables us to represent video contents with visual and semantic features simultaneously. Thirdly, we construct a temporally constrained video segment affinity matrix and use a partially near duplicate image discovery technique to cluster visually and semantically consistent video frames together. These frame clusters can then be condensed to form an informative and compact summary of the video. We will present experimental results to show the effectiveness of the semantic attributes in assisting the visual features in video summarization and our new technique achieves state-of-the-art performance. Ke Sun 0006, Jiasong Zhu, Zhuo Lei, Xianxu Hou, Qian Zhang 0018, Jiang Duan, Guoping Qiu |
ICME | 7 |
| 2017 | A Color Prediction System for Interactive Drawing Based Image Retrieval on Mobile DevicesabstractQuery suggestion plays a key role in improving the usability of image search. Textual Query Suggestion, widely used in existing search text-based image retrieval engines, is able to suggest a list of textual query terms based on users' query input. This paper presents a color prediction system dedicated to interactive drawing based image retrieval system on mobile devices. Usually, such systems retrieve the results after a complete sketch is submitted as input. We propose a color prediction system that recommends colors in the palette and predict drawings to assist users in creating queries. The predicted drawings can be fed to the retrieval system so it can retrieve results as soon as a user draws a stroke on the canvas. We conduct user studies to compare our system with a general drawing based image retrieval interface and show that our system can help users retrieve target image with an improved efficiency. Guoping Qiu |
ISM | 2 |
| 2017 | Deep Feature Consistent Variational AutoencoderabstractWe present a novel method for constructing Variational Autoencoder (VAE). Instead of using pixel-by-pixel loss, we enforce deep feature consistency between the input and the output of a VAE, which ensures the VAE's output to preserve the spatial correlation characteristics of the input, thus leading the output to have a more natural visual appearance and better perceptual quality. Based on recent deep learning works such as style transfer, we employ a pre-trained deep convolutional neural network (CNN) and use its hidden features to define a feature perceptual loss for VAE training. Evaluated on the CelebA face dataset, we show that our model produces better results than other methods in the literature. We also show that our method can produce latent vectors that can capture the semantic information of face expressions and can be used to achieve state-of-the-art performance in facial attribute prediction. Xianxu Hou, LinLin Shen, Ke Sun 0006, Guoping Qiu |
WACV | 4 |
| 2016 | Clustering Symmetric Positive Definite Matrices on the Riemannian Manifolds
Ligang Zheng, Guoping Qiu, Jiwu Huang |
ACCV (1) | 2 |
| 2016 | A classification method for estimating the illuminant of an imageabstractIdentifying the light source of an image is important for image processing tasks such as colour correction and white point balancing. This is also known as colour constancy in computer vision. This paper presents a novel clustering classification colour constancy framework (the 4C method). Based on the assumption that similar illuminants will result in similar white point colours, we first use a clustering algorithm to group similar white point colours of the training samples into the same cluster. We then treat the images in the same cluster as belonging to the same illumination source and each cluster as one class of illuminants. The colour constancy problem, i.e., that of estimating the unknown illuminant of an image, becomes that of identifying which illuminant class (cluster) the images illuminant falling into. To achieve this, we derive an effective colour feature representation of the image and use a general classification algorithm to classify the image into one of the illuminant classes (clusters). We present experimental results on publicly available testing datasets and show that our new method is competitive to state of the art. Guoping Qiu |
VCIP | 2 |
| 2016 | Crowd density estimation based on rich features and random projection forestabstractCurrent state of the art crowd density estimation methods are based on computationally expensive Gaussian process regression or Ridge regression models which can only handle a small number of features. In many computer vision applications, it has been empirically shown that a richer set of image features can lead to enhanced performances. In this paper, we reason that using more image features could potentially boost the performances of crowd detection and thus propose to employ much extensive and richer feature sets for crowd density estimation. To achieve computational efficiency and scalability, we use random forest as the regression model whose tree structure is intrinsically fast and scalable. Unlike traditional approaches to random forest construction, we embed random projection in the tree nodes to simultaneously combat the curse of dimensionality and to introduce randomness in the tree construction (we call this Random Projection Forest) thus making our new method very efficient and effective. Experimental results on two public pedestrian detection video databases show that our new method achieves state of the art performances that are superior to those of previously published regression techniques. Bolei Xu, Guoping Qiu |
WACV | 2 |
| 2016 | Habitat image annotation with low-level features, medium-level knowledge and location information
Mercedes Torres, Guoping Qiu |
Multim. Syst. | 2 |
| 2015 | A High Dynamic Range Microscopic Video System
Chi Zheng, Salvador Garcia Bernal, Guoping Qiu |
ICIG (1) | 3 |
| 2015 | Optimal Pricing for the Competitive and Evolutionary Cloud Market
Bolei Xu, Tao Qin 0001, Guoping Qiu, Tie-Yan Liu |
IJCAI | 3 |
| 2015 | A Comparison of Five HSV Color Selection Interfaces for Mobile Painting Search
Guoping Qiu, Natasha Alechina, Sarah Atkinson |
INTERACT (2) | 2 |
| 2015 | Bundling Centre for Landmark Image DiscoveryabstractThis paper introduces a novel method to efficiently discover/cluster landmark images in large image collection. We consider each cluster as a combination of several sub-clusters, which is composed of images taken from different view points of the identical landmark. For each sub-cluster, we find its local centre represented by a group of similar images, and define it as the bundling centre (BC). We therefore start the image discovery/clustering by identifying the BCs and accomplish the task by efficiently growing and merging those sub-clusters represented by different BCs. In our proposed method, we use min-hash based method to build a sparse graph so as to avoid the time-consuming full-scale exhaustive pairwise image matching. Based on the information provided by the sparse graph, BCs are identified as local dense neighbors sharing high intra-similarity. We have also proposed a weighted voting method to efficiently grow these BCs with high accuracy. More importantly, the fixed local centres can ensure each sub-cluster contains identical landmark and generate result with high precision. In addition, compared to a single representative (iconic) image, the group of similar images obtained by each BC can provide more comprehensive cluster information and, thus, overcome the problem of low recall caused by information lost during visual word quantization. We present experimental results on two landmark datasets and show that, without query expansion, our method can boost landmark image discovery/clustering performances of state of the art techniques. Qian Zhang 0018, Guoping Qiu |
ICMR | 2 |
| 2015 | A Preliminary Examination of the User Behavior in Query-by-Drawing Portrait Painting Search on Mobile DevicesabstractAlthough many researchers have studied the user behavior of using text-based information search engine, less is known about search pattern for mobile content-based image search. We developed a Query-by-Drawing (QbD) mobile application, and conducted a user study on it to explore the search behavior of painting search by drawing on the touchscreen phone. Based on the resulting drawings and video-logs of drawing procedures on three task conditions, we analyzed the patterns of query formulation and query modification. We further examined the effects of user characteristic and task type on the search strategy when using our mobile application. The results elicited some guidelines for mobile QbD image search interface design and informed the potential improvements of our application. Guoping Qiu, Natasha Alechina, Sarah Atkinson |
MoMM | 2 |
| 2014 | Fast and accurate Nearest Neighbor search in the manifolds of symmetric positive definite matricesabstractIn this paper, we present a fast and accurate Nearest Neighbor (NN) search method in the Riemannian manifolds formed by a kind of structured data - symmetric positive definite (SPD) matrices. We use an ensemble of vocabulary trees based on hierarchical k-means clustering and query these trees to find the NN candidates in sub-linear time. As generating these vocabulary trees with widely used affine-invariant Riemannian metric (AIRM) will be very time-demanding, we propose to use the second-order approximation to AIRM (SOA-AIRM). We evaluate the proposed NN search algorithm in the application scenario of near-duplicate image detection in a large database. Experimental results demonstrate that the proposed method significantly outperforms state of the art techniques in terms of both accuracy and speed. Ligang Zheng, Guoping Qiu, Jiwu Huang, Jiang Duan |
ICASSP | 2 |
| 2014 | A portable real-time high-dynamic range video systemabstractUsing a single sensor to capture videos of scenes containing very high and very low luminance contents at the same time is still a challenge. An alternative solution is to use multiple sensors and specially designed optics. Such approach also faces some significant technical challenges including the manufacture of precision optics, accurate alignment of different sensors, recovering high dynamic range video frames and rendering them for display in real-time. In this work, we present a new real-time high dynamic range video system which is able to produce high quality live video up to 30 frames per second. The innovative features of the system include novel optics, new high dynamic range radiance map recovery methods and GPU based real-time HDR video processing and display. A video recorded live by the new system can be found in this vimeo link: http://vimeo.com/97595588. We will bring the system to the conference and perform live demo. Salvador Garcia Bernal, Guoping Qiu |
ICIP | 2 |
| 2014 | Crowdsourcing based radio map anomalous event detection system for calibration-on-demandabstractCalibrations for Wi-Fi Indoor Positioning Systems (IPS) are expensive and time-consuming. Recalibrations are necessary periodically to compensate the drifts of radio map. Many crowdsourcing-based approaches have been proposed to calibrate the radio map strategically with user-contributed data. However, these approaches still suffer from several drawbacks. Calibration-on-demand is desirable to reduce the maintenance cost and to enhance the reliability of the Wi-Fi IPS. Wi-Fi IPS degrades only when the radio map changes significantly and lastingly. Such changes are caused by radio map anomalous events. Examples of these events include changes in indoor environment, existence of spurious data from the sensing system and unexpected failure in utility infrastructures. This paper presents a theoretical and experimental study of a crowdsourcing online Wi-Fi radio map anomalous event detection system. Detection of these events will enhance the operation intelligence of Wi-Fi IPS and enable high-level context-aware services. This system is consisted of an outlier detector for abnormal signal identification and an event discriminator for determination of event occurrence. Real world smartphone data and simulated data are used to test the performance of the system. Initial results show that the system is able to detect over 92% of simulated events and 87.5% of real world events with a low false alarm rate. Guoping Qiu, Yupeng Gao, Xiong Fang, Andy Chang, Chuen-Yu Chan |
IPIN | 2 |
| 2014 | Can People Finger-draw Color-sketches from Memory for Painting Search on Mobile Phone?abstractFor the case of people desire to find the previously-seen painting but only have vague memory of the painting, we designed and built a mobile phone application to enable people to search for paintings by drawing rough color sketches. Three-phase memory studies -- with 15-minute delay, 1-week delay and 1-month delay -were conducted to explore if people could draw from their visual memory and the resulting drawing were useful for search. Seventeen participants were involved in three memory studies. The experiment results implied that most of participants could draw usable rough color sketches from their memory as painting queries even one month after viewing. Our research also demonstrated that users could improve their performance if they got more familiar with the functions and usage of our application. Sarah Atkinson, Guoping Qiu, Natasha Alechina |
MoMM | 3 |
| 2014 | A label ranking method based on Gaussian mixture model
Yangming Zhou, Yangguang Liu, Xiao Zhi Gao 0001, Guoping Qiu |
Knowl. Based Syst. | 4 |
| 2014 | Face hallucination based on sparse local-pixel structureabstractIn this paper, we propose a face-hallucination method, namely face hallucination based on sparse local-pixel structure. In our framework, a high resolution (HR) face is estimated from a single frame low resolution (LR) face with the help of the facial dataset. Unlike many existing face-hallucination methods such as the from local-pixel structure to global image super-resolution method (LPS-GIS) and the super-resolution through neighbor embedding, where the prior models are learned by employing the least-square methods, our framework aims to shape the prior model using sparse representation. Then this learned prior model is employed to guide the reconstruction process. Experiments show that our framework is very flexible, and achieves a competitive or even superior performance in terms of both reconstruction error and visual quality. Our method still exhibits an impressive ability to generate plausible HR facial images based on their sparse local structures. Cheng Cai, Guoping Qiu, Kin-Man Lam 0001 |
Pattern Recognit. | 3 |
| 2014 | A Novel Polar Space Random Field Model for the Detection of Glandular StructuresabstractIn this paper, we propose a novel method to detect glandular structures in microscopic images of human tissue. We first convert the image from Cartesian space to polar space and then introduce a novel random field model to locate the possible boundary of a gland. Next, we develop a visual feature-based support vector regressor to verify if the detected contour corresponds to a true gland. And finally, we combine the outputs of the random field and the regressor to form the GlandVision algorithm for the detection of glandular structures. Our approach can not only detect the existence of the gland, but also can accurately locate it with pixel accuracy. In the experiments, we treat the task of detecting glandular structures as object (gland) detection and segmentation problems respectively. The results indicate that our new technique outperforms state-of-the-art computer vision algorithms in respective fields. Hao Fu 0001, Guoping Qiu, Jie Shu, Mohammad Ilyas |
IEEE Trans. Medical Imaging | 2 |
| 2014 | Fast Near-Duplicate Image Detection Using Uniform Randomized TreesabstractIndexing structure plays an important role in the application of fast near-duplicate image detection, since it can narrow down the search space. In this article, we develop a cluster of uniform randomized trees (URTs) as an efficient indexing structure to perform fast near-duplicate image detection. The main contribution in this article is that we introduce “uniformity” and “randomness” into the indexing construction. The uniformity requires classifying the object images into the same scale subsets. Such a decision makes good use of the two facts in near-duplicate image detection, namely: (1) the number of categories is huge; (2) a single category usually contains only a small number of images. Therefore, the uniform distribution is very beneficial to narrow down the search space and does not significantly degrade the detection accuracy. The randomness is embedded into the generation of feature subspace and projection direction, improveing the flexibility of indexing construction. The experimental results show that the proposed method is more efficient than the popular locality-sensitive hashing and more stable and flexible than the traditional KD-tree. Yanqiang Lei, Guoping Qiu, Ligang Zheng, Jiwu Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2013 | Information-Based Scale Saliency Methods with Wavelet Sub-band Energy Density Descriptors
Anh Cat Le Ngo, Li-Minn Ang, Guoping Qiu, Kah Phooi Seng |
ACIIDS (2) | 3 |
| 2013 | Multiscale Discriminant Saliency for Visual Attention
Anh Cat Le Ngo, Li-Minn Ang, Guoping Qiu, Kah Phooi Seng |
ICCSA (1) | 3 |
| 2013 | Recovering High Dynamic Range Radiance Maps from Photographs Revisited: A Simple and Important FixabstractAn influential technique introduced by Debevec and Malik has become the de facto standard for recovering high dynamic range radiance maps from photographs and has been widely used in research and commercial systems for over a decade. However, we have discovered an important defect in the original algorithm that will make this technique often fail to produce reasonable results in the extremely bright or dark regions of a scene. In this paper, we introduce a novel technique to correct this defect. Instead of the original algorithm where only pixel values from the photographs are used to guide the synthesis of the high dynamic range radiance map, we explicitly incorporate the shutter speed information of the camera. At each spatial pixel location, we estimate a "suitable shutter" that will make that location best exposed. A pixel's contribution to the high dynamic range radiance value is not only a function of its value but also depends on the difference between shutter speed used to take the pixel and the estimated "suitable shutter" of that pixel. We also show that this new idea can be successfully used to directly fuse differently exposed photographs into a single low dynamic range image for display in conventional low dynamic range devices. Yujie Mei, Guoping Qiu |
ICIG | 2 |
| 2013 | A Semi-automatic Image Analysis Tool for Biomarker Detection in Immunohistochemistry AnalysisabstractDigitized tissue slide analysis allows the Pathologists to use computer assisted image analysis technology to reduce time cost and increase the accuracy of diagnosis. In this paper, we present an ease to use semi-automatic tool which can detect and separate stains in tissue samples correctly. This statistical model based tool has been applied to detecting different kinds of stains. Multiple evaluation processes are implemented to demonstrate the robustness, accuracy and usefulness of this tool. Experimental results show that the tool performs significantly better than established popular methods such as color deconvolution and CMYK in stained color separation. Jie Shu, Guoping Qiu, Mohammad Ilyas |
ICIG | 2 |
| 2013 | Multi-scale visual attention & saliency modelling with decision theory
Anh Cat Le Ngo, Li-Minn Ang, Guoping Qiu, Kah Phooi Seng |
ICIP | 3 |
| 2013 | Habitat classification using random forest based image annotationabstractHabitat classification is an important ecological activity used to monitor environmental biodiversity. Current classification techniques rely heavily on human surveyors and are laborious, time consuming, expensive and subjective. In this paper, we approach habitat classification as an automatic image annotation problem. We have developed a novel method for annotating ground-taken photographs with the habitats present in them using random projection forests. For this purpose, we have collected and manually annotated a geo-referenced habitat image database with over 1000 ground photographs. We compare the use of two different types of input (blocks within images and the whole images) to classify habitats. We also compare our approach with a popular random forest implementation. Results show that our approach has a lower error rate and it is able to classify three habitats (Woodland and scrub, Grassland and marsh, and Miscellaneous) with a high recall. Mercedes Torres, Guoping Qiu |
ICIP | 2 |
| 2013 | Interactive skin condition recognitionabstractIt is believed that there are between 1000 to 2000 skin conditions, and about 20% are difficult to diagnose. An intelligent system capable of making accurate diagnosis not only helps patients in places where access to health services are scarce, but also benefits typical general practitioners who have received minimal dermatology training. In this paper, we introduce a challenging dataset developed by gathering 2309 images from 44 different skin conditions, and collecting answers to simple perceptual questions from 361 “Amazon Mechanical Turk” workers. We also propose a method based on random forest technology that combines visual features of the skin lesion images with user provided answers to achieve promising recognition rates. We believe that our solution can be potentially improved and installed on smart phones and tablets to enhance quality of life in patients across the world. Orod Razeghi, Qian Zhang 0018, Guoping Qiu |
ICME | 3 |
| 2013 | Tree partition voting min-hash for partial duplicate image discoveryabstractDiscovering partially duplicated images such as those of the same scenes, buildings or objects taken from different angles, distances and vantage points can be very useful in applications such as managing large image repositories and image search on the Internet. In this paper, we present a novel technique for partial duplicate image discovery. The new technique, termed tree partition voting min-hash (TmH), first partitions interest points within an image based on their geometric or photometric (appearance) properties using a spatial partition tree data structure and then finds potential partial duplicate images through a traditional partition min-hash (PmH) method [1]. We have developed a k-d tree partition min-hash (kdTmH) and a random projection tree partition min-hash (rpTmH) technique and have also developed a weighted voting algorithm for improving the similarity measure of a pair of hashing sketches. We present experimental results on 3 datasets and show that TmH significantly outperforms PmH in terms of recall and precision performances without increasing complexity and that the new voting algorithms performs better than sketch matching techniques in the literature. Qian Zhang 0018, Hao Fu 0001, Guoping Qiu |
ICME | 3 |
| 2013 | Integrating low-level and semantic features for object consistent segmentation
Hao Fu 0001, Guoping Qiu |
Neurocomputing | 2 |
| 2012 | GlandVision: A Novel Polar Space Random Field Model for Glandular Biological Structure DetectionabstractIn this paper, we propose a novel method for detecting glandular structures in microscopic images of human tissue. We first transform the image from Cartesian space to polar space and introduce a novel random field model with an efficient inference strategy that uses two simple chain graphs to approximate a circular graph to infer possible boundary of a gland. We then develop a visual feature based support vector regressor (SVR) to verify if the inferred contour corresponds to a true gland. And finally, we combine the outputs of the random field and the regressor to form the GlandVision algorithm for the detection of glandular structures. In the experiments, we treat the task of detecting glandular structures as object (gland) proposal, detection and segmentation problems respectively and show that our new technique outperforms state of the art computer vision algorithms developed for generic objects. Hao Fu 0001, Guoping Qiu, Mohammad Ilyas, Jie Shu |
BMVC | 2 |
| 2012 | CamBlend: an object focused collaboration toolabstractCamBlend is a new focus-in-context panoramic video collaboration system designed to facilitate the interaction with and around objects in a lightweight, flexible package. As well as the ability to view very high resolution local and remote video that covers a full 180#176; field of view, the system contains a number of tools which facilitate bi-directional pointing between two remote spaces. In the first quasi-naturalistic exploratory study on a focus-in-context video system, we show a number of unique object referencing behaviours, including un-intentional or 'implicit' pointing and a number of scenarios where this was advantageous. Additionally the study highlighted some of the problems inherent in aligning between screen-based and real-world perspectives. James Norris, Holger Schnädelbach, Guoping Qiu |
CHI | 3 |
| 2012 | Random Forest for Image Annotation
Hao Fu 0001, Qian Zhang 0018, Guoping Qiu |
ECCV (6) | 3 |
| 2012 | Visual saliency based on fast nonparametric multidimensional entropy estimationabstractBottom-up visual saliency can be computed through information theoretic models but existing methods face significant computational challenges. Whilst nonparametric methods suffer from the curse of dimensionality problem and are computationally expensive, parametric approaches have the difficulty of determining the shape parameters of the distribution models. This paper makes two contributions to information theoretic based visual saliency models. First, we formulate visual saliency as center surround conditional entropy which gives a direct and intuitive interpretation of the center surround mechanism under the information theoretic framework. Second, and more importantly, we introduce a fast nonparametric multidimensional entropy estimation solution to make information theoretic-based saliency models computationally tractable and practicable in realtime applications. We present experimental results on publicly available eye-tracking image databases to demonstrate that the proposed method is competitive to state of the art. Anh Cat Le Ngo, Guoping Qiu, Geoff Underwood, Li-Minn Ang, Kah Phooi Seng |
ICASSP | 2 |
| 2012 | Efficient coarse-to-fine near-duplicate image detection in riemannian manifoldabstractThis paper presents an efficient coarse-to-fine strategy for near duplicate image detection in a Riemannian space. At the coarse level, we use the faster but less accurate log-Euclidean Riemannian metric to search the entire database to retrieve a subset of the images that are likely to contain the near duplicates of the querying image; and at the fine level, we use the more accurate but computationally more demanding affine-invariant Riemannian metric to search the coarse level results to accurately identify near-duplicates. We present experimental results to show that the new coarse to fine strategy can be over 20 times faster than existing techniques using affine-invariant Riemannian metric without sacrificing accuracy. Ligang Zheng, Guoping Qiu, Jiwu Huang |
ICASSP | 2 |
| 2012 | GPU-accelerated local tone-mapping for high dynamic range imagesabstractThis paper presents a very fast local tone mapping method for displaying high dynamic range (HDR) images. Though local tone mapping operators produce better local contrast and details, they are usually slow. We have solved this problem by designing a highly parallel algorithm, which can be easily implemented on a Graphics Processing Unit (GPU) to harvest high computational efficiency. At the same time, the proposed method mimics the local adaption mechanism of the human visual system and thus gives good results for a wide variety of images. Qiyuan Tian, Jiang Duan, Guoping Qiu |
ICIP | 3 |
| 2012 | Fast semantic image retrieval based on random forestabstractThis paper introduces random forest as a computational and data structure paradigm for fusing low-level visual features and high-level semantic concepts for image retrieval. We use visual features to split the tree nodes and use the image labels to supervise the splitting to make images located at the same tree node share similar semantic concepts as well as visual similarities. We exploit such a random forest and define the semantic neighbor set (SNS) of a given image as the union of all images in the leaf nodes that this image falls onto. From SNS we further define the semantic similarity measure (SSM) between two images as the number of trees in which they share the same leaf nodes within a SNS. With SNS and SSM, example-based image retrieval becomes that of first finding the SNS of the querying image and then ranking the images according to the SSMs between the querying image and images in its SNS. We also show that the new technique can be adapted for keyword-based semantic image retrieval. The inherent efficient tree data structure leads to fast solutions. We will present experimental results to show the effectiveness of this new semantic image retrieval technique. Hao Fu 0001, Guoping Qiu |
ACM Multimedia | 2 |
| 2012 | Supervised learning and anti-learning of colorectal cancer classes and survival rates from cellular biology parametersabstractIn this paper, we describe a dataset relating to cellular and physical conditions of patients who are operated upon to remove colorectal tumours. This data provides a unique insight into immunological status at the point of tumour removal, tumour classification and post-operative survival. Attempts are made to learn relationships between attributes (physical and immunological) and the resulting tumour stage and survival. Results for conventional machine learning approaches can be considered poor, especially for predicting tumour stages for the most important types of cancer. This poor performance is further investigated and compared with a synthetic, dataset based on the logical exclusive-OR function and it is shown that there is a significant level of “anti-learning” present in all supervised methods used and this can be explained by the highly dimensional, complex and sparsely representative dataset. For predicting the stage of cancer from the immunological attributes, anti-learning approaches outperform a range of popular algorithms. Chris M. Roadknight, Uwe Aickelin, Guoping Qiu, John Scholefield, Lindy Durrant |
SMC | 3 |
| 2012 | Near-Duplicate Image Detection in a Visually Salient Riemannian SpaceabstractThis paper presents a framework for near-duplicate image detection in a visually salient Riemannian space. A visual saliency model is first used to identify salient regions of the image and then the salient region covariance matrix (SCOV) of various image features is computed. SCOV, which lies in a Riemannian manifold, is used as a robust and compact image content descriptor. An efficient coarse-to-fine Riemannian (CTOFR) image search strategy has been developed to improve efficiency while maintaining accuracy. CTOFR first uses a computationally fast but less accurate log-Euclidean Riemannian metric to do a coarse level search of the entire database and retrieve a subset of likely targets and then uses a computationally expensive but more accurate affine-invariant Riemannian metric to search the returns from the coarse search. We present experimental results to demonstrate that SCOV is a very compact, robust, and discriminative descriptor which is competitive to other state-of-the-art descriptors for near-duplicate image and video detection. We show that CTOFR can yield significant speedups over traditional full search methods without sacrificing accuracy, and that the larger the database the higher the speedup factor. Ligang Zheng, Yanqiang Lei, Guoping Qiu, Jiwu Huang |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2011 | Feature Combination beyond Basic ArithmeticsabstractKernel-based feature combination techniques such as Multiple Kernel Learning use arithmetical operations to linearly combine different kernels. We have observed that the kernel distributions of different features are usually very different. We argue that the similarity distributions amongst the data points for a given dataset should not change with their representation features and propose the concept of relative kernel distribution invariance (RKDI). We have developed a very simple histogram matching based technique to achieve RKDI by transforming the kernels to a canonical distribution. We have performed extensive experiments on various computer vision and machine learning datasets and show that calibrating the kernels to an empirically chosen canonical space before they are combined can always achieve a performance gain over state-of-art methods. As histogram matching is a remarkably simple and robust technique, the new method is universally applicable to kernel-based feature combination. Hao Fu 0001, Guoping Qiu, Hangen He |
BMVC | 2 |
| 2011 | Local contrast stretch based tone mapping for high dynamic range imagesabstractThis paper presents a local tone mapping method to render high dynamic range images on conventional displays. We adaptively stretch contrast in local regions to reproduce local contrast. In order to avoid halos, we use bilateral filtering to smooth the image prior to the contrast stretching operation. Our method is fast and easy to use, and the experiment results show that the technique can produce good results on a variety of high dynamic range images. Jiang Duan, Wenpeng Dong, Guoping Qiu |
CIMSIVP | 4 |
| 2011 | Integrating Low-level and Semantic Features for Object Consistent SegmentationabstractThe aim of semantic segmentation is to assign each pixel a semantic label. Numerous methods for semantic segmentation have been proposed in recent years and most of them chose pixel or super pixel as the processing primitives. However, as the information contained in a pixel or a super pixel is not discriminative enough, the outputs of these algorithms are usually not object consistent. To tackle this problem, we introduce the concept of object-like regions as a new and higher level processing primitive. We first experimentally showed that using object-like regions as processing primitives can boost semantic segmentation accuracy, and then proposed a novel method to produce object-like regions by integrating state-of art low-level segmentation algorithms with typical semantic segmentation algorithms through a novel semantic feature feedback mechanism. We present experimental results on the publicly available image understanding database MSRC21 and show that the new method can achieve state of the art semantic segmentation results with far fewer processing primitives. Hao Fu 0001, Guoping Qiu |
ICIG | 2 |
| 2011 | Saliency Modulated High Dynamic Range Image Tone MappingabstractThis paper presents a new high dynamic range image tone mapping technique - saliency modulated tone mapping (SMTM). The HDR image is not directly viewable and dynamic range compression will unavoidably loose information. A saliency map analyzes the visual importance of the regions and can therefore direct the tone mapping operators to preserve the visual conspicuity of the regions that should more likely attract visual attention. In SMTM, we have developed a very fast algorithm to first compute the visual saliency map of the high dynamic range radiance map and then directly use the saliency of the local regions to control the local tone mapping curve such that highly salient regions will have their details and contrast better protected so as to remain salient and attract visual attention in the tone mapped display. We present experimental results to show that SMTM provides competitive performances to state of the art tone mapping techniques in rending visually pleasing low dynamic range displays. We also show that SMTM is better able to preserve the visual saliency of the HDR image and that SMTM renders high saliency regions to stand out to attract observers attention. Yujie Mei, Guoping Qiu, Kin-Man Lam 0001 |
ICIG | 2 |
| 2011 | Salient covariance for near-duplicate image and video detectionabstractThis paper introduces the covariance matrix of visually salient image features as a compact and robust descriptor for near duplicate image and video copy detection. We make two novel contributions. We first present a fast method for computing information theoretic based visual saliency maps using a data independent fast transform to replace the conventional data dependent computationally demanding transforms. We then introduce salient covariance (SCOV) - the covariance matrix of various image features within the visually salient regions and use SCOV for near duplicate image and video copy detection. We present experimental results to show that our new fast visual saliency computation technique improves efficiency without compromising performances. We demonstrate that SCOV is a very compact and robust feature for near duplicate image and video copy detection. Compared to popular features such as GIST, SCOV is not only more robust against various manipulations but also can be over 20 times more compact whilst achieving the same or better performances. Ligang Zheng, Guoping Qiu, Jiwu Huang, Hao Fu 0001 |
ICIP | 2 |
| 2011 | A Hybrid Probabilistic Model for Unified Collaborative and Content-Based Image TaggingabstractThe increasing availability of large quantities of user contributed images with labels has provided opportunities to develop automatic tools to tag images to facilitate image search and retrieval. In this paper, we present a novel hybrid probabilistic model (HPM) which integrates low-level image features and high-level user provided tags to automatically tag images. For images without any tags, HPM predicts new tags based solely on the low-level image features. For images with user provided tags, HPM jointly exploits both the image features and the tags in a unified probabilistic framework to recommend additional tags to label the images. The HPM framework makes use of the tag-image association matrix (TIAM). However, since the number of images is usually very large and user-provided tags are diverse, TIAM is very sparse, thus making it difficult to reliably estimate tag-to-tag co-occurrence probabilities. We developed a collaborative filtering method based on nonnegative matrix factorization (NMF) for tackling this data sparsity issue. Also, an L1 norm kernel method is used to estimate the correlations between image features and semantic concepts. The effectiveness of the proposed approach has been evaluated using three databases containing 5,000 images with 371 tags, 31,695 images with 5,587 tags, and 269,648 images with 5,018 tags, respectively. William Kwok-Wai Cheung, Guoping Qiu, Xiangyang Xue 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2011 | From Local Pixel Structure to Global Image Super-Resolution: A New Face Hallucination FrameworkabstractWe have developed a new face hallucination framework termed from local pixel structure to global image super-resolution (LPS-GIS). Based on the assumption that two similar face images should have similar local pixel structures, the new framework first uses the input low-resolution (LR) face image to search a face database for similar example high-resolution (HR) faces in order to learn the local pixel structures for the target HR face. It then uses the input LR face and the learned pixel structures as priors to estimate the target HR face. We present a three-step implementation procedure for the framework. Step 1 searches the database for K example faces that are the most similar to the input, and then warps the K example images to the input using optical flow. Step 2 uses the warped HR version of the K example faces to learn the local pixel structures for the target HR face. An effective method for learning local pixel structures from an individual face, and an adaptive procedure for fusing the local pixel structures of different example faces to reduce the influence of warping errors, have been developed. Step 3 estimates the target HR face by solving a constrained optimization problem by means of an iterative procedure. Experimental results show that our new method can provide good performances for face hallucination, both in terms of reconstruction error and visual quality; and that it is competitive with existing state-of-the-art methods. Kin-Man Lam 0001, Guoping Qiu, Tingzhi Shen |
IEEE Trans. Image Process. | 3 |
| 2010 | Recent advances in high dynamic range imaging technologyabstractRecently, visual representations using high dynamic range (HDR) images become increasingly popular, with advancement of technologies for increasing the dynamic range of image. HDR image is expected to be used in wide-ranging applications such as digital cinema, digital photography and next generation broadcast, because of its high quality and its powerful expression ability. HDR imaging technologies will spread its sphere of influence in imaging industry. In this paper, we review the state-of-the-art studies and the trends of the HDR imaging, in terms of the following three points: (1) HDR imaging sensor and HDR image generation techniques as image acquisition technologies, (2) encode method of HDR images for efficient transmission and storage, (3) human visual system issues associated with reproduction of HDR image. Yukihiro Bandoh, Guoping Qiu, Masahiro Okuda, Scott Daly, Til Aach, Oscar C. Au |
ICIP | 2 |
| 2010 | A hierarchical algorithm for image multi-labelingabstractThis paper presents an efficient two-stage method for multi-class image labeling. We first propose a simple label-filtering algorithm (LFA), which can remove most of the irrelevant labels for a query image while the potential labels are maintained. With a small population of potential labels left, we then apply the Naive-Bayes Nearest-Neighbor (NBNN) classifier as the second stage of our algorithm to identify the labels for the query image. This approach has been evaluated on the Corel database, and compared to existing algorithms. Experiment results show that our proposed algorithm can achieve a promising result, as it outperforms existing algorithms. Jiwei Hu, Kin-Man Lam 0001, Guoping Qiu |
ICIP | 3 |
| 2010 | Learning local pixel structure for face hallucinationabstractIn this paper, we present a novel learning-based face hallucination method based on the assumption that similar faces will have similar local pixel structures. We use the low- resolution (LR) input face to search a database for K example faces that are the most similar to the input and align them with the input accordingly. The local pixel structures of the target high-resolution (HR) image are learned from those warped HR example faces in a neighbor embedding manner, and a total variation (TV) constraint is employed to aid the learning of all pixels' embedding weights. The learned local pixel structures are then used as constraints to reconstruct a HR version of the input face. Experimental results show that the method performs well in terms of both reconstruction error and visual quality. Kin-Man Lam 0001, Guoping Qiu, Tingzhi Shen |
ICIP | 3 |
| 2010 | Tone mapping HDR images using optimization: A general frameworkabstractThis paper presents a novel tone mapping framework. First, we introduce a tone mapping fidelity principle which explicitly stipulates that tone-mapped image data should not only be visually enhanced but should also stay faithful to the original image. Second, this principle naturally translates tone mapping into a constrained optimization problem where a two-term cost function, one measures the difference between the tone mapped image and a visually enhanced version of the image, and the other measures the difference between the tone mapped image and the original image, is optimized. The relative weightings of the two terms in the cost function not only offers an insightful and simple mechanism to control the appearance of the tone mapped image but also enables the introduction of spatially varying or uniform weighting functions thus unifying local and global tone mapping in a single framework. We present results of tone mapping high dynamic range (HDR) images and low dynamic range JPEG images to demonstrate the effectiveness of the new tone mapping framework. Guoping Qiu, Yujie Mei, Kin-Man Lam 0001 |
ICIP | 1 |
| 2010 | Tone-mapping high dynamic range images by novel histogram adjustment
Jiang Duan, Marco Bressan 0003, Christopher R. Dance, Guoping Qiu |
Pattern Recognit. | 4 |
| 2010 | Interactive imaging and vision - Ideas, algorithms and applications
Guoping Qiu, Pong C. Yuen |
Pattern Recognit. | 1 |
| 2010 | JPEG error analysis and its applications to digital image forensicsabstractJPEG is one of the most extensively used image formats. Understanding the inherent characteristics of JPEG may play a useful role in digital image forensics. In this paper, we introduce JPEG error analysis to the study of image forensics. The main errors of JPEG include quantization, rounding, and truncation errors. Through theoretically analyzing the effects of these errors on single and double JPEG compression, we have developed three novel schemes for image forensics including identifying whether a bitmap image has previously been JPEG compressed, estimating the quantization steps of a JPEG image, and detecting the quantization table of a JPEG image. Extensive experimental results show that our new methods significantly outperform existing techniques especially for the images of small sizes. We also show that the new method can reliably detect JPEG image blocks which are as small as 8 × 8 pixels and compressed with quality factors as high as 98. This performance is important for analyzing and locating small tampered regions within a composite image. Weiqi Luo 0001, Jiwu Huang, Guoping Qiu |
IEEE Trans. Inf. Forensics Secur. | 3 |
| 2009 | Example-based image super-resolution with class-specific predictors
Kin-Man Lam 0001, Guoping Qiu, Lansun Shen, Suyu Wang |
J. Vis. Commun. Image Represent. | 3 |
| 2009 | Object motion detection using information theoretic spatio-temporal saliency
Chang Liu 0108, Pong C. Yuen, Guoping Qiu |
Pattern Recognit. | 3 |
| 2008 | Learning a discriminative sparse tri-value transformabstractSimple binary patterns have been successfully used for extracting feature representations for visual object classification. In this paper, we present a method to learn a set of discriminative tri-value patterns for projecting high dimensional raw visual inputs into a low dimensional subspace for tasks such as face detection. Unlike previous methods that use predefined simple transform bases to generate tens of thousands features first and then use machine learning to select the most useful features, our method attempts to learn discriminative transform bases directly. Since it would be extremely hard to develop analytical solutions, we define an objective function that can be solved using simulated annealing. To reduce the search space, we impose sparseness and smoothness constraints on the transform bases. Experimental results demonstrate that our method is effective and provides an alternative approach to effective visual object classification. Zhenhua Qu, Guoping Qiu, Pong C. Yuen |
ICPR | 2 |
| 2008 | Collaborative and content-based image labelingabstractMany on-line photo sharing systems allow users to tag their images so as to support semantic image search. In this paper, we study how one can take advantages of the already-tagged images to (semi-)automate the labeling of newly uploaded ones. In particular, we propose a hybrid approach for the prediction where user-provided tags and image visual contents are fused under a unified probabilistic framework. Kernel smoothing and collaborative filtering techniques are explored for improving the accuracy of the probabilistic models estimation. By comparing with some state-of-the-art content-based image labeling methods, we have empirically shown that 1) the proposed method can achieve comparable tag prediction accuracy when there is no user-provided tag, and that 2) it can significantly boost the prediction accuracy if the user can provide just a few tags. William Kwok-Wai Cheung, Xiangyang Xue 0001, Guoping Qiu |
ICPR | 4 |
| 2008 | A Novel Method for Block Size Forensics Based on Morphological Operations
Weiqi Luo 0001, Jiwu Huang, Guoping Qiu |
IWDW | 3 |
| 2008 | Erratum to "Learning to display high dynamic range images": [Pattern Recognition 40 (10) 2641-2655]
Guoping Qiu, Jiang Duan, Graham D. Finlayson |
Pattern Recognit. | 1 |
| 2008 | Classification in an informative sample subspace
Guoping Qiu, Jianzhong Fang |
Pattern Recognit. | 1 |
| 2007 | A Novel Method for Detecting Cropped and Recompressed Image BlockabstractOne of the most common practices in image tampering involves cropping a patch from a source and pasting it onto a target. In this paper, we present a novel method for the detection of such tampering operations in JPEG images. The lossy JPEG compression introduces inherent blocking artifacts into the image and our method exploits such artifacts to serve as a 'watermark' for the detection of image tampering. We develop the blocking artifact characteristics matrix (BACM) and show that, for the original JPEG images, the BACM exhibits regular symmetrical shape; for images that are cropped from another JPEG image and re-saved as JPEG images, the regular symmetrical property of the BACM is destroyed. We fully exploit this property of the BACM and derive representation features from the BACM to train a support vector machine (SVM) classifier for recognizing whether an image is an original JPEG image or it has been cropped from another JPEG image and re-saved as a JPEG image. We present experiment results to show the efficacy of our method. Weiqi Luo 0001, Zhenhua Qu, Jiwu Huang, Guoping Qiu |
ICASSP (2) | 4 |
| 2007 | Display HDR Image using a Gain MapabstractIn this paper, we present a novel method for the display of high dynamic range images. The new method first computes a gain map image using a computational approach inspired by a machine learning algorithm and sums the gain map and the original image together; it then linearly scales the sum image to fit the dynamic range of the display devices. Results are presented to demonstrate the effectiveness of this new method and it is also shown that the new approach is an effective method for enhancing standard (8 bits/pixl) images. Jian Guan 0004, Guoping Qiu |
ICIP (3) | 2 |
| 2007 | An Information Theoretic Model of Spatiotemporal Visual SaliencyabstractThis paper presents a principled and practical method for the computation of visual saliency of spatiotemporal events in full motion videos. Based on the assumption that uniqueness or informative-ness correlates with saliency, our model predicts the saliency of a spatiotemporal event based on the information it contains. To compute the uniqueness of the spatiotemporal events, we model the joint spatial and temporal conditional probability distributions of the spatiotemporal events and compute their spatiotemporal saliencies in a natural and integrated framework. To make the information theoretic model practical, we have developed methods to simplify the model and computational process. Testing results on several video sequences demonstrate that our model is effective in predicting visually salient spatiotemporal events and is comparable to state of the art. It is expected that our principled and practical model will find widespread applications in multimedia content analysis and processing. Guoping Qiu, Xiaodong Gu 0005, Zhibo Chen 0001, Quqing Chen |
ICME | 1 |
| 2007 | A Human Vision System based Flash Picture Coding Method for Video CodingabstractFlash light due to photographing widely appears in video sequences especially in those obtained from news interviews, conferences and sports matches. When a flash picture is encoded, the intensity changes drastically and non-uniformly so that motion estimation can not find a well-matching block in reference pictures. Accordingly, much more bits are generated than the neighboring non-flash pictures. In this paper, a novel flash picture coding method which represents a new concept of video coding is proposed. Based on human vision system (HVS) property, rather than encoding the original flash picture, it encodes an artificial flash picture which is made up of an artificial non-flash (or de-flashed) picture and some parameters for flash effect modeling. Experiments show the proposed flash picture coding method greatly reduces the coded bits of the flash picture while still provides a competitive subjective quality. Quqing Chen, Zhengang Nie, Zhibo Chen 0001, Xiaodong Gu 0005, Guoping Qiu |
ISCAS | 5 |
| 2007 | Improving Video Coding at Scene Cuts using Attention based Adaptive Bit AllocationabstractExisting video coding methods can cause visual quality and buffer occupancy to fluctuate significantly at scene cuts. To address this problem, we have developed a novel visual attention based adaptive bit allocation method. We first perform scene cut detection to extract frames in the vicinities of dramatic scene changes; we then perform visual saliency analysis on those frames to grade the macro-blocks according to their visual importance; and finally we devise a visual attention based adaptive bit allocation scheme which assigns more bits to visually salient blocks and fewer bits to visually less important blocks. We will present experimental results which demonstrate that at scene cut areas, coding quality in terms of PSNR of our method are both higher and much smoother than those of existing coding methods and the buffer occupancy of our method is also much more consistent and has less fluctuation. Our method is compatible with other rate control schemes and can be easily implemented to improve existing video coding standards such as MPEG-2, H.264/AVC, and others. Zhibo Chen 0001, Guoping Qiu, Lihua Zhu, Quqing Chen, Xiaodong Gu 0005 |
ISCAS | 2 |
| 2007 | Special issue on high dynamic range imaging
Guoping Qiu, Erik Reinhard, Graham D. Finlayson |
J. Vis. Commun. Image Represent. | 1 |
| 2007 | Learning to display high dynamic range images
Guoping Qiu, Jiang Duan, Graham D. Finlayson |
Pattern Recognit. | 1 |
| 2007 | Visual guided navigation for image retrieval
Guoping Qiu, Jeremy Morris, Xunli Fan |
Pattern Recognit. | 1 |
| 2006 | Semi-supervised Learning based on Bayesian Networks and Optimization for Interactive Image RetrievalabstractIn this paper, we present a novel interactive image retrieval technique using semi-supervised learning. Recently, Guan and Qiu [8, 9] have shown that by constructing a Bayesian Network where the nodes represent the (continuous) class membership scores and arcs represent the dependence relations of the data points, the (semi-supervised) classification problem can be formulated as a quadratic optimization problem; and by using the labeled data as linear constraints, the optimization problem yields a large, sparse system of linear equations which can be solved very efficiently using standard methods. In this work, we show that this semi-supervised learning method can be naturally adopted as a computational tool to incorporate users feedbacks for interactive image retrieval. We present experimental results to show the effectiveness of our new interactive image retrieval method. We also show that semisupervised learning can have advantages over supervised and unsupervised learning in image retrieval applications. 1 Mai Yang, Jian Guan 0004, Guoping Qiu, Kin-Man Lam 0001 |
BMVC | 3 |
| 2005 | Color by linear neighborhood embeddingabstractThis paper presents a new semiautomatic method for adding colors to monochrome images. The new method, termed color by linear neighborhood embedding (CLNE) first computes the geometry of local image patches in the monochrome image and then transmits the geometry to the chrominance. Following computational methods of the locally linear embedding (LLE) algorithm, we compute the geometry of the monochrome image by solving a constrained least squares problem and embed the computed geometry in the chrominance images by solving a linearly constrained quadratic optimization problem, both in closed forms. We also show that the computed geometry of different color channels are highly correlated thus empirically supporting methods that embed the geometry of the monochrome in the chrominance for colorization. Guoping Qiu, Jian Guan 0004 |
ICIP (3) | 1 |
| 2005 | A New Approach to Estimating Hidden Message Length in Stochastic Modulation Steganography
Jiwu Huang, Guoping Qiu |
IWDW | 3 |
| 2005 | Spectral Images and Features Co-Clustering with Application to Content-based Image RetrievalabstractIn this paper, we present a spectral graph partitioning method for the co-clustering of images and features. We present experimental results, which show that spectral co-clustering has computational advantages over traditional k-means algorithm, especially when the dimensionalities of feature vectors are high. In the context of image clustering, we also show that spectral co-clustering gives better performances. We advocate that the images and features co-clustering framework offers new opportunities for developing advanced image database management technology and illustrate a possible scheme for exploiting the co-clustering results for developing a novel content-based image retrieval method Jian Guan 0004, Guoping Qiu, Xiangyang Xue 0001 |
MMSP | 2 |
| 2005 | A new key frame representation for video segment retrievalabstractIn this paper, we propose an optimal key frame representation scheme based on global statistics for video shot retrieval. Each pixel in this optimal key frame is constructed by considering the probability of occurrence of those pixels at the corresponding pixel position among the frames in a video shot. Therefore, this constructed key frame is called temporally maximum occurrence frame (TMOF), which is an optimal representation of all the frames in a video shot. The retrieval performance of this representation scheme is further improved by considering the k pixel values with the largest probabilities of occurrence and the highest peaks of the probability distribution of occurrence at each pixel position for a video shot. The corresponding schemes are called k-TMOF and k-pTMOF, respectively. These key frame representation schemes are compared to other histogram-based techniques for video shot representation and retrieval. In the experiments, three video sequences in the MPEG-7 content set were used to evaluate the performances of the different key frame representation schemes. Experimental results show that our proposed representations outperform the alpha-trimmed average histogram for video retrieval. Kin-Wai Sze, Kin-Man Lam 0001, Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2004 | Novel histogram processing for colour image enhancementabstractIn practice, the histogram equalization often produces images with unnatural appearances and visually disturbing artefacts. One of the reasons for these unwanted effects is that the histogram equalization attempts to force the output image to have a uniform pixel distribution regardless of what the original image's pixel distribution may be. In this paper, we present a novel histogram processing algorithm which takes into account the original image's pixel distribution in the equalization process. The method uses a single parameter to control the degree of contrast enhancement to ensure that the output have an enhanced appearance which is also faithful to that of the original image and is free of unwanted visually disturbing artefacts. We first develop the algorithm for the luminance channel and then extend the method to the colour components. We present experimental results to demonstrate the better performances of our new method over established methods in the literature. Jiang Duan, Guoping Qiu |
ICIG | 2 |
| 2004 | CVPIC image retrieval based on block colour co-occurance matrix and pattern histogramabstractCompressed domain image processing techniques are becoming increasingly important. Compressed domain retrieval allows the calculation of image features and hence content-based image retrieval (CBIR) to be performed directly on the compressed data without the need for decoding it beforehand. The colour visual pattern image coding (CVPIC) technique represents a compression algorithm where the compressed form is directly meaningful. Based on CVPIC, we introduce a compressed domain retrieval algorithm that makes immediate use of the fact that colour and pattern information is readily available in the CVPIC domain. Colour features are exploited by building a block co-occurance matrix of colour indices while shape information is represented through pattern histograms. Combining these two types of descriptors results in an efficient and effective image retrieval method that even outperforms popular pixel-based algorithms such as colour histograms, colour coherence vectors and colour correlograms. Gerald Schaefer, Simon Lieutaud, Guoping Qiu |
ICIP | 3 |
| 2004 | From sensory coding to scene classificationabstractWe present a storage-efficient and computationally fast method for rapid navigation/browsing through large image repositories and for content-based image retrieval. In the developed system, multiple resolution, orientation achromatic, and opponent chromatic channels are sequentially encoded by a maximal information sensory encoding model, which conveniently and effectively indexes the images into a binary tree data structure. Content-based image retrieval, database navigation, and image browsing are done very efficiently and rapidly by manipulating the n-bit binary keys in the binary tree data structure. We present experimental results to demonstrate the effectiveness of our method. Guoping Qiu, Jeremy Morris, Xunli Fan |
MMSP | 1 |
| 2004 | Embedded colour image coding for content-based retrieval
Guoping Qiu |
J. Vis. Commun. Image Represent. | 1 |
| 2004 | Compressing histogram representations for automatic colour photo categorization
Guoping Qiu, Xia Feng, Jianzhong Fang |
Pattern Recognit. | 1 |
| 2003 | Face Detection Based on Multiple Regression and Recognition Support Vector MachinesabstractThis paper presents a novel approach to face detection. A potential face pattern is first filtered by a Gaussian derivative filter bank to generate a set of derivative images, which are then transformed by the Angular Radial Transform (ART) to form a compact set of representation feature vectors. Using these feature vectors for face detection is based on a two level multiple support vector machines (SVMs) strategy. At the first level, a separate SVM is trained for each derivative image to indicate the presence/absence of a face in the input based on the features from that derivative image alone. These SVMs are trained as binary classifiers but used for regression in the sense that they output continuous values. At the second level, a single SVM takes the outputs of the first level SVMs as input to make the overall and final decision to determine whether the current input is a face or nonface pattern. Experimental results are presented to demonstrate the effectiveness of the method. 1 Jianzhong Fang, Guoping Qiu |
BMVC | 2 |
| 2003 | Appearance indexingabstractAlthough it is very hard to quantify the visual impression of an image, we conjecture that the overall visual appearance of an image may be the combined effect of the appearances of image patches of various shapes and sizes. We can imagine that all possible appearances of image patches of all possible shapes and sizes form a conceptual appearance space. Each point in the appearance space therefore corresponds to a certain combination of the object parameters (shape, size, surface texture, pose, orientation, etc) and the imaging conditions (illuminating source, viewing angle, sensor response etc.). Using real sensor data and unsupervised learning, statistically most representative appearance prototypes can be found to approximate the appearance space. Statistics of these appearance prototypes present in an image therefore characterize the image content, which in turn can be used to perform tasks such as content-based image retrieval. Guoping Qiu |
ICASSP (3) | 1 |
| 2003 | Human face detection using angular radial transform and support vector machinesabstractThis paper presents a new face detection method. For a potential face pattern, a histogram equalized intensity map and a local intensity variance map are created to normalize the pattern photometrically. We then view these two maps as geometric shapes and apply the angular radial transform (ART) to derive a compact representation of the pattern. The ART transform coefficients are then used as input to a support vector machine (SVM) to determine the presence or absence of a face in the pattern. We also develop a SVM based skin color detection technique as a preprocessing step and only search image regions that contain sufficient large number of skin pixels thus greatly enhancing the detection speed. Experimental results are presented to demonstrate the effectiveness of the new method. Jianzhong Fang, Guoping Qiu |
ICIP (1) | 2 |
| 2003 | Color photo categorization using compressed histograms and support vector machinesabstractIn this paper, an efficient method using various histogram-based (high-dimensional) image content descriptors for automatically classifying general color photos into relevant categories is presented. Principal component analysis (PCA) is used to project the original high dimensional histograms onto their eigenspaces. Lower dimensional eigenfeatures are then used to train support vector machines (SVMs) to classify images into their categories. Experimental results show that even though different descriptors perform differently, they are all highly redundant. It is shown that the dimensionality of all these descriptors, regardless of their performances, can be significantly reduced without affecting classification accuracy. Such scheme would be useful when it is used in an interactive setting for relevant feedback in content-based image retrieval, where low dimensional content descriptors enable fast online learning and reclassification of results. Xia Feng, Jianzhong Fang, Guoping Qiu |
ICIP (3) | 3 |
| 2003 | Scene cut detection using the colored pattern appearance modelabstractIn this paper, we propose to use the colored pattern appearance model (CPAM) as a content representation for video scene break detection. This model represents a scene by means of global statistics of the local visual appearance, and was originally motivated by studies in human color vision. The performance of this method is compared to several histogram-based approaches. An adaptive thresholding technique, namely entropic thresholding, is applied to determine the respective optimal threshold values for each of the approaches. In the experiments, the two video sequences in the MPEG-7 content set are used to evaluate the performances of the CPAM and the histogram-based methods. Experimental results show that our proposed model outperforms other histogram-based approaches in scene break detection. Kin-Wai Sze, Kin-Man Lam 0001, Guoping Qiu |
ICIP (2) | 3 |
| 2003 | Color image indexing using BTCabstractThis paper presents a new application of a well-studied image coding technique, namely block truncation coding (BTC). It is shown that BTC can not only be used for compressing color images, it can also be conveniently used for content-based image retrieval from image databases. From the BTC compressed stream (without performing decoding), we derive two image content description features, one termed the block color co-occurrence matrix (BCCM) and the other block pattern histogram (BPH). We use BCCM and BPH to compute the similarity measures of images for content-based image retrieval applications. Experimental results are presented which demonstrate that BCCM and BPH are comparable to similar state of the art techniques. Guoping Qiu |
IEEE Trans. Image Process. | 1 |
| 2003 | Frequency layered color indexing for content-based image retrievalabstractImage patches of different spatial frequencies are likely to have different perceptual significance as well as reflect different physical properties. Incorporating such concept is helpful to the development of more effective image retrieval techniques. We introduce a method which separates an image into layers, each of which retains only pixels in areas with similar spatial frequency characteristics and uses simple low-level features to index the layers individually. The scheme associates indexing features with perceptual and physical significance thus implicitly incorporating high level knowledge into low level features. We present a computationally efficient implementation of the method, which enhances the power and at the same time retains the simplicity and elegance of basic color indexing. Experimental results are presented to demonstrate the effectiveness of the method. Guoping Qiu, Kin-Man Lam 0001 |
IEEE Trans. Image Process. | 1 |
| 2002 | Indexing chromatic and achromatic patterns for content-based colour image retrieval
Guoping Qiu |
Pattern Recognit. | 1 |
| 2001 | Detection Algorithm of Particle Contamination in Reticle Images with Continuous Wavelet TransformabstractThis paper presents an inspection method of particle contamination for semiconductor reticles using continuous wavelet transform. Particle defect is considered as a singularity in the reticle image, and wavelet transform is applied to detect such an event. By taking the local maxima of wavelet transform as suspected defects, the candidate pixels under inspection reduce to a small fraction of the whole image. From the evolution of wavelet coefficients across scales, two features are extracted to identify detects among the suspected defects, Lipschitz Exponent (L.E.) and smoothing factor. With the rules of LE and smoothing factor trained from the image database, the defects can be detected with high accuracy. From the test of synthetic reticle images with defects at different size, it is concluded that the maximum scale of wavelet under which the defect is still visible can not be over about 8 times as much as the defect size. The simulation of a real reticle image and synthetic defects shows the effectiveness of the present method. Chaoquan Chen, Guoping Qiu |
BMVC | 2 |
| 2001 | Spectral and Spatial Invariant Image Retrieval using Scene Structural MatrixabstractWe introduce Scene Structural Matrix (SSM), a novel image content descriptor and its application to invariant image retrieval. The SSM captures the overall structural characteristics of the scene by indexing the geometric features of the image. We employ a binary image tree (bintree) to partition the image and from which we derive multiscale geometric structural descriptors of the image. We have applied the SSM to contentbased image retrieval from image databases. Experimental results show that SSM is particularly effective in retrieving images with strong structural features, such as landscape photographs. We will show that SSM is robust against spatial and spectral distortions thus making it superior to current state of the art techniques such as colour correlogram in certain applications. We will also show that images retrieved by the SSM are more relevant than those returned by colour correlogram and colour histogram. 1. Guoping Qiu, S. Sudirman |
BMVC | 1 |
| 2001 | Color image coding, indexing and retrieval using binary space partitioning treeabstractThis paper presents a unified approach to colour image coding, content-based indexing, and retrieval for database applications. The binary space partitioning (BSP) tree, traditionally used in gray scale image coding (Wu 1992, Radha et al. 1996) is extended to represent colour images. The BSP tree, hence the structure, of the image is explicitly coded. A method is developed to compute the similarities of images based on their BSP tree representations. In image database applications, the images in the database are coded by BSP tree to achieve a good balance between storage efficiency and easy manipulation of image data. Content-based image querying is performed in the compressed bit streams by comparing the BSP tree of the query image with those of the images in the database. Guoping Qiu, S. Sudirman |
ICIP (1) | 1 |
| 2001 | Constraint adaptive segmentation for color image coding and content-based retrievalabstractWe present a constraint adaptive image segmentation technique designed to achieve the combined purposes of color image coding/compression, indexing and content-based retrieval. An image is segmented into homogeneous squared regions of variable sizes such that it can be encoded very efficiently. From the segmented image, we derive an effective and efficient image content description feature termed the region and color co-occurrence matrix (RACOM) as image index for content-based indexing and retrieval. Experimental results are presented to demonstrate that RACOM is a very effective image content description feature and has comparable performance to state of the art methods, colour correlogram and MPEG7 colour structure histogram, in content-based image retrieval from large image databases. Guoping Qiu |
MMSP | 1 |
| 2001 | Scene structural matrix for image indexing and retrievalabstractWe present a new image indexing and retrieval method for image database applications. We introduce a novel image content description feature termed scene structural matrix (SSM). The SSM captures the overall structural information of the scene by indexing the geometric features of the image. We also introduce a constraint adaptive image segmentation method termed middle cut (MC) to partition the image and derive the multi-scale geometric structures of the image. Experimental results show that SSM is particularly effective in retrieving images with strong structural features, such as landscape photographs. It is also shown that SSM is robust against spatial and spectral distortions thus making it superior to the traditional colour histogram, colour correlogram and MPEG7-color structure histogram in certain applications. Guoping Qiu, S. Sudirman |
MMSP | 1 |
| 2001 | Image coding using a colored pattern appearance model
Guoping Qiu |
VCIP | 1 |
| 2000 | Interresolution Look-up Table for Improved Spatial Magnification of Image
Guoping Qiu |
J. Vis. Commun. Image Represent. | 1 |
| 2000 | MLP for adaptive postprocessing block-coded imagesabstractA new technique based on the multilayer perceptron (MLP) neural network is proposed for blocking-artifact removal in block-coded images. The new method is based on the concept of learning-by-examples. The compressed image and its original uncompressed version are used to train the neural networks. In the developed scheme, inter-block slopes of the compressed image are used as input, the difference between the original uncompressed and the compressed image is used as desired output for training the networks. Blocking-artifact removal is realized by adding the neural network's outputs to the compressed image. The new technique has been applied to process JPEG compressed images. Experimental results show significant improvements in both visual quality and peak signal-to-noise ratio. It is also shown the present method is comparable to other state of the art techniques for quality enhancement in block-coded images. Guoping Qiu |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 1999 | Pattern colour separable image coding for content based indexingabstractA colour image coding system based on the pattern colour separable (PCS) theory of human colour vision is developed. The PCS image coding system codes an image using three separate channels: pattern, colour and strength. The pattern channel corresponds to the visual pathway that is sensitive to spatial frequency information, the colour channel corresponds to the visual pathway that is sensitive to the colour information and the strength channel corresponds to the visual pathway that is sensitive to the strength of the visual stimulus. Decoding of the image is based on a neural network trained in a self-organizing manner. Compression performances of the system are shown to be very good. Query by image contents is a successful approach to image database indexing and retrieval. Because image data take up huge amount of storage space and network bandwidth, compressing the image is a standard practice. Given that images are compressed/coded in a database, it mould be more efficient if the compressed data can also be directly used as indices for content-based query. We show that the contents (shapes and colours) of the images are embedded in the compressed data of the PCS system and therefore can be readily and efficiently used as indices for content based image indexing. Guoping Qiu |
MMSP | 1 |
| 1999 | A progressively predictive image pyramid for efficient lossless codingabstractA low entropy pyramidal image data structure suited for lossless coding and progressive transmission is proposed in this work. The new coder, called the progressively predictive pyramid (PPP) is based on the well-known Laplacian pyramid. By introducing inter-resolution predictors into the original Laplacian pyramid, we show that the entropy level in the original pyramid can be reduced significantly. To take full advantage of progressive transmission, a scheme is introduced to create the predictor adaptively, thus eliminating the need to transmit the predictor and reducing the coding overheads. A method for designing the predictor is presented. Numerical results show that PPP is superior to traditional approaches to pyramid generation in the sense that the pyramids generated by PPP always have significantly lower entropy values. Guoping Qiu |
IEEE Trans. Image Process. | 1 |
| 1998 | Multi-grid Edge Models for Magnifying Digital Images
Guoping Qiu |
ACCV (2) | 1 |
| 1994 | Improved clustering using deterministic annealing with a gradient descent technique
Guoping Qiu, Martin R. Varley, Trevor J. Terrell |
Pattern Recognit. Lett. | 1 |
| 1994 | Functional optimization properties of median filteringabstractIn this letter, we use a new approach for studying the properties of median filtering. Specifically, using threshold decomposition, it is shown that median filtering operation minimizes a two-term cost function of the output state of the median filter. The first term of the cost function measures the smoothness between the median filter output and its neighbor points within the operation window, and the second term measures the discrepancy between the filter output and its original signal. The results from this study have provided us with a new tool to analyze and understand some of the properties of the median filtering operation, including weighted median filtering.> Guoping Qiu |
IEEE Signal Process. Lett. | 1 |