VLDB 2026 Research / reviewers in the wild / expert
Paul L. Rosin
dblp:r/PaulLRosin
· DBLP profile ↗
249ranked-venue papers
75as first author
55since 2021 · last 2026
0000-0002-4965-3884ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 154 · 33 first-author · 37 since 2021Artificial intelligence and machine learning · 129 · 57 first-author · 21 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Human-computer interaction and ubiquitous computing · 7 · 2 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Temporal Inconsistency Guidance for Super-resolution Video Quality AssessmentabstractAs super-resolution (SR) techniques introduce unique distortions that fundamentally differ from those caused by traditional degradation processes (e.g., compression), there is an increasing demand for specialized video quality assessment (VQA) methods tailored to SR-generated content. One critical factor affecting perceived quality is temporal inconsistency, which refers to irregularities between consecutive frames. However, existing VQA approaches rarely quantify this phenomenon or explicitly investigate its relationship with human perception. Moreover, SR videos exhibit amplified inconsistency levels as a result of enhancement processes. In this paper, we propose Temporal Inconsistency Guidance for Super-resolution Video Quality Assessment (TIG-SVQA) that underscores the critical role of temporal inconsistency in guiding the quality assessment of SR videos. We first design a perception-oriented approach to quantify frame-wise temporal inconsistency. Based on this, we introduce the Inconsistency Highlighted Spatial Module, which localizes inconsistent regions at both coarse and fine scales. Inspired by the human visual system, we further develop an Inconsistency Guided Temporal Module that performs progressive temporal feature aggregation: (1) a consistency-aware fusion stage in which a visual memory capacity block adaptively determines the information load of each temporal segment based on inconsistency levels, and (2) an informative filtering stage for emphasizing quality-related features. Extensive experiments on both single-frame and multi-frame SR video scenarios demonstrate that our method significantly outperforms state-of-the-art VQA approaches. Xiaoyuan Yang 0003, Weide Liu, Xin Jin 0014, Xu Jia 0012, Yukun Lai, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
AAAI | 7 |
| 2026 | Cross-Modal Interaction for Multi-Dimensional AI-Generated Image Quality Assessment
Minghao Zou, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
QoMEX | 3 |
| 2026 | Line Art Colorization with Offset Prior-based Diffusion ModelabstractReference-based line art video colorization colorizes the target line art according to reference images, which is an essential stage for the cartoon production workflow. However, the manual colorization process is time-consuming and repetitive, making automatic video colorization highly desirable. Existing cartoon colorization methods struggle with domain misalignment between the reference and line art images and the loss of details caused by compression into a low-dimensional space in the existing video diffusion models, reducing colorization quality. In this paper, we propose an Offset Prior-based Diffusion Model (OPDM) for cartoon video colorization, which utilizes the powerful generation capability of the diffusion model and cross-domain matching priors to generate high-quality colorization results. Specifically, we design a simple and effective Offset-Adapter that leverages the idea of sampling offsets in deformable convolution to estimate the cross-domain spatial offset features between the target line arts and reference images. We further introduce a new training strategy that combines forward diffusion and reverse denoising in the training stage to ensure content consistency. Experiments on a public cartoon dataset and our newly constructed long cartoon video dataset demonstrate that our proposed method outperforms the existing state-of-the-art line art coloring methods. Code is available at https://github.com/xzh976/OPDM. Yukun Lai, Paul L. Rosin |
WACV | 5 |
| 2026 | Foreword to the Special Section on CAD/Graphics 2023
Joaquim Jorge 0001, Shi-Min Hu 0001, Paul L. Rosin, Yiyu Cai |
Comput. Graph. | 3 |
| 2026 | EHIN: Early-aware hierarchical interaction network for weakly-supervised referring image segmentation
Anqing Chen, Wanli Ma 0001, Weide Liu, Yakun Ju, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
Neurocomputing | 8 |
| 2026 | FaceEditor: Text-driven and mask-constrained face attribute editing
Lin Zhang 0041, Weiliang Meng, Paul L. Rosin, Yukun Lai, Yaonan Wang 0001 |
Pattern Recognit. | 5 |
| 2026 | P3C-DNet: Pseudo-Groundtruth Contrastive Learning With Color Calibration Dehazing NetworkabstractExisting dehazing methods primarily rely on synthetic hazy images for supervised learning. While effective on synthetic datasets, these methods often struggle to generalize to real-world hazy images, leading to issues such as color distortion and incomplete haze removal. Moreover, their limited adaptability to real-world datasets and inability to handle complex haze scenarios remain significant challenges. To address these limitations, we propose a novel unsupervised framework P3C-DNet (Pseudo-groundtruth Contrastive learning with Color Calibration Dehazing Network). Our P3C-DNet introduces a Pseudo-groundtruth image generation strategy through the Pseudo-groundtruth Contrastive Supervision (PCS) module, which overcomes the lack of real haze-free training data by generating high-quality Pseudo-groundtruth images. To further refine the dehazing process, we incorporate a codebook-based image coding and matching mechanism that aligns Pseudo-groundtruth images with hazy inputs, enhancing the accuracy and detail of the dehazed outputs. To address the prevalent issue of color distortion, especially in complex environments, our P3C-DNet integrates a Dynamic Color Restoration Block (DCRB) to ensure visual quality and color consistency in the dehazed results. Experimental evaluations demonstrate that our P3C-DNet achieves superior performance in haze removal, color fidelity, and detail preservation, significantly outperforming existing methods and setting a new benchmark for real-world dehazing tasks. Ze Ouyang, Weiliang Meng, Paul L. Rosin, Yukun Lai, Yaonan Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2026 | CLIP-Hand: CLIP-based regressor for hand pose estimation and mesh recovery
Feng Zhou 0007, Shuang Ji, Pei Shen, Ju Dai, JunJun Pan, Yukun Lai, Paul L. Rosin |
Vis. Comput. | 7 |
| 2025 | Chat-Driven 3D Human Pose and Shape Editing with Large Language ModelsabstractGenerating and creating humanoid 3D models has received increasing attention recently due to its fundamental support for many high-level 3D applications. Although automatic 3D pose and shape reconstruction methods have achieved promising results, there are still some failure cases due to self-occlusions, viewpoint changes, and the complexity of human pose articulations. In this paper, we propose a novel way to leverage Large Language Models (LLMs) to interactively reconstruct human pose and shape based on a Skinned Multi-Person Linear (SMPL) model. We construct a mapping table to fine-tune an LLM, enabling it to understand user inputs better and output the positional information of joint points. Additionally, a simple neural network is adopted to regress the shape cues of the SMPL. We demonstrate a gallery of results of numerous poses and shapes. We validate our method via numerical evaluations, user studies, and comparisons to manually posed characters and previous work. Feng Zhou 0007, Ju Dai, Mengxiao Zhu 0004, Yongmei Zhang, Yukun Lai, Paul L. Rosin |
ICASSP | 7 |
| 2025 | LGA-Net: Learning Local and Global Affinities for Sparse Scribble Based Image Colorization
Hongjin Lyu, Bo Li 0023, Paul L. Rosin, Yukun Lai |
ICCV | 3 |
| 2025 | MCHM25: Multimedia Computing for Health and MedicineabstractRecent years have witnessed an unprecedented growth of multimodal data in healthcare, ranging from distributed sensors and medical imaging devices (MRI, CT, X-rays) to digital health platforms that integrate audio, video, 3D geometry, and clinical text. The increasing availability of such data presents significant opportunities for computer-aided diagnosis and intelligent healthcare solutions, yet also poses substantial challenges in multimodal integration, large-scale analysis, and real-world deployment. The 2nd International Workshop on Multimedia Computing for Health and Medicine (MCHM'25), held in conjunction with ACM Multimedia 2025, focuses on advanced multimedia computing techniques, including mobile and hardware solutions, for tackling real-world problems in healthcare. The workshop brings together researchers and practitioners in multimedia computing, artificial intelligence, and medicine to explore emerging methods, applications, and systems that have a direct impact on human health. Wei Zhou 0021, Hadi Amirpour, Li Yu 0004, Jungong Han, Richang Hong, Paul L. Rosin |
ACM Multimedia | 6 |
| 2025 | CCDb+: Enhanced Annotations and Multi-Modal Benchmark for Natural Dyadic ConversationsabstractBackchannel signals play a critical role in social interaction, expressing attentiveness, agreement, and emotion in both human and human-agent conversations. However, few multi-modal databases exist in this area due to the complexity of categorisation and the high cost of precise timing, especially in naturalistic dyadic conversations. To address these challenges, we introduce CCDb+ (Cardiff Conversation Database +) an enhanced version of CCDb, with 25 newly annotated conversations and corrections to 14 previously annotated conversations, along with thorough consistency checks to ensure annotation reliability. Additionally, we propose a multi-modal process for backchannel detection as a baseline, showing that both visual and acoustic cues contribute significantly to understanding backchannel behaviour. Recognising that backchannel signals often intersect with other social cues, we introduce several detection sub-tasks-such as smile, nodding, and agreement-with baseline results for each. Finally, we demonstrate multi-modal paradigms for nuanced signals like nodding and thinking. The database and associated annotations are publicly available at https://huggingface.co/datasets/CardiffVisualComputing/CCDb. Yukun Lai, Paul L. Rosin |
ACM Multimedia | 3 |
| 2025 | Canonical pose reconstruction from single depth image for 3D non-rigid pose recovery on limited datasets
Fahd Alhamazani, Paul L. Rosin, Yukun Lai |
Comput. Graph. | 2 |
| 2025 | A three-level benchmark dataset for spatial and temporal forensic analysis of videos
Naheed Akhtar, Mubbashar Saddique, Paul L. Rosin, Xianfang Sun, Muhammad Hussain 0001, Zulfiqar Habib |
Mach. Vis. Appl. | 3 |
| 2025 | Image Manipulation Quality AssessmentabstractImage quality assessment (IQA) and its computational models play a vital role in modern computer vision applications. Research has traditionally focused on signal distortions arising during image compression and transmission, and their impact on perceived image quality. However, little attention is paid to image manipulation that alters an image using various filters. With the prevalence of image manipulation in real-life scenarios, it is critical to understand how humans perceive filter-altered images and to develop reliable IQA models capable of automatically assessing the quality of filtered images. In this paper, we build a new IQA database for filter-altered images, comprised of 360 images manipulated by various filters. To ensure the subjective IQA faithfully reflects human visual perception, we conduct a fully-controlled psychovisual experiment. Building upon the ground truth, we propose an innovative deep learning-based no-reference IQA (NR-IQA) model named IMQA that can accurately predict the perceived quality of filter-altered images. This model involves constructing an image filtering-aware module to learn discriminatory features for filter-altered images; and fuses these features with the representations generated by an image quality-aware module. Experimental results demonstrate the superior performance of the proposed IMQA model. Xinbo Wu, Jianxun Lou, Wan'an Liu, Paul L. Rosin, Gualtiero Colombo 0001, Stuart M. Allen, Roger M. Whitaker, Hantao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Perception-Oriented Bidirectional Attention Network for Image Super-Resolution Quality AssessmentabstractMany super-resolution (SR) algorithms have been proposed to increase image resolution. However, full-reference (FR) image quality assessment (IQA) metrics for comparing and evaluating different SR algorithms are limited. In this work, we propose the Perception-oriented Bidirectional Attention Network (PBAN) for image SR FR-IQA, which is composed of three modules: an image encoder module, a perception-oriented bidirectional attention (PBA) module, and a quality prediction module. First, we encode the input images for feature representations. Inspired by the characteristics of the human visual system, we then construct the perception-oriented PBA module. Specifically, different from existing attention-based SR IQA methods, we conceive a Bidirectional Attention to bidirectionally construct visual attention to distortion, which is consistent with the generation and evaluation processes of SR images. To further guide the quality assessment towards the perception of distorted information, we propose Grouped Multi-scale Deformable Convolution, enabling the proposed method to adaptively perceive distortion. Moreover, we design Sub-information Excitation Convolution to direct visual perception to both sub-pixel and sub-channel attention. Finally, the quality prediction module is exploited to integrate quality-aware features and regress quality scores. Extensive experiments demonstrate that our proposed PBAN outperforms state-of-the-art quality assessment methods. Xiaoyuan Yang 0003, Guanghui Yue 0001, Jun Fu 0007, Qiuping Jiang, Xu Jia 0012, Paul L. Rosin, Hantao Liu, Wei Zhou 0021 |
IEEE Trans. Image Process. | 7 |
| 2025 | Distortion-Induced Saliency Shifts in VideoabstractVisual saliency modelling is of fundamental importance in modern video processing and its applications. Our previous eye-tracking study revealed that signal distortions caused by editing, compression, or transmission alter gaze patterns and consequently induce saliency shifts in both spatial and temporal domains. Saliency shifts provide crucial insights into viewers’ behavioural responses to video distortions, facilitating the perception-based optimisation of video algorithms. However, the spatio-temporal saliency shifts and their measurable effects on perception related applications remain largely unexplored. In this paper, we first investigate the measurement of distortion-induced saliency shifts (DSS) in videos and analyse DSS behaviours as functions of video content, time order and critical distortion disruption. Second, based on our findings, we construct three vision models to quantitatively simulate distinct DSS behaviours and integrate them into a comprehensive DSS behaviour model. Finally, we demonstrate that the computational DSS model can enhance emerging video technologies. Xinbo Wu, Jianxun Lou, Zhengyan Dong, Fan Zhang 0017, Paul L. Rosin, Hantao Liu |
IEEE Trans. Multim. | 5 |
| 2025 | Asymmetric Information Enhanced Mapping Framework for Multirobot Exploration Based on Deep Reinforcement LearningabstractDespite significant advancements in multirobot technologies, efficiently and collaboratively exploring an unknown environment remains a major challenge. In this paper, we propose AIM-Mapping, an Asymmetric InforMation enhanced Mapping framework based on deep reinforcement learning. The framework fully leverages the privileged information to help construct the environmental representation as well as the supervised signal in an asymmetric actor-critic training framework. Specifically, privileged information is used to evaluate exploration performance through an asymmetric feature representation module and a mutual information evaluation module. The decision-making network employs the trained feature encoder to extract structural information of the environment and integrates it with a topological map constructed based on geometric distance. By leveraging this topological map representation, we apply topological graph matching to assign corresponding boundary points to each robot as long-term goal points. We conduct experiments in both iGibson simulation environments and real-world scenarios. The results demonstrate that the proposed method achieves significant performance improvements compared to existing approaches. Jiyu Cheng, Junhui Fan, Xiaolei Li 0003, Paul L. Rosin, Yibin Li 0001, Wei Zhang 0021 |
IEEE Trans. Robotics | 4 |
| 2025 | AttentionPainter: An Efficient and Adaptive Stroke Predictor for Scene PaintingabstractStroke-based Rendering (SBR) aims to decompose an input image into a sequence of parameterized strokes, which can be rendered into a painting that resembles the input image. Recently, Neural Painting methods that utilize deep learning and reinforcement learning models to predict the stroke sequences have been developed, but suffer from longer inference time or unstable training. To address these issues, we propose AttentionPainter, an efficient and adaptive model for single-step neural painting. First, we propose a novel scalable stroke predictor, which predicts a large number of stroke parameters within a single forward process, instead of the iterative prediction of previous Reinforcement Learning or auto-regressive methods, which makes AttentionPainter faster than previous neural painting methods. To further increase the training efficiency, we propose a Fast Stroke Stacking algorithm, which brings 13 times acceleration for training. Moreover, we propose Stroke-density Loss, which encourages the model to use small strokes for detailed information, to help improve the reconstruction quality. Finally, we design a Stroke Diffusion Model as an application of AttentionPainter, which conducts the denoising process in the stroke parameter space and facilitates stroke-based inpainting and editing applications helpful for human artists' design. Extensive experiments show that AttentionPainter outperforms the state-of-the-art neural painting methods. Yizhe Tang, Yue Wang 0020, Ran Yi 0002, Xin Tan 0002, Lizhuang Ma, Yukun Lai, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 8 |
| 2025 | Stacked deep fusion GAN for enhanced text-to-image generation
Yaqi Sun, Paul L. Rosin, Yukun Lai |
Vis. Comput. | 3 |
| 2024 | SCD: Statistical Color Distribution-Based Objective Image Colorization Quality Assessment
Hongjin Lyu, Hareeharan Elangovan, Paul L. Rosin, Yukun Lai |
CGI (1) | 3 |
| 2024 | SuperSVG: Superpixel-Based Scalable Vector Graphics SynthesisabstractSVG (Scalable Vector Graphics) is a widely used graphics format that possesses excellent scalability and editability. Image vectorization, which aims to convert raster images to SVGs, is an important yet challenging problem in computer vision and graphics. Existing image vectorization methods either suffer from low reconstruction accuracy for complex images or require long computation time. To address this issue, we propose SuperSVG, a superpixel-based vectorization model that achieves fast and high-precision image vectorization. Specifically, we decompose the input image into superpixels to help the model focus on areas with similar colors and textures. Then, we propose a two-stage self-training framework, where a coarse-stage model is employed to reconstruct the main structure and a refinement-stage model is used for enriching the details. Moreover, we propose a novel dynamic path warping loss to help the refinement-stage model to inherit knowledge from the coarse-stage model. Extensive qualitative and quantitative experiments demonstrate the superior performance of our method in terms of reconstruction accuracy and inference time compared to state-of-the-art approaches. The code is available in https://github.com/sjtuplayer/SuperSVG. Ran Yi 0002, Baihong Qian, Jiangning Zhang, Paul L. Rosin, Yukun Lai |
CVPR | 5 |
| 2024 | Deep Generative Model based Rate-Distortion for Image Downscaling AssessmentabstractIn this paper, we propose Image Downscaling Assessment by Rate-Distortion (IDA-RD), a novel measure to quantitatively evaluate image downscaling algorithms. In contrast to image-based methods that measure the quality of downscaled images, ours is process-based that draws ideas from rate-distortion theory to measure the distortion incurred during downscaling. Our main idea is that downscaling and super-resolution (SR) can be viewed as the encoding and decoding processes in the rate-distortion model, respectively, and that a downscaling algorithm that preserves more details in the resulting low-resolution (LR) images should lead to less dis-torted high-resolution (HR) images in SR. In other words, the distortion should increase as the downscaling algorithm deteriorates. However, it is non-trivial to measure this distortion as it requires the SR algorithm to be blind and stochastic. Our key insight is that such requirements can be met by re-cent SR algorithms based on deep generative models that can find all matching HR images for a given LR image on their learned manifolds. Extensive experimental results show the effectiveness of our IDA-RD measure. Our code is available at: https://github.com/Byronliang8/Ida-Rd Yuanbang Liang, Bhavesh Garg, Paul L. Rosin, Yipeng Qin |
CVPR | 3 |
| 2024 | AHRNET: Attention and Heatmap-Based Regressor for Hand Pose Estimation and Mesh RecoveryabstractEstimating 3D hand pose and recovering the full hand surface mesh from a single RGB image is a challenging task due to self-occlusions, viewpoint changes, and the complexity of hand articulations. In this paper, we propose a novel framework that combines an attention mechanism with heatmap regression to accurately and efficiently predict 3D joint locations and reconstruct the hand mesh. We adopt a pooling attention module that learns to focus on relevant regions in the input image to extract better features for handling occlusions, while greatly reducing the computational cost. The multi-scale 2D heatmaps provide spatial constraints to guide the 3D vertex predictions. By exploiting the complementary strengths of sparse 2D supervision and dense mesh regression, our method accurately reconstructs hand meshes with realistic details. Extensive experiments on standard benchmarks demonstrate that the proposed method efficiently improves the performance of 3D hand pose estimation and mesh recovery. The reproducible recipes are available at https://github.com/SDiannn/AHRNET-Heatmap. Feng Zhou 0007, Pei Shen, Ju Dai, Yukun Lai, Paul L. Rosin |
ICASSP | 7 |
| 2024 | SAMVG: A Multi-Stage Image Vectorization Model with the Segment-Anything ModelabstractVector graphics are widely used in graphical designs and have received more and more attention. However, unlike raster images which can be easily obtained, acquiring high-quality vector graphics, typically through automatically converting from raster images, remains a significant challenge, especially for more complex images such as photos or artworks. In this paper, we propose SAMVG, a multi-stage model to vectorize raster images into SVG (Scalable Vector Graphics). Firstly, SAMVG uses general image segmentation provided by the Segment-Anything Model and uses a novel filtering method to identify the best dense segmentation map for the entire image. Secondly, SAMVG then identifies missing components and adds more detailed components to the SVG. Through a series of extensive experiments, we demonstrate that SAMVG can produce high quality SVGs in any domain while requiring less computation time and complexity compared to previous state-of-the-art methods. Haokun Zhu, Juang Ian Chong, Ran Yi 0002, Yukun Lai, Paul L. Rosin |
ICASSP | 6 |
| 2024 | Burnsnet: Burn Region Segmentation Network From Color Images With Two-Way CNNabstractBurn injury is a serious health issue leading to several thousands of annual fatalities. The color image-based automated burns diagnostic and assessment methods hold the potential for timely diagnosis and treatment. However, the research is limited in this domain which remains a major challenge. In this work, we explore and address the complex task of burn region segmentation in color images of burn patients. We present a semantic segmentation network that has two parallel sub-networks: a spatial-stream network for extracting low-level features and a contextual-stream network for generating a larger receptive field. Our network utilizes the pre-trained ResNet101 network, global average pooling, and instance normalization for better encoding and fusion of the network outputs. This dual-stream approach optimizes the performance in situations where data scarcity poses a challenge, facilitating robust semantic segmentation despite limited training samples. We prepared a pixel-wise labeled dataset for burn region segmentation and the experimental results on this dataset show that our proposed network outperforms several state-of-the-art semantic segmentation methods. Our method achieved mIOU and Matthews’ correlation coefficient (MCC) of 74.3% and 81.7%, respectively, approximately 4.5% higher than the second-best performing method. The Extended Burn Image Segmentation (EBIS) dataset and our model are available at https://github.com/VEDAs-Lab/EBIS Joohi Chauhan, Paul L. Rosin, Puneet Goyal |
ICIP | 2 |
| 2024 | Knowledge Distillation for Road Detection Based on Cross-Model Semi-Supervised LearningabstractThe advancement of knowledge distillation has played a crucial role in enabling the transfer of knowledge from larger teacher models to smaller and more efficient student models, and is particularly beneficial for online and resource-constrained applications. The effectiveness of the student model heavily relies on the quality of the distilled knowledge received from the teacher. Given the accessibility of unlabelled remote sensing data, semi-supervised learning has become a prevalent strategy for enhancing model performance. However, relying solely on semi-supervised learning with smaller models may be insufficient due to their limited capacity for feature extraction. This limitation restricts their ability to exploit training data. To address this issue, we propose an integrated approach that combines knowledge distillation and semi-supervised learning methods. This hybrid approach leverages the robust capabilities of large models to effectively utilise large unlabelled data whilst subsequently providing the small student model with rich and informative features for enhancement. The proposed semi-supervised learning-based knowledge distillation (SSLKD) approach demonstrates a notable improvement in the performance of the student model, in the application of road segmentation surpassing the effectiveness of traditional semi-supervised learning methods. Wanli Ma 0001, Oktay Karakus, Paul L. Rosin |
IGARSS | 3 |
| 2024 | AesStyler: Aesthetic Guided Universal Style TransferabstractRecent studies have shown impressive progress in universal style transfer which can integrate arbitrary styles into content images. However, existing approaches struggle with low aesthetics and disharmonious patterns in the final results. To address this problem, we propose AesStyler, a novel Aesthetic Guided Universal Style Transfer method. Specifically, our approach introduces the aesthetic assessment model, trained on a dataset with human-assessed aesthetic scores, into the universal style transfer task to accurately capture aesthetic features that universally resonate with human aesthetic preferences. Unlike previous methods which only consider aesthetics of specific style images, we propose to build a Universal Aesthetic Codebook (UAC) to harness universal aesthetic features that encapsulate the global aspects of aesthetics. Aesthetic features are fed into a novel Universal and Style-specific Aesthetic-Guided Attention (USAesA) module to guide the style transfer process. USAesA empowers our model to integrate the aesthetic attributes of both universal and style-specific aesthetic features with style features and facilitates the fusion of these aesthetically enhanced style features with content features. Extensive experiments and user studies have demonstrated that our approach generates aesthetically more harmonious and pleasing results than the state-ofthe- art methods, both aesthetic-free and aesthetic-aware. The code is available at: https://github.com/zwandering/AesStyler. Ran Yi 0002, Haokun Zhu, Yukun Lai, Paul L. Rosin |
ACM Multimedia | 5 |
| 2024 | Exploiting Inter-Sample Affinity for Knowability-Aware Universal Domain Adaptation
Yifan Wang 0020, Lin Zhang 0041, Ran Song 0001, Hongliang Li 0001, Paul L. Rosin, Wei Zhang 0021 |
Int. J. Comput. Vis. | 5 |
| 2024 | TranSalNet+: Distortion-aware saliency predictionabstractPredicting the saliency of images affected by distortion is a challenging but emerging research problem. Given a distorted image, we wish to accurately predict saliency as perceived by humans. A recent distortion-aware saliency benchmark – the CUDAS database – reveals the inadequacy of existing saliency models in handling distorted images. In this paper, we devise a deep learning Distortion-Aware Saliency Module (DASM) that enables capturing saliency features related to image distortions, and integrates this module into a saliency prediction architecture. To achieve the high expressive capability of DASM using supervised learning, we create a dedicated dataset that draws upon a large-scale saliency dataset and machine-generated image quality assessments . Experimental results demonstrate the superior performance of the proposed model in predicting the saliency of distorted images. Jianxun Lou, Xinbo Wu, Padraig Corcoran, Paul L. Rosin, Hantao Liu |
Neurocomputing | 4 |
| 2024 | Cross-lingual font style transfer with full-domain convolutional attention
Tian-le Ji, Paul L. Rosin, Yukun Lai, Weiliang Meng, Yaonan Wang 0001 |
Pattern Recognit. | 3 |
| 2024 | HairManip: High quality hair manipulation via hair element disentangling
Lin Zhang 0041, Paul L. Rosin, Yukun Lai, Yaonan Wang 0001 |
Pattern Recognit. | 3 |
| 2024 | Ship Landmark: An Informative Ship Image Annotation and Its ApplicationsabstractVisual perception of ships has been attracting increasing attention in the fields of computer vision and ocean engineering. Despite the extensive work related to landmark detection of common objects, the role of landmarks in ship perception has been overlooked. In this paper, we aim to fill this gap by focusing on ship landmarks. Specifically, we give a comprehensive analysis of both the physical structure and deep features of ships, which finds that highlighted areas in feature maps correspond with structurally significant parts of ships. By summarizing the locations of such areas in ships, we define 20 ship landmarks and build the Ship Landmark Dataset (SLAD), the first ship dataset with landmark annotations. We also provide a benchmark for ship landmark detection by evaluating state-of-the-art landmark detection methods on the newly built SLAD. Moreover, we showcased several applications of ship landmarks, including ship recognition, ship image generation, key area detection for ships, and ship detection. Project web page:https://vsislab.github.io/Ships_VSIS/. Mingxin Zhang 0006, Qian Zhang 0076, Ran Song 0001, Paul L. Rosin, Wei Zhang 0021 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2023 | Towards Artistic Image Aesthetics Assessment: a Large-scale Dataset and a New MethodabstractImage aesthetics assessment (IAA) is a challenging task due to its highly subjective nature. Most of the current studies rely on large-scale datasets (e.g., AVA and AADB) to learn a general model for all kinds of photography images. However, little light has been shed on measuring the aesthetic quality of artistic images, and the existing datasets only contain relatively few artworks. Such a defect is a great obstacle to the aesthetic assessment of artistic images. To fill the gap in the field of artistic image aesthetics assessment (AIAA), we first introduce a large-scale AIAA dataset: Boldbrush Artistic Image Dataset (BAlD), which consists of 60,337 artistic images covering various art forms, with more than 360,000 votes from online users. We then propose a new method, SAAN (Style-specific Art Assessment Network), which can effectively extract and utilize style-specific and generic aesthetic information to evaluate artistic images. Experiments demonstrate that our proposed approach outperforms existing lAA methods on the proposed BAlD dataset according to quantitative comparisons. We believe the proposed dataset and method can serve as a foundation for future AIAA works and inspire more research in this field. Dataset and code are available at: https://github.com/Dreemurr-T/BAID.git Ran Yi 0002, Haoyuan Tian, Yukun Lai, Paul L. Rosin |
CVPR | 5 |
| 2023 | Confidence Guided Semi-Supervised Learning in Land Cover ClassificationabstractSemi-supervised learning has been well developed to help reduce the cost of manual labelling by exploiting a large quantity of unlabelled data. Especially in the application of land cover classification, pixel-level manual labelling in large-scale imagery is labour-intensive, time-consuming and expensive. However, existing semi-supervised learning methods pay limited attention to the quality of pseudo-labels during training even though the quality of training data is one of the critical factors determining network performance. In order to fill this gap, we develop a confidence-guided semi-supervised learning (CGSSL) approach to make use of high-confidence pseudo labels and reduce the negative effect of low-confidence ones for land cover classification. Meanwhile, the proposed semi-supervised learning approach uses multiple network architectures to increase the diversity of pseudo labels. The proposed semi-supervised learning approach significantly improves the performance of land cover classification compared to the classic semi-supervised learning methods and even outperforms fully supervised learning with a complete set of labelled imagery of the benchmark Potsdam land cover dataset. Wanli Ma 0001, Oktay Karakus, Paul L. Rosin |
IGARSS | 3 |
| 2023 | 3DCascade-GAN: Shape completion from single-view depth imagesabstractDepth images can be easily acquired using depth cameras. However, these images only contain partial information about the shape due to unavoidable self-occlusion. Thanks to the availability of large datasets of shapes, it is possible to use a learning-based approach to produce complete shapes from single depth images. State-of-the-art generative adversarial network (GAN) architectures can produce reasonable results. However, the use of relatively local convolutions restricts GAN architectures from producing globally plausible shapes. In this study, we develop a novel dynamic latent code selection mechanism in which the model learns to select only important codes from the latent space. Furthermore, a novel 3D self-attention (3DSA) layer is introduced that is able to capture non-local relationships across the 3D space. We further design a GAN architecture that uses a multistage encoder–decoder to recover the shape, where our 3DSA layer is introduced to the discriminator to help attend to global features, which stabilizes the model learning and encourages shape refinement, making our reconstruction more structurally plausible. Through extensive experiments, we demonstrate that our method outperforms other state-of-the-art methods for single depth image 3D reconstruction. Fahd Alhamazani, Yukun Lai, Paul L. Rosin |
Comput. Graph. | 3 |
| 2023 | 3D Visual Saliency: An Independent Perceptual Measure or a Derivative of 2D Image Saliency?abstractWhile 3D visual saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and has been well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art 3D visual saliency methods remain poor at predicting human fixations. Cues emerging prominently from these experiments suggest that 3D visual saliency might associate with 2D image saliency. This paper proposes a framework that combines a Generative Adversarial Network and a Conditional Random Field for learning visual saliency of both a single 3D object and a scene composed of multiple 3D objects with image saliency ground truth to 1) investigate whether 3D visual saliency is an independent perceptual measure or just a derivative of image saliency and 2) provide a weakly supervised method for more accurately predicting 3D visual saliency. Through extensive experiments, we not only demonstrate that our method significantly outperforms the state-of-the-art approaches, but also manage to answer the interesting and worthy question proposed within the title of this paper. Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Quality Metric Guided Portrait Line Drawing Generation From Unpaired Training DataabstractFace portrait line drawing is a unique style of art which is highly abstract and expressive. However, due to its high semantic constraints, many existing methods learn to generate portrait drawings using paired training data, which is costly and time-consuming to obtain. In this paper, we propose a novel method to automatically transform face photos to portrait drawings using unpaired training data with two new features; i.e., our method can (1) learn to generate high quality portrait drawings in multiple styles using a single network and (2) generate portrait drawings in a "new style" unseen in the training data. To achieve these benefits, we (1) propose a novel quality metric for portrait drawings which is learned from human perception, and (2) introduce a quality loss to guide the network toward generating better looking portrait drawings. We observe that existing unpaired translation methods such as CycleGAN tend to embed invisible reconstruction information indiscriminately in the whole drawings due to significant information imbalance between the photo and portrait drawing domains, which leads to important facial features missing. To address this problem, we propose a novel asymmetric cycle mapping that enforces the reconstruction information to be visible and only embedded in the selected facial regions. Along with localized discriminators for important facial regions, our method well preserves all important facial features in the generated drawings. Generator dissection further explains that our model learns to incorporate face semantic information during drawing generation. Extensive experiments including a user study show that our model outperforms state-of-the-art methods. Ran Yi 0002, Yong-Jin Liu 0001, Yukun Lai, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | 3D Face Reconstruction and Gaze Tracking in the HMD for Virtual InteractionabstractWith the rapid development of virtual reality (VR) technology, VR headsets, a.k.a. Head-Mounted Displays (HMDs), are widely available, allowing immersive 3D content to be viewed. A natural need for truly immersive VR is to allow bidirectional communication: the user should be able to interact with the virtual world using facial expressions and eye gaze, in addition to traditional means of interaction. The typical application scenario includes VR virtual conferencing and virtual roaming, where ideally users are able to see other users’ expressions and have eye contact with them in the virtual world. In addition, eye gaze also provides a natural means of interaction with virtual objects. Despite significant achievements in recent years for reconstruction of 3D faces from RGB or RGB-D images, it remains a challenge to reliably capture and reconstruct 3D facial expressions including eye gaze when the user is wearing an HMD, because the majority of the face is occluded, especially those areas around the eyes which are essential for recognizing facial expressions and eye gaze. In this paper, we introduce a novel real-time system that is able to capture and reconstruct 3D faces wearing HMDs, and robustly recover eye gaze. We further propose a novel method to map eye gaze directions to the 3D virtual world, which provides a novel and useful interactive mode in VR. We compare our method with state-of-the-art techniques both qualitatively and quantitatively, and demonstrate the effectiveness of our system using live capture. Yukun Lai, Shihong Xia, Paul L. Rosin, Lin Gao 0004 |
IEEE Trans. Multim. | 4 |
| 2022 | Analysis of Video Quality Induced Spatio-Temporal Saliency ShiftsabstractHuman viewers’ eye movements reflect their perceptual responses to visual signals. Previous research has shown that distortions in videos cause spatio-temporal gaze shifts, which means gaze behaviour is related to video quality perception. It would be highly beneficial to understand gaze behaviour of viewing videos of varying perceived quality. However, little is known about the interactions between gaze, video content and distortions. In this paper, based on our eye-tracking database for video quality (SVQ160), we perform systematic analyses to reveal the impact of video content (VC) and time order (TO) on gaze shifts. Findings and quantitative methods for gaze behaviour can be used to develop advanced video quality metrics and video processing algorithms. Xinbo Wu, Zhengyan Dong, Fan Zhang 0017, Paul L. Rosin, Hantao Liu |
ICIP | 4 |
| 2022 | Learning to Predict 3D Mesh SaliencyabstractMesh saliency, which measures the perceptual importance of different regions on a mesh, benefits a wide range of applications. However, existing mesh saliency models are largely built with hard-coded formulae, which cannot capture true human perception. Some existing techniques utilise indirect measures to capture user perception (e.g., mouse clicks), which can be unreliable. In this work, we collect eye-tracking data for 3D objects seen from different views, and develop an optimisation-based approach to fusing heat-maps captured from individual views to form consistent saliency maps on meshes. To predict mesh saliency on a new shape, we further develop a learning-based approach that regresses local surface characteristics based on a set of input features. Experimental results show that our learning-based method achieves better performance than state-of-the-art methods for unseen shapes. We will make our dataset publicly available. Dalia A. ALfarasani, Thomas Sweetman, Yukun Lai, Paul L. Rosin |
ICPR | 4 |
| 2022 | SHREC'21: Quantifying shape complexity
Mazlum Ferhat Arslan, Alexandros Haridis, Paul L. Rosin, Sibel Tari, Charlotte Brassey, James D. Gardiner, Asli Gençtav, Murat Genctav |
Comput. Graph. | 3 |
| 2022 | Foreword to the special issue on 3D object retrieval 2021 workshop (3DOR2021)
Silvia Biasotti, Roberto M. Dyke, Yukun Lai, Paul L. Rosin, Remco C. Veltkamp |
Comput. Graph. | 4 |
| 2022 | NPRportrait 1.0: A three-level benchmark for non-photorealistic rendering of portraitsabstractRecently, there has been an upsurge of activity in image-based non-photorealistic rendering (NPR), and in particular portrait image stylisation, due to the advent of neural style transfer (NST). However, the state of performance evaluation in this field is poor, especially compared to the norms in the computer vision and machine learning communities. Unfortunately, the task of evaluating image stylisation is thus far not well defined, since it involves subjective, perceptual, and aesthetic aspects. To make progress towards a solution, this paper proposes a new structured, three-level, benchmark dataset for the evaluation of stylised portrait images. Rigorous criteria were used for its construction, and its consistency was validated by user studies. Moreover, a new methodology has been developed for evaluating portrait stylisation algorithms, which makes use of the different benchmark levels as well as annotations provided by user studies regarding the characteristics of the faces. We perform evaluation for a wide variety of image stylisation methods (both portrait-specific and general purpose, and also both traditional NPR approaches and NST) using the new benchmark dataset. Paul L. Rosin, Yukun Lai, David Mould, Ran Yi 0002, Itamar Berger, Lars Doyle, Seungyong Lee 0001, Chuan Li 0001, Yong-Jin Liu 0001, Amir Semmo, Ariel Shamir, Minjung Son 0001, Holger Winnemöller |
Comput. Vis. Media | 1 |
| 2022 | Scale-aware network with modality-awareness for RGB-D indoor semantic segmentation
Feng Zhou 0007, Yukun Lai, Paul L. Rosin, Fengquan Zhang |
Neurocomputing | 3 |
| 2022 | Preface
Shi-Min Hu 0001, Paul L. Rosin, Tian-Jia Shao |
J. Comput. Sci. Technol. | 2 |
| 2022 | Cross-validation of a semantic segmentation network for natural history collection specimensabstractAbstract Semantic segmentation has been proposed as a tool to accelerate the processing of natural history collection images. However, developing a flexible and resilient segmentation network requires an approach for adaptation which allows processing different datasets with minimal training and validation. This paper presents a cross-validation approach designed to determine whether a semantic segmentation network possesses the flexibility required for application across different collections and institutions. Consequently, the specific objectives of cross-validating the semantic segmentation network are to (a) evaluate the effectiveness of the network for segmenting image sets derived from collections different from the one in which the network was initially trained on; and (b) test the adaptability of the segmentation network for use in other types of collections. The resilience to data variations from different institutions and the portability of the network across different types of collections are required to confirm its general applicability. The proposed validation method is tested on the Natural History Museum semantic segmentation network, designed to process entomological microscope slides. The proposed semantic segmentation network is evaluated through a series of cross-validation experiments designed to test using data from two types of collections: microscope slides (from three institutions) and herbarium sheets (from seven institutions). The main contribution of this work is the method, software and ground truth sets created for this cross-validation as they can be reused in testing similar segmentation proposals in the context of digitization of natural history collections. The cross-validation of segmentation methods should be a required step in the integration of such methods into image processing workflows for natural history collections. Abraham Nieva de la Hidalga, Paul L. Rosin, Xianfang Sun, Laurence Livermore, James Durrant, James Turner, Mathias Dillen, Alicia Musson, Sarah Phillips, Quentin Groom, Alex R. Hardisty |
Mach. Vis. Appl. | 2 |
| 2022 | LiTMNet: A deep CNN for efficient HDR image reconstruction from a single LDR image
Guotao Wu, Ran Song 0001, Mingxin Zhang 0006, Xiaolei Li 0003, Paul L. Rosin |
Pattern Recognit. | 5 |
| 2022 | Learning on 3D Meshes With Laplacian Encoding and Poolingabstract3D models are commonly used in computer vision and graphics. With the wider availability of mesh data, an efficient and intrinsic deep learning approach to processing 3D meshes is in great need. Unlike images, 3D meshes have irregular connectivity, requiring careful design to capture relations in the data. To utilize the topology information while staying robust under different triangulations, we propose to encode mesh connectivity using Laplacian spectral analysis, along with mesh feature aggregation blocks (MFABs) that can split the surface domain into local pooling patches and aggregate global information amongst them. We build a mesh hierarchy from fine to coarse using Laplacian spectral clustering, which is flexible under isometric transformations. Inside the MFABs there are pooling layers to collect local information and multi-layer perceptrons to compute vertex features of increasing complexity. To obtain the relationships among different clusters, we introduce a Correlation Net to compute a correlation matrix, which can aggregate the features globally by matrix multiplication with cluster features. Our network architecture is flexible enough to be used on meshes with different numbers of vertices. We conduct several experiments including shape segmentation and classification, and our method outperforms state-of-the-art algorithms for these tasks on the ShapeNet and COSEG datasets. Yi-Ling Qiao, Lin Gao 0004, Jie Yang 0038, Paul L. Rosin, Yukun Lai, Xilin Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2022 | A review of image and video colorization: From analogies to deep learningabstractImage colorization is a classic and important topic in computer graphics, where the aim is to add color to a monochromatic input image to produce a colorful result. In this survey, we present the history of colorization research in chronological order and summarize popular algorithms in this field. Early work on colorization mostly focused on developing techniques to improve the colorization quality. In the last few years, researchers have considered more possibilities such as combining colorization with NLP (natural language processing) and focused more on industrial applications. To better control the color, various types of color control are designed, such as providing reference images or color-scribbles. We have created a taxonomy of the colorization methods according to the input type, divided into grayscale, sketch-based and hybrid. The pros and cons are discussed for each algorithm, and they are compared according to their main characteristics. Finally, we discuss how deep learning, and in particular Generative Adversarial Networks (GANs), has changed this field. Jia-Qi Zhang, You-You Zhao, Paul L. Rosin, Yukun Lai, Lin Gao 0004 |
Vis. Informatics | 4 |
| 2021 | Large-Capacity Image Steganography Based on Invertible Neural NetworksabstractMany attempts have been made to hide information in images, where one main challenge is how to increase the payload capacity without the container image being detected as containing a message. In this paper, we propose a large-capacity Invertible Steganography Network (ISN) for image steganography. We take steganography and the recovery of hidden images as a pair of inverse problems on image domain transformation, and then introduce the forward and backward propagation operations of a single invertible network to leverage the image embedding and extracting problems. Sharing all parameters of our single ISN architecture enables us to efficiently generate both the container image and the revealed hidden image(s) with high quality. Moreover, in our architecture the capacity of image steganography is significantly improved by naturally increasing the number of channels of the hidden image branch. Comprehensive experiments demonstrate that with this significant improvement of the steganography payload capacity, our ISN achieves state-of-the-art in both visual and quantitative comparisons. Shao-Ping Lu, Paul L. Rosin |
CVPR | 4 |
| 2021 | Mesh Saliency: An Independent Perceptual Measure or a Derivative of Image Saliency?abstractWhile mesh saliency aims to predict regional importance of 3D surfaces in agreement with human visual perception and is well researched in computer vision and graphics, latest work with eye-tracking experiments shows that state-of-the-art mesh saliency methods remain poor at predicting human fixations. Cues emerging prominently from these experiments suggest that mesh saliency might associate with the saliency of 2D natural images. This paper proposes a novel deep neural network for learning mesh saliency using image saliency ground truth to 1) investigate whether mesh saliency is an independent perceptual measure or just a derivative of image saliency and 2) provide a weakly supervised method for more accurately predicting mesh saliency. Through extensive experiments, we not only demonstrate that our method outperforms the current state-of-the-art mesh saliency method by 116% and 21% in terms of linear correlation coefficient and AUC respectively, but also reveal that mesh saliency is intrinsically related with both image saliency and object categorical information. Codes are available at https://github.com/rsong/MIMO-GAN. Ran Song 0001, Wei Zhang 0021, Yitian Zhao, Yonghuai Liu, Paul L. Rosin |
CVPR | 5 |
| 2021 | Line Drawings for Face Portraits From Photos Using Global and Local Structure Based GANsabstractDespite significant effort and notable success of neural style transfer, it remains challenging for highly abstract styles, in particular line drawings. In this paper, we propose APDrawingGAN++, a generative adversarial network (GAN) for transforming face photos to artistic portrait drawings (APDrawings), which addresses substantial challenges including highly abstract style, different drawing techniques for different facial features, and high perceptual sensitivity to artifacts. To address these, we propose a composite GAN architecture that consists of local networks (to learn effective representations for specific facial features) and a global network (to capture the overall content). We provide a theoretical explanation for the necessity of this composite GAN structure by proving that any GAN with a single generator cannot generate artistic styles like APDrawings. We further introduce a classification-and-synthesis approach for lips and hair where different drawing styles are used by artists, which applies suitable styles for a given input. To capture the highly abstract art form inherent in APDrawings, we address two challenging operations-(1) coping with lines with small misalignments while penalizing large discrepancy and (2) generating more continuous lines-by introducing two novel loss terms: one is a novel distance transform loss with nonlinear mapping and the other is a novel line continuity loss, both of which improve the line quality. We also develop dedicated data augmentation and pre-training to further improve results. Extensive experiments, including a user study, show that our method outperforms state-of-the-art methods, both qualitatively and quantitatively. Ran Yi 0002, Mengfei Xia, Yong-Jin Liu 0001, Yukun Lai, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | DeepFaceEditing: deep face generation and editing with disentangled geometry and appearance controlabstractRecent facial image synthesis methods have been mainly based on conditional generative models. Sketch-based conditions can effectively describe the geometry of faces, including the contours of facial components, hair structures, as well as salient edges (e.g., wrinkles) on face surfaces but lack effective control of appearance, which is influenced by color, material, lighting condition, etc. To have more control of generated results, one possible approach is to apply existing disentangling works to disentangle face images into geometry and appearance representations. However, existing disentangling methods are not optimized for human face editing, and cannot achieve fine control of facial details such as wrinkles. To address this issue, we propose DeepFaceEditing, a structured disentanglement framework specifically designed for face images to support face generation and editing with disentangled control of geometry and appearance. We adopt a local-to-global approach to incorporate the face domain knowledge: local component images are decomposed into geometry and appearance representations, which are fused consistently using a global fusion module to improve generation quality. We exploit sketches to assist in extracting a better geometry representation, which also supports intuitive geometry editing via sketching. The resulting method can either extract the geometry and appearance representations from face images, or directly extract the geometry representation from face sketches. Such representations allow users to easily edit and synthesize face images, with decoupled control of their geometry and appearance. Both qualitative and quantitative evaluations show the superior detail and appearance control abilities of our method compared to state-of-the-art methods. Feng-Lin Liu, Yukun Lai, Paul L. Rosin, Chunpeng Li, Hongbo Fu 0001, Lin Gao 0004 |
ACM Trans. Graph. | 4 |
| 2021 | Mesh Saliency via Weakly Supervised Classification-for-Saliency CNNabstractRecently, effort has been made to apply deep learning to the detection of mesh saliency. However, one major barrier is to collect a large amount of vertex-level annotation as saliency ground truth for training the neural networks. Quite a few pilot studies showed that this task is difficult. In this work, we solve this problem by developing a novel network trained in a weakly supervised manner. The training is end-to-end and does not require any saliency ground truth but only the class membership of meshes. Our Classification-for-Saliency CNN (CfS-CNN) employs a multi-view setup and contains a newly designed two-channel structure which integrates view-based features of both classification and saliency. It essentially transfers knowledge from 3D object classification to mesh saliency. Our approach significantly outperforms the existing state-of-the-art methods according to extensive experimental results. Also, the CfS-CNN can be directly used for scene saliency. We showcase two novel applications based on scene saliency to demonstrate its utility. Ran Song 0001, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Unpaired Portrait Drawing Generation via Asymmetric Cycle MappingabstractPortrait drawing is a common form of art with high abstraction and expressiveness. Due to its unique characteristics, existing methods achieve decent results only with paired training data, which is costly and time-consuming to obtain.In this paper, we address the problem of automatic transfer from face photos to portrait drawings with unpaired training data. We observe that due to the significant imbalance of information richness between photos and drawings, existing unpaired transfer methods such as CycleGAN tends to embed invisible reconstruction information indiscriminately in the whole drawings, leading to important facial features partially missing in drawings. To address this problem, we propose a novel asymmetric cycle mapping that enforces the reconstruction information to be visible (by a truncation loss) and only embedded in selective facial regions (by a relaxed forward cycle-consistency loss). Along with localized discriminators for the eyes, nose and lips, our method well preserves all important facial features in the generated portrait drawings. By introducing a style classifier and taking the style vector into account, our method can learn to generate portrait drawings in multiple styles using a single network. Extensive experiments show that our model outperforms state-of-the-art methods. Ran Yi 0002, Yong-Jin Liu 0001, Yukun Lai, Paul L. Rosin |
CVPR | 4 |
| 2020 | Siamese Graph Convolution Network for Face Sketch Recognition: An application using Graph structure for face photo-sketch recognitionabstractIn this paper, we present a novel Siamese graph convolution network (GCN) for face sketch recognition. To build a graph from an image, we utilize a deep learning method to detect the image edges, and then use a superpixel method to segment the edge image. Each segmented superpixel region is taken as a node, and each pair of adjacent regions forms an edge of the graph. Graphs from both a face sketch and a face photo are input into the Siamese GCN for recognition. A deep graph matching method is used to share messages between cross-modal graphs in this model. Experiments show that the GCN can obtain high performance on several face photo-sketch datasets, including seen and unseen face photo-sketch datasets. It is also shown that the model performance based on the graph structure representation of the data using the Siamese GCN is more stable than a Siamese CNN model. Liang Fan, Xianfang Sun, Paul L. Rosin |
ICPR | 3 |
| 2020 | SHREC'20: Shape correspondence with non-isometric deformations abstractEstimating correspondence between two shapes continues to be a challenging problem in geometry processing. Most current methods assume deformation to be near-isometric, however this is often not the case. For this paper, a collection of shapes of different animals has been curated, where parts of the animals (e.g., mouths, tails & ears) correspond yet are naturally non-isometric. Ground-truth correspondences were established by asking three specialists to independently label corresponding points on each of the models with respect to a previously labelled reference model. We employ an algorithmic strategy to select a single point for each correspondence that is representative of the proposed labels. A novel technique that characterises the sparsity and distribution of correspondences is employed to measure the performance of ten shape correspondence methods. Roberto M. Dyke, Yukun Lai, Paul L. Rosin, Stefano Zappalà, Seana Dykes, Daoliang Guo, Kun Li 0001, Riccardo Marin, Simone Melzi, Jing-Yu Yang 0002 |
Comput. Graph. | 3 |
| 2020 | SHREC 2020: Multi-domain protein shape retrieval challenge
Florent Langenfeld, Yuxu Peng, Yukun Lai, Paul L. Rosin, Tunde Aderinwale, Genki Terashi, Charles Christoffer, Daisuke Kihara, Halim Benhabiles, Karim Hammoudi, Adnane Cabani, Féryal Windal, Mahmoud Melkemi, Andrea Giachetti 0001, Stelios K. Mylonas, Apostolos Axenopoulos, Petros Daras, Ekpo Otu, Matthieu Montès |
Comput. Graph. | 4 |
| 2020 | 3D computational modeling and perceptual analysis of kinetic depth effectsabstractHumans have the ability to perceive kinetic depth effects , i.e., to perceived 3D shapes from 2D projections of rotating 3D objects. This process is based on a variety of visual cues such as lighting and shading effects. However, when such cues are weak or missing, perception can become faulty, as demonstrated by the famous silhouette illusion example of the spinning dancer . Inspired by this, we establish objective and subjective evaluation models of rotated 3D objects by taking their projected 2D images as input. We investigate five different cues: ambient luminance, shading, rotation speed, perspective, and color difference between the objects and background. In the objective evaluation model, we first apply 3D reconstruction algorithms to obtain an objective reconstruction quality metric, and then use quadratic stepwise regression analysis to determine weights of depth cues to represent the reconstruction quality. In the subjective evaluation model, we use a comprehensive user study to reveal correlations with reaction time and accuracy, rotation speed, and perspective. The two evaluation models are generally consistent, and potentially of benefit to inter-disciplinary research into visual perception and 3D reconstruction. Mengyao Cui 0001, Shao-Ping Lu, Miao Wang 0004, Yongliang Yang 0002, Yukun Lai, Paul L. Rosin |
Comput. Vis. Media | 6 |
| 2020 | Adaptive gradient-based block compressive sensing with sparsity for noisy images
Paul L. Rosin, Yukun Lai, Jinhua Zheng, Yaonan Wang 0001 |
Multim. Tools Appl. | 2 |
| 2020 | Subspace Clustering via Good NeighborsabstractFinding the informative subspaces of high-dimensional datasets is at the core of numerous applications in computer vision, where spectral-based subspace clustering is arguably the most widely studied method due to its strong empirical performance. Such algorithms first compute an affinity matrix to construct a self-representation for each sample using other samples as a dictionary. Sparsity and connectivity of the self-representation play important roles in effective subspace clustering. However, simultaneous optimization of both factors is difficult due to their conflicting nature, and most existing methods are designed to address only one factor. In this paper, we propose a post-processing technique to optimize both sparsity and connectivity by finding good neighbors. Good neighbors induce key connections among samples within a subspace and not only have large affinity coefficients but are also strongly connected to each other. We reassign the coefficients of the good neighbors and eliminate other entries to generate a new coefficient matrix. We show that the few good neighbors can effectively recover the subspace, and the proposed post-processing step of finding good neighbors is complementary to most existing subspace clustering algorithms. Experiments on five benchmark datasets show that the proposed algorithm performs favorably against the state-of-the-art methods with negligible additional computation cost. Jufeng Yang, Jie Liang 0007, Kai Wang 0001, Paul L. Rosin, Ming-Hsuan Yang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2020 | Measuring Shapes with Desired Convex PolygonsabstractIn this paper we have developed a family of shape measures. All the measures from the family evaluate the degree to which a shape looks like a predefined convex polygon. A quite new approach in designing object shape based measures has been applied. In most cases such measures were defined by exploiting some shape properties. Such properties are optimized (e.g., maximized or minimized) by certain shapes and based on this, the new shape measures were defined. An illustrative example might be the shape circularity measure derived by exploiting the well-known result that the circle has the largest area among all the shapes with the same perimeter. Of course, there are many more such examples (e.g., ellipticity, linearity, elongation, and squareness measures are some of them). There are different approaches as well. In the approach applied here, no desired property is needed and no optimizing shape has to be found. We start from a desired convex polygon, and develop the related shape measure. The method also allows a tuning parameter. Thus, there is a new 2-fold family of shape measures, dependent on a predefined convex polygon, and a tuning parameter, that controls the measure's behavior. The measures obtained range over the interval (0,1] and pick the maximal possible value, equal to 1, if and only if the measured shape coincides with the selected convex polygon that was used to develop the particular measure. All the measures are invariant with respect to translations, rotations, and scaling transformations. An extension of the method leads to a family of new shape convexity measures. Jovisa D. Zunic, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2020 | Structure-Preserving Neural Style TransferabstractState-of-the-art neural style transfer methods have demonstrated amazing results by training feed-forward convolutional neural networks or using an iterative optimization strategy. The image representation used in these methods, which contains two components: style representation and content representation, is typically based on high-level features extracted from pretrained classification networks. Because the classification networks are originally designed for object recognition, the extracted features often focus on the central object and neglect other details. As a result, the style textures tend to scatter over the stylized outputs and disrupt the content structures. To address this issue, we present a novel image stylization method that involves an additional structure representation. Our structure representation, which considers two factors: i) the global structure represented by the depth map and ii) the local structure details represented by the image edges, effectively reflects the spatial distribution of all the components in an image as well as the structure of dominant objects respectively. Experimental results demonstrate that our method achieves an impressive visual effectiveness, which is particularly significant when processing images sensitive to structure distortion, e.g. images containing multiple objects potentially at different depths, or dominant objects with clear structures. Ming-Ming Cheng, Xiao-Chang Liu, Shao-Ping Lu, Yukun Lai, Paul L. Rosin |
IEEE Trans. Image Process. | 6 |
| 2020 | Sparse Graph Regularized Mesh Color Edit PropagationabstractMesh color edit propagation aims to propagate the color from a few color strokes to the whole mesh, which is useful for mesh colorization, color enhancement and color editing, etc. Compared with image edit propagation, luminance information is not available for 3D mesh data, so the color edit propagation is more difficult on 3D meshes than images, with far less research carried out. This paper proposes a novel solution based on sparse graph regularization. Firstly, a few color strokes are interactively drawn by the user, and then the color will be propagated to the whole mesh by minimizing a sparse graph regularized nonlinear energy function. The proposed method effectively measures geometric similarity over shapes by using a set of complementary multiscale feature descriptors, and effectively controls color bleeding via a sparse ℓ1 optimization rather than quadratic minimization used in existing work. The proposed framework can be applied for the task of interactive mesh colorization, mesh color enhancement and mesh color editing. Extensive qualitative and quantitative experiments show that the proposed method outperforms the state-of-the-art methods. Bo Li 0023, Yukun Lai, Paul L. Rosin |
IEEE Trans. Image Process. | 3 |
| 2020 | WSCNet: Weakly Supervised Coupled Networks for Visual Sentiment Classification and DetectionabstractAutomatic assessment of sentiment from visual content has gained considerable attention with the increasing tendency of expressing opinions online. In this paper, we solve the problem of visual sentiment analysis, which is challenging due to the high-level abstraction in the recognition process. Existing methods based on convolutional neural networks learn sentiment representations from the holistic image, despite the fact that different image regions can have different influence on the evoked sentiment. In this paper, we introduce a weakly supervised coupled convolutional network (WSCNet). Our method is dedicated to automatically selecting relevant soft proposals given weak annotations (e.g., global image labels), thereby significantly reducing the annotation burden, and encompasses the following contributions. First, the proposed WSCNet detects a sentiment-specific soft map by training a fully convolutional network with the cross spatial pooling strategy in the detection branch. Second, both the holistic and localized information are utilized by coupling the sentiment map with deep features as semantic vector in the classification branch. The sentiment detection and classification branches are integrated into a unified deep framework optimized in an end-to-end manner. Extensive experiments demonstrate that the proposed WSCNet outperforms the state-of-the-art results on seven benchmark datasets. Dongyu She, Jufeng Yang, Ming-Ming Cheng, Yukun Lai, Paul L. Rosin, Liang Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2020 | Self-Paced Balance Learning for Clinical Skin Disease RecognitionabstractClass imbalance is a challenging problem in many classification tasks. It induces biased classification results for minority classes that contain less training samples than others. Most existing approaches aim to remedy the imbalanced number of instances among categories by resampling the majority and minority classes accordingly. However, the imbalanced level of difficulty of recognizing different categories is also crucial, especially for distinguishing samples with many classes. For example, in the task of clinical skin disease recognition, several rare diseases have a small number of training samples, but they are easy to diagnose because of their distinct visual properties. On the other hand, some common skin diseases, e.g., eczema, are hard to recognize due to the lack of special symptoms. To address this problem, we propose a self-paced balance learning (SPBL) algorithm in this paper. Specifically, we introduce a comprehensive metric termed the complexity of image category that is a combination of both sample number and recognition difficulty. First, the complexity is initialized using the model of the first pace, where the pace indicates one iteration in the self-paced learning paradigm. We then assign each class a penalty weight that is larger for more complex categories and smaller for easier ones, after which the curriculum is reconstructed by rearranging the training samples. Consequently, the model can iteratively learn discriminative representations via balancing the complexity in each pace. Experimental results on the SD-198 and SD-260 benchmark data sets demonstrate that the proposed SPBL algorithm performs favorably against the state-of-the-art methods. We also demonstrate the effectiveness of the SPBL algorithm's generalization capacity on various tasks, such as indoor scene image recognition and object classification. Jufeng Yang, Jie Liang 0007, Xiaoxiao Sun 0002, Ming-Ming Cheng, Paul L. Rosin, Liang Wang 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2020 | Distinction of 3D Objects and Scenes via Classification Network and Markov Random FieldabstractAn importance measure of 3D objects inspired by human perception has a range of applications since people want computers to behave like humans in many tasks. This paper revisits a well-defined measure, distinction of 3D surface mesh, which indicates how important a region of a mesh is with respect to classification. We develop a method to compute it based on a classification network and a Markov Random Field (MRF). The classification network learns view-based distinction by handling multiple views of a 3D object. Using a classification network has an advantage of avoiding the training data problem which has become a major obstacle of applying deep learning to 3D object understanding tasks. The MRF estimates the parameters of a linear model for combining the view-based distinction maps. The experiments using several publicly accessible datasets show that the distinctive regions detected by our method are not just significantly different from those detected by methods based on handcrafted features, but more consistent with human perception. We also compare it with other perceptual measures and quantitatively evaluate its performance in the context of two applications. Furthermore, due to the view-based nature of our method, we are able to easily extend mesh distinction to 3D scenes containing multiple objects. Ran Song 0001, Yonghuai Liu, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 3 |
| 2020 | Automatic semantic style transfer using deep convolutional neural networks and soft masks
Paul L. Rosin, Yukun Lai, Yaonan Wang 0001 |
Vis. Comput. | 2 |
| 2019 | APDrawingGAN: Generating Artistic Portrait Drawings From Face Photos With Hierarchical GANsabstractSignificant progress has been made with image stylization using deep learning, especially with generative adversarial networks (GANs). However, existing methods fail to produce high quality artistic portrait drawings. Such drawings have a highly abstract style, containing a sparse set of continuous graphical elements such as lines, and so small artifacts are much more exposed than for painting styles. Moreover, artists tend to use different strategies to draw different facial features and the lines drawn are only loosely related to obvious image features. To address these challenges, we propose APDrawingGAN, a novel GAN based architecture that builds upon hierarchical generators and discriminators combining both a global network (for images as a whole) and local networks (for individual facial regions). This allows dedicated drawing strategies to be learned for different facial features. Since artists' drawings may not have lines perfectly aligned with image features, we develop a novel loss to measure similarity between generated and artists' drawings based on distance transforms, leading to improved strokes in portrait drawing. To train APDrawingGAN, we construct an artistic drawing dataset containing high-resolution portrait photos and corresponding professional artistic drawings. Extensive experiments, including a user study, show that APDrawingGAN produces significantly better artistic drawings than state-of-the-art methods. Ran Yi 0002, Yong-Jin Liu 0001, Yukun Lai, Paul L. Rosin |
CVPR | 4 |
| 2019 | Pose2Seg: Detection Free Human Instance SegmentationabstractThe standard approach to image instance segmentation is to perform the object detection first, and then segment the object from the detection bounding-box. More recently, deep learning methods like Mask R-CNN perform them jointly. However, little research takes into account the uniqueness of the "human" category, which can be well defined by the pose skeleton. Moreover, the human pose skeleton can be used to better distinguish instances with heavy occlusion than using bounding-boxes. In this paper, we present a brand new pose-based instance segmentation framework for humans which separates instances based on human pose, rather than proposal region detection. We demonstrate that our pose-based framework can achieve better accuracy than the state-of-art detection-based approach on the human instance segmentation problem, and can moreover better handle occlusion. Furthermore, there are few public datasets containing many heavily occluded humans along with comprehensive annotations, which makes this a challenging problem seldom noticed by researchers. Therefore, in this paper we introduce a new benchmark "Occluded Human (OCHuman)", which focuses on occluded humans with comprehensive annotations including bounding-box, human pose and instance masks. This dataset contains 8110 detailed annotated human instances within 4731 images. With an average 0.67 MaxIoU for each person, OCHuman is the most complex and challenging dataset related to human instance segmentation. Through this dataset, we want to emphasize occlusion as a challenging problem for researchers to study. Song-Hai Zhang, Ruilong Li, Paul L. Rosin, Zixi Cai, Dingcheng Yang, Hao-Zhi Huang 0001, Shi-Min Hu 0001 |
CVPR | 4 |
| 2019 | Scoot: A Perceptual Metric for Facial SketchesabstractWhile it is trivial for humans to quickly assess the perceptual similarity between two images, the underlying mechanism are thought to be quite complex. Despite this, the most widely adopted perceptual metrics today, such as SSIM and FSIM, are simple, shallow functions, and fail to consider many factors of human perception. Recently, the facial modeling community has observed that the inclusion of both structure and texture has a significant positive benefit for face sketch synthesis (FSS). But how perceptual are these so-called “perceptual features”? Which elements are critical for their success? In this paper, we design a perceptual metric, called Structure Co-Occurrence Texture (Scoot), which simultaneously considers the block-level spatial structure and co-occurrence texture statistics. To test the quality of metrics, we propose three novel meta-measures based on various reliable properties. Extensive experiments verify that our Scoot metric exceeds the performance of prior work. Besides, we built the first largest scale (152k judgments) human-perception-based sketch database that can evaluate how well a metric consistent with human perception. Our results suggest that “spatial structure” and “co-occurrence texture” are two generally applicable perceptual features in face sketch synthesis. Deng-Ping Fan, Shengchuan Zhang, Yu-Huan Wu, Yun Liu 0011, Ming-Ming Cheng, Bo Ren 0003, Paul L. Rosin, Rongrong Ji |
ICCV | 7 |
| 2019 | Non-rigid registration under anisotropic deformationsabstractNon-rigid registration of deformed 3D shapes is a challenging and fundamental task in geometric processing, which aims to non-rigidly deform a source shape into alignment with a target shape. Current state-of-the-art methods assume deformations to be near-isometric. This assumption does not reflect real-world conditions, for example in large-scale deformation, where moderate anisotropic deformations (e.g., stretches) are common. In this paper we propose two significant changes to a typical registration pipeline to address such challenging deformations. First, we introduce a method to estimate anisotropic non-isometric deformations and incorporate this into an iterative non-rigid registration pipeline. Second, we compute additional correspondences in non-isometrically deforming regions using reliable correspondences as landmarks and prune inconsistent correspondences. We compare the performance of our proposed algorithm to several state-of-the-art methods using existing benchmarks. Experimental results show that our method outperforms existing methods. Roberto M. Dyke, Yukun Lai, Paul L. Rosin, Gary K. L. Tam |
Comput. Aided Geom. Des. | 3 |
| 2019 | BING: Binarized normed gradients for objectness estimation at 300fpsabstractTraining a generic objectness measure to produce object proposals has recently become of significant interest. We observe that generic objects with well-defined closed boundaries can be detected by looking at the norm of gradients, with a suitable resizing of their corresponding image windows to a small fixed size. Based on this observation and computational reasons, we propose to resize the window to 8 × 8 and use the norm of the gradients as a simple 64D feature to describe it, for explicitly training a generic objectness measure. We further show how the binarized version of this feature, namely binarized normed gradients (BING), can be used for efficient objectness estimation, which requires only a few atomic operations (e.g., add, bitwise shift, etc.). To improve localization quality of the proposals while maintaining efficiency, we propose a novel fast segmentation method and demonstrate its effectiveness for improving BING’s localization performance, when used in multi-thresholding straddling expansion (MTSE) post-processing. On the challenging PASCAL VOC2007 dataset, using 1000 proposals per image and intersection-over-union threshold of 0.5, our proposal method achieves a 95.6% object detection rate and 78.6% mean average best overlap in less than 0.005 second per image. Ming-Ming Cheng, Yun Liu 0011, Wen-Yan Lin, Paul L. Rosin, Philip Torr 0001 |
Comput. Vis. Media | 5 |
| 2019 | A multi-scale topological shape model for single and multiple component shapes
Padraig Corcoran, Jovisa D. Zunic, Paul L. Rosin |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Edge-texture feature-based image forgery detection with cross-dataset evaluation
Khurshid Asghar, Xianfang Sun, Paul L. Rosin, Mubbashar Saddique, Muhammad Hussain 0001, Zulfiqar Habib |
Mach. Vis. Appl. | 3 |
| 2019 | Automatic Example-Based Image Colorization Using Location-Aware Cross-Scale MatchingabstractGiven a reference colour image and a destination grayscale image, this paper presents a novel automatic colourisation algorithm that transfers colour information from the reference image to the destination image. Since the reference and destination images may contain content at different or even varying scales (due to changes of distance between objects and the camera), existing texture matching based methods can often perform poorly. We propose a novel cross-scale texture matching method to improve the robustness and quality of the colourisation results. Suitable matching scales are considered locally, which are then fused using global optimisation that minimises both the matching errors and spatial change of scales. The minimisation is efficiently solved using a multi-label graph-cut algorithm. Since only low-level texture features are used, texture matching based colourisation can still produce semantically incorrect results, such as meadow appearing above the sky. We consider a class of semantic violation where the statistics of up-down relationships learnt from the reference image are violated and propose an effective method to identify and correct unreasonable colourisation. Finally, a novel nonlocal ℓ1 optimisation framework is developed to propagate high confidence micro-scribbles to regions of lower confidence to produce a fully colourised image. Qualitative and quantitative evaluations show that our method outperforms several state-of-the-art methods. Bo Li 0023, Yukun Lai, Matthew John, Paul L. Rosin |
IEEE Trans. Image Process. | 4 |
| 2019 | Simultaneous Subspace Clustering and Cluster Number Estimating Based on Triplet RelationshipabstractIn this paper, we propose a unified framework to discover the number of clusters and group the data points into different clusters using subspace clustering simultaneously. Real data distributed in a high-dimensional space can be disentangled into a union of low-dimensional subspaces, which can benefit various applications. To explore such intrinsic structure, state-of-the-art subspace clustering approaches often optimize a self-representation problem among all samples, to construct a pairwise affinity graph for spectral clustering. However, a graph with pairwise similarities lacks robustness for segmentation, especially for samples which lie on the intersection of two subspaces. To address this problem, we design a hyper-correlation-based data structure termed as the triplet relationship, which reveals high relevance and local compactness among three samples. The triplet relationship can be derived from the self-representation matrix, and be utilized to iteratively assign the data points to clusters. Based on the triplet relationship, we propose a unified optimizing scheme to automatically calculate clustering assignments. Specifically, we optimize a model selection reward and a fusion reward by simultaneously maximizing the similarity of triplets from different clusters while minimizing the correlation of triplets from the same cluster. The proposed algorithm also automatically reveals the number of clusters and fuses groups to avoid over-segmentation. Extensive experimental results on both synthetic and real-world datasets validate the effectiveness and robustness of the proposed method. Jie Liang 0007, Jufeng Yang, Ming-Ming Cheng, Paul L. Rosin, Liang Wang 0001 |
IEEE Trans. Image Process. | 4 |
| 2019 | Image-driven unsupervised 3D model co-segmentation
Paul L. Rosin, Xianfang Sun, Jianguo Xiao, Zhouhui Lian |
Vis. Comput. | 2 |
| 2018 | FLIC: Fast Linear Iterative Clustering With Active SearchabstractIn this paper, we reconsider the clustering problem for image over-segmentation from a new perspective. We propose a novel search algorithm named “active search” which explicitly considers neighboring continuity. Based on this search method, we design a back-and-forth traversal strategy and a "joint" assignment and update step to speed up the algorithm. Compared to earlier works, such as Simple Linear Iterative Clustering (SLIC) and its follow-ups, who use fixed search regions and perform the assignment and the update step separately, our novel scheme reduces the iteration number before convergence, as well as improves boundary sensitivity of the over-segmentation results. Extensive evaluations on the Berkeley segmentation benchmark verify that our method outperforms competing methods under various evaluation metrics. In particular, lowest time cost is reported among existing methods (approximately 30 fps for a 481321 image on a single CPU core). To facilitate the development of over-segmentation, the code will be publicly available. Jiaxing Zhao, Bo Ren 0003, Qibin Hou, Ming-Ming Cheng, Paul L. Rosin |
AAAI | 5 |
| 2018 | Weakly Supervised Coupled Networks for Visual Sentiment AnalysisabstractAutomatic assessment of sentiment from visual content has gained considerable attention with the increasing tendency of expressing opinions on-line. In this paper, we solve the problem of visual sentiment analysis using the high-level abstraction in the recognition process. Existing methods based on convolutional neural networks learn sentiment representations from the holistic image appearance. However, different image regions can have a different influence on the intended expression. This paper presents a weakly supervised coupled convolutional network with two branches to leverage the localized information. The first branch detects a sentiment specific soft map by training a fully convolutional network with the cross spatial pooling strategy, which only requires image-level labels, thereby significantly reducing the annotation burden. The second branch utilizes both the holistic and localized information by coupling the sentiment map with deep features for robust classification. We integrate the sentiment detection and classification branches into a unified deep framework and optimize the network in an end-to-end manner. Extensive experiments on six benchmark datasets demonstrate that the proposed method performs favorably against the state-of-the-art methods for visual sentiment analysis. Jufeng Yang, Dongyu She, Yukun Lai, Paul L. Rosin, Ming-Hsuan Yang 0001 |
CVPR | 4 |
| 2018 | Clinical Skin Lesion Diagnosis Using Representations Inspired by Dermatologist CriteriaabstractThe skin is the largest organ in human body. Around 30%-70% of individuals worldwide have skin related health problems, for whom effective and efficient diagnosis is necessary. Recently, computer aided diagnosis (CAD) systems have been successfully applied to the recognition of skin cancers in dermatoscopic images. However, little work has concentrated on the commonly encountered skin diseases in clinical images captured by easily-accessed cameras or mobile phones. Meanwhile, for a CAD system, the representations of skin lesions are required to be understandable for dermatologists so that the predictions are convincing. To address this problem, we present effective representations inspired by the accepted dermatological criteria for diagnosing clinical skin lesions. We demonstrate that the dermatological criteria are highly correlated with measurable visual components. Accordingly, we design six medical representations considering different criteria for the recognition of skin lesions, and construct a diagnosis system for clinical skin disease images. Experimental results show that the proposed medical representations can not only capture the manifestations of skin lesions effectively, and consistently with the dermatological criteria, but also improve the prediction performance with respect to the state-of-the-art methods based on uninterpretable features. Jufeng Yang, Xiaoxiao Sun 0002, Jie Liang 0007, Paul L. Rosin |
CVPR | 4 |
| 2018 | Real-Time 3D Face Reconstruction and Gaze Tracking for Virtual RealityabstractWith the rapid development of virtual reality (VR) technology, VR glasses, a.k.a. Head-Mounted Displays (HMDs) are widely available, allowing immersive 3D content to be viewed. A natural need for truly immersive VR is to allow bidirectional communication: the user should be able to interact with the virtual world using facial expressions and eye gaze, in addition to traditional means of interaction. Typical application scenarios include VR virtual conferencing and virtual roaming, where ideally users are able to see other users' expressions and have eye contact with them in the virtual world. Despite significant achievements in recent years for reconstruction of 3D faces from RGB or RGB- D images, it remains a challenge to reliably capture and reconstruct 3D facial expressions including eye gaze when the user is wearing VR glasses, because the majority of the face is occluded, especially those areas around the eyes which are essential for recognizing facial expressions and eye gaze. In this paper, we introduce a novel real-time system that is able to capture and reconstruct 3D faces wearing HMDs and robustly recover eye gaze. We demonstrate the effectiveness of our system using live capture and more results are shown in the accompanying video. Lin Gao 0004, Yukun Lai, Paul L. Rosin, Shihong Xia |
VR | 4 |
| 2018 | An evaluation of canonical forms for non-rigid 3D shape retrievalabstractCanonical forms attempt to factor out a non-rigid shape’s pose, giving a pose-neutral shape. This opens up the possibility of using methods originally designed for rigid shape retrieval for the task of non-rigid shape retrieval. We extend our recent benchmark for testing canonical form algorithms. Our new benchmark is used to evaluate a greater number of state-of-the-art canonical forms, on five recent non-rigid retrieval datasets, within two different retrieval frameworks. A total of fifteen different canonical form methods are compared. We find that the difference in retrieval accuracy between different canonical form methods is small, but varies significantly across different datasets. We also find that efficiency is the main difference between the methods. David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Zhi-Quan Cheng, Zhouhui Lian, Sipin Nie, Longcun Jin, Gil Shamai, Yusuf Sahillioglu, Ladislav Kavan |
Graph. Model. | 4 |
| 2018 | Biharmonic deformation transfer with automatic key point selection
Jie Yang 0038, Lin Gao 0004, Yukun Lai, Paul L. Rosin, Shihong Xia |
Graph. Model. | 4 |
| 2018 | FLIC: Fast linear iterative clustering with active searchabstractIn this paper, we reconsider the clustering problem for image over-segmentation from a new perspective. We propose a novel search algorithm called “active search” which explicitly considers neighbor continuity. Based on this search method, we design a back-and-forth traversal strategy and a joint assignment and update step to speed up the algorithm. Compared to earlier methods, such as simple linear iterative clustering (SLIC) and its variants, which use fixed search regions and perform the assignment and the update steps separately, our novel scheme reduces the number of iterations required for convergence, and also provides better boundaries in the over-segmentation results. Extensive evaluation using the Berkeley segmentation benchmark verifies that our method outperforms competing methods under various evaluation metrics. In particular, our method is fastest, achieving approximately 30 fps for a 481 × 321 image on a single CPU core. To facilitate further research, our code is made publicly available. Jiaxing Zhao, Bo Ren 0003, Qibin Hou, Ming-Ming Cheng, Paul L. Rosin |
Comput. Vis. Media | 5 |
| 2018 | Disconnectedness: A new moment invariant for multi-component shapes
Jovisa D. Zunic, Paul L. Rosin, Vladimir Ilic |
Pattern Recognit. | 2 |
| 2018 | Robust Virtual Unrolling of Historical Parchment XMT ImagesabstractWe develop a framework to virtually unroll fragile historical parchment scrolls, which cannot be physically unfolded via a sequence of X-ray tomographic slices, thus providing easy access to those parchments whose contents have remained hidden for centuries. The first step is to produce a topologically correct segmentation, which is challenging as the parchment layers vary significantly in thickness, contain substantial interior textures and can often stick together in places. For this purpose, our method starts with linking the broken layers in a slice using the topological structure propagated from its previous processed slice. To ensure topological correctness, we identify fused regions by detecting junction sections, and then match them using global optimization efficiently solved by the blossom algorithm, taking into account the shape energy of curves separating fused layers. The fused layers are then separated using as-parallel-as-possible curves connecting junction section pairs. To flatten the segmented parchment, pixels in different frames need to be put into alignment. This is achieved via a dynamic programming-based global optimization, which minimizes the total matching distances and penalizes stretches. Eventually, the text of the parchment is revealed by ink projection. We demonstrate the effectiveness of our approach using challenging real-world data sets, including the water damaged fifteenth century Bressingham scroll. Chang Liu 0009, Paul L. Rosin, Yukun Lai, Weiduo Hu |
IEEE Trans. Image Process. | 2 |
| 2018 | Dynamic Match Kernel With Deep Convolutional Features for Image RetrievalabstractFor image retrieval methods based on bag of visual words, much attention has been paid to enhancing the discriminative powers of the local features. Although retrieved images are usually similar to a query in minutiae, they may be significantly different from a semantic perspective, which can be effectively distinguished by convolutional neural networks (CNN). Such images should not be considered as relevant pairs. To tackle this problem, we propose to construct a dynamic match kernel by adaptively calculating the matching thresholds between query and candidate images based on the pairwise distance among deep CNN features. In contrast to the typical static match kernel which is independent to the global appearance of retrieved images, the dynamic one leverages the semantical similarity as a constraint for determining the matches. Accordingly, we propose a semantic-constrained retrieval framework by incorporating the dynamic match kernel, which focuses on matched patches between relevant images and filters out the ones for irrelevant pairs. Furthermore, we demonstrate that the proposed kernel complements recent methods, such as hamming embedding, multiple assignment, local descriptors aggregation, and graph-based re-ranking, while it outperforms the static one under various settings on off-the-shelf evaluation metrics. We also propose to evaluate the matched patches both quantitatively and qualitatively. Extensive experiments on five benchmark data sets and large-scale distractors validate the merits of the proposed method against the state-of-the-art methods for image retrieval. Jufeng Yang, Jie Liang 0007, Kai Wang 0001, Paul L. Rosin, Ming-Hsuan Yang 0001 |
IEEE Trans. Image Process. | 5 |
| 2018 | Visual Sentiment Prediction Based on Automatic Discovery of Affective RegionsabstractAutomatic assessment of sentiment from visual content has gained considerable attention with the increasing tendency of expressing opinions via images and videos online. This paper investigates the problem of visual sentiment analysis, which involves a high-level abstraction in the recognition process. While most of the current methods focus on improving holistic representations, we aim to utilize the local information, which is inspired by the observation that both the whole image and local regions convey significant sentiment information. We propose a framework to leverage affective regions, where we first use an off-the-shelf objectness tool to generate the candidates, and employ a candidate selection method to remove redundant and noisy proposals. Then, a convolutional neural network (CNN) is connected with each candidate to compute the sentiment scores, and the affective regions are automatically discovered, taking the objectness score as well as the sentiment score into consideration. Finally, the CNN outputs from local regions are aggregated with the whole images to produce the final predictions. Our framework only requires image-level labels, thereby significantly reducing the annotation burden otherwise required for training. This is especially important for sentiment analysis since sentiment can be abstract, and labeling affective regions is too subjective and labor-consuming. Extensive experiments show that the proposed algorithm outperforms the state-of-the-art approaches on eight popular benchmark datasets. Jufeng Yang, Dongyu She, Ming Sun 0006, Ming-Ming Cheng, Paul L. Rosin, Liang Wang 0001 |
IEEE Trans. Multim. | 5 |
| 2018 | Automatic unpaired shape deformation transferabstractTransferring deformation from a source shape to a target shape is a very useful technique in computer graphics. State-of-the-art deformation transfer methods require either point-wise correspondences between source and target shapes, or pairs of deformed source and target shapes with corresponding deformations. However, in most cases, such correspondences are not available and cannot be reliably established using an automatic algorithm. Therefore, substantial user effort is needed to label the correspondences or to obtain and specify such shape sets. In this work, we propose a novel approach to automatic deformation transfer between two unpaired shape sets without correspondences. 3D deformation is represented in a high-dimensional space. To obtain a more compact and effective representation, two convolutional variational autoencoders are learned to encode source and target shapes to their latent spaces. We exploit a Generative Adversarial Network (GAN) to map deformed source shapes to deformed target shapes, both in the latent spaces, which ensures the obtained shapes from the mapping are indistinguishable from the target shapes. This is still an under-constrained problem, so we further utilize a reverse mapping from target shapes to source shapes and incorporate cycle consistency loss, i.e. applying both mappings should reverse to the input shape. This VAE-Cycle GAN (VC-GAN) architecture is used to build a reliable mapping between shape spaces. Finally, a similarity constraint is employed to ensure the mapping is consistent with visual similarity, achieved by learning a similarity neural network that takes the embedding vectors from the source and target latent spaces and predicts the light field distance between the corresponding shapes. Experimental results show that our fully automatic method is able to obtain high-quality deformation transfer results with unpaired data sets, comparable or better than existing methods where strict correspondences are required. Lin Gao 0004, Jie Yang 0038, Yi-Ling Qiao, Yukun Lai, Paul L. Rosin, Weiwei Xu 0003, Shihong Xia |
ACM Trans. Graph. | 5 |
| 2017 | An open-data, agent-based model of alcohol related crimeabstractThe allocation of resources to challenge city centre violent crime traditionally relies on historical data to identify hot-spots. The usefulness of such data-driven approaches is limited when historical data is scarce or unavailable (e.g. planning of a new city) or insufficiently representative (e.g. does not account for novel events, such as Olympic Games). In some cities, crime data is not systematically accumulated at all. We present a graph-constrained agent based simulation model of alcohol-related violent crime that is capable of predicting areas of likely violent crime without requiring any historical data. The only inputs to our simulation are publicly available geographical data, which makes our method immediately applicable to a wide range of tasks, such as optimal city planning, police patrol optimisation, devising alcohol licensing policies. In experiments, we evaluate our model and demonstrate agreement of our model's predictions on where and when violence will occur with real-world violent crime data. Analyses indicate that our agent based model may be able to make a significant contribution to attempts to prevent violence through deterrence or by design. Joseph Redfern, Kirill A. Sidorov, Paul L. Rosin, Simon C. Moore, Padraig Corcoran, David Marshall 0001 |
AVSS | 3 |
| 2017 | 4D Analysis of Facial Ageing Using Dynamic FeaturesabstractFacial ageing analysis based on 4D data (3D plus time) is much more robust to pose changes and illumination variations than using 2D image and video. The purpose of this investigation was to measure the effects of age and gender related facial changes using dynamic 3D facial scans. Experiments were carried out on the subjects, who were divided into two groups by age (15-30 years and 31-60 years). Each group was further subdivided by gender. 3D scans of the subjects were processed to extract facial features which were tracked through the duration of the data capture. Subsequently, a set of dynamic features were computed from these facial features, as well as static features for comparison. Two-way multivariate analysis of variance (MANOVA) of these features demonstrated that statistically significant age and gender related differences could be detected. We show that 3D facial dynamics provide more useful information than static features for the characterisation of smiles. Khtam Al-Meyah, David Marshall 0001, Paul L. Rosin |
KES | 3 |
| 2017 | Developing and applying a benchmark for evaluating image stylization
David Mould, Paul L. Rosin |
Comput. Graph. | 2 |
| 2017 | Practical automatic background substitution for live videoabstractIn this paper we present a novel automatic background substitution approach for live video. The objective of background substitution is to extract the foreground from the input video and then combine it with a new background. In this paper, we use a color line model to improve the Gaussian mixture model in the background cut method to obtain a binary foreground segmentation result that is less sensitive to brightness differences. Based on the high quality binary segmentation results, we can automatically create a reliable trimap for alpha matting to refine the segmentation boundary. To make the composition result more realistic, an automatic foreground color adjustment step is added to make the foreground look consistent with the new background. Compared to previous approaches, our method can produce higher quality binary segmentation results, and to the best of our knowledge, this is the first time such an automatic and integrated background substitution system has been proposed which can run in real time, which makes it practical for everyday applications. Hao-Zhi Huang 0001, Xiaonan Fang 0001, Yufei Ye 0001, Song-Hai Zhang, Paul L. Rosin |
Comput. Vis. Media | 5 |
| 2017 | Example-based image colorization via automatic feature selection and fusion
Bo Li 0023, Yukun Lai, Paul L. Rosin |
Neurocomputing | 3 |
| 2017 | Intelligent Visual Media Processing: When Graphics Meets Vision
Ming-Ming Cheng, Qibin Hou, Song-Hai Zhang, Paul L. Rosin |
J. Comput. Sci. Technol. | 4 |
| 2017 | Detecting violent and abnormal crowd activity using temporal analysis of grey level co-occurrence matrix (GLCM)-based texture measuresabstractThe severity of sustained injury resulting from assault-related violence can be minimised by reducing detection time. However, it has been shown that human operators perform poorly at detecting events found in video footage when presented with simultaneous feeds. We utilise computer vision techniques to develop an automated method of abnormal crowd detection that can aid a human operator in the detection of violent behaviour. We observed that behaviour in city centre environments often occurs in crowded areas, resulting in individual actions being occluded by other crowd members. We propose a real-time descriptor that models crowd dynamics by encoding changes in crowd texture using temporal summaries of grey level co-occurrence matrix features. We introduce a measure of inter-frame uniformity and demonstrate that the appearance of violent behaviour changes in a less uniform manner when compared to other types of crowd behaviour. Our proposed method is computationally cheap and offers real-time description. Evaluating our method using a privately held CCTV dataset and the publicly available Violent Flows, UCF Web Abnormality and UMN Abnormal Crowd datasets, we report a receiver operating characteristic score of 0.9782, 0.9403, 0.8218 and 0.9956, respectively. Kaelon Lloyd, Paul L. Rosin, David Marshall 0001, Simon C. Moore |
Mach. Vis. Appl. | 2 |
| 2017 | Example-Based Image Colorization Using Locality Consistent Sparse RepresentationabstractImage colorization aims to produce a natural looking color image from a given gray-scale image, which remains a challenging problem. In this paper, we propose a novel example-based image colorization method exploiting a new locality consistent sparse representation. Given a single reference color image, our method automatically colorizes the target gray-scale image by sparse pursuit. For efficiency and robustness, our method operates at the superpixel level. We extract low-level intensity features, mid-level texture features, and high-level semantic features for each superpixel, which are then concatenated to form its descriptor. The collection of feature vectors for all the superpixels from the reference image composes the dictionary. We formulate colorization of target superpixels as a dictionary-based sparse reconstruction problem. Inspired by the observation that superpixels with similar spatial location and/or feature representation are likely to match spatially close regions from the reference image, we further introduce a locality promoting regularization term into the energy formulation, which substantially improves the matching consistency and subsequent colorization results. Target superpixels are colorized based on the chrominance information from the dominant reference superpixels. Finally, to further improve coherence while preserving sharpness, we develop a new edge-preserving filter for chrominance channels with the guidance from the target gray-scale image. To the best of our knowledge, this is the first work on sparse pursuit image colorization from single reference images. Experimental results demonstrate that our colorization method outperforms the state-of-the-art methods, both visually and quantitatively using a user study. Bo Li 0023, Fuchen Zhao, Zhuo Su 0001, Xiangguo Liang, Yukun Lai, Paul L. Rosin |
IEEE Trans. Image Process. | 6 |
| 2016 | Skeleton-based canonical forms for non-rigid 3D shape retrievalabstractThe retrieval of non-rigid 3D shapes is an important task. A common technique is to simplify this problem to a rigid shape retrieval task by producing a bending-invariant canonical form for each shape in the dataset to be searched. It is common for these techniques to attempt to “unbend” a shape by applying multidimensional scaling (MDS) to the distances between points on the mesh, but this leads to unwanted local shape distortions. We instead perform the unbending on the skeleton of the mesh, and use this to drive the deformation of the mesh itself. This leads to computational speed-up, and reduced distortion of local shape detail. We compare our method against other canonical forms: our experiments show that our method achieves state-of-the-art retrieval accuracy in a recent canonical forms benchmark, and only a small drop in retrieval accuracy over the state-of-the-art in a second recent benchmark, while being significantly faster. David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin |
Comput. Vis. Media | 3 |
| 2016 | Shape Retrieval of Non-rigid 3D Human Modelsabstract3D models of humans are commonly used within computer graphics and vision, and so the ability to distinguish between body shapes is an important shape retrieval problem. We extend our recent paper which provided a benchmark for testing non-rigid 3D shape retrieval algorithms on 3D human models. This benchmark provided a far stricter challenge than previous shape benchmarks. We have added 145 new models for use as a separate training set, in order to standardise the training data used and provide a fairer comparison. We have also included experiments with the FAUST dataset of human scans. All participants of the previous benchmark study have taken part in the new tests reported here, many providing updated results using the new data. In addition, further participants have also taken part, and we provide extra analysis of the retrieval results. A total of 25 different shape retrieval methods are compared. David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Zhouhui Lian, Masaki Aono, A. Ben Hamza, Alexander M. Bronstein, Michael M. Bronstein, S. Bu, Umberto Castellani, S. Cheng, Valeria Garro, Andrea Giachetti 0001, Afzal Godil, Luca Isaia, Henry Johan, Long Lai, Bo Li 0013, Chenfeng Li, Hai-Sheng Li 0002, Roee Litman, Yijuan Lu, Li Sun 0004, Gary K. L. Tam, Atsushi Tatsuma, Jianbo Ye |
Int. J. Comput. Vis. | 3 |
| 2016 | Improving Shape from Shading with Interactive Tabu Search
Jing Wu 0004, Paul L. Rosin, Xianfang Sun, Ralph R. Martin |
J. Comput. Sci. Technol. | 2 |
| 2016 | Combining cellular automata and local binary patterns for copy-move forgery detection
Dijana Tralic, Sonja Grgic, Xianfang Sun, Paul L. Rosin |
Multim. Tools Appl. | 4 |
| 2016 | Measuring linearity of curves in 2D and 3D
Paul L. Rosin, Jovanka Pantovic, Jovisa D. Zunic |
Pattern Recognit. | 1 |
| 2015 | Towards 4D Coupled Models of Conversational Facial Expression InteractionsabstractIn this paper we introduce a novel approach for building 4D coupled statistical models of conversational facial expression interactions. To build these coupled models we use 3D AAMs for feature extraction, 4D polynomial fitting for sequence representation, and concatenated feature vectors of frontchannel-backchannel interactions (with offset values) for the coupled model. Using a coupled model of conversation smile interactions, we predicted each sequence’s backchannel signal. In a subsequent experiment, human observers rated predicted sequences as highly similar to the originals. Our results demonstrate the usefulness of coupled models as powerful tools to analyse and synthesise key aspects of conversational interactions, including conversation timings, backchannel responses to frontchannel signals, and the spatial and temporal dynamics of conversational facial expression interactions. Jason Vandeventer, Lukas Gräser, Magdalena Rychlowska, Paul L. Rosin, David Marshall 0001 |
BMVC | 4 |
| 2015 | Improved DSIFT Descriptor Based Copy-Rotate-Move Forgery Detection
Ali Retha Hasoon Khayeat, Xianfang Sun, Paul L. Rosin |
PSIVT | 3 |
| 2015 | Feature Neighbourhood Mutual Information for multi-modal image registration: An application to eye fundus imaging
Philip A. Legg, Paul L. Rosin, David Marshall 0001, James E. Morgan |
Pattern Recognit. | 2 |
| 2015 | Euclidean-distance-based canonical forms for non-rigid 3D shape retrievalabstractRetrieval of 3D shapes is a challenging problem, especially for non-rigid shapes. One approach giving favourable results uses multidimensional scaling (MDS) to compute a canonical form for each mesh, after which rigid shape matching can be applied. However, a drawback of this method is that it requires geodesic distances to be computed between all pairs of mesh vertices. Due to the super-quadratic computational complexity, canonical forms can only be computed for low-resolution meshes. We suggest a linear time complexity method for computing a canonical form, using Euclidean distances between pairs of a small subset of vertices. This approach has comparable retrieval accuracy but lower time complexity than using global geodesic distances, allowing it to be used on higher resolution meshes, or for more meshes to be considered within a time budget. David Pickup, Xianfang Sun, Paul L. Rosin, Ralph R. Martin |
Pattern Recognit. | 3 |
| 2014 | An Efficient Approach to Correspondences between Multiple Non-Rigid PartsabstractAbstract Identifying multiple deformable parts on meshes and establishing dense correspondences between them are tasks of fundamental importance to computer graphics, with applications to e.g. geometric edit propagation and texture transfer. Much research has considered establishing correspondences between non‐rigid surfaces, but little work can both identify similar multiple deformable partsandhandle partial shape correspondences. This paper addresses two related problems, treating them as a whole: (i) identifying similar deformable parts on a mesh, related by anon‐rigidtransformation to a given query part, and (ii) establishing dense point correspondences automatically between such parts. We show that simple and efficient techniques can be developed if we make the assumption that these parts locally undergo isometric deformation. Our insight is that similar deformable parts are suggested by large clusters of point correspondences that are isometrically consistent. Once such parts are identified,densepoint correspondences can be obtained by an iterative propagation process. Our techniques are applicable to models with arbitrary topology. Various examples demonstrate the effectiveness of our techniques. Gary K. L. Tam, Ralph R. Martin, Paul L. Rosin, Yukun Lai |
Comput. Graph. Forum | 3 |
| 2014 | Use of non-photorealistic rendering and photometric stereo in making bas-reliefs from photographs
Jing Wu 0004, Ralph R. Martin, Paul L. Rosin, Xianfang Sun, Yukun Lai, Christian Wallraven |
Graph. Model. | 3 |
| 2014 | Facial expression recognition in dynamic sequences: An integrated approach
Hui Fang 0003, Neil Mac Parthaláin, Andrew J. Aubrey, Gary K. L. Tam, Rita Borgo, Paul L. Rosin, Phil W. Grant, David Marshall 0001, Min Chen 0001 |
Pattern Recognit. | 6 |
| 2014 | Virtual unrolling and information recovery from scanned scrolled historical documents
Oksana Samko, Yukun Lai, David Marshall 0001, Paul L. Rosin |
Pattern Recognit. | 4 |
| 2014 | Scan integration as a labelling problem
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
Pattern Recognit. | 4 |
| 2014 | Efficient Circular ThresholdingabstractOtsu's algorithm for thresholding images is widely used, and the computational complexity of determining the threshold from the histogram is O(N) where N is the number of histogram bins. When the algorithm is adapted to circular rather than linear histograms then two thresholds are required for binary thresholding. We show that, surprisingly, it is still possible to determine the optimal threshold in O(N) time. The efficient optimal algorithm is over 300 times faster than traditional approaches for typical histograms and is thus particularly suitable for real-time applications. We further demonstrate the usefulness of circular thresholding using the adapted Otsu criterion for various applications, including analysis of optical flow data, indoor/outdoor image classification, and non-photorealistic rendering. In particular, by combining circular Otsu feature with other colour/texture features, a 96.9% correct rate is obtained for indoor/outdoor classification on the well known IITM-SCID2 data set, outperforming the state-of-the-art result by 4.3%. Yukun Lai, Paul L. Rosin |
IEEE Trans. Image Process. | 2 |
| 2014 | Mesh saliency via spectral processingabstractWe propose a novel method for detecting mesh saliency, a perceptually-based measure of the importance of a local region on a 3D surface mesh. Our method incorporates global considerations by making use of spectral attributes of the mesh, unlike most existing methods which are typically based on local geometric cues. We first consider the properties of the log-Laplacian spectrum of the mesh. Those frequencies which show differences from expected behaviour capture saliency in the frequency domain. Information about these frequencies is considered in the spatial domain at multiple spatial scales to localise the salient features and give the final salient areas. The effectiveness and robustness of our approach are demonstrated by comparisons to previous approaches on a range of test models. The benefits of the proposed method are further evaluated in applications such as mesh simplification, mesh segmentation, and scan integration, where we show how incorporating mesh saliency can provide improved results. Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
ACM Trans. Graph. | 4 |
| 2014 | Diffusion pruning for rapidly and robustly selecting global correspondences using local isometryabstractFinding correspondences between two surfaces is a fundamental operation in various applications in computer graphics and related fields. Candidate correspondences can be found by matching local signatures, but as they only consider local geometry, many are globally inconsistent. We provide a novel algorithm to prune a set of candidate correspondences to those most likely to be globally consistent. Our approach can handle articulated surfaces, and ones related by a deformation which is globally nonisometric, provided that the deformation is locally approximately isometric. Our approach uses an efficient diffusion framework, and only requires geodesic distance calculations in small neighbourhoods, unlike many existing techniques which require computation of global geodesic distances. We demonstrate that, for typical examples, our approach provides significant improvements in accuracy, yet also reduces time and memory costs by a factor of several hundred compared to existing pruning techniques. Our method is furthermore insensitive to holes, unlike many other methods. Gary K. L. Tam, Ralph R. Martin, Paul L. Rosin, Yukun Lai |
ACM Trans. Graph. | 3 |
| 2014 | Artistic rendering enhancing global structure
Yukun Lai, Paul L. Rosin |
Vis. Comput. | 2 |
| 2013 | Measuring Linearity of Curves
Jovisa D. Zunic, Jovanka Pantovic, Paul L. Rosin |
ICPRAM | 3 |
| 2013 | Making bas-reliefs from photographs of human faces
Jing Wu 0004, Ralph R. Martin, Paul L. Rosin, Xianfang Sun, Frank C. Langbein, Yukun Lai, David Marshall 0001 |
Comput. Aided Des. | 3 |
| 2013 | Artistic minimal rendering with lines and blocks
Paul L. Rosin, Yukun Lai |
Graph. Model. | 1 |
| 2013 | Visualizing Natural Image StatisticsabstractNatural image statistics is an important area of research in cognitive sciences and computer vision. Visualization of statistical results can help identify clusters and anomalies as well as analyze deviation, distribution, and correlation. Furthermore, they can provide visual abstractions and symbolism for categorized data. In this paper, we begin our study of visualization of image statistics by considering visual representations of power spectra, which are commonly used to visualize different categories of images. We show that they convey a limited amount of statistical information about image categories and their support for analytical tasks is ineffective. We then introduce several new visual representations, which convey different or more information about image statistics. We apply ANOVA to the image statistics to help select statistically more meaningful measurements in our design process. A task-based user evaluation was carried out to compare the new visual representations with the conventional power spectra plots. Based on the results of the evaluation, we made further improvement of visualizations by introducing composite visual representations of image statistics. Hui Fang 0003, Gary K. L. Tam, Rita Borgo, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, Christian Wallraven, Douglas W. Cunningham, David Marshall 0001, Min Chen 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2013 | Abstract Art by Shape ClassificationabstractThis paper shows that classifying shapes is a tool useful in nonphotorealistic rendering (NPR) from photographs. Our classifier inputs regions from an image segmentation hierarchy and outputs the "best" fitting simple shape such as a circle, square, or triangle. Other approaches to NPR have recognized the benefits of segmentation, but none have classified the shape of segments. By doing so, we can create artwork of a more abstract nature, emulating the style of modern artists such as Matisse and other artists who favored shape simplification in their artwork. The classifier chooses the shape that "best" represents the region. Since the classifier is trained by a user, the "best shape" has a subjective quality that can over-ride measurements such as minimum error and more importantly captures user preferences. Once trained, the system is fully automatic, although simple user interaction is also possible to allow for differences in individual tastes. A gallery of results shows how this classifier contributes to NPR from images by producing abstract artwork. Yi-Zhe Song, David Pickup, Chuan Li 0001, Paul L. Rosin, Peter Hall 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2013 | Registration of 3D Point Clouds and Meshes: A Survey from Rigid to NonrigidabstractThree-dimensional surface registration transforms multiple three-dimensional data sets into the same coordinate system so as to align overlapping components of these sets. Recent surveys have covered different aspects of either rigid or nonrigid registration, but seldom discuss them as a whole. Our study serves two purposes: 1) To give a comprehensive survey of both types of registration, focusing on three-dimensional point clouds and meshes and 2) to provide a better understanding of registration from the perspective of data fitting. Registration is closely related to data fitting in which it comprises three core interwoven components: model selection, correspondences and constraints, and optimization. Study of these components 1) provides a basis for comparison of the novelties of different techniques, 2) reveals the similarity of rigid and nonrigid registration in terms of problem representations, and 3) shows how overfitting arises in nonrigid registration and the reasons for increasing interest in intrinsic techniques. We further summarize some practical issues of registration which include initializations and evaluations, and discuss some of our own observations, insights and foreseeable research trends. Gary K. L. Tam, Zhi-Quan Cheng, Yukun Lai, Frank C. Langbein, Yonghuai Liu, David Marshall 0001, Ralph R. Martin, Xianfang Sun, Paul L. Rosin |
IEEE Trans. Vis. Comput. Graph. | 9 |
| 2013 | 3D point of interest detection via spectral irregularity diffusion
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
Vis. Comput. | 4 |
| 2012 | Measuring Linearity of Closed Curves and Connected Compound Curves
Paul L. Rosin, Jovanka Pantovic, Jovisa D. Zunic |
ACCV (3) | 1 |
| 2012 | A new convexity measurement for 3D meshesabstractThis paper presents a novel convexity measurement for 3D meshes. The new convexity measure is calculated by minimizing the ratio of the summed area of valid regions in a mesh's six views, which are projected on faces of the bounding box whose edges are parallel to the coordinate axes, to the sum of three orthogonal projected areas of the mesh. The complete definition, theoretical analysis, and a computing algorithm of our convexity measure are explicitly described. This paper also proposes a new 3D shape descriptor CD (i.e., Convexity Distribution) based on the distribution of above-mentioned ratios, which are computed by randomly rotating the mesh around its center, to better describe the object's convexity-related properties compared to existing convexity measurements. Our experiments not only show that the proposed convexity measure corresponds well with human intuition, but also demonstrate the effectiveness of the new convexity measure and the new shape descriptor by significantly improving the performance of other methods in the application of 3D shape retrieval. Zhouhui Lian, Afzal Godil, Paul L. Rosin, Xianfang Sun |
CVPR | 3 |
| 2012 | Saliency-guided integration of multiple scansabstractWe present a novel method to integrate multiple 3D scans captured from different viewpoints. Saliency information is used to guide the integration process. The multi-scale saliency of a point is specifically designed to reflect its sensitivity to registration errors. Then scans are partitioned into salient and non-salient regions through an Markov Random Field (MRF) framework where neighbourhood consistency is incorporated to increase the robustness against potential scanning errors. We then develop different schemes to discriminatively integrate points in the two regions. For the points in salient regions which are more sensitive to registration errors, we employ the Iterative Closest Point algorithm to compensate the local registration error and find the correspondences for the integration. For the points in non-salient regions which are less sensitive to registration errors, we integrate them via an efficient and effective point-shifting scheme. A comparative study shows that the proposed method delivers improved surface integration. Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
CVPR | 4 |
| 2012 | Hybrid phoneme based clustering approach for audio driven facial animationabstractWe consider the problem of producing accurate facial animation corresponding to a given input speech signal. A popular technique previously used for Audio Driven Facial Animation is to build a joint audio-visual model using Active Appearance Models (AAMs) to represent possible facial variations and Hidden Markov Models (HMMs) to select the correct appearance based on the input audio. However there are several questions that remained unanswered. In particular the choice of clustering technique and the choice of the number of clusters in the HMM may have significant influence over the quality of the produced videos. We have investigated a range of clustering techniques in order to improve the quality of the HMM produced, and proposed a new structure based on using Gaussian Mixture Models (GMMs) to model each phoneme separately. We compared our approach to several alternatives using a public dataset of 300 phonetically labeled sentences spoken by a single person and found that our approach produces more accurate animation. In addition, we use a hybrid approach where the training data is phonetically labeled thus producing a model with better separation of phonemes, but test audio data is not labeled, thus making our approach for generating facial animation less laborious and fully automatic. Benjamin Havell, Paul L. Rosin, Saeid Sanei, Andrew J. Aubrey, David Marshall 0001, Yulia Hicks |
ICASSP | 2 |
| 2012 | Conditional random field-based mesh saliencyabstractWe propose a new method for detecting mesh saliency, a reflection of perception-based regional importance for 3D meshes. The basic idea is to incorporate the Conditional Random Field (CRF) framework with a saliency detection process. We first produce a multi-scale representation for a mesh. Then, a CRF is designed to robustly detect salient regions utilising neighbourhood consistency. By inferring the CRF via belief propagation algorithm, we actually make use of the global statistic information in the saliency detection process. Experimental results demonstrate the robustness and the effectiveness of the proposed method. Ran Song 0001, Yonghuai Liu, Yitian Zhao, Ralph R. Martin, Paul L. Rosin |
ICIP | 5 |
| 2011 | Segmentation of Parchment Scrolls for Virtual UnrollingabstractIn this paper we introduce a framework for the segmentation of scanned scrolled parchments, based on a novel graph cut based approach with an additional shape prior, in combination with anisotropic diffusion and geometry-constrained postprocessing. This problem has not been investigated by the computer vision community properly yet due to the parchment scanning technology novelty, and is extremely important for effective data recovery from historical scrolled documents whose content is inaccessible due to the deterioration of the parchment. To date, parchment segmentation has required user interaction, which is very time consuming for such data. We demonstrate with real examples how our algorithm is able to solve the major problem for scrolled parchment analysis, namely segment connected layers, and process the data without user interaction. Oksana Samko, Yukun Lai, David Marshall 0001, Paul L. Rosin |
BMVC | 4 |
| 2011 | Shape Description by Bending Invariant Moments
Paul L. Rosin |
CAIP (1) | 1 |
| 2011 | Mapping and manipulating facial dynamicsabstractThis paper describes a novel approach to building models of temporal dynamics for facial animation with applications in performing perceptual testing of trustworthiness. A vital component of the system is a method to bring two image sequences into temporal alignment. Our approach is to project the two sequences into face space (built using shape models [1]) and apply dynamic time warping (DTW). However, the variability in the sequences causes the standard DTW algorithm to perform poorly on our data, and so we have overcome this by extending DTW in the following ways: 1) the signal magnitudes are augmented by incorporating derivatives [2], and a scheme for estimating weights in the cost function is proposed, 2) the set of sequences is used to build a graph, with nodes representing sequences and edges indicating the cost of applying the extended DTW to align pairs of sequences; better alignments between sequences can now be found by traversing the minimum cost path through the graph. Once all signals are aligned to a common temporal reference it is straightforward to map the temporal dynamics from one face to another. A remapped face is synthesised using the new trajectory in face space to drive an active appearance model [1]. Furthermore, the common temporal reference allows us to build a statistical model of the dynamics. This can be used to both identify dynamics of interest and also to manipulate the dynamics, e.g. to reduce or exaggerate facial dynamics. Andrew J. Aubrey, Vedran Kajic, Ivana Cingovska, Paul L. Rosin, David Marshall 0001 |
FG | 4 |
| 2011 | MRF-based automatic image ordering and its application to mosaicingabstractA fast and robust auto-sorting method for image ordering based on Markov Random Fields (MRF) is proposed. We present a specific MRF model for the ordering problem and use pairwise phase correlation for the formulation. The MRF is inferred by a modified belief propagation (BP) method. Experimental results prove that the new method can reorder a disorganised collection of images without human input, prior information or restrictions, as just the first stage of a multi stage mosaicing process, but also provides information that can be used to guide a mosaicing process in order to reduce both local mismatch and global error accumulation. Ran Song 0001, Yonghuai Liu, Yitian Zhao, Ralph R. Martin, Paul L. Rosin |
ICASSP | 5 |
| 2011 | Visualization of Time-Series Data in Parameter Space for Understanding Facial DynamicsabstractAbstract Over the past decade, computer scientists and psychologists have made great efforts to collect and analyze facial dynamics data that exhibit different expressions and emotions. Such data is commonly captured as videos and are transformed into feature‐based time‐series prior to any analysis. However, the analytical tasks, such as expression classification, have been hindered by the lack of understanding of the complex data space and the associated algorithm space. Conventional graph‐based time‐series visualization is also found inadequate to support such tasks. In this work, we adopt a visual analytics approach by visualizing the correlation between the algorithm space and our goal – classifying facial dynamics. We transform multiple feature‐based time‐series for each expression in measurement space to a multi‐dimensional representation in parameter space. This enables us to utilize parallel coordinates visualization to gain an understanding of the algorithm space, providing a fast and cost‐effective means to support the design of analytical algorithms. Gary K. L. Tam, Hui Fang 0003, Andrew J. Aubrey, Phil W. Grant, Paul L. Rosin, David Marshall 0001, Min Chen 0001 |
Comput. Graph. Forum | 5 |
| 2011 | Measuring linearity of open planar curve segments
Jovisa D. Zunic, Paul L. Rosin |
Image Vis. Comput. | 2 |
| 2011 | Orientation and anisotropy of multi-component shapes from boundary information
Paul L. Rosin, Jovisa D. Zunic |
Pattern Recognit. | 1 |
| 2011 | Fast Rule Identification and Neighborhood Selection for Cellular AutomataabstractCellular automata (CA) with given evolution rules have been widely investigated, but the inverse problem of extracting CA rules from observed data is less studied. Current CA rule extraction approaches are both time consuming and inefficient when selecting neighborhoods. We give a novel approach to identifying CA rules from observed data and selecting CA neighborhoods based on the identified CA model. Our identification algorithm uses a model linear in its parameters and gives a unified framework for representing the identification problem for both deterministic and probabilistic CA. Parameters are estimated based on a minimum variance criterion. An incremental procedure is applied during CA identification to select an initial coarse neighborhood. Redundant cells in the neighborhood are then removed based on parameter estimates, and the neighborhood size is determined using the Bayesian information criterion. Experimental results show the effectiveness of our algorithm and that it outperforms other leading CA identification algorithms. Xianfang Sun, Paul L. Rosin, Ralph R. Martin |
IEEE Trans. Syst. Man Cybern. Part B | 2 |
| 2010 | MRF Labeling for Multi-view Range Image Integration
Ran Song 0001, Yonghuai Liu, Ralph R. Martin, Paul L. Rosin |
ACCV (2) | 4 |
| 2010 | Image processing using 3-state cellular automata
Paul L. Rosin |
Comput. Vis. Image Underst. | 1 |
| 2010 | Rectilinearity of 3D Meshes
Zhouhui Lian, Paul L. Rosin, Xianfang Sun |
Int. J. Comput. Vis. | 2 |
| 2010 | A Hu moment invariant as a shape circularity measure
Jovisa D. Zunic, Kaoru Hirota, Paul L. Rosin |
Pattern Recognit. | 3 |
| 2010 | Assessing the Uniqueness and Permanence of Facial Actions for Use in Biometric ApplicationsabstractAlthough the human face is commonly used as a physiological biometric, very little work has been done to exploit the idiosyncrasies of facial motions for person identification. In this paper, we investigate theuniquenessandpermanenceof facial actions to determine whether these can be used as a behavioral biometric. Experiments are carried out using 3-D video data of participants performing a set of very short verbal and nonverbal facial actions. The data have been collected over long time intervals to assess the variability of the subjects' emotional and physical conditions. Quantitative evaluations are performed for both the identification and the verification problems; the results indicate that emotional expressions (e.g., smile and disgust) are not sufficiently reliable for identity recognition in real-life situations, whereas speech-related facial movements show promising potential. Lanthao Benedikt, Darren Cosker, Paul L. Rosin, David Marshall 0001 |
IEEE Trans. Syst. Man Cybern. Part A | 3 |
| 2009 | A Robust Solution to Multi-modal Image Registration by Combining Mutual Information with Multi-scale Derivatives
Philip A. Legg, Paul L. Rosin, David Marshall 0001, James E. Morgan |
MICCAI (1) | 2 |
| 2009 | Rapid and effective segmentation of 3D models using random walks
Yukun Lai, Shi-Min Hu 0001, Ralph R. Martin, Paul L. Rosin |
Comput. Aided Geom. Des. | 4 |
| 2009 | Real-time content-aware image resizing
Hua Huang 0001, TianNan Fu, Paul L. Rosin, Chun Qi |
Sci. China Ser. F Inf. Sci. | 3 |
| 2009 | Expressive line drawings of human faces from range images
Yuezhu Huang, Ralph R. Martin, Paul L. Rosin, Xiangxu Meng, Chenglei Yang |
Sci. China Ser. F Inf. Sci. | 3 |
| 2009 | Noise analysis and synthesis for 3D laser depth scanners
Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Frank C. Langbein |
Graph. Model. | 2 |
| 2009 | An Alternative Approach to Computing Shape Orientation with an Application to Compound Shapes
Jovisa D. Zunic, Paul L. Rosin |
Int. J. Comput. Vis. | 2 |
| 2009 | Edge-Aware Level Set Diffusion and Bilateral Filtering Reconstruction for Image Magnification
Hua Huang 0001, Paul L. Rosin, Chun Qi |
J. Comput. Sci. Technol. | 3 |
| 2009 | A simple method for detecting salient regions
Paul L. Rosin |
Pattern Recognit. | 1 |
| 2009 | Classification of pathological shapes using convexity measures
Paul L. Rosin |
Pattern Recognit. Lett. | 1 |
| 2009 | Bas-Relief Generation Using Adaptive Histogram EqualizationabstractAn algorithm is presented to automatically generate bas-reliefs based on adaptive histogram equalization (AHE), starting from an input height field. A mesh model may alternatively be provided, in which case a height field is first created via orthogonal or perspective projection. The height field is regularly gridded and treated as an image, enabling a modified AHE method to be used to generate a bas-relief with a user-chosen height range. We modify the original image-contrast-enhancement AHE method to use gradient weights also to enhance the shape features of the bas-relief. To effectively compress the height field, we limit the height-dependent scaling factors used to compute relative height variations in the output from height variations in the input; this prevents any height differences from having too great effect. Results of AHE over different neighborhood sizes are averaged to preserve information at different scales in the resulting bas-relief. Compared to previous approaches, the proposed algorithm is simple and yet largely preserves original shape features. Experiments show that our results are, in general, comparable to and in some cases better than the best previously published methods. Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Frank C. Langbein |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2008 | Shapes Fit For PurposeabstractThis paper is about shape fitting to regions that segment an image and some applications that rely on the abstraction that offers. The novelty lies in three areas: (1) we fit a shape drawn from a selection of shape families, not just one class of shape, using a supervised classifier; (2) We use results from the classifier to match photographs and artwork of particular objects using a few qualitative shapes, which overcomes the significant differences between photographs and paintings; (3) We further use the shape classifier to process photographs into abstract synthetic art which, so far as we know, is novel too. Thus we use our shape classier in both discriminative (matching) and generative (image synthesis) tasks. We conclude the level of abstraction offered by our shape classifier is novel and useful. 1 Anupriya Balikai, Paul L. Rosin, Yi-Zhe Song, Peter Hall 0001 |
BMVC | 2 |
| 2008 | Facial Dynamics in Biometric IdentificationabstractThis paper investigates the use of facial gestures for identity recognition. This is the first time that such a quantitative evaluation is conducted, comparing the analyses of 2D versus 3D dynamic data of verbal and nonverbal facial actions. Suitable data processing and feature extraction methods are examined, then a number of pattern matching techniques including the Fréchet distance, Correlation Coefficients, Hidden-Markov Models, Dynamic Time Warping and its derived forms are compared, in light of which an improved algorithm is proposed. Finally, a face recognition prototype using facial dynamics is built, achieving an Equal Error Rate EER=1.6%. 1 Lanthao Benedikt, Vedran Kajic, Darren Cosker, Paul L. Rosin, David Marshall 0001 |
BMVC | 4 |
| 2008 | Fast mesh segmentation using random walksabstract3D mesh models are now widely available for use in various applications. The demand for automatic model analysis and understanding is ever increasing. Mesh segmentation is an important step towards model understanding, and acts as a useful tool for different mesh processing applications, e.g. reverse engineering and modeling by example. We extend a random walk method used previously for image segmentation to give algorithms for both interactive and automatic mesh segmentation. This method is extremely efficient, and scales almost linearly with increasing number of faces. For models of moderate size, interactive performance is achieved with commodity PCs. It is easy-to-implement, robust to noise in the mesh, and yields results suitable for downstream applications for both graphical and engineering models. Yukun Lai, Shi-Min Hu 0001, Ralph R. Martin, Paul L. Rosin |
Symposium on Solid and Physical Modeling | 4 |
| 2008 | Noise in 3D laser range scanner dataabstractThis paper discusses noise in range data measured by a Konica Minolta Vivid 910 scanner. Previous papers considering denoising 3D mesh data have often used artificial data comprising Gaussian noise, which is independently distributed at each mesh point. Measurements of an accurately machined, almost planar test surface indicate that real scanner data does not have such properties. An initial characterisation of real scanner noise for this test surface shows that the errors are not quite Gaussian, and more importantly, exhibit significant short range correlation. This analysis yields a simple model for generating noise with similar characteristics. We also examine the effect of two typical mesh denoising algorithms on the real noise present in the test data. The results show that new denoising algorithms are required to effectively remove real scanner noise. Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Frank C. Langbein |
Shape Modeling International | 2 |
| 2008 | Random walks for feature-preserving mesh denoising
Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Frank C. Langbein |
Comput. Aided Geom. Des. | 2 |
| 2008 | A two-component rectilinearity measure
Paul L. Rosin |
Comput. Vis. Image Underst. | 1 |
| 2007 | Measuring the Orientability of Shapes
Paul L. Rosin |
CAIP | 1 |
| 2007 | A Definition for Orientation for Multiple Component Shapes
Jovisa D. Zunic, Paul L. Rosin |
CAIP | 2 |
| 2007 | Random walks for mesh denoisingabstractThis paper considers an approach to mesh denoising based on the concept of random walks. The proposed method consists of two stages: a face normal filtering procedure, followed by a vertex position updating procedure which integrates the denoised face normals in a least-squares sense. Face normal filtering is performed by weighted averaging of normals in a neighbourhood. The weights are based on the probability of arriving at a given neighbour after a random walk of a virtual particle starting at a given face of the mesh and moving a fixed number of steps. The probability of a particle stepping from its current face to a given neighboring face is determined by the angle between the two face normals, using a Gaussian distribution whose width is adaptively adjusted to enhance the feature-preserving property of the algorithm. The vertex position updating procedure uses the conjugate gradient algorithm for speed of convergence. Analysis and experiments show that random walks of different step lengths yield similar denoising results. In particular, iterative application of a one-step random walk in a progressive manner effectively preserves detailed features while denoising the mesh very well. We observe that this approach is faster than many other feature-preserving mesh denoising algorithms. Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Frank C. Langbein |
Symposium on Solid and Physical Modeling | 2 |
| 2007 | Evaluating Harker and O'Leary's distance approximation for ellipse fitting
Paul L. Rosin |
Pattern Recognit. Lett. | 1 |
| 2007 | Fast and Effective Feature-Preserving Mesh DenoisingabstractWe present a simple and fast mesh denoising method, which can remove noise effectively, while preserving mesh features such as sharp edges and corners. The method consists of two stages. Firstly, noisy face normals are filtered iteratively by weighted averaging of neighboring face normals. Secondly, vertex positions are iteratively updated to agree with the denoised face normals. The weight function used during normal filtering is much simpler than that used in previous similar approaches, being simply a trimmed quadratic. This makes the algorithm both fast and simple to implement. Vertex position updating is based on the integration of surface normals using a least-squares error criterion. Like previous algorithms, we solve the least-squares problem by gradient descent, but whereas previous methods needed user input to determine the iteration step size, we determine it automatically. In addition, we prove the convergence of the vertex position updating approach. Analysis and experiments show the advantages of our proposed method over various earlier surface denoising methods. Xianfang Sun, Paul L. Rosin, Ralph R. Martin, Frank C. Langbein |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2006 | Shape Orientability
Jovisa D. Zunic, Paul L. Rosin, Lazar Kopanja |
ACCV (2) | 2 |
| 2006 | Segmenting reliefs on triangle meshesabstractSculptural reliefs are widely used in various industries for purposes such as applying brands to packaging and decorating porcelain. In order to easily apply reliefs to CAD models, it is often desirable to reverse-engineer previously designed and manufactured reliefs. 3D scanners can generate triangle meshes from objects with reliefs; however, previous mesh segmentation work has not considered the particular problem of separation of reliefs from background. We consider here the specific case of segmenting a simple relief delimited by a single outer contour, which lies on a smooth, slowly varying background. Generally, such reliefs meet the surrounding surface in a small step, enabling us to devise a specific method for such relief segmentation.We find the boundary between the background and the relief using an adaptive snake. It starts at a simple user-drawn contour, and is driven inwards by a collapsing force until it matches the relief's boundary. Our method is insensitive to the choice of the initial contour. The snake's limiting position is controlled by a feature energy term designed to find a step. A refinement strategy is then used to drive the snake into concavities of the relief contour.We demonstrate operation of our algorithm using real scanned models with different relief contour shapes and triangle meshes with different resolutions. Shenglan Liu 0005, Ralph R. Martin, Frank C. Langbein, Paul L. Rosin |
Symposium on Solid and Physical Modeling | 4 |
| 2006 | A symmetric convexity measure
Paul L. Rosin, Christine L. Mumford |
Comput. Vis. Image Underst. | 1 |
| 2006 | A model of diatom shape and texture for analysis, synthesis and identification
Yulia Hicks, David Marshall 0001, Paul L. Rosin, Ralph R. Martin, David G. Mann, S. J. M. Droop |
Mach. Vis. Appl. | 3 |
| 2006 | Selection of the optimal parameter value for the Isomap algorithm
Oksana Samko, David Marshall 0001, Paul L. Rosin |
Pattern Recognit. Lett. | 3 |
| 2006 | Training Cellular Automata for Image ProcessingabstractExperiments were carried out to investigate the possibility of training cellular automata (CA) to perform several image processing tasks. Even if only binary images are considered, the space of all possible rule sets is still very large, and so the training process is the main bottleneck of such an approach. In this paper, the sequential floating forward search method for feature selection was used to select good rule sets for a range of tasks, namely noise filtering (also applied to grayscale images using threshold decomposition), thinning, and convex hulls. Various objective functions for driving the search were considered. Several modifications to the standard CA formulation were made (the B-rule and two-cycle CAs), which were found, in some cases, to improve performance. Paul L. Rosin |
IEEE Trans. Image Process. | 1 |
| 2006 | On the Orientability of ShapesabstractThe orientation of a shape is a useful quantity, and has been shown to affect performance of object recognition in the human visual system. Shape orientation has also been used in computer vision to provide a properly oriented frame of reference, which can aid recognition. However, for certain shapes, the standard moment-based method of orientation estimation fails. We introduce as a new shape feature shape orientability, which defines the degree to which a shape has distinct (but not necessarily unique) orientation. A new method is described for measuring shape orientability, and has several desirable properties. In particular, unlike the standard moment-based measure of elongation, it is able to differentiate between the varying levels of orientability of n-fold rotationally symmetric shapes. Moreover, the new orientability measure is simple and efficient to compute (for an n-gon we describe an O(n) algorithm). Jovisa D. Zunic, Paul L. Rosin, Lazar Kopanja |
IEEE Trans. Image Process. | 2 |
| 2005 | Measuring rectilinearity
Paul L. Rosin, Jovisa D. Zunic |
Comput. Vis. Image Underst. | 1 |
| 2005 | Toward Perceptually Realistic Talking Heads: Models, Methods, and McGurkabstractMotivated by the need for an informative, unbiased, and quantitative perceptual method for the evaluation of a talking head we are developing, we propose a new test based on the “McGurk Effect.” Our approach helps to identify strengths and weaknesses in visual--speech synthesis algorithms for talking heads and facial animations, in general, and uses this insight to guide further development. We also evaluate the behavioral quality of our facial animations in comparison to real-speaker footage and demonstrate our tests by applying them to our current speech-driven facial animation system. Darren Cosker, David Marshall 0001, Paul L. Rosin, Susan Paddock, Simon K. Rushton |
ACM Trans. Appl. Percept. | 3 |
| 2004 | Proceedings of the 13th British Machine Vision Conference
Paul L. Rosin, David Marshall 0001 |
Image Vis. Comput. | 1 |
| 2004 | A New Convexity Measure for PolygonsabstractAbstract-Convexity estimators are commonly used in the analysis of shape. In this paper, we define and evaluate a new convexity measure for planar regions bounded by polygons. The new convexity measure can be understood as a "boundary-based" measure and in accordance with this it is more sensitive to measured boundary defects than the so called "area-based" convexity measures. When compared with the convexity measure defined as the ratio between the Euclidean perimeter of the convex hull of the measured shape and the Euclidean perimeter of the measured shape then the new convexity measure also shows some advantages-particularly for shapes with holes. The new convexity measure has the following desirable properties: 1) the estimated convexity is always a number from (0, 1], 2) the estimated convexity is 1 if and only if the measured shape is convex, 3) there are shapes whose estimated convexity is arbitrarily close to 0, 4) the new convexity measure is invariant under similarity transformations, and 5) there is a simple and fast procedure for computing the new convexity measure. Jovisa D. Zunic, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2004 | Measuring sigmoidality
Paul L. Rosin |
Pattern Recognit. | 1 |
| 2004 | Agent-based computer vision
Paul L. Rosin, Omer F. Rana |
Pattern Recognit. | 1 |
| 2003 | Measuring Sigmoidality
Paul L. Rosin |
CAIP | 1 |
| 2003 | Measuring shape: ellipticity, rectangularity, and triangularity
Paul L. Rosin |
Mach. Vis. Appl. | 1 |
| 2003 | Rectilinearity Measurements for PolygonsabstractThe paper introduces a shape measure intended to describe the extent to which a closed polygon is rectilinear. Other than somewhat obvious measures of rectilinearity (e.g., the sum of the differences of each corner's angle from multiples of 90/spl deg/), there has been little work in deriving a measure that is straightforward to compute, is invariant under scale, rotation, and translation, and corresponds with the intuitive notion of rectilinear shapes. There are applications in a number of different areas of computer vision and photogrammetry. Rectilinear structures often correspond to human-made objects and are therefore justified as attentional cues for further processing. For instance, in aerial image processing and reconstruction, where building footprints are often rectilinear on the local ground plane, building structures, once recognized as rectilinear, can be matched to corresponding shapes in other views for stereo reconstruction. Perceptual grouping algorithms may seek to complete shapes based on the assumption that the object in question is rectilinear. Using the proposed measure, such systems can verity this assumption. Jovisa D. Zunic, Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2003 | Assessing the behaviour of polygonal approximation algorithms
Paul L. Rosin |
Pattern Recognit. | 1 |
| 2003 | Superellipse fitting to partial data
Paul L. Rosin |
Pattern Recognit. | 2 |
| 2003 | Shape measures for image retrieval
George Gagaudakis, Paul L. Rosin |
Pattern Recognit. Lett. | 2 |
| 2003 | Comments on "ground from figure discrimination"
Paul L. Rosin |
Pattern Recognit. Lett. | 1 |
| 2003 | Evaluation of global image thresholding for change detection
Paul L. Rosin, Efstathios V. Ioannidis |
Pattern Recognit. Lett. | 1 |
| 2002 | Modelling life cycle related and individual shape variation in biological specimensabstractThe main purpose of this research is to develop methods for automatic identification of biological specimens in digital photographs and drawings held in a database. Incorporation of taxonomic drawings into a visual indexing system has not been attempted to date. Diatoms are a single cell microscopic algae that provide a particularly suitable case study. Identification of diatoms is a challenging task due to the huge number of the species, blurred boundaries between species, and life cycle related shape changes. A novel model based on principal curves representing the life cycle related shape variation of a number of diatom species has been developed. Our model is suitable for reconstruction purposes, allowing us to produce drawings of a variety of diatom shapes, thus providing a link between the photographs and drawings. We present the classification results of photographed and drawn specimens based on the model and compare our results to another recent system for diatom identification. Finally, given a diatom specimen, we are able not only to identify the species it belongs to but also to pinpoint the stage in the life cycle it represents. Yulia Hicks, David Marshall 0001, Ralph R. Martin, Paul L. Rosin, Micha Bayer, David G. Mann |
BMVC | 4 |
| 2002 | A Convexity Measurement for PolygonsabstractConvexity estimators are commonly used in the analysis of shape. In this paper we define and evaluate a new easily computable measure of convexity for polygons. Let P be an arbitrary polygon. If 7 ) (P, c) denotes the perimeter in the sense of l metrics of the polygon obtained by the rotation of P by angle c with the origin as the center of the applied rotation, and if 7)2 (R(P, c)) is the Euclidean perimeter of the minimal rectangle R(P, c) having the edges parallel to coordinate axes which includes such a rotated polygon P, then we show that C(P) defined as C(P) = min 7)2(R(P,c)) a C[0,2rr] 7)1 (P, c) can be used as an estimate for the convexity of P. Several desirable properties of C (P) are proved, as well. Jovisa D. Zunic, Paul L. Rosin |
BMVC | 2 |
| 2002 | A Rectilinearity Measurement for Polygons
Jovisa D. Zunic, Paul L. Rosin |
ECCV (2) | 2 |
| 2002 | Automatic landmarking for building biological shape modelsabstractWe present a new method for automatic landmark extraction from the contours of biological specimens. Our ultimate goal is to enable automatic identification of biological specimens in photographs and drawings held in a database. We propose to use active appearance models for visual indexing of both photographs and drawings. Automatic landmark extraction will assist us in building the models. We describe the results of using our method on drawings and photographs of examples of diatoms, and present an active shape model built using automatically extracted data. Yulia Hicks, David Marshall 0001, Ralph R. Martin, Paul L. Rosin, Micha Bayer, David G. Mann |
ICIP (2) | 4 |
| 2002 | Multimodal retinal imaging: new strategies for the detection of glaucomaabstractGlaucoma is a serious worldwide disease whose treatment can be improved by early detection. As part of a new clinical approach this paper introduces some preliminary studies in the computerised detection of the disease. In particular we consider the problems of registering 3D laser data with a digital image. We introduce a new method based on windowed mutual information and show that it performs better than the standard mutual information technique. Paul L. Rosin, David Marshall 0001, James E. Morgan |
ICIP (3) | 1 |
| 2002 | Thresholding for Change Detection
Paul L. Rosin |
Comput. Vis. Image Underst. | 1 |
| 2002 | Incorporating shape into histograms for CBIR
George Gagaudakis, Paul L. Rosin |
Pattern Recognit. | 2 |
| 2001 | Shape measures for image retrievalabstractOne of the main goals in content based image retrieval (CBIR) is to incorporate shape into the process in a reliable manner. In order to overcome the difficulties of directly obtaining shape information (in particular avoiding region segmentation) we develop several shape measures that tackle the problem in an indirect manner, requiring only a minimal amount of segmentation. A histogram-based scheme is then used, maintaining low complexity with high efficiency and robustness. The obtained results showed that the synergy of the shape measures worked out providing an improvement over the colour histogram. George Gagaudakis, Paul L. Rosin |
ICIP (2) | 2 |
| 2001 | Unimodal thresholding
Paul L. Rosin |
Pattern Recognit. | 1 |
| 2001 | Robust pixel unmixingabstractPixel unmixing is commonly performed by employing a least squared (LS) error criterion, making it sensitive to outliers. As an alternative, the least median of squares (LMedS) method is proposed. Not only is it extremely robust, but it is efficient and straightforward both to implement and use. Paul L. Rosin |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2000 | Using CBIR and Pathfinder Networks for Image Database VisualizationabstractDigital images and videos have an increasingly important role in today's telecommunication and our everyday life in modern information society. In this paper, we explore the synergy between content-based image retrieval (CBIR) techniques and Pathfinder networks. Salient image features, based on colour, texture, edge orientation, and edge distance, are extracted. The structural modelling capabilities of Pathfinder are then applied to simplify and visualise the strongest interrelationships in the image database. George Gagaudakis, Paul L. Rosin, Chaomei Chen |
ICPR | 2 |
| 2000 | Measuring Shape: Ellipticity, Rectangularity, and TriangularityabstractObject classification often operates by making decisions based on the values of several shape properties measured from the image. The paper describes and tests several algorithms for calculating ellipticity, rectangularity, and triangularity shape descriptors. Paul L. Rosin |
ICPR | 1 |
| 2000 | Content-Based Image VisualizationabstractThe proliferation of content based image retrieval techniques has highlighted the need to understand the relationship between image clustering based on low-level image features and image clustering made by human users. In conventional image retrieval systems, images are typically characterized by a range of features such as color, texture, and shape. However, little is known to what extent these low-level features can be effectively combined with information visualization techniques such that users may explore images in a digital library according to visual similarities. The authors compare and analyze a number of Pathfinder networks of images generated based on such features. Salient structures of images are visualized according to features extracted from color, texture, and shape orientation. Implications for visualizing and constructing hypermedia systems are discussed. Chaomei Chen, George Gagaudakis, Paul L. Rosin |
IV | 3 |
| 2000 | Fitting SuperellipsesabstractIn the literature, methods for fitting superellipses to data tend to be computationally expensive due to the nonlinear nature of the problem. This paper describes and tests several fitting techniques which provide different trade-offs between efficiency and accuracy. In addition, we describe various alternative error of fit measures that can be applied by most superellipse fitting methods. Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2000 | Shape partitioning by convexityabstractThe partitioning of two dimensional (2-D) shapes into subparts is an important component of shape analysis. The paper defines a formulation of convexity as a criterion of good part decomposition. Its appropriateness is validated by applying it to some simple shapes as well as against showing its close correspondence with Hoffman and Singh's (1997) part saliency factors. Paul L. Rosin |
IEEE Trans. Syst. Man Cybern. Part A | 1 |
| 1999 | Incorporating Shape into Histograms for CBIRabstractThis paper describes an indexing system for use in Content Based Image Retrieval. The standard colour histogram approach is simple, efficient, and robust. However, it does not include shape information, which leads to problems (e.g., many-to-many mappings). To remedy this, we use additional features in an attempt to incorporate shape and textual information to the index key. Our experiments showed that the combination of colour, texture, distance and orientation histograms gave approximately 10% improvement of recall over the standard colour histogram. George Gagaudakis, Paul L. Rosin |
BMVC | 2 |
| 1999 | Shape Partitioning by ConvexityabstractThe partitioning of 2D shapes into subparts is an important component of shape analysis. This paper defines a formulation of convexityasa criterion of good part decomposition. It's appropriateness is validated by applying it to some simple shapes as well as against showing its close correspondence with Hoffman and Singh's part saliency factors. Paul L. Rosin |
BMVC | 1 |
| 1999 | A survey and comparison of traditional piecewise circular approximations to the ellipse
Paul L. Rosin |
Comput. Aided Geom. Des. | 1 |
| 1999 | Further Five-Point Fit Ellipse Fitting
Paul L. Rosin |
Graph. Model. Image Process. | 1 |
| 1999 | Measuring Corner Properties
Paul L. Rosin |
Comput. Vis. Image Underst. | 1 |
| 1999 | Measuring rectangularity
Paul L. Rosin |
Mach. Vis. Appl. | 1 |
| 1999 | Robust pose estimationabstractStandard least-squares (LS) methods for pose estimation of objects are sensitive to outliers which can occur due to mismatches. Even a single mismatch can severely distort the estimated pose. This paper describes a least-median of squares (LMedS) approach to estimating pose using point matches. It is both robust (resistant to up to 50% outliers) and efficient (linear in the number of points). The basic algorithm is then extended to improve performance in the presence of two types of noise: 1) type I which perturbs all data values by small amounts (e.g., Gaussian) and 2) type II which can corrupt a few data values by large amounts. Paul L. Rosin |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 1998 | Assessing the Behaviour of Polygonal Approximation AlgorithmsabstractIn a recent paper we described a method for assessing the accuracy of polygonal approximation algorithms [16]. Here we develop several measures to assess the stability of such approximation algorithms under variations in their scale parameters. A monotonicity index is introduced that can be applied to analyse the change in the approximation error or the number of line segments against increasing scale. A consistency index quantifies the variation in results produced at the same scale by an algorithm (but with different input parameter values). Finally, the previously developed accuracy figure of merit is calculated and averaged over 21 test curves for different parameter values to obtain more reliable scores. 1 Paul L. Rosin |
BMVC | 1 |
| 1998 | Thresholding for Change Detection
Paul L. Rosin |
ICCV | 1 |
| 1998 | Ellipse Fitting Using Orthogonal Hyperbolae and Stirling's Oval
Paul L. Rosin |
Graph. Model. Image Process. | 1 |
| 1998 | The effects of data filtering on neural network learning
Paul L. Rosin, Freddy Fierens |
Neurocomputing | 1 |
| 1998 | Refining Region EstimatesabstractA method for improving the segmentation of images is presented. It involves taking an initial segmentation provided by some other means, and modifying the region boundaries depending on the estimated region models until an equilibrium is reached. The advantages of this technique are: (1) no parameters are required, (2) it is invariant under constant scalings of the image intensities, and (3) it is relatively insensitive to the position and topology of the initial segmentation. Examples are given of its application to single and multi-scale intensity images, textured images, range images and multi-band satellite images. Paul L. Rosin |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1998 | Guest editorial
Dimitris C. Dracopoulos, Paul L. Rosin |
Neural Comput. Appl. | 2 |
| 1998 | Determining local natural scales of curves
Paul L. Rosin |
Pattern Recognit. Lett. | 1 |
| 1997 | Measuring Corner Properties
Paul L. Rosin |
BMVC | 1 |
| 1997 | Thresholding for Change Detection
Paul L. Rosin |
BMVC | 1 |
| 1997 | Further Five Point Fit Ellipse Fitting
Paul L. Rosin |
BMVC | 1 |
| 1997 | Edges: saliency measures and automatic thresholding
Paul L. Rosin |
Mach. Vis. Appl. | 1 |
| 1997 | Techniques for Assessing Polygonal Approximations of CurvesabstractGiven the enormous number of available methods for finding polygonal approximations to curves techniques are required to assess different algorithms. Some of the standard approaches are shown to be unsuitable if the approximations contain varying numbers of lines. Instead, we suggest assessing an algorithm's results relative to an optimal polygon, and describe a measure which combines the relative fidelity and efficiency of a curve segmentation. We use this measure to compare the application of 23 algorithms to a curve first used by Teh and Chin (1989); their integral square errors (ISEs) are assessed relative to the optimal ISE. In addition, using an example of pose estimation, it is shown how goal-directed evaluation can be used to select an appropriate assessment criterion. Paul L. Rosin |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1997 | Response to KanataniabstractAbstract — We discuss the advantages and disadvantages of two approaches to model selection: the information theo-retic method suggested by Kanatani [3] and others, and our heuristic sequential selection method [9]. Paul L. Rosin, Geoff A. W. West |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1996 | Analysing Error of Fit Functions for EllipsesabstractWe describe several established error of fit (EOF) functions for use in the least square fitting of ellipses, and introduce a further four new EOFs.Four measures are given for assessing the suitability of such EOFs, quantifying their linearity, curvature bias, asymmetry, and overall goodness.These measures enable a better understanding to be gained of the individual merits of the EOF functions. Paul L. Rosin |
BMVC | 1 |
| 1996 | Techniques for Assessing Polygonal Approximations of CurvesabstractGiven the enormous number of available methods for finding polygonal approximations to curves techniques are required to assess different algorithms. Some of the standard approaches are shown to be unsuitable if the approximations contain varying numbers of lines. Instead, we suggest assessing an algorithm's results relative to an optimal polygon, and describe a measure which combines the relative fidelity and efficiency of a curve segmentation. We use this measure to compare the application of fifteen algorithms to a curve first used by Teh and Chin [31]; their ISEs are assessed relative to the optimal ISE. In addition, using an example of pose estimation, it is shown how goal-directed evaluation can be used to select an appropriate assessment criterion. Paul L. Rosin |
BMVC | 1 |
| 1996 | Augmenting Corner Descriptors
Paul L. Rosin |
CVGIP Graph. Model. Image Process. | 1 |
| 1996 | Assessing Error of Fit Functions for Ellipses
Paul L. Rosin |
CVGIP Graph. Model. Image Process. | 1 |
| 1996 | Analysing error of fit functions for ellipses
Paul L. Rosin |
Pattern Recognit. Lett. | 1 |
| 1995 | Image Difference Threshold Strategies and Shadow DetectionabstractThe paper considers two problems associated with the detection and classification of motion in image sequences obtained from a static camera. Motion is detected by differencing a reference and the ``current'' image frame, and therefore requires a suitable reference image and the selection of an appropriate detection threshold. Several threshold selection methods are investigated, and an algorithm based on hysteresis thresholding is shown to give acceptably good results over a number of test image sets. The second part of the paper examines the problem of detecting shadow regions within the image which are associated with the object motion. This is based on the notion of a shadow as a semi-transparent region in the image which retains a (reduced contrast) representation of the underlying surface pattern, texture or grey value. The method uses a region growing algorithm which uses a growing criterion based on a fixed attenuation of the photometric gain over the shadow region, in comparison to the reference image. Paul L. Rosin, Tim J. Ellis |
BMVC | 1 |
| 1995 | Salience Distance Transforms
Paul L. Rosin, Geoff A. W. West |
CVGIP Graph. Model. Image Process. | 1 |
| 1995 | Dynamic Threshold Determination by Local and Global Edge Evaluation
Svetha Venkatesh, Paul L. Rosin |
CVGIP Graph. Model. Image Process. | 2 |
| 1995 | Early Image Representation by Slope Districts
Paul L. Rosin |
J. Vis. Commun. Image Represent. | 1 |
| 1995 | Nonparametric Segmentation of Curves into Various RepresentationsabstractThis paper describes and demonstrates the operation and performance of an algorithm for segmenting connected points into a combination of representations such as lines, circular, elliptical and superelliptical arcs, and polynomials. The algorithm has a number of interesting properties including being scale invariant, nonparametric, general purpose, and efficient. Paul L. Rosin, Geoff A. W. West |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1994 | Grouping Curved LinesabstractThis paper examines the problem of automatically grouping image curves. In contrast, most previous work has been restricted to points and straight lines. Some of the computational aspects of the groupings of continuation, parallelism, and proximity are analysed, and the issues of neighbourhoods, combinatorix, and multiple scales are discussed. Paul L. Rosin |
BMVC | 1 |
| 1994 | Determining local natural scales of curvesabstractAn alternative to representing curves at a single scale or a fixed number of multiple scales is to represent them only at their natural (i.e. most significant) scales. This allows all the important information concerning the different sized structures contained in the curve to be explicitly represented without the overhead of redundant representations of the curve. This paper describes several approaches to determining the local natural scales of curves. That is, various possibly overlapping sections of the curve should be represented at certain scales depending on their shape. The merits and drawbacks of the techniques are described, and the results of implementing one of them are shown. Paul L. Rosin |
ICPR (1) | 1 |
| 1994 | Non-Parametric Multiscale Curve SmoothingabstractLowe8 demonstrated a method for automatically segmenting and smoothing image curves by varying degrees. It was intended to remove noise and unnecessary fine detail, aiding subsequent processing such as grouping and matching. An alternative technique is described in this paper that is based on recursively subdividing the curve into alternative sets of sections. Rather than use thresholds on the values of curvature and its derivatives to determine the segmentation and degree of smoothing our technique is driven by three qualitative measures: (1) a criterion for selecting potential breakpoints, (2) a criterion for determining the amount of smoothing for curve sections, (3) a significance measure that determines which sections form the best selection. The advantages of the technique are robustness, scale invariance, and the absence of parameters. Paul L. Rosin |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 1993 | Multi-Scale Salience Distance TransformsabstractThe distance transform has been proposed for use in computer vision for a number of applications such as matching and skeletonisation. This paper proposes two things: (1) a multi-scale distance transform to overcome the need to choose edge thresholds and scale and (2) the addition of various saliency factors such as edge strength, length and curvature to the basic distance transform to improve its effectiveness. Results are presented for applications of matching and snake fitting. 1 Paul L. Rosin, Geoff A. W. West |
BMVC | 1 |
| 1993 | Multiscale Representation and Matching of Curves Using Codons
Paul L. Rosin |
CVGIP Graph. Model. Image Process. | 1 |
| 1993 | Extracting natural scales using fourier descriptors
Paul L. Rosin, Svetha Venkatesh |
Pattern Recognit. | 1 |
| 1993 | Acquiring information from cues
Paul L. Rosin |
Pattern Recognit. Lett. | 1 |
| 1993 | Ellipse fitting by accumulating five-point fits
Paul L. Rosin |
Pattern Recognit. Lett. | 1 |
| 1993 | A note on the least squares fitting of ellipses
Paul L. Rosin |
Pattern Recognit. Lett. | 1 |
| 1992 | Multistage Combined Ellipse and Line Detection
Geoff A. W. West, Paul L. Rosin |
BMVC | 2 |
| 1992 | Representing curves at their natural scales
Paul L. Rosin |
Pattern Recognit. | 1 |
| 1992 | Early image representation using regions defined by maximum gradient paths between singular points
Paul L. Rosin, Alan C. F. Colchester, David J. Hawkes |
Pattern Recognit. | 1 |
| 1992 | Detection and verification of surfaces of revolution by perceptual grouping
Paul L. Rosin, Geoff A. W. West |
Pattern Recognit. Lett. | 1 |
| 1991 | Detecting and Classifying Intruders in Image Sequences
Paul L. Rosin |
BMVC | 1 |
| 1991 | Extracting surfaces of revolution by perceptual grouping of ellipsesabstractEllipses seen in an image may be the 2-D projection of 3-D circles from the scene. Given this assumption, ellipses can be grouped into perceptual groups from which inferences about the 3-D structure of objects can be made. Methods are proposed for extracting groupings corresponding to surfaces of revolution. A Hough transform approach is used for grouping, after which the confidence in the plausibility of the perceptual group is improved by detecting symmetry groupings.> Paul L. Rosin, Geoff A. W. West |
CVPR | 1 |
| 1991 | Frame-based system for image interpretation
Paul L. Rosin, Tim J. Ellis |
Image Vis. Comput. | 1 |
| 1991 | Techniques for segmenting image curves into meaningful descriptions
Geoff A. W. West, Paul L. Rosin |
Pattern Recognit. | 2 |
| 1990 | Perceptual grouping of circular arcs under projectionabstractTheories and techniques of shape constancy and shape from contour often assert that ellipses seen in an image are the 2D projection of 3D circles from the scene. Given this assumption, several methods for grouping ellipses are described. From these perceptual groups inferences about the 3D structure of their generating circles are made. This will be useful for model invokation. The world (both natural and man-made) is constrained by physical laws to a finite number of basic patterns. Certain three dimensional relationships between features in a scene give rise to two dimensional relationships that are relatively viewpoint invariant and can be distinguished as non-accidental. Thus, those regularities observable in the image arise not by accident, but are projections of real regularities in the scene. The Gestalt movement proposed Paul L. Rosin, Geoff A. W. West |
BMVC | 1 |
| 1990 | Segmenting curves into elliptic arcs and straight linesabstractA method is described for segmenting edge data into a combination of straight lines and elliptic arcs. The two-stage process first segments the data into straight line segments. Ellipses are then fitted to the line data. This is much faster than curve fitting directly to pixel data since the lines provide a great reduction in data. Segmentation is performed in the paradigm suggested by D.G. Lowe (1987). A measure of significance is defined that produces a scale-invariant description and allows the replacement of sequences of line segments by ellipses without requiring any thresholds. A method for fitting ellipses to arbitrary curves, essential for this algorithm, has been developed, based on an iterative Kalman filter. This is guaranteed to produce an elliptical fit even though the best conic fit may be a hyperbola or parabola.> Paul L. Rosin, Geoff A. W. West |
ICCV | 1 |
| 1989 | Segmentation of edges into lines and arcs
Paul L. Rosin, Geoff A. W. West |
Image Vis. Comput. | 1 |