VLDB 2026 Research / reviewers in the wild / expert
John P. Collomosse
dblp:72/4504 · also John Philip Collomosse
· DBLP profile ↗
99ranked-venue papers
15as first author
28since 2021 · last 2026
0000-0003-3580-4685ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 79 · 13 first-author · 23 since 2021Artificial intelligence and machine learning · 52 · 6 first-author · 19 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 5 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 since 2021Security and privacy · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | NRGMark: Localized Watermarking for Energy Transparency in ImagesabstractWe present NRGMark, a region-based image watermarking framework to embed provenance metadata into composite graphic designs such as posters. NRGMark enables imperceptible watermarking of distinct visual elements each carrying independent metadata on aspects like environmental impact, such as the energy consumption associated with generative AI (GenAI) use. NRGMark extends image watermark encoder-decoder models by incorporating an object localization network to detect and decode multiple watermarked regions within a document, even under image transformations and physical print–scan degradation. NRGMark interoperates with several watermarking techniques and the emerging C2PA open standard for media provenance to encode environmental impact metadata. We demonstrate NRGMark on both synthetic and real-world design layouts, illustrating its potential to support energy transparency in the age of GenAI. Shruti Agarwal, Élie Michel, Vishal Asnani, Tania Mathern, John P. Collomosse |
WACV | 5 |
| 2026 | Fact-Checking With Contextual Narratives: Leveraging Retrieval-Augmented LLMs for Social Media AnalysisabstractFact-checking systems have gained traction as scalable solutions, yet they often face challenges such as handling diverse evidence sources, integrating multimodal data, and presenting comprehensive narratives. In this work, we propose cluster-based retrieval augmented verification with explanation (CRAVE), a novel framework that integrates retrieval-augmented large language models (LLMs) with clustering techniques to address multimodal misinformation on social media. The framework is designed to process multimodal inputs (text and images) and iteratively refine evidence through agent-based mechanisms. We validated the framework on multiple real-world and synthetic datasets, showing that breaking up evidence into narrative clusters improves both retrieval precision, clustering quality, and judgment accuracy, showcasing its potential as a robust decision-support tool for fact-checkers. Arka Ujjal Dey, Muhammad Junaid Awan, Georgia Channing, Christian Schröder de Witt, John P. Collomosse |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2025 | MAGNET: Augmenting Generative Decoders with Representation Learning and Infilling CapabilitiesabstractSavya Khosla, Aditi Tiwari, Kushal Kafle, Simon Jenni, Handong Zhao, John Collomosse, Jing Shi. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Savya Khosla, Aditi Tiwari, Kushal Kafle, Simon Jenni, Handong Zhao, John P. Collomosse, Jing Shi 0005 |
ACL (1) | 6 |
| 2025 | Content Authenticities: A Discussion on the Values of Provenance Data for Creatives and Their AudiencesabstractThe proliferation of AI-generated digital content has intensified the user demand for accurate provenance information to ensure content authenticity.Technical advancements now provide tools to make the digital media content supply chain more transparent through the use of provenance data.This paper foregrounds the importance of understanding how the situated nature of user-content engagement influences perceptions and uses of this data.Insights from a workshop with experts in the creative media sector suggest that, as the adoption of provenance data becomes more common, users need richer and more nuanced information.We suggest that analyzing the increasing demand for content authenticity through the lens of multiple "authenticities", each reflecting different user needs and contexts, can help identify and address the needs for, and uses of, provenance data by creators and audiences alike. Caterina Moruzzi, Ella Tallyn, Frances Liddell, Billy Dixon, John P. Collomosse, Chris Elsden |
Creativity & Cognition | 5 |
| 2025 | Multitwine: Multi-Object Compositing with Text and Layout ControlabstractWe introduce the first generative model capable of simultaneous multi-object compositing, guided by both text and layout. Our model allows for the addition of multiple objects within a scene, capturing a range of interactions from simple positional relations (e.g., next to, in front of) to complex actions requiring reposing (e.g., hugging, playing guitar). When an interaction implies additional props, like ‘taking a selfie’, our model autonomously generates these supporting objects. By jointly training for compositing and subject-driven generation, also known as customization, we achieve a more balanced integration of textual and visual inputs for text-driven object compositing. As a result, we obtain a versatile model with state-of-the-art performance in both tasks. We further present a data generation pipeline leveraging visual and language models to effortlessly synthesize multimodal, aligned training data. Gemma Canet Tarrés, Zhe Lin 0001, He Zhang 0004, Andrew Gilbert, John P. Collomosse, Soo Ye Kim |
CVPR | 6 |
| 2025 | TrustMark: Robust Watermarking and Watermark Removal for Arbitrary Resolution Images
Tu Bui, Shruti Agarwal, John P. Collomosse |
ICCV | 3 |
| 2025 | DiffTell: A High-Quality Dataset for Describing Image Manipulation Changes
Zonglin Di, Jing Shi 0005, Hao Tan 0002, Alexander Black 0001, John P. Collomosse, Yang Liu 0018 |
ICCV | 6 |
| 2025 | On the Coexistence and Ensembling of WatermarksabstractWatermarking, the practice of embedding imperceptible information into media such as images, videos, audio, and text, is essential for intellectual property protection, content provenance and attribution. The growing complexity of digital ecosystems necessitates watermarks for different uses to be embedded in the same media. However, to detect and decode all watermarks, they need to coexist well with one another. We perform the first study of coexistence of deep image watermarking methods and, contrary to intuition, we find that various open-source watermarks can coexist with only minor impacts on image quality and decoding robustness. The coexistence of watermarks also opens the avenue for ensembling watermarking methods. We show how ensembling can increase the overall message capacity and enable new trade-offs between capacity, accuracy, robustness and image quality, without needing to retrain the base models. Aleksandar Petrov, Shruti Agarwal, Philip Torr 0001, Adel Bibi, John P. Collomosse |
NeurIPS | 5 |
| 2025 | CLASS: Conditional Latent Architecture for Search and Synthesis of Design LayoutsabstractWe propose CLASS; a novel unified model for the syn-thesis and search for design layouts, two tasks that are often handled separately by prior works. We propose to learn a compact and coherent latent feature of a layout supporting joint search and synthesis. This allows vari-ous operations such style-conditioned layout generation, la-tent space manipulation and provides seamless integration of search and synthesis for an effective design workflow. We train CLASS with a dual decoder: a new transformer-based layout-conditioned decoder and a CNN-based raster decoder. The latent-conditioned decoder explicitly conditions upon a latent vector while generating a layout in an auto-regressive fashion. We train CLASS under variational framework which in conjunction with a raster-decoder en-hances the latent representation improving both generation and retrieval performances. We show the effectiveness of CLASS on the RICO and PubLayNet benchmarks, and demonstrate that CLASS is capable of high-quality synthe-sis from scratch, as well as performing self-completion, in-terpolation, project between design layouts, whilst achieving close to or better than state-of-the-art search performance. Dipu Manandhar, Paul Guerrero 0001, John P. Collomosse |
WACV | 4 |
| 2024 | VIXEN: Visual Text Comparison Network for Image Difference CaptioningabstractWe present VIXEN - a technique that succinctly summarizes in text the visual differences between a pair of images in order to highlight any content manipulation present. Our proposed network linearly maps image features in a pairwise manner, constructing a soft prompt for a pretrained large language model. We address the challenge of low volume of training data and lack of manipulation variety in existing image difference captioning (IDC) datasets by training on synthetically manipulated images from the recent InstructPix2Pix dataset generated via prompt-to-prompt editing framework. We augment this dataset with change summaries produced via GPT-3. We show that VIXEN produces state-of-the-art, comprehensible difference captions for diverse image contents and edit types, offering a potential mitigation against misinformation disseminated via manipulated image content. Code and data are available at http://github.com/alexblck/vixen Alexander Black 0001, Jing Shi 0005, Tu Bui, John P. Collomosse |
AAAI | 5 |
| 2024 | ProMark: Proactive Diffusion Watermarking for Causal AttributionabstractGenerative AI (GenAI) is transforming creative work-flows through the capability to synthesize and manipulate images via high-level prompts. Yet creatives are not well supported to receive recognition or reward for the use of their content in GenAI training. To this end, we propose ProMark, a causal attribution technique to attribute a synthetically generated image to its training data concepts like objects, motifs, templates, artists, or styles. The concept information is proactively embedded into the input training images using imperceptible watermarks, and the diffusion models (unconditional or conditional) are trained to retain the corresponding watermarks in generated images. We show that we can embed as many as 216unique water-marks into the training data, and each training image can contain more than one watermark. ProMark can maintain image quality whilst outperforming correlation-based attribution. Finally, several qualitative examples are presented, providing the confidence that the presence of the watermark conveys a causative relationship between training data and synthetic images. Vishal Asnani, John P. Collomosse, Tu Bui, Xiaoming Liu 0002, Shruti Agarwal |
CVPR | 2 |
| 2024 | FineMatch: Aspect-Based Fine-Grained Image and Text Mismatch Detection and Correction
Hang Hua, Jing Shi 0005, Kushal Kafle, Simon Jenni, Daoan Zhang, John P. Collomosse, Scott Cohen, Jiebo Luo 0001 |
ECCV (9) | 6 |
| 2024 | Thinking Outside the BBox: Unconstrained Generative Object Compositing
Gemma Canet Tarrés, Zhe Lin 0001, Jianming Zhang 0001, Dan Ruta, Andrew Gilbert, John P. Collomosse, Soo Ye Kim |
ECCV (62) | 8 |
| 2024 | SegGuard: Defending Scene Segmentation Against Adversarial Patch AttackabstractAdversarial Patch Attacks (APAs) induce prediction errors by inserting carefully crafted regions into images. This paper presents the first defence against APAs for deep networks that perform semantic segmentation of scenes. We show that a conditional generator can be trained to produce patches on demand targeting specific classes and achieving superior performance versus conventional pixel-optimised patch attacks. We then leverage this generator along with the segmentation network as part of a generative adversarial network, which trains the model to ignore the adversarial patches produced by the generator, while simultaneously training the generator to produce updated patches to attack the fine-tuned network. We show that our process confers strong protection against adversarial patches, and that this protection generalises to traditional pixel-optimised adversarial patches. Thomas Gittings, Steve A. Schneider, John P. Collomosse |
ICIP | 3 |
| 2023 | Layout Representation Learning with Spatial and Structural HierarchiesabstractWe present a novel hierarchical modeling method for layout representation learning, the core of design documents (e.g., user interface, poster, template). Existing works on layout representation often ignore element hierarchies, which is an important facet of layouts, and mainly rely on the spatial bounding boxes for feature extraction. This paper proposes a Spatial-Structural Hierarchical Auto-Encoder (SSH-AE) that learns hierarchical representation by treating a hierarchically annotated layout as a tree format. On the one side, we model SSH-AE from both spatial (semantic views) and structural (organization and relationships) perspectives, which are two complementary aspects to represent a layout. On the other side, the semantic/geometric properties are associated at multiple resolutions/granularities, naturally handling complex layouts. Our learned representations are used for effective layout search from both spatial and structural similarity perspectives. We also newly involve the tree-edit distance (TED) as an evaluation metric to construct a comprehensive evaluation protocol for layout similarity assessment, which benefits a systematic and customized layout search. We further present a new dataset of POSTER layouts which we believe will be useful for future layout research. We show that our proposed SSH-AE outperforms the existing methods achieving state-of-the-art performance on two benchmark datasets. Code is available at github.com/yueb17/SSH-AE. Dipu Manandhar, John P. Collomosse, Yun Fu 0001 |
AAAI | 4 |
| 2023 | Audio-Visual Contrastive Learning with Temporal Self-SupervisionabstractWe propose a self-supervised learning approach for videos that learns representations of both the RGB frames and the accompanying audio without human supervision. In contrast to images that capture the static scene appearance, videos also contain sound and temporal scene dynamics. To leverage the temporal and aural dimension inherent to videos, our method extends temporal self-supervision to the audio-visual setting and integrates it with multi-modal contrastive objectives. As temporal self-supervision, we pose playback speed and direction recognition in both modalities and propose intra- and inter-modal temporal ordering tasks. Furthermore, we design a novel contrastive objective in which the usual pairs are supplemented with additional sample-dependent positives and negatives sampled from the evolving feature space. In our model, we apply such losses among video clips and between videos and their temporally corresponding audio clips. We verify our model design in extensive ablation experiments and evaluate the video and audio representations in transfer experiments to action recognition and retrieval on UCF101 and HMBD51, audio classification on ESC50, and robust video fingerprinting on VGG-Sound, with state-of-the-art results. Simon Jenni, Alexander Black 0001, John P. Collomosse |
AAAI | 3 |
| 2023 | SceneComposer: Any-Level Semantic Image SynthesisabstractWe propose a new framework for conditional image synthesis from semantic layouts of any precision levels, ranging from pure text to a 2D semantic canvas with precise shapes. More specifically, the input layout consists of one or more semantic regions with free-form text descriptions and adjustable precision levels, which can be set based on the desired controllability. The framework naturally reduces to text-to-image (T2I) at the lowest level with no shape information, and it becomes segmentation-to-image (S2I) at the highest level. By supporting the levels in-between, our framework is flexible in assisting users of different drawing expertise and at different stages of their creative workflow. We introduce several novel techniques to address the challenges coming with this new setup, including a pipeline for collecting training data; a precision-encoded mask pyramid and a text feature map representation to jointly encode precision level, semantics, and composition information; and a multi-scale guided diffusion model to synthesize images. To evaluate the proposed method, we collect a test dataset containing user-drawn layouts with diverse scenes and styles. Experimental results show that the proposed method can generate high-quality images following the layout at given precision, and compares favorably against existing methods. Project page https://zengxianyu.github.io/scenec/ Yu Zeng 0001, Zhe Lin 0001, Jianming Zhang 0001, Qing Liu 0017, John P. Collomosse, Jason Kuen, Vishal M. Patel |
CVPR | 5 |
| 2023 | VADER: Video Alignment Differencing and RetrievalabstractWe propose VADER, a spatio- temporal matching, alignment, and change summarization method to help fight misinformation spread via manipulated videos. VADER matches and coarsely aligns partial video fragments to candidate videos using a robust visual descriptor and scalable search over adaptively chunked video content. A transformer- based alignment module then refines the temporal localization of the query fragment within the matched video. A space- time comparator module identifies regions of manipulation between aligned content, invariant to any changes due to any residual temporal misalignments or artifacts arising from non- editorial changes of the content. Robustly matching video to a trusted source enables conclusions to be drawn on video provenance, enabling informed trust decisions on content encountered. Code and data are available at https://github.com/AlexBlck/vader Alexander Black 0001, Simon Jenni, Tu Bui, Md. Mehrab Tanjim, Stefano Petrangeli, Ritwik Sinha, Viswanathan (Vishy) Swaminathan, John P. Collomosse |
ICCV | 8 |
| 2023 | Scene designer: compositional sketch-based image retrieval with contrastive learning and an auxiliary synthesis task
Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir Ponti |
Multim. Tools Appl. | 3 |
| 2022 | RepMix: Representation Mixing for Robust Attribution of Synthesized Images
Tu Bui, Ning Yu 0006, John P. Collomosse |
ECCV (14) | 3 |
| 2022 | CoGS: Controllable Generation and Search from Sketch and Style
Cusuh Ham, Gemma Canet Tarrés, Tu Bui, James Hays, Zhe Lin 0001, John P. Collomosse |
ECCV (16) | 6 |
| 2022 | StyleBabel: Artistic Style Tagging and CaptioningabstractWe present StyleBabel, a unique open access dataset of natural language captions and free-form tags describing the artistic style of over 135K digital artworks, collected via a novel participatory method from experts studying at specialist art and design schools. StyleBabel was collected via an iterative method, inspired by ‘Grounded Theory’: a qualitative approach that enables annotation while co-evolving a shared language for fine-grained artistic style attribute description. We demonstrate several downstream tasks for StyleBabel, adapting the recent ALADIN architecture for fine-grained style similarity, to train cross-modal embeddings for: 1) free-form tag generation; 2) natural language description of artistic style; 3) fine-grained text search of style. To do so, we extend ALADIN with recent advances in Visual Transformer (ViT) and cross-modal representation learning, achieving a state of the art accuracy in fine-grained style retrieval. Dan Ruta, Andrew Gilbert, Pranav Aggarwal, Naveen Marri, Ajinkya Kale, Jo Briggs, Chris Speed, Hailin Jin, Baldo Faieta, Alex Filipkowski, Zhe Lin 0001, John P. Collomosse |
ECCV (8) | 12 |
| 2022 | TAPESTRY: A De-Centralized Service for Trusted Interaction OnlineabstractWe present a novel de-centralised service for proving the provenance of online digital identity, exposed as an assistive tool to help non-expert users make better decisions about whom to trust online. Our service harnesses the digital personhood (DP); the longitudinal and multi-modal signals created through users’ lifelong digital interactions, as a basis for evidencing the provenance of identity. We describe how users may exchange trust evidence derived from their DP, in a granular and privacy-preserving manner, with other users in order to demonstrate coherence and longevity in their behaviour online. This is enabled through a novel secure infrastructure combining hybrid on- and off-chain storage combined with deep learning for DP analytics and visualization. We show how our tools enable users to make more effective decisions on whether to trust unknown third parties online, and also to spot behavioural deviations in their own social media footprints indicative of account hijacking. Daniel Cooper, John P. Collomosse, Constantin Catalin Dragan, Mark Manulis, Jamie Steane, Arthi Kanchana Manohar, Jo Briggs, Helen S. Jones, Wendy Moncur |
IEEE Trans. Serv. Comput. | 3 |
| 2021 | Magic Layouts: Structural Prior for Component Detection in User Interface Designs
Dipu Manandhar, Hailin Jin, John P. Collomosse |
CVPR | 3 |
| 2021 | OSCAR-Net: Object-centric Scene Graph Attention for Image AttributionabstractImages tell powerful stories but cannot always be trusted. Matching images back to trusted sources (attribution) enables users to make a more informed judgment of the images they encounter online. We propose a robust image hashing algorithm to perform such matching. Our hash is sensitive to manipulation of subtle, salient visual details that can substantially change the story told by an image. Yet the hash is invariant to benign transformations (changes in quality, codecs, sizes, shapes, etc.) experienced by images during online redistribution. Our key contribution is OSCAR-Net1(Object-centric Scene Graph Attention for Image Attribution Network); a robust image hashing model inspired by recent successes of Transformers in the visual domain. OSCAR-Net constructs a scene graph representation that attends to fine-grained changes of every object’s visual appearance and their spatial relationships. The network is trained via contrastive learning on a dataset of original and manipulated images yielding a state of the art image hash for content fingerprinting that scales to millions of images. Eric Nguyen, Tu Bui, Viswanathan (Vishy) Swaminathan, John P. Collomosse |
ICCV | 4 |
| 2021 | ALADIN: All Layer Adaptive Instance Normalization for Fine-grained Style SimilarityabstractWe present ALADIN (All Layer AdaIN); a novel architecture for searching images based on the similarity of their artistic style. Representation learning is critical to visual search, where distance in the learned search embedding reflects image similarity. Learning an embedding that discriminates fine-grained variations in style is hard, due to the difficulty of defining and labelling style. ALADIN takes a weakly supervised approach to learning a representation for fine-grained style similarity of digital artworks, leveraging BAM-FG, a novel large-scale dataset of user generated content groupings gathered from the web. ALADIN sets a new state of the art accuracy for style-based visual search over both coarse labelled style data (BAM) and BAM-FG; a new 2.62 million image dataset of 310,000 fine-grained style groupings also contributed by this work. Dan Ruta, Saeid Motiian, Baldo Faieta, Zhe Lin 0001, Hailin Jin, Alex Filipkowski, Andrew Gilbert, John P. Collomosse |
ICCV | 8 |
| 2021 | Compositional Sketch SearchabstractWe present an algorithm for searching image collections using free-hand sketches that describe the appearance and relative positions of multiple objects1Sketch based image retrieval (SBIR) methods predominantly match queries containing a single, dominant object invariant to its position within an image. Our work exploits drawings as a concise and intuitive representation for specifying entire scene compositions. We train a convolutional neural network (CNN) to encode masked visual features from sketched objects, pooling these into a spatial descriptor encoding the spatial relationships and appearances of objects in the composition. Training the CNN backbone as a Siamese network under triplet loss yields a metric search embedding for measuring compositional similarity which may be efficiently leveraged for visual search by applying product quantization. Alexander Black 0001, Tu Bui, Long Mai, Hailin Jin, John P. Collomosse |
ICIP | 5 |
| 2021 | Neural architecture search for deep image prior
Kary Ho, Andrew Gilbert, Hailin Jin, John P. Collomosse |
Comput. Graph. | 4 |
| 2020 | Vax-a-Net: Training-Time Defence Against Adversarial Patch Attacks
Thomas Gittings, Steve A. Schneider, John P. Collomosse |
ACCV (4) | 3 |
| 2020 | DeepVoxels++: Enhancing the Fidelity of Novel View Synthesis from 3D Voxel Embeddings
Tong He 0002, John P. Collomosse, Hailin Jin, Stefano Soatto |
ACCV (1) | 2 |
| 2020 | Semantic Estimation of 3D Body Shape and Pose using Minimal Cameras
Andrew Gilbert, Matthew Trumble, Adrian Hilton 0001, John P. Collomosse |
BMVC | 4 |
| 2020 | Sketchformer: Transformer-Based Representation for Sketched StructureabstractSketchformer is a novel transformer-based representation for encoding free-hand sketches input in a vector form, i.e. as a sequence of strokes. Sketchformer effectively addresses multiple tasks: sketch classification, sketch based image retrieval (SBIR), and the reconstruction and interpolation of sketches. We report several variants exploring continuous and tokenized input representations, and contrast their performance. Our learned embedding, driven by a dictionary learning tokenization scheme, yields state of the art performance in classification and image retrieval tasks, when compared against baseline representations driven by LSTM sequence to sequence architectures: SketchRNN and derivatives. We show that sketch reconstruction and interpolation are improved significantly by the Sketchformer embedding for complex sketches with longer stroke sequences. Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir Ponti |
CVPR | 3 |
| 2020 | Learning Structural Similarity of User Interface Layouts Using Graph Networks
Dipu Manandhar, Dan Ruta, John P. Collomosse |
ECCV (22) | 3 |
| 2020 | Geo-PIFu: Geometry and Pixel Aligned Implicit Functions for Single-view Human ReconstructionabstractWe propose Geo-PIFu, a method to recover a 3D mesh from a monocular color image of a clothed person. Our method is based on a deep implicit function-based representation to learn latent voxel features using a structure-aware 3D U-Net, to constrain the model in two ways: first, to resolve feature ambiguities in query point encoding, second, to serve as a coarse human shape proxy to regularize the high-resolution mesh and encourage global shape regularity. We show that, by both encoding query points and constraining global shape using latent voxel features, the reconstruction we obtain for clothed human meshes exhibits less shape distortion and improved surface details compared to competing methods. We evaluate Geo-PIFu on a recent human mesh public dataset that is 10x larger than the private commercial dataset used in PIFu and previous derivative work. On average, we exceed the state of the art by 42.7% reduction in Chamfer and Point-to-Surface Distances, and 19.4% reduction in normal estimation errors. Tong He 0002, John P. Collomosse, Hailin Jin, Stefano Soatto |
NeurIPS | 2 |
| 2020 | Real-Time Multi-person Motion Capture from Multi-view Video and IMUsabstractAbstract A real-time motion capture system is presented which uses input from multiple standard video cameras and inertial measurement units (IMUs). The system is able to track multiple people simultaneously and requires no optical markers, specialized infra-red cameras or foreground/background segmentation, making it applicable to general indoor and outdoor scenarios with dynamic backgrounds and lighting. To overcome limitations of prior video or IMU-only approaches, we propose to use flexible combinations of multiple-view, calibrated video and IMU input along with a pose prior in an online optimization-based framework, which allows the full 6-DoF motion to be recovered including axial rotation of limbs and drift-free global position. A method for sorting and assigning raw input 2D keypoint detections into corresponding subjects is presented which facilitates multi-person tracking and rejection of any bystanders in the scene. The approach is evaluated on data from several indoor and outdoor capture environments with one or more subjects and the trade-off between input sparsity and tracking performance is discussed. State-of-the-art pose estimation performance is obtained on the Total Capture (mutli-view video and IMU) and Human 3.6M (multi-view video) datasets. Finally, a live demonstrator for the approach is presented showing real-time capture, solving and character animation using a light-weight, commodity hardware setup. Charles Malleson, John P. Collomosse, Adrian Hilton 0001 |
Int. J. Comput. Vis. | 2 |
| 2020 | Deep learning with wearable based heart rate variability for prediction of mental and general health
Louise V. Coutts, David Plans, Alan W. Brown, John P. Collomosse |
J. Biomed. Informatics | 4 |
| 2020 | Tamper-Proofing Video With Hierarchical Attention Autoencoder Hashing on BlockchainabstractWe present ARCHANGEL; a novel distributed ledger based system for assuring the long-term integrity of digital video archives. First, we introduce a novel deep network architecture using a hierarchical attention autoencoder (HAAE) to compute temporal content hashes (TCHs) from minutes or hour-long audio-visual streams. Our TCHs are sensitive to accidental or malicious content modification (tampering). The focus of our self-supervised HAAE is to guard against content modification such as frame truncation or corruption but ensure invariance against format shift (i.e. codec change). This is necessary due to the curatorial requirement for archives to format shift video over time to ensure future accessibility. Second, we describe how the TCHs (and the models used to derive them) are secured via a proof-of-authority blockchain distributed across multiple independent archives. We report on the efficacy of ARCHANGEL within the context of a trial deployment in which the national government archives of the United Kingdom, United States of America, Estonia, Australia and Norway participated. Tu Bui, Daniel Cooper, John P. Collomosse, Mark Bell, Alex Green 0002, John Sheridan, Jez Higgins, Arindra Das, Jared Keller 0001, Olivier Thereaux |
IEEE Trans. Multim. | 3 |
| 2020 | Inpainting of Wide-Baseline Multiple Viewpoint VideoabstractWe describe a non-parametric algorithm for multiple-viewpoint video inpainting. Uniquely, our algorithm addresses the domain of wide baseline multiple-viewpoint video (MVV) with no temporal look-ahead in near real time speed. A Dictionary of Patches (DoP) is built using multi-resolution texture patches reprojected from geometric proxies available in the alternate views. We dynamically update the DoP over time, and a Markov Random Field optimisation over depth and appearance is used to resolve and align a selection of multiple candidates for a given patch, this ensures the inpainting of large regions in a plausible manner conserving both spatial and temporal coherence. We demonstrate the removal of large objects (e.g., people) on challenging indoor and outdoor MVV exhibiting cluttered, dynamic backgrounds and moving cameras. Andrew Gilbert, Matthew Trumble, Adrian Hilton 0001, John P. Collomosse |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2019 | Robust Synthesis of Adversarial Visual Examples Using a Deep Image Prior
Thomas Gittings, Steve A. Schneider, John P. Collomosse |
BMVC | 3 |
| 2019 | LiveSketch: Query Perturbations for Guided Sketch-Based Visual SearchabstractLiveSketch is a novel algorithm for searching large image collections using hand-sketched queries. LiveSketch tackles the inherent ambiguity of sketch search by creating visual suggestions that augment the query as it is drawn, making query specification an iterative rather than one-shot process that helps disambiguate users' search intent. Our technical contributions are: a triplet convnet architecture that incorporates an RNN based variational autoencoder to search for images using vector (stroke-based) queries; real-time clustering to identify likely search intents (and so, targets within the search embedding); and the use of backpropagation from those targets to perturb the input stroke sequence, so suggesting alterations to the query in order to guide the search. We show improvements in accuracy and time-to-task over contemporary baselines using a 67M image corpus. John P. Collomosse, Tu Bui, Hailin Jin |
CVPR | 1 |
| 2019 | An Internal Learning Approach to Video InpaintingabstractWe propose a novel video inpainting algorithm that simultaneously hallucinates missing appearance and motion (optical flow) information, building upon the recent 'Deep Image Prior' (DIP) that exploits convolutional network architectures to enforce plausible texture in static images. In extending DIP to video we make two important contributions. First, we show that coherent video inpainting is possible without a priori training. We take a generative approach to inpainting based on internal (within-video) learning without reliance upon an external corpus of visual data to train a one-size-fits-all model for the large space of general videos. Second, we show that such a framework can jointly generate both appearance and flow, whilst exploiting these complementary modalities to ensure mutual consistency. We show that leveraging appearance statistics specific to each video achieves visually plausible results whilst handling the challenging problem of long-term consistency. Long Mai, Hailin Jin, Ning Xu 0007, John P. Collomosse |
ICCV | 6 |
| 2019 | Fusing Visual and Inertial Sensors with Semantics for 3D Human Pose EstimationabstractWe propose an approach to accurately estimate 3D human pose by fusing multi-viewpoint video (MVV) with inertial measurement unit (IMU) sensor data, without optical markers, a complex hardware setup or a full body model. Uniquely we use a multi-channel 3D convolutional neural network to learn a pose embedding from visual occupancy and semantic 2D pose estimates from the MVV in a discretised volumetric probabilistic visual hull. The learnt pose stream is concurrently processed with a forward kinematic solve of the IMU data and a temporal model (LSTM) exploits the rich spatial and temporal long range dependencies among the solved joints, the two streams are then fused in a final fully connected layer. The two complementary data sources allow for ambiguities to be resolved within each sensor modality, yielding improved accuracy over prior methods. Extensive evaluation is performed with state of the art performance reported on the popular Human 3.6M dataset (Ionescu et al. in Intell IEEE Trans Pattern Anal Mach 36(7):1325–1339, 2014), the newly released TotalCapture dataset and a challenging set of outdoor videos TotalCaptureOutdoor. We release the new hybrid MVV dataset (TotalCapture) comprising of multi-viewpoint video, IMU and accurate 3D skeletal joint ground truth derived from a commercial motion capture system. The dataset is available online at http://cvssp.org/data/totalcapture/ . Andrew Gilbert, Matthew Trumble, Charles Malleson, Adrian Hilton 0001, John P. Collomosse |
Int. J. Comput. Vis. | 5 |
| 2018 | Deep Manifold Alignment for Mid-Grain Sketch Based Image Retrieval
Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse |
ACCV (3) | 4 |
| 2018 | Disentangling Structure and Aesthetics for Style-Aware Image CompletionabstractContent-aware image completion or in-painting is a fundamental tool for the correction of defects or removal of objects in images. We propose a non-parametric in-painting algorithm that enforces both structural and aesthetic (style) consistency within the resulting image. Our contributions are two-fold: (1) we explicitly disentangle image structure and style during patch search and selection to ensure a visually consistent look and feel within the target image. (2) we perform adaptive stylization of patches to conform the aesthetics of selected patches to the target image, so harmonizing the integration of selected patches into the final composition. We show that explicit consideration of visual style during in-painting delivers excellent qualitative and quantitative results across the varied image styles and content, over the Places2 scene photographic dataset and a challenging new in-painting dataset of artwork derived from BAM! Andrew Gilbert, John P. Collomosse, Hailin Jin, Brian L. Price |
CVPR | 2 |
| 2018 | ARCHANGEL: Trusted Archives of Digital Public DocumentsabstractWe present ARCHANGEL; a decentralised platform for ensuring the long-term integrity of digital documents stored within public archives. Document integrity is fundamental to public trust in archives. Yet currently that trust is built upon institutional reputation --- trust at face value in a centralised authority, like a national government archive or University. ARCHANGEL proposes a shift to a technological underscoring of that trust, using distributed ledger technology (DLT) to cryptographically guarantee the provenance, immutability and so the integrity of archived documents. We describe the ARCHANGEL architecture, and report on a prototype of that architecture build over the Ethereum infrastructure. We report early evaluation and feedback of ARCHANGEL from stakeholders in the research data archives space. John P. Collomosse, Tu Bui, John Sheridan, Alex Green 0002, Mark Bell, Jamie Fawcett, Jez Higgins, Olivier Thereaux |
DocEng | 1 |
| 2018 | Volumetric Performance Capture from Minimal Camera Viewpoints
Andrew Gilbert, Marco Volino, John P. Collomosse, Adrian Hilton 0001 |
ECCV (11) | 3 |
| 2018 | Deep Autoencoder for Combined Human Pose Estimation and Body Model Upscaling
Matthew Trumble, Andrew Gilbert, Adrian Hilton 0001, John P. Collomosse |
ECCV (10) | 4 |
| 2018 | TAPESTRY: Visualizing Interwoven Identities for Trust ProvenanceabstractIn this paper we report our study involving an early prototype of TAPESTRY, a service to support people and businesses to connect safely online through the use of a Machine Learning generated visualization. Establishing the veracity of the person or business behind a pseudonomized identity, online, is a challenge for many people. In the burgeoning digital economy, finding ways to support good decision-making in potentially risky online exchanges is of vital importance. In this paper, we propose a Machine Learning method to extract temporal patterns from data on individuals' behavioral norms in their online activity. This monitors and communicates the coherence of these activities to others, especially those who are about to disclose personal information to the individual, in a visualization. We report findings from a user trial that examined how people accessed and interpreted the TAPESTRY visualization to inform their decisions on who to back in a mock crowdfunding campaign to evaluate its efficacy. The study proved the protocol of the Machine Learning method and qualitative insights are informing iterations of the visualization design to enhance user experience and support understanding. John P. Collomosse, Arthi Kanchana Manohar, Jo Briggs, Jamie Steane |
VizSEC | 2 |
| 2018 | Sketching out the details: Sketch-based image retrieval using convolutional neural networks with multi-stage regressionabstractWe propose and evaluate several deep network architectures for measuring the similarity between sketches and photographs, within the context of the sketch based image retrieval (SBIR) task.We study the ability of our networks to generalize across diverse object categories from limited training data, and explore in detail strategies for weight sharing, pre-processing, data augmentation and dimensionality reduction.In addition to a detailed comparative study of network configurations, we contribute by describing a hybrid multi-stage training network that exploits both contrastive and triplet networks to exceed state of the art performance on several SBIR benchmarks by a significant margin.Datasets and models are available at www.cvssp.org. Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse |
Comput. Graph. | 4 |
| 2017 | Real-Time Full-Body Motion Capture from Video and IMUsabstractA real-time full-body motion capture system is presented which uses input from a sparse set of inertial measurement units (IMUs) along with images from two or more standard video cameras and requires no optical markers or specialized infra-red cameras. A real-time optimization-based framework is proposed which incorporates constraints from the IMUs, cameras and a prior pose model. The combination of video and IMU data allows the full 6-DOF motion to be recovered including axial rotation of limbs and drift-free global position. The approach was tested using both indoor and outdoor captured data. The results demonstrate the effectiveness of the approach for tracking a wide range of human motion in real time in unconstrained indoor/outdoor scenes. Charles Malleson, Andrew Gilbert, Matthew Trumble, John P. Collomosse, Adrian Hilton 0001, Marco Volino |
3DV | 4 |
| 2017 | Total Capture: 3D Human Pose Estimation Fusing Video and Inertial Sensors
Matthew Trumble, Andrew Gilbert, Charles Malleson, Adrian Hilton 0001, John P. Collomosse |
BMVC | 5 |
| 2017 | Sketched Visual Narratives for Image and Video SearchabstractThe internet is transforming into a visual medium; over 80% of the internet is forecast to be visual content by 2018, and most of this content will be consumed on mobile devices featuring a touch-screen as their primary interface. Gestural interaction, such as sketch, presents an intuitive way to interact with these devices. Imagine a Google image search in which you specify your query by sketching the desired image with your finger, rather than (or in addition to) describing it with text words. Sketch offers an orthogonal perspective on visual search - enabling concise specification of appearance (via sketch) in addition to semantics (via text). In this talk, John Collomosse will present a summary of his group's work on the use of free-hand sketches for the visual search and manipulation of images and video. He will begin by describing a scalable system for sketch based search of multi-million image databases, based upon their Gradient Field HOG (GF-HOG) descriptor. He will then describe how deep learning can be used to enhance performance of the retrieval. Imagine a product catalogue in which you sketched, say an engineering part, rather than using a text or serial numbers to find it? John will then describe how scalable search of video can be similarly achieved, through the depiction of sketched visual narratives that depict not only objects but also their motion (dynamics) as a constraint to find relevant video clips. The work presented in this talk has been supported by the EPSRC and AHRC between 2012-2016. John P. Collomosse |
DocEng | 1 |
| 2017 | Sketching with Style: Visual Search with Sketches and Aesthetic ContextabstractWe propose a novel measure of visual similarity for image retrieval that incorporates both structural and aesthetic (style) constraints. Our algorithm accepts a query as sketched shape, and a set of one or more contextual images specifying the desired visual aesthetic. A triplet network is used to learn a feature embedding capable of measuring style similarity independent of structure, delivering significant gains over previous networks for style discrimination. We incorporate this model within a hierarchical triplet network to unify and learn a joint space from two discriminatively trained streams for style and structure. We demonstrate that this space enables, for the first time, style-constrained sketch search over a diverse domain of digital artwork comprising graphics, paintings and drawings. We also briefly explore alternative query modalities. John P. Collomosse, Tu Bui, Kimberly Wilber, Hailin Jin |
ICCV | 1 |
| 2017 | BAM! The Behance Artistic Media Dataset for Recognition Beyond PhotographyabstractComputer vision systems are designed to work well within the context of everyday photography. However, artists often render the world around them in ways that do not resemble photographs. Artwork produced by people is not constrained to mimic the physical world, making it more challenging for machines to recognize.,,This work is a step toward teaching machines how to categorize images in ways that are valuable to humans. First, we collect a large-scale dataset of contemporary artwork from Behance, a website containing millions of portfolios from professional and commercial artists. We annotate Behance imagery with rich attribute labels for content, emotions, and artistic media. Furthermore, we carry out baseline experiments to show the value of this dataset for artistic style prediction, for improving the generality of existing object classifiers, and for the study of visual domain adaptation. We believe our Behance Artistic Media dataset will be a good starting point for researchers wishing to study artistic imagery and relevant problems. This dataset can be found at https://bam-dataset.org/. Kimberly Wilber, Hailin Jin, Aaron Hertzmann, John P. Collomosse, Serge J. Belongie |
ICCV | 5 |
| 2017 | Compact descriptors for sketch-based image retrieval using a triplet loss convolutional neural networkabstractWe present an efficient representation for sketch based image retrieval (SBIR) derived from a triplet loss convolutional neural network (CNN). We treat SBIR as a cross-domain modelling problem, in which a depiction invariant embedding of sketch and photo data is learned by regression over a siamese CNN architecture with half-shared weights and modified triplet loss function. Uniquely, we demonstrate the ability of our learned image descriptor to generalise beyond the categories of object present in our training data, forming a basis for general cross-category SBIR. We explore appropriate strategies for training, and for deriving a compact image descriptor from the learned representation suitable for indexing data on resource constrained e. g. mobile devices. We show the learned descriptors to outperform state of the art SBIR on the defacto standard Flickr15k dataset using a significantly more compact (56 bits per image, i. e. ≈ 105KB total) search index than previous methods. Datasets and models are available from the CVSSP datasets server at www.cvssp.org. Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse |
Comput. Vis. Image Underst. | 4 |
| 2016 | Evolutionary data purification for social media classificationabstractWe present a novel algorithm for the semantic labeling of photographs shared via social media. Such imagery is diverse, exhibiting high intra-class variation that demands large training data volumes to learn representative classifiers. Unfortunately image annotation at scale is noisy resulting in errors in the training corpus that confound classifier accuracy. We show how evolutionary algorithms may be applied to select a 'purified' subset of the training corpus to optimize classifier performance. We demonstrate our approach over a variety of image descriptors (including deeply learned features) and support vector machines. Stuart James, John P. Collomosse |
ICPR | 2 |
| 2015 | Font finder: Visual recognition of typeface in printed documentsabstractWe describe a novel algorithm for visually identifying the font used in a scanned printed document. Our algorithm requires no pre-recognition of characters in the string (i. e. optical character recognition). Gradient orientation features are collected local the character boundaries, and quantized into a hierarchical Bag of Visual Words representation. Following stop-word analysis, classification via logistic regression (LR) of the codebooked features yields per-character probabilities which are combined across the string to decide the posterior for each font. We achieve 93.4% accuracy over a 1000 font database of scanned printed text comprising Latin characters. Tu Bui, John P. Collomosse |
ICIP | 2 |
| 2015 | Foreword: Special section on visual media production
John P. Collomosse, Peter Hall 0001 |
Comput. Graph. | 1 |
| 2015 | 4D Model Flow: Precomputed Appearance Alignment for Real-time 4D Video InterpolationabstractWe introduce the concept of 4D model flow for the precomputed alignment of dynamic surface appearance across 4D video sequences of different motions reconstructed from multi-view video. Precomputed 4D model flow allows the efficient parametrization of surface appearance from the captured videos, which enables efficient real-time rendering of interpolated 4D video sequences whilst accurately reproducing visual dynamics, even when using a coarse underlying geometry. We estimate the 4D model flow using an image-based approach that is guided by available geometry proxies. We propose a novel representation in surface texture space for efficient storage and online parametric interpolation of dynamic appearance. Our 4D model flow overcomes previous requirements for computationally expensive online optical flow computation for data-driven alignment of dynamic surface appearance by precomputing the appearance alignment. This leads to an efficient rendering technique that enables the online interpolation between 4D videos in real time, from arbitrary viewpoints and with visual quality comparable to the state of the art. Dan Casas, Christian Richardt, John P. Collomosse, Christian Theobalt, Adrian Hilton 0001 |
Comput. Graph. Forum | 3 |
| 2015 | Comprehensible Video ThumbnailsabstractAbstract We present the Comprehensible Video Thumbnail; an automatically generated visual précis that summarizes salient objects and their dynamics within a video clip. Salient moving objects are detected within clips using a novel stochastic sampling technique that identifies, clusters and then tracks regions exhibiting affine motion coherence within the clip. Tracks are analyzed to determine salient instants at which motion and/or appearance changes significantly, and the resulting objects arranged in a stylized composition optimized to reduce visual clutter and enhance understanding of scene content through classification and depiction of motion type and trajectory. The result is an object‐level visual gist of the clip, obtained with full automation and depicting content and motion with greater descriptive power that prior approaches. We demonstrate these benefits through a user study in which the comprehension of our video thumbnails is compared to the state of the art over a wide variety of sports footage. Jongdae Kim, Charles Gray, Paul Asente, John P. Collomosse |
Comput. Graph. Forum | 4 |
| 2015 | Hybrid Skeletal-Surface Motion Graphs for Character Animation from 4D Performance CaptureabstractWe present a novel hybrid representation for character animation from 4D Performance Capture (4DPC) data which combines skeletal control with surface motion graphs. 4DPC data are temporally aligned 3D mesh sequence reconstructions of the dynamic surface shape and associated appearance from multiple-view video. The hybrid representation supports the production of novel surface sequences which satisfy constraints from user-specified key-frames or a target skeletal motion. Motion graph path optimisation concatenates fragments of 4DPC data to satisfy the constraints while maintaining plausible surface motion at transitions between sequences. Space-time editing of the mesh sequence using a learned part-based Laplacian surface deformation model is performed to match the target skeletal motion and transition between sequences. The approach is quantitatively evaluated for three 4DPC datasets with a variety of clothing styles. Results for key-frame animation demonstrate production of novel sequences that satisfy constraints on timing and position of less than 1% of the sequence duration and path length. Evaluation of motion-capture-driven animation over a corpus of 130 sequences shows that the synthesised motion accurately matches the target skeletal motion. The combination of skeletal control with the surface motion graph extends the range and style of motion which can be produced while maintaining the natural dynamics of shape and appearance from the captured performance. Peng Huang 0001, Margara Tejera, John P. Collomosse, Adrian Hilton 0001 |
ACM Trans. Graph. | 3 |
| 2014 | Admixed portrait: reflections on being online as a new parentabstractThis Pictorial documents the process of designing a device as an intervention within a field study of new parents. The device was deployed in participating parents' homes to invite reflection on their everyday experiences of portraying self and others through social media in their transition to parenthood. Diego Trujillo-Pisanty, Abigail Durrant, Sarah Martindale, Stuart James, John P. Collomosse |
Conference on Designing Interactive Systems | 5 |
| 2014 | Optimal Representation of Multiple View Video
Marco Volino, Dan Casas, John P. Collomosse, Adrian Hilton 0001 |
BMVC | 3 |
| 2014 | A particle filtering approach to salient video object localizationabstractWe describe a novel fully automatic algorithm for identifying salient objects in video based on their motion. Spatially coherent clusters of optical flow vectors are sampled to generate estimates of affine motion parameters local to super-pixels identified within each frame. These estimates, combined with spatial data, form coherent point distributions in a 5D solution space corresponding to objects or parts there-of. These distributions are temporally denoised using a particle filtering approach, and clustered to estimate the position and motion parameters of salient moving objects in the clip. We demonstrate localization of salient object/s in a variety of clips exhibiting moving and cluttered backgrounds. Charles Gray, Stuart James, John P. Collomosse, Paul Asente |
ICIP | 3 |
| 2014 | Incremental transfer learning for object recognition in streaming videoabstractWe present a new incremental learning framework for realtime object recognition in video streams. ImageNet is used to bootstrap a set of one-vs-all incrementally trainable SVMs which are updated by user annotation events during streaming. We adopt an inductive transfer learning (ITL) approach to warp the video feature space to the ImageNet feature space, so enabling the incremental updates. Uniquely, the transformation used for the ITL warp is also learned incrementally using the same update events. We demonstrate a semi-automated video logging (SAVL) system using our incrementally learned ITL approach and show this to outperform existing SAVL which uses non-incremental transfer learning. Jongdae Kim, John P. Collomosse |
ICIP | 2 |
| 2014 | Wide Baseline Multi-view Video Matting Using a Hybrid Markov Random FieldabstractWe describe a novel framework for segmenting a time- and view-coherent foreground matte sequence from synchronised multiple view video. We construct a Markov Random Field (MRF) comprising links between super pixels corresponded across views, and links between super pixels and their constituent pixels. Texture, colour and disparity cues are incorporated to model foreground appearance. We solve using a multi-resolution iterative approach enabling an eight view high definition (HD) frame to be processed in less than a minute. Furthermore we incorporate a temporal diffusion process introducing a prior on the MRF using information propagated from previous frames, and a facility for optional user correction. The result is a set of temporally coherent mattes solved for simultaneously across views for each frame, exploiting similarities across views and time. Tinghuai Wang, John P. Collomosse, Adrian Hilton 0001 |
ICPR | 2 |
| 2014 | ReEnact: Sketch based Choreographic Design from Archival Dance FootageabstractWe describe a novel system for synthesising video choreography using sketched visual storyboards comprising human poses (stick men) and action labels. First, we describe an algorithm for searching archival dance footage using sketched pose. We match using an implicit representation of pose parsed from a mix of challenging low and high fidelity footage. In a training pre-process we learn a mapping between a set of exemplar sketches and corresponding pose representations parsed from the video, which are generalized at query-time to enable retrieval over previously unseen frames, and over additional unseen videos. Second, we describe how a storyboard of sketched poses, interspersed with labels indicating connecting actions, may be used to drive the synthesis of novel video choreography from the archival footage. Stuart James, Manuel J. Fonseca, John P. Collomosse |
ICMR | 3 |
| 2014 | 4D video textures for interactive character appearanceabstractAbstract 4D Video Textures (4DVT) introduce a novel representation for rendering video‐realistic interactive character animation from a database of 4D actor performance captured in a multiple camera studio. 4D performance capture reconstructs dynamic shape and appearance over time but is limited to free‐viewpoint video replay of the same motion. Interactive animation from 4D performance capture has so far been limited to surface shape only. 4DVT is the final piece in the puzzle enabling video‐realistic interactive animation through two contributions: a layered view‐dependent texture map representation which supports efficient storage, transmission and rendering from multiple view video capture; and a rendering approach that combines multiple 4DVT sequences in a parametric motion space, maintaining video quality rendering of dynamic surface appearance whilst allowing high‐level interactive control of character motion and viewpoint. 4DVT is demonstrated for multiple characters and evaluated both quantitatively and through a user‐study which confirms that the visual quality of captured video is maintained. The 4DVT representation achieves >90% reduction in size and halves the rendering cost. Dan Casas, Marco Volino, John P. Collomosse, Adrian Hilton 0001 |
Comput. Graph. Forum | 3 |
| 2014 | TouchCut: Fast image and video segmentation using single-touch interaction
Tinghuai Wang, John P. Collomosse |
Comput. Vis. Image Underst. | 3 |
| 2014 | Guest Editorial: Tracking, Detection and Segmentation
Richard Bowden, John P. Collomosse, Krystian Mikolajczyk |
Int. J. Comput. Vis. | 2 |
| 2013 | Learnable Stroke Models for Example-based Portrait PaintingabstractWe present a novel algorithm for stylizing photographs into portrait paintings comprised of curved brush strokes. Rather than drawing upon a prescribed set of heuristics to place strokes, our system learns a flexible model of artistic style by analyzing training data from a human artist. Given a training pair — a source image and painting of that image—a non-parametric model of style is learned by observing the geometry and tone of brush strokes local to image features. A Markov Random Field (MRF) enforces spatial coherence of style parameters. Style models local to facial features are learned using a semantic segmentation of the input face image, driven by a combination of an Active Shape Model and Graph-cut. We evaluate style transfer between a variety of training and test images, demonstrating a wide gamut of learned brush and shading styles. Tinghuai Wang, John P. Collomosse, Andrew Hunter, Darryl Greig |
BMVC | 2 |
| 2013 | Markov random fields for sketch based video retrievalabstractWe describe a new system for searching video databases using free-hand sketched queries. Our query sketches depict both object appearance and motion, and are annotated with keywords that indicate the semantic category of each object. We parse space-time volumes from video to form graph representation, which we match to sketches under a Markov Random Field (MRF) optimization. The MRF energy function is used to rank videos for relevance and contains unary, pairwise and higher-order potentials that reflect the colour, shape, motion and type of sketched objects. We evaluate performance over a dataset of 500 sports footage clips. Rui Hu 0007, Stuart James, Tinghuai Wang, John P. Collomosse |
ICMR | 4 |
| 2013 | A performance evaluation of gradient field HOG descriptor for sketch based image retrieval
Rui Hu 0007, John P. Collomosse |
Comput. Vis. Image Underst. | 2 |
| 2013 | Virtual Volumetric Graphics on Commodity Displays Using 3D Viewer Tracking
Charles Malleson, John P. Collomosse |
Int. J. Comput. Vis. | 2 |
| 2013 | State of the "Art": A Taxonomy of Artistic Stylization Techniques for Images and VideoabstractThis paper surveys the field of nonphotorealistic rendering (NPR), focusing on techniques for transforming 2D input (images and video) into artistically stylized renderings. We first present a taxonomy of the 2D NPR algorithms developed over the past two decades, structured according to the design characteristics and behavior of each technique. We then describe a chronology of development from the semiautomatic paint systems of the early nineties, through to the automated painterly rendering systems of the late nineties driven by image gradient analysis. Two complementary trends in the NPR literature are then addressed, with reference to our taxonomy. First, the fusion of higher level computer vision and NPR, illustrating the trends toward scene analysis to drive artistic abstraction and diversity of style. Second, the evolution of local processing approaches toward edge-aware filtering for real-time stylization of images and video. The survey then concludes with a discussion of open challenges for 2D NPR identified in recent NPR symposia, including topics such as user and aesthetic evaluation. Jan Eric Kyprianidis, John P. Collomosse, Tinghuai Wang, Tobias Isenberg 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2012 | Topic based pose relevance learning in dance archivesabstractThis paper improves spatial pyramid kernel (SPK) and proposes a relevance learning approach to compare performer's poses in a large dance archive, the NRCD collection1. Domain knowledge of Choreutics is exploited to define pose topics and a selection operator is developed for pose topic matching. The visual structure descriptor of self similarity (SSF) is extended to hierarchical self similarity (HSSF) to keep shape context. The framework of Bag-of-Visual Words (BOVW) is applied to encode as well as to speed up the matching on pose topics/topic combinations. This alleviates the complexity in limb allocation which is infeasible in our data. Extensive experiments show that the new approach outperforms the original SPK in both precision and robustness. Reede Ren, John P. Collomosse, Joemon M. Jose |
CIKM | 2 |
| 2012 | Touchcut: Single-touch object segmentation driven by level set methodsabstractIn this paper, we propose an object segmentation algorithm driven by minimal user interactions. Compared to previous user-guided systems, our system can cut out the desired object in a given image with only a single finger touch minimizing user effort. The proposed model harnesses both edge and region based local information in an adaptive manner as well as geometric cues implied by the user-input to achieve fast and robust segmentation in a level set framework. We demonstrate the advantages of our method in terms of computational efficiency and accuracy comparing qualitatively and quantitatively with graph cut based techniques. Tinghuai Wang, John P. Collomosse |
ICASSP | 3 |
| 2012 | Annotated Free-Hand Sketches for Video Retrieval Using Object Semantics and Motion
Rui Hu 0007, Stuart James, John P. Collomosse |
MMM | 3 |
| 2012 | Skeletons from sketches of dancing posesabstractThe contribution of this paper is a sketch parser able to recognize the several components of a skeleton described using the drawing of a stick-man. We describe the sketch parser in detail, and briefly outline how it is applied to form the front-end of a sketch based retrieval system capable of searching for human poses in archival dance footage. Manuel J. Fonseca, Stuart James, John P. Collomosse |
VL/HCC | 3 |
| 2012 | Visual Sentences for Pose Retrieval Over Low-Resolution Cross-Media Dance CollectionsabstractWe describe a system for matching human posture (pose) across a large cross-media archive of dance footage spanning nearly 100 years, comprising digitized photographs and videos of rehearsals and performances. This footage presents unique challenges due to its age, quality and diversity. We propose a forest-like pose representation combining visual structure (self-similarity) descriptors over multiple scales, without explicitly detecting limb positions which would be infeasible for our data. We explore two complementary multi-scale representations, applying passage retrieval and latent Dirichlet allocation (LDA) techniques inspired by the text retrieval domain, to the problem of pose matching. The result is a robust system capable of quickly searching large cross-media collections for similarity to a visually specified query pose. We evaluate over a cross-section of the UK National Research Centre for Dance's (UK-NRCD), and the Siobhan Davies Replay's (SDR) digital dance archives, using visual queries supplied by dance professionals. We demonstrate significant performance improvements over two base-lines: classical single and multi-scale bag of visual words (BoVW) and spatial pyramid kernel (SPK) matching . Reede Ren, John P. Collomosse |
IEEE Trans. Multim. | 2 |
| 2012 | Probabilistic Motion Diffusion of Labeling Priors for Coherent Video SegmentationabstractWe present a robust algorithm for temporally coherent video segmentation. Our approach is driven by multi-label graph cut applied to successive frames, fusing information from the current frame with an appearance model and labeling priors propagated forwarded from past frames. We propagate using a novel motion diffusion model, producing a per-pixel motion distribution that mitigates against cumulative estimation errors inherent in systems adopting “hard” decisions on pixel motion at each frame. Further, we encourage spatial coherence by imposing label consistency constraints within image regions (super-pixels) obtained via a bank of unsupervised frame segmentations, such as mean-shift. We demonstrate quantitative improvements in accuracy over state-of-the-art methods on a variety of sequences exhibiting clutter and agile motion, adopting the Berkeley methodology for our comparative evaluation. Tinghuai Wang, John P. Collomosse |
IEEE Trans. Multim. | 2 |
| 2011 | A bag-of-regions approach to sketch-based image retrievalabstractThis paper presents a system for retrieving photographs using free-hand sketched queries. Regions are extracted from each image by gathering nodes of a hierarchical image segmentation into a bag-of-regions (BoR) representation. The BoR represents object shape at multiple scales, encoding shape even in the presence of adjacent clutter. We extract a shape representation from each region, using the Gradient Field HoG (GF-HOG) descriptor which enables direct comparison with the sketched query. The retrieval pipeline yields significant performance improvements over the previous GF-HOG results reliant on single-scale Canny edge maps, and over leading descriptors (SIFT, SSIM) for visual search. In addition, our system enables localization of the sketched object within matching images. Rui Hu 0007, Tinghuai Wang, John P. Collomosse |
ICIP | 3 |
| 2011 | A BOVW Based Query Generative Model
Reede Ren, John P. Collomosse, Joemon M. Jose |
MMM (1) | 2 |
| 2011 | Special section on Non-Photorealistic Animation and Rendering (NPAR) 2010
John P. Collomosse, Tobias Isenberg 0001 |
Comput. Graph. | 1 |
| 2011 | Stylized ambient displays of digital media collections
Tinghuai Wang, John P. Collomosse, Rui Hu 0007, David Slatter, Darryl Greig, Phil Cheatle |
Comput. Graph. | 2 |
| 2010 | Gradient field descriptor for sketch based retrieval and localizationabstractWe present an image retrieval system driven by free-hand sketched queries depicting shape. We introduce Gradient Field HoG (GF-HOG) as a depiction invariant image descriptor, encapsulating local spatial structure in the sketch and facilitating efficient codebook based retrieval. We show improved retrieval accuracy over 3 leading descriptors (Self Similarity, SIFT, HoG) across two datasets (Flickr160, ETHZ extended objects), and explain how GF-HOG can be combined with RANSAC to localize sketched objects within relevant images. We also demonstrate a prototype sketch driven photo montage application based on our system. Rui Hu 0007, Mark Barnard, John P. Collomosse |
ICIP | 3 |
| 2010 | Multi-label propagation for coherent video segmentation and artistic stylizationabstractWe present a new algorithm for segmenting video frames into temporally stable colored regions, applying our technique to create artistic stylizations (e.g. cartoons and paintings) from real video sequences. Our approach is based on a multi-label graph cut applied to successive frames, in which the color data term and label priors are incrementally updated and propagated over time. We demonstrate coherent segmentation and stylization over a variety of home videos. Tinghuai Wang, Jean-Yves Guillemaut, John P. Collomosse |
ICIP | 3 |
| 2010 | Motion-sketch Based Video Retrieval Using a Trellis Levenshtein DistanceabstractWe present a fast technique for retrieving video clips using free-hand sketched queries. Visual keypoints within each video are detected and tracked to form short trajectories, which are clustered to form a set of space-time tokens summarising video content. A Viterbi process matches a space-time graph of tokens to a description of colour and motion extracted from the query sketch. Inaccuracies in the sketched query are ameliorated by computing path cost using a Levenshtein (edit) distance. We evaluate over datasets of sports footage. Rui Hu 0007, John P. Collomosse |
ICPR | 2 |
| 2009 | Storyboard sketches for Content Based Video RetrievalabstractWe present a novel Content Based Video Retrieval (CBVR) system, driven by free-hand sketch queries depicting both objects and their movement (via dynamic cues; streak-lines and arrows). Our main contribution is a probabilistic model of video clips (based on Linear Dynamical Systems), leading to an algorithm for matching descriptions of sketched objects to video. We demonstrate our model fitting to clips under static and moving camera conditions, exhibiting linear and oscillatory motion. We evaluate retrieval on two real video data sets, and on a video data set exhibiting controlled variation in shape, color, motion and clutter. John P. Collomosse, Graham McNeill |
ICCV | 1 |
| 2009 | Mobile augmented reality based 3D snapshotsabstractWe describe a mobile augmented reality application that is based on 3D snapshotting using multiple photographs. Optical square markers provide the anchor for reconstructed virtual objects in the scene. A novel approach based on pixel flow highly improves tracking performance. This dual tracking approach also allows for a new single-button user interface metaphor for moving virtual objects in the scene. The development of the AR viewer was accompanied by user studies confirming the chosen approach. Peter Keitler, Frieder Pankratz, Björn Schwerdtfeger, Daniel Pustka, Wolf Rödiger, Gudrun Klinker, Christian Rauch 0005, Anup Chathoth, John P. Collomosse, Yi-Zhe Song |
ISMAR | 9 |
| 2008 | Free-hand sketch grouping for video retrievalabstractWe present an algorithm for extracting object descriptions from free-hand sketches of remembered scenes, drawn as video retrieval queries. Our sketches depict scene content, as well as indicators of motion. We report an exploratory study investigating how people sketch to depict recalled events. We incorporate several observations from this study into the design of a novel sketch parsing algorithm. We draw upon a temporal HMM classifier to recognise common pictograms, and graph-cut to identify more general objects. John P. Collomosse, Graham McNeill, Leon Adam Watts |
ICPR | 1 |
| 2007 | RTcams: A New Perspective on Nonphotorealistic Rendering from PhotographsabstractAbstract-We introduce a simple but versatile camera model that we call the Rational Tensor Camera (RTcam). RTcams are well principled mathematically and provably subsume several important contemporary camera models in both computer graphics and vision; their generality is one contribution. They can be used alone or compounded to produce more complicated visual effects. In this paper, we apply RTcams to generate synthetic artwork with novel perspective effects from real photographs. Existing Nonphotorealistic Rendering from Photographs (NPRP) is constrained to the projection inherent in the source photograph, which is most often linear. RTcams lift this restriction and so contribute to NPRP via multiperspective projection. This paper describes RTcams, compares them to contemporary alternatives, and discusses how to control them in practice. Illustrative examples are provided throughout. Peter Hall 0001, John P. Collomosse, Yi-Zhe Song, Peiyi Shen, Chuan Li 0001 |
IEEE Trans. Vis. Comput. Graph. | 2 |
| 2006 | Video motion analysis for the synthesis of dynamic cues and Futurist art
John P. Collomosse, Peter Hall 0001 |
Graph. Model. | 1 |
| 2005 | Video Paintbox: The fine art of video painting
John P. Collomosse, Peter Hall 0001 |
Comput. Graph. | 1 |
| 2005 | Rendering cartoon-style motion cues in post-production video
John P. Collomosse, David Rowntree, Peter Hall 0001 |
Graph. Model. | 1 |
| 2005 | Stroke Surfaces: Temporally Coherent Artistic Animations from VideoabstractThe contribution of this paper is a novel framework for synthesizing nonphotorealistic animations from real video sequences. We demonstrate that, through automated mid-level analysis of the video sequence as a spatiotemporal volume--a block of frames with time as the third dimension--we are able to generate animations in a wide variety of artistic styles, exhibiting a uniquely high degree of temporal coherence. In addition to rotoscoping, matting, and novel temporal effects unique to our method, we demonstrate the extension of static nonphotorealistic rendering (NPR) styles to video, including painterly, sketchy, and cartoon shading. We demonstrate how this novel coherent shading framework may be combined with our earlier motion emphasis work to produce a comprehensive "Video Paintbox" capable of rendering complete cartoon-styled animations from video clips. John P. Collomosse, David Rowntree, Peter Hall 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2004 | A Mid-Level Description of Video, with Application to Non-photorealistic AnimationabstractThe contribution of this paper is a novel spatiotemporal description of real video sequences. Our description comprises a set of surfaces that separate objects in the video, and an accompanying database which describes objects in the video. We explain how to automatically process video into our description, and use it to solve a long-standing problem in Computer Graphics; that of automatically producing non-photorealistic (NPR) animations from video sequences. We show empirically that our description is highly compact relative to alternative coding schemes, and that our NPR animation technique out performs the current state-of-the-art in automatic video painting by about an order of magnitude, in terms of temporal coherence. 1 John P. Collomosse, Peter Hall 0001 |
BMVC | 1 |
| 2003 | Video Analysis for Cartoon-like Special EffectsabstractIn recent years the Vision community has shown interest in processing images and video for use by the entertainment industries. Typical applications include 3D reconstruction of models, and rendering graphics models into video. This paper broadly aligns with that trend, but differs in that we process video to emphasise motion in Cartoon-like styles, in which moving objects deform in defiance of physical laws, and leave trailing marks of one kind or another in their wake. We provide an introduction to the effects real animators use, and show how a judicious choice of standard processing techniques, supplemented by novel methods, can be used to achieve convincing results. We illustrate the robustness of our method using several video sequences, ranging in content from simple oscillatory to articulated motion, under both static and moving camera conditions. John P. Collomosse, David Rowntree, Peter Hall 0001 |
BMVC | 1 |
| 2003 | Cubist Style Rendering from PhotographsabstractThe contribution of the paper is a novel nonphotorealistic rendering (NPR) technique, influenced by the style of Cubist art. Specifically, we are motivated by artists such as Picasso and Braque, who produced art work by composing elements of a scene taken from multiple points of view; paradoxically, such compositions convey a sense of motion without assuming temporal dependence between views. Our method accepts a set of two-dimensional images as input and produces a Cubist style painting with minimal user interaction. We use salient features identified within the image set, such as eyes, noses, and mouths, as compositional elements; we believe the use of such features to be a unique contribution to NPR. Before composing features into a final image, we geometrically distort them to produce the more angular forms common in Cubist art. Finally, we render the composition to give a painterly effect, using an automatic algorithm. This paper describes our method, illustrating the application of our algorithm with a gallery of images. We conclude with a critical appraisal and suggest the use of "high-level" features is of interest to NPR. John P. Collomosse, Peter Hall 0001 |
IEEE Trans. Vis. Comput. Graph. | 1 |