Tu Bui

dblp:172/9628 · DBLP profile ↗
← Back
20ranked-venue papers
7as first author
11since 2021 · last 2025
0000-0001-6622-9703ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 17 · 6 first-author · 10 since 2021Artificial intelligence and machine learning · 14 · 4 first-author · 9 since 2021Databases, data management, data science and information retrieval · 1
YearPublicationVenuePosition
2025 TrustMark: Robust Watermarking and Watermark Removal for Arbitrary Resolution Images
Tu Bui, Shruti Agarwal, John P. Collomosse
ICCV1
2025 A Closer Look at Multimodal Representation Collapse
abstract
We aim to develop a fundamental understanding of modality collapse, a recently observed empirical phenomenon wherein models trained for multimodal fusion tend to rely only on a subset of the modalities, ignoring the rest. We show that modality collapse happens when noisy features from one modality are entangled, via a shared set of neurons in the fusion head, with predictive features from another, effectively masking out positive contributions from the predictive features of the former modality and leading to its collapse. We further prove that cross-modal knowledge distillation implicitly disentangles such representations by freeing up rank bottlenecks in the student encoder, denoising the fusion-head outputs without negatively impacting the predictive features from either modality. Based on the above findings, we propose an algorithm that prevents modality collapse through explicit basis reallocation, with applications in dealing with missing modalities. Extensive experiments on multiple multimodal benchmarks validate our theoretical claims. Project page: https://abhrac.github.io/mmcollapse/.
Abhra Chaudhuri, Anjan Dutta 0001, Tu Bui, Serban Georgescu
ICML3
2024 VIXEN: Visual Text Comparison Network for Image Difference Captioning
abstract
We present VIXEN - a technique that succinctly summarizes in text the visual differences between a pair of images in order to highlight any content manipulation present. Our proposed network linearly maps image features in a pairwise manner, constructing a soft prompt for a pretrained large language model. We address the challenge of low volume of training data and lack of manipulation variety in existing image difference captioning (IDC) datasets by training on synthetically manipulated images from the recent InstructPix2Pix dataset generated via prompt-to-prompt editing framework. We augment this dataset with change summaries produced via GPT-3. We show that VIXEN produces state-of-the-art, comprehensible difference captions for diverse image contents and edit types, offering a potential mitigation against misinformation disseminated via manipulated image content. Code and data are available at http://github.com/alexblck/vixen
Alexander Black 0001, Jing Shi 0005, Tu Bui, John P. Collomosse
AAAI4
2024 ATLANTIS: A Framework for Automated Targeted Language-guided Augmentation Training for Robust Image Search
Inderjeet Singh 0001, Roman Vainshtein, Alon Zolfi, Asaf Shabtai, Tu Bui, Jonathan Brokman, Omer Hofman, Fumiyoshi Kasahara, Kentaro Tsuji, Hisashi Kojima
BMVC5
2024 ProMark: Proactive Diffusion Watermarking for Causal Attribution
abstract
Generative AI (GenAI) is transforming creative work-flows through the capability to synthesize and manipulate images via high-level prompts. Yet creatives are not well supported to receive recognition or reward for the use of their content in GenAI training. To this end, we propose ProMark, a causal attribution technique to attribute a synthetically generated image to its training data concepts like objects, motifs, templates, artists, or styles. The concept information is proactively embedded into the input training images using imperceptible watermarks, and the diffusion models (unconditional or conditional) are trained to retain the corresponding watermarks in generated images. We show that we can embed as many as 216unique water-marks into the training data, and each training image can contain more than one watermark. ProMark can maintain image quality whilst outperforming correlation-based attribution. Finally, several qualitative examples are presented, providing the confidence that the presence of the watermark conveys a causative relationship between training data and synthetic images.
Vishal Asnani, John P. Collomosse, Tu Bui, Xiaoming Liu 0002, Shruti Agarwal
CVPR3
2023 VADER: Video Alignment Differencing and Retrieval
abstract
We propose VADER, a spatio- temporal matching, alignment, and change summarization method to help fight misinformation spread via manipulated videos. VADER matches and coarsely aligns partial video fragments to candidate videos using a robust visual descriptor and scalable search over adaptively chunked video content. A transformer- based alignment module then refines the temporal localization of the query fragment within the matched video. A space- time comparator module identifies regions of manipulation between aligned content, invariant to any changes due to any residual temporal misalignments or artifacts arising from non- editorial changes of the content. Robustly matching video to a trusted source enables conclusions to be drawn on video provenance, enabling informed trust decisions on content encountered. Code and data are available at https://github.com/AlexBlck/vader
Alexander Black 0001, Simon Jenni, Tu Bui, Md. Mehrab Tanjim, Stefano Petrangeli, Ritwik Sinha, Viswanathan (Vishy) Swaminathan, John P. Collomosse
ICCV3
2023 Scene designer: compositional sketch-based image retrieval with contrastive learning and an auxiliary synthesis task
Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir Ponti
Multim. Tools Appl.2
2022 RepMix: Representation Mixing for Robust Attribution of Synthesized Images
Tu Bui, Ning Yu 0006, John P. Collomosse
ECCV (14)1
2022 CoGS: Controllable Generation and Search from Sketch and Style
Cusuh Ham, Gemma Canet Tarrés, Tu Bui, James Hays, Zhe Lin 0001, John P. Collomosse
ECCV (16)3
2021 OSCAR-Net: Object-centric Scene Graph Attention for Image Attribution
abstract
Images tell powerful stories but cannot always be trusted. Matching images back to trusted sources (attribution) enables users to make a more informed judgment of the images they encounter online. We propose a robust image hashing algorithm to perform such matching. Our hash is sensitive to manipulation of subtle, salient visual details that can substantially change the story told by an image. Yet the hash is invariant to benign transformations (changes in quality, codecs, sizes, shapes, etc.) experienced by images during online redistribution. Our key contribution is OSCAR-Net1(Object-centric Scene Graph Attention for Image Attribution Network); a robust image hashing model inspired by recent successes of Transformers in the visual domain. OSCAR-Net constructs a scene graph representation that attends to fine-grained changes of every object’s visual appearance and their spatial relationships. The network is trained via contrastive learning on a dataset of original and manipulated images yielding a state of the art image hash for content fingerprinting that scales to millions of images.
Eric Nguyen, Tu Bui, Viswanathan (Vishy) Swaminathan, John P. Collomosse
ICCV2
2021 Compositional Sketch Search
abstract
We present an algorithm for searching image collections using free-hand sketches that describe the appearance and relative positions of multiple objects1Sketch based image retrieval (SBIR) methods predominantly match queries containing a single, dominant object invariant to its position within an image. Our work exploits drawings as a concise and intuitive representation for specifying entire scene compositions. We train a convolutional neural network (CNN) to encode masked visual features from sketched objects, pooling these into a spatial descriptor encoding the spatial relationships and appearances of objects in the composition. Training the CNN backbone as a Siamese network under triplet loss yields a metric search embedding for measuring compositional similarity which may be efficiently leveraged for visual search by applying product quantization.
Alexander Black 0001, Tu Bui, Long Mai, Hailin Jin, John P. Collomosse
ICIP2
2020 Sketchformer: Transformer-Based Representation for Sketched Structure
abstract
Sketchformer is a novel transformer-based representation for encoding free-hand sketches input in a vector form, i.e. as a sequence of strokes. Sketchformer effectively addresses multiple tasks: sketch classification, sketch based image retrieval (SBIR), and the reconstruction and interpolation of sketches. We report several variants exploring continuous and tokenized input representations, and contrast their performance. Our learned embedding, driven by a dictionary learning tokenization scheme, yields state of the art performance in classification and image retrieval tasks, when compared against baseline representations driven by LSTM sequence to sequence architectures: SketchRNN and derivatives. We show that sketch reconstruction and interpolation are improved significantly by the Sketchformer embedding for complex sketches with longer stroke sequences.
Leo Sampaio Ferraz Ribeiro, Tu Bui, John P. Collomosse, Moacir Ponti
CVPR2
2020 Tamper-Proofing Video With Hierarchical Attention Autoencoder Hashing on Blockchain
abstract
We present ARCHANGEL; a novel distributed ledger based system for assuring the long-term integrity of digital video archives. First, we introduce a novel deep network architecture using a hierarchical attention autoencoder (HAAE) to compute temporal content hashes (TCHs) from minutes or hour-long audio-visual streams. Our TCHs are sensitive to accidental or malicious content modification (tampering). The focus of our self-supervised HAAE is to guard against content modification such as frame truncation or corruption but ensure invariance against format shift (i.e. codec change). This is necessary due to the curatorial requirement for archives to format shift video over time to ensure future accessibility. Second, we describe how the TCHs (and the models used to derive them) are secured via a proof-of-authority blockchain distributed across multiple independent archives. We report on the efficacy of ARCHANGEL within the context of a trial deployment in which the national government archives of the United Kingdom, United States of America, Estonia, Australia and Norway participated.
Tu Bui, Daniel Cooper, John P. Collomosse, Mark Bell, Alex Green 0002, John Sheridan, Jez Higgins, Arindra Das, Jared Keller 0001, Olivier Thereaux
IEEE Trans. Multim.1
2019 LiveSketch: Query Perturbations for Guided Sketch-Based Visual Search
abstract
LiveSketch is a novel algorithm for searching large image collections using hand-sketched queries. LiveSketch tackles the inherent ambiguity of sketch search by creating visual suggestions that augment the query as it is drawn, making query specification an iterative rather than one-shot process that helps disambiguate users' search intent. Our technical contributions are: a triplet convnet architecture that incorporates an RNN based variational autoencoder to search for images using vector (stroke-based) queries; real-time clustering to identify likely search intents (and so, targets within the search embedding); and the use of backpropagation from those targets to perturb the input stroke sequence, so suggesting alterations to the query in order to guide the search. We show improvements in accuracy and time-to-task over contemporary baselines using a 67M image corpus.
John P. Collomosse, Tu Bui, Hailin Jin
CVPR2
2018 Deep Manifold Alignment for Mid-Grain Sketch Based Image Retrieval
Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse
ACCV (3)1
2018 ARCHANGEL: Trusted Archives of Digital Public Documents
abstract
We present ARCHANGEL; a decentralised platform for ensuring the long-term integrity of digital documents stored within public archives. Document integrity is fundamental to public trust in archives. Yet currently that trust is built upon institutional reputation --- trust at face value in a centralised authority, like a national government archive or University. ARCHANGEL proposes a shift to a technological underscoring of that trust, using distributed ledger technology (DLT) to cryptographically guarantee the provenance, immutability and so the integrity of archived documents. We describe the ARCHANGEL architecture, and report on a prototype of that architecture build over the Ethereum infrastructure. We report early evaluation and feedback of ARCHANGEL from stakeholders in the research data archives space.
John P. Collomosse, Tu Bui, John Sheridan, Alex Green 0002, Mark Bell, Jamie Fawcett, Jez Higgins, Olivier Thereaux
DocEng2
2018 Sketching out the details: Sketch-based image retrieval using convolutional neural networks with multi-stage regression
abstract
We propose and evaluate several deep network architectures for measuring the similarity between sketches and photographs, within the context of the sketch based image retrieval (SBIR) task.We study the ability of our networks to generalize across diverse object categories from limited training data, and explore in detail strategies for weight sharing, pre-processing, data augmentation and dimensionality reduction.In addition to a detailed comparative study of network configurations, we contribute by describing a hybrid multi-stage training network that exploits both contrastive and triplet networks to exceed state of the art performance on several SBIR benchmarks by a significant margin.Datasets and models are available at www.cvssp.org.
Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse
Comput. Graph.1
2017 Sketching with Style: Visual Search with Sketches and Aesthetic Context
abstract
We propose a novel measure of visual similarity for image retrieval that incorporates both structural and aesthetic (style) constraints. Our algorithm accepts a query as sketched shape, and a set of one or more contextual images specifying the desired visual aesthetic. A triplet network is used to learn a feature embedding capable of measuring style similarity independent of structure, delivering significant gains over previous networks for style discrimination. We incorporate this model within a hierarchical triplet network to unify and learn a joint space from two discriminatively trained streams for style and structure. We demonstrate that this space enables, for the first time, style-constrained sketch search over a diverse domain of digital artwork comprising graphics, paintings and drawings. We also briefly explore alternative query modalities.
John P. Collomosse, Tu Bui, Kimberly Wilber, Hailin Jin
ICCV2
2017 Compact descriptors for sketch-based image retrieval using a triplet loss convolutional neural network
abstract
We present an efficient representation for sketch based image retrieval (SBIR) derived from a triplet loss convolutional neural network (CNN). We treat SBIR as a cross-domain modelling problem, in which a depiction invariant embedding of sketch and photo data is learned by regression over a siamese CNN architecture with half-shared weights and modified triplet loss function. Uniquely, we demonstrate the ability of our learned image descriptor to generalise beyond the categories of object present in our training data, forming a basis for general cross-category SBIR. We explore appropriate strategies for training, and for deriving a compact image descriptor from the learned representation suitable for indexing data on resource constrained e. g. mobile devices. We show the learned descriptors to outperform state of the art SBIR on the defacto standard Flickr15k dataset using a significantly more compact (56 bits per image, i. e. ≈ 105KB total) search index than previous methods. Datasets and models are available from the CVSSP datasets server at www.cvssp.org.
Tu Bui, Leo Sampaio Ferraz Ribeiro, Moacir Ponti, John P. Collomosse
Comput. Vis. Image Underst.1
2015 Font finder: Visual recognition of typeface in printed documents
abstract
We describe a novel algorithm for visually identifying the font used in a scanned printed document. Our algorithm requires no pre-recognition of characters in the string (i. e. optical character recognition). Gradient orientation features are collected local the character boundaries, and quantized into a hierarchical Bag of Visual Words representation. Following stop-word analysis, classification via logistic regression (LR) of the codebooked features yields per-character probabilities which are combined across the string to decide the posterior for each font. We achieve 93.4% accuracy over a 1000 font database of scanned printed text comprising Latin characters.
Tu Bui, John P. Collomosse
ICIP1