EDBT 2026 Demo / reviewers in the wild / expert
Ewa Kijak
dblp:09/6343
· DBLP profile ↗
45ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0003-0232-1097ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 26 · 3 first-author · 7 since 2021Artificial intelligence and machine learning · 22 · 9 since 2021Databases, data management, data science and information retrieval · 3 · 1 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Forensic Cost of Watermark RemovalabstractCurrent watermark removal methods are evaluated on two axes: attack success rate and perceptual quality. We show that this is insufficient. While state-of-the-art attacks successfully degrade the watermark signal without visible distortion, they leave distinct statistical artifacts that betray the removal attempt. We name this overlooked axis Watermark Removal Detection (WRD) and demonstrate that a modern classifier trained on these artifacts achieves state-of-the-art detection rates at 10− 3 FPR across every removal method tested. No existing attack accounts for this forensic leakage. We benchmark leading watermarking schemes against standard removal pipelines under the extended evaluation triple of attack success, perceptual quality, and forensic detectability, and find that no current method balances all three. Our results establish forensic stealthiness as a necessary requirement for watermark removal. Gautier Evennou, Ewa Kijak |
IH&MMSec | 2 |
| 2026 | Text-Aided Domain Adaptation for CLIP-like models and application to challenging domain shifts
Louis Hémadou, Héléna Vorobieva, Ewa Kijak, Frédéric Jurie |
Comput. Vis. Image Underst. | 3 |
| 2025 | Adapting Without Seeing: Text-Aided Domain Adaptation for Adapting CLIP-like Models to Novel DomainsabstractThis paper addresses the challenge of adapting large vision models, such as CLIP, to domain shifts in image classification tasks. While these models, pre-trained on vast datasets like LAION 2B, offer powerful visual representations, they may struggle when applied to domains significantly different from their training data, such as industrial applications. We introduce TADA, a Text-Aided Domain Adaptation method that adapts the visual representations of these models to new domains without requiring target domain images. TADA leverages verbal descriptions of the domain shift to capture the differences between the pre-training and target domains. Our method integrates seamlessly with fine-tuning strategies, including prompt learning methods. We demonstrate TADA’s effectiveness in improving the performance of large vision models on domain-shifted data, achieving state-of-the-art results on benchmarks like DomainNet. Louis Hémadou, Héléne Vorobieva, Ewa Kijak, Frédéric Jurie |
ICASSP | 3 |
| 2025 | Reframing Image Difference Captioning with BLIP2IDC and Synthetic AugmentationabstractThe rise of the generative models quality during the past years enabled the generation of edited variations of images at an important scale. To counter the harmful effects of such technology, the Image Difference Captioning (IDC) task aims to describe the differences between two images. While this task is successfully handled for simple 3D rendered images, it struggles on real-world images. The reason is twofold: the training data-scarcity, and the difficulty to capture fine-grained differences between complex images. To address those issues, we propose in this paper a simple yet effective framework to both adapt existing image captioning models to the IDC task and augment IDC datasets. We introduce BLIP2IDC, an adaptation of BLIP2 to the IDC task at low computational cost, and show it outperforms two-streams approaches by a significant margin on real-world IDC datasets. We also propose to use synthetic augmentation to improve the performance of IDC models in an agnostic fashion. We show that our synthetic augmentation strategy provides high quality data, leading to a challenging new dataset well-suited for IDC named Syned11The code, weights and dataset are available at https://github.com/gautierevn/BLIP2IDC. Gautier Evennou, Antoine Chaffin, Vivien Chappelier, Ewa Kijak |
WACV | 4 |
| 2024 | Distinctive Image Captioning: Leveraging Ground Truth Captions in Clip Guided Reinforcement LearningabstractTraining image captioning models using teacher forcing results in very generic samples, whereas more distinctive captions can be very useful in retrieval applications or to produce alternative texts describing images for accessibility. Reinforcement Learning (RL) allows to use cross-modal retrieval similarity score between the generated caption and the input image as reward to guide the training, leading to more distinctive captions. Recent studies show that pre-trained cross-modal retrieval models can be used to provide this reward, completely eliminating the need for reference captions. However, we argue in this paper that Ground Truth (GT) captions can still be useful in this RL framework. We propose a new image captioning model training strategy that makes use of GT captions in different ways. Firstly, they can be used to train a simple MLP discriminator that serves as a regularization to prevent reward hacking and ensures the fluency of generated captions, resulting in a textual GAN setup extended for multimodal inputs. Secondly, they can serve as strong baselines when added to the pool of captions used to compute the proposed contrastive reward to reduce the variance of gradient estimate. Thirdly, they can serve as additional trajectories in the RL strategy, resulting in a teacher forcing loss weighted by the similarity of the GT to the image. This objective acts as an additional learning signal grounded to the distribution of the GT captions. Experiments on MS COCO demonstrate the interest of the proposed training strategy to produce distinctive captions while maintaining high writing quality. Antoine Chaffin, Ewa Kijak, Vincent Claveau |
ICIP | 2 |
| 2023 | Graph-Based Feature Learning from Image Markers
Isabela Borlido Barcelos, Leonardo de Melo Joao, Zenilton Kleber Gonçalves do Patrocínio Jr., Ewa Kijak, Alexandre X. Falcão, Silvio Jamil Ferzoli Guimarães |
CIARP | 4 |
| 2023 | Embedding Space Interpolation Beyond Mini-Batch, Beyond Pairs and Beyond ExamplesabstractMixup refers to interpolation-based data augmentation, originally motivated as a way to go beyond empirical risk minimization (ERM). Its extensions mostly focus on the definition of interpolation and the space (input or feature) where it takes place, while the augmentation process itself is less studied. In most methods, the number of generated examples is limited to the mini-batch size and the number of examples being interpolated is limited to two (pairs), in the input space.
We make progress in this direction by introducing MultiMix, which generates an arbitrarily large number of interpolated examples beyond the mini-batch size and interpolates the entire mini-batch in the embedding space. Effectively, we sample on the entire convex hull of the mini-batch rather than along linear segments between pairs of examples.
On sequence data, we further extend to Dense MultiMix. We densely interpolate features and target labels at each spatial location and also apply the loss densely. To mitigate the lack of dense labels, we inherit labels from examples and weight interpolation factors by attention as a measure of confidence.
Overall, we increase the number of loss terms per mini-batch by orders of magnitude at little additional cost. This is only possible because of interpolating in the embedding space. We empirically show that our solutions yield significant improvement over state-of-the-art mixup methods on four different benchmarks, despite interpolation being only linear. By analyzing the embedding space, we show that the classes are more tightly clustered and uniformly spread over the embedding space, thereby explaining the improved behavior. Shashanka Venkataramanan, Ewa Kijak, Laurent Amsaleg, Yannis Avrithis |
NeurIPS | 2 |
| 2023 | HDR-LFNet: Inverse tone mapping using fusion network
Mathieu Chambe, Ewa Kijak, Zoltán Miklós 0001, Olivier Le Meur, Rémi Cozot, Kadi Bouatouch |
Comput. Graph. | 2 |
| 2023 | Graph-based image gradients aggregated with random forests
Raquel Almeida 0001, Ewa Kijak, Simon Malinowski, Zenilton Kleber Gonçalves do Patrocínio Jr., Arnaldo de Albuquerque Araújo, Silvio Jamil Ferzoli Guimarães |
Pattern Recognit. Lett. | 2 |
| 2022 | AlignMixup: Improving Representations By Interpolating Aligned FeaturesabstractMixup is a powerful data augmentation method that in-terpolates between two or more examples in the input or feature space and between the corresponding target labels. However, how to best interpolate images is not well defined. Recent mixup methods overlay or cut-and-paste two or more objects into one image, which needs care in selecting regions. Mixup has also been connected to autoencoders, because often autoencoders generate an image that continuously deforms into another. However, such images are typically of low quality. In this work, we revisit mixup from the deformation perspective and introduce AligtiMixup, where we geometrically align two images in the feature space. The correspondences allow us to interpolate between two sets of features, while keeping the locations of one set. Interestingly, this retains mostly the geometry or pose of one image and the appearance or texture of the other. We also show that an autoencoder can still improve representation learning under mixup, without the classifier ever seeing decoded images. AlignMixup outperforms state-of-the-art mixup methods on five different benchmarks. Code available at https://github.com/shashankvkt/AlignMixup_CVPR22.git Shashanka Venkataramanan, Ewa Kijak, Laurent Amsaleg, Yannis Avrithis |
CVPR | 2 |
| 2022 | It Takes Two to Tango: Mixup for Deep Metric Learning
Shashanka Venkataramanan, Bill Psomas, Ewa Kijak, Laurent Amsaleg, Konstantinos Karantzalos, Yannis Avrithis |
ICLR | 3 |
| 2022 | Generative Cooperative Networks for Natural Language GenerationabstractGenerative Adversarial Networks (GANs) have known a tremendous success for many continuous generation tasks, especially in the field of image generation. However, for discrete outputs such as language, optimizing GANs remains an open problem with many instabilities, as no gradient can be properly back-propagated from the discriminator output to the generator parameters. An alternative is to learn the generator network via reinforcement learning, using the discriminator signal as a reward, but such a technique suffers from moving rewards and vanishing gradient problems. Finally, it often falls short compared to direct maximum-likelihood approaches. In this paper, we introduce Generative Cooperative Networks, in which the discriminator architecture is cooperatively used along with the generation policy to output samples of realistic texts for the task at hand. We give theoretical guarantees of convergence for our approach, and study various efficient decoding schemes to empirically achieve state-of-the-art results in two main NLG tasks. Sylvain Lamprier, Thomas Scialom, Antoine Chaffin, Vincent Claveau, Ewa Kijak, Jacopo Staiano, Benjamin Piwowarski |
ICML | 5 |
| 2022 | Generating Artificial Texts as Substitution or Complement of Training DataabstractThe quality of artificially generated texts has considerably improved with the advent of transformers. The question of using these models to generate learning data for supervised learning tasks naturally arises, especially when the original language resource cannot be distributed, or when it is small. In this article, this question is explored under 3 aspects: (i) are artificial data an efficient complement? (ii) can they replace the original data when those are not available or cannot be distributed for confidentiality reasons? (iii) can they improve the explainability of classifiers? Different experiments are carried out on classification tasks - namely sentiment analysis on product reviews and Fake News detection - using artificially generated data by fine-tuned GPT-2 models. The results show that such artificial data can be used in a certain extend but require pre-processing to significantly improve performance. We also show that bag-of-words approaches benefit the most from such data augmentation. Vincent Claveau, Antoine Chaffin, Ewa Kijak |
LREC | 3 |
| 2022 | PPL-MCTS: Constrained Textual Generation Through Discriminator-Guided MCTS DecodingabstractLarge language models (LM) based on Transformers allow to generate plausible long texts. In this paper, we explore how this generation can be further controlled at decoding time to satisfy certain constraints (e.g. being non-toxic, conveying certain emotions, using a specific writing style, etc.) without fine-tuning the LM.Precisely, we formalize constrained generation as a tree exploration process guided by a discriminator that indicates how well the associated sequence respects the constraint. This approach, in addition to being easier and cheaper to train than fine-tuning the LM, allows to apply the constraint more finely and dynamically. We propose several original methods to search this generation tree, notably the Monte Carlo Tree Search (MCTS) which provides theoretical guarantees on the search efficiency, but also simpler methods based on re-ranking a pool of diverse sequences using the discriminator scores. These methods are evaluated, with automatic and human-based metrics, on two types of constraints and languages: review polarity and emotion control in French and English. We show that discriminator-guided MCTS decoding achieves state-of-the-art results without having to tune the language model, in both tasks and languages. We also demonstrate that other proposed decoding methods based on re-ranking can be really effective when diversity among the generated propositions is encouraged. Antoine Chaffin, Vincent Claveau, Ewa Kijak |
NAACL-HLT | 3 |
| 2022 | Which Discriminator for Cooperative Text Generation?abstractLanguage models generate texts by successively predicting probability distributions for next tokens given past ones. A growing field of interest tries to leverage external information in the decoding process so that the generated texts have desired properties, such as being more natural, non toxic, faithful, or having a specific writing style. A solution is to use a classifier at each generation step, resulting in a cooperative environment where the classifier guides the decoding of the language model distribution towards relevant texts for the task at hand. In this paper, we examine three families of (transformer-based) discriminators for this specific task of cooperative decoding: bidirectional, left-to-right and generative ones. We evaluate the pros and cons of these different types of discriminators for cooperative generation, exploring respective accuracy on classification tasks along with their impact on the resulting sample quality and computational performances. We also provide the code of a batched implementation of the powerful cooperative decoding strategy used for our experiments, the Monte Carlo Tree Search, working with each discriminator for Natural Language Generation. Antoine Chaffin, Thomas Scialom, Sylvain Lamprier, Jacopo Staiano, Benjamin Piwowarski, Ewa Kijak, Vincent Claveau |
SIGIR | 6 |
| 2021 | Detecting Human-Object Interaction with Mixed SupervisionabstractHuman object interaction (HOI) detection is an important task in image understanding and reasoning. It is in a form of HOI triplet 〈human, verb, object〉, requiring bounding boxes for human and object, and action between them for the task completion. In other words, this task requires strong supervision for training that is however hard to procure. A natural solution to overcome this is to pursue weakly-supervised learning, where we only know the presence of certain HOI triplets in images but their exact location is unknown. Most weakly-supervised learning methods do not make provision for leveraging data with strong supervision, when they are available; and indeed a naive combination of this two paradigms in HOI detection fails to make contributions to each other. In this regard we propose a mixed-supervised HOI detection pipeline: thanks to a specific design of momentum-independent learning that learns seamlessly across these two types of supervision. Moreover, in light of the annotation insufficiency in mixed supervision, we introduce an HOI element swapping technique to synthesize diverse and hard negatives across images and improve the robustness of the model. Our method is evaluated on the challenging HICO-DET dataset. It performs close to or even better than many fully-supervised methods by using a mixed amount of strong and weak annotations; furthermore, it outperforms representative state of the art weakly- and fully-supervised methods under the same supervision. Suresh Kirthi Kumaraswamy, Miaojing Shi, Ewa Kijak |
WACV | 3 |
| 2019 | Combining convolutional side-outputs for road image segmentationabstractImage segmentation consists in creating partitions within an image into meaningful areas and objects. It can be used in scene understanding and recognition, in fields like biology, medicine, robotics, satellite imaging, amongst others. In this work we take advantage of the learned model in a deep architecture, by extracting side-outputs at different layers of the network for the task of image segmentation. We study the impact of the amount of side-outputs and evaluate strategies to combine them. A post-processing filtering based on mathematical morphology idempotent functions is also used in order to remove some undesirable noises. Experiments were performed on the publicly available KITTI Road Dataset for image segmentation. Our comparison shows that the use of multiples side outputs can increase the overall performance of the network, making it easier to train and more stable when compared with a single output in the end of the network. Also, for a small number of training epochs (500), we achieved a competitive performance when compared to the best algorithm in KITTI Evaluation Server. Felipe A. L. Reis, Raquel Almeida 0001, Ewa Kijak, Simon Malinowski, Silvio Jamil Ferzoli Guimarães, Zenilton Kleber Gonçalves do Patrocínio Jr. |
IJCNN | 3 |
| 2018 | Active Learning to Assist Annotation of Aerial Images in Environmental SurveysabstractNowadays, remote sensing technologies greatly ease environmental assessment using aerial images. Such data are most often analyzed by a manual operator, leading to costly and non scalable solutions. In the fields of both machine learning and image processing, many algorithms have been developed to fasten and automate this complex task. Their main common assumption is the need to have prior ground truth available. However, for field experts or engineers, manually labeling the objects requires a time-consuming and tedious process. Restating the labeling issue as a binary classification one, we propose a method to assist the costly annotation task by introducing an active learning process, considering a query-by-group strategy. Assuming that a comprehensive context may be required to assist the annotator with the labeling task of a single instance, the labels of all the instances of an image are indeed queried. A score based on instances distribution is defined to rank the images for annotation and an appropriate retraining step is derived to simultaneously reduce the interaction cost and improve the classifier performances at each iteration. A numerical study on real images is conducted to assess the algorithm performances. It highlights promising results regarding the classification rate along with the chosen re-training strategy and the number of interactions with the user. Mathieu Laroze, Romain Dambreville, Chloé Friguet, Ewa Kijak, Sébastien Lefèvre |
CBMI | 4 |
| 2018 | Context-Aware Forgery Localization in Social-Media Images: A Feature-Based Approach EvaluationabstractIn this paper, we study context-aware methods to localize tamperings in images from social media. The problem is defined as a comparison between image pairs: an near-duplicate image retrieved from the network and a tampered version. We propose a method based on local features matching, followed by a kernel density estimation, that we compare to recent similar approaches. The proposed approaches are evaluated on two dedicated datasets containing a variety of representative tamperings in images from social media, with difficult examples. Context-aware methods are proven to be better than blind image forensics approach. However, the evaluation allows to analyze the strengths and weaknesses of the contextual-based methods on realistic datasets. Cédric Maigrot, Ewa Kijak, Vincent Claveau |
ICIP | 2 |
| 2018 | Fusion-Based Multimodal Detection of Hoaxes in Social NetworksabstractSocial networks make it possible to share information rapidly and massively. Yet, one of their major drawback comes from the absence of verification of the piece of information, especially with viral messages. This is the issue addressed by the participants of the Verification Multimedia Use task of Mediaeval 2016. They used several approaches and clues from different modalities (text, image, social information). In this paper, we explore the interest of combining and merging these approaches in order to evaluate the predictive power of each modality and to make the most of their potential complementarity. Cédric Maigrot, Vincent Claveau, Ewa Kijak |
WI | 3 |
| 2017 | Strategies to Select Examples for Active Learning with Conditional Random Fields
Vincent Claveau, Ewa Kijak |
CICLing (1) | 2 |
| 2017 | Unsupervised Part Learning for Visual RecognitionabstractPart-based image classification aims at representing categories by small sets of learned discriminative parts, upon which an image representation is built. Considered as a promising avenue a decade ago, this direction has been neglected since the advent of deep neural networks. In this context, this paper brings two contributions: first, this work proceeds one step further compared to recent part-based models (PBM), focusing on how to learn parts without using any labeled data. Instead of learning a set of parts per class, as generally performed in the PBM literature, the proposed approach both constructs a partition of a given set of images into visually similar groups, and subsequently learns a set of discriminative parts per group in a fully unsupervised fashion. This strategy opens the door to the use of PBM in new applications where labeled data are typically not available, such as instance-based image retrieval. Second, this paper shows that despite the recent success of end-to-end models, explicit part learning can still boost classification performance. We experimentally show that our learned parts can help building efficient image representations, which outperform state-of-the art Deep Convolutional Neural Networks (DCNN) on both classification and retrieval tasks. Ronan Sicre, Yannis Avrithis, Ewa Kijak, Frédéric Jurie |
CVPR | 3 |
| 2016 | Direct vs. indirect evaluation of distributional thesauriabstractWith the success of word embedding methods in various Natural Language Processing tasks, all the field of distributional semantics has experienced a renewed interest. Beside the famous word2vec, recent studies have presented efficient techniques to build distributional thesaurus; in particular, Claveau et al. (2014) have already shown that Information Retrieval (IR) tools and concepts can be successfully used to build a thesaurus. In this paper, we address the problem of the evaluation of such thesauri or embedding models and compare their results. Through several experiments and by evaluating directly the results with reference lexicons, we show that the recent IR-based distributional models outperform state-of-the-art systems such as word2vec. Following the work of Claveau and Kijak (2016), we use IR as an applicative framework to indirectly evaluate the generated thesaurus. Here again, this task-based evaluation validates the IR approach used to build the thesaurus. Moreover, it allows us to compare these results with those from the direct evaluation framework used in the literature. The observed differences bring these evaluation habits into question. Vincent Claveau, Ewa Kijak |
COLING | 2 |
| 2016 | Distributional Thesauri for Information Retrieval and vice versa
Vincent Claveau, Ewa Kijak |
LREC | 2 |
| 2016 | Partial least squares for face hashing
Cassio E. dos Santos, Ewa Kijak, Guillaume Gravier, William Robson Schwartz |
Neurocomputing | 2 |
| 2015 | Beyond MSER: Maximally Stable Regions using Tree of ShapesabstractThis article explores the application of a tree-based feature extraction algorithm for the widely-used MSER features, and proposes a Tree of Shapes based detector of Maximally Stable Regions. Changing an underlying component tree in the algorithm allows considering alternative properties and pixel orderings for extracting the Maximally Stable Regions. Differences introduced to the region structure with changing the underlying tree are discussed, as well as the spatial organization of the detected regions imposed by using a self-dual image representation for detection. Performance evaluation is carried out on a standard matching benchmark in terms of repeat ability and matching score under different image transformations, as well as in a large scale image retrieval setup, measuring Mean Average Precision. The proposed de-scriptor is compared to the standard MSER implementation as well as a tree-based MSER implementation, achieving competitive results in the matching setup and outperforming the baseline MSER in the retrieval experiments. Petra Bosilj, Ewa Kijak, Sébastien Lefèvre |
BMVC | 2 |
| 2015 | Short local descriptors from 2D connected pattern spectraabstractWe propose a local region descriptor based on connected pattern spectra, and combined with normalized central moments. The descriptors are calculated for MSER regions of the image, and their performance compared against SIFT. The MSER regions were chosen because they can be efficiently selected by constructing a max-tree, a structure used to calculate both descriptors and region moments. Experiments on the UCID database show an improvement over SIFT in two out of five experimental setups, and comparable performance in two other experiments. The new descriptors are only half the size of SIFT, resulting in 4 times faster query times when performing exact search on descriptor index built from 262 images. Petra Bosilj, Ewa Kijak, Michael H. F. Wilkinson, Sébastien Lefèvre |
ICIP | 2 |
| 2014 | Improving distributional thesauri by exploring the graph of neighbors
Vincent Claveau, Ewa Kijak, Olivier Ferret |
COLING | 2 |
| 2014 | Generating and using probabilistic morphological resources for the biomedical domain
Vincent Claveau, Ewa Kijak |
LREC | 2 |
| 2012 | Face recognition using Co-occurrence Histograms of Oriented GradientsabstractRecently, Histogram of Oriented Gradient (HOG) is applied in face recognition. In this paper, we apply Co-occurrence of Oriented Gradient (CoHOG), which is an extension of HOG, on the face recognition problem. Some weighted functions for magnitude gradient are tested. We also proposed a weighted approach for CoHOG, where a weight value is set for each subregion of face image. Numerical experiments performed on Yale and ORL datasets show that 1) CoHOG has recognition accuracy higher than HOG; 2) using gradient magnitude in CoHOG improves recognition results; and 3) weighted CoHOG approach improves accuracy recognition rate. The recognition results using CoHOG are competitive with some of the state of the art methods. This proves the effectiveness of CoHOG descriptor for face recognition. Thanh-Toan Do, Ewa Kijak |
ICASSP | 2 |
| 2012 | Enlarging hacker's toolbox: Deluding image recognition by attacking keypoint orientationsabstractContent-Based Image Retrieval Systems (CBIRS) used in forensics related contexts require very good image recognition capabilities. Whereas the robustness criterion has been extensively covered by Computer Vision or Multimedia literature, none of these communities explored the security of CBIRS. Recently, preliminary studies have shown real systems can be deluded by applying transformations to images that are very specific to the SIFT local description scheme commonly used for recognition. The work presented in this paper adds one strategy for attacking images, and somehow enlarges the box of tools hackers can use for deluding systems. This paper shows how the orientation of keypoints can be tweaked, which in turn lowers matches since this deeply changes the final SIFT feature vectors. The method learns what visual patch should be applied to change the orientation of keypoints thanks to an SVM-based process. Experiments with a database made of 100,000 real world images confirms the effectiveness of this keypoint-orientation attacking scheme. Thanh-Toan Do, Ewa Kijak, Laurent Amsaleg, Teddy Furon |
ICASSP | 2 |
| 2012 | Security-oriented picture-in-picture visual modificationsabstractThe performance of Content-Based Image Retrieval Systems (CBIRS) is typically evaluated via benchmarking their capacity to match images despite various generic distortions such as crops, rescalings or Picture in Picture (PiP) attacks, which are the most challenging. Distortions are made in a very generic manner, by applying a set of transformations that are completely independent from the systems later performing recognition tasks. Recently, studies have shown that exploiting the finest details of the various techniques used in a CBIRS offers the opportunity to create distortions that dramatically reduce the recognition performance. Such a security perspective is taken in this paper. Instead of creating generic PiP distortions, it proposes a creation scheme able to delude the recognition capabilities of a CBIRS that is representative of state of the art techniques as it relies on SIFT, high-dimensional k-nearest neighbors searches and geometrical robustification steps. Experiments using 100,000 real-world images confirm the effectiveness of these security-oriented PiP visual modifications. Thanh-Toan Do, Ewa Kijak, Laurent Amsaleg, Teddy Furon |
ICMR | 2 |
| 2011 | Image compression using the Iteration-Tuned and Aligned DictionaryabstractWe present a new, block-based image codec based on sparse representations using a learned, structured dictionary called the Iteration Tuned and Aligned Dictionary (ITAD). The question of selecting the number of atoms used in the representation of each image block is addressed with a new, global (image-wide), rate-distortion-based sparsity selection criterion. We show experimentally that our codec outperforms JPEG2000 in both quantitative evaluations (by 0.9 dB to 4 dB) and qualitative evaluations. Joaquin Zepeda, Christine Guillemot, Ewa Kijak |
ICASSP | 3 |
| 2010 | Approximate nearest neighbors using sparse representationsabstractA new method is introduced that makes use of sparse image representations to search for approximate nearest neighbors (ANN) under the normalized inner-product distance. The approach relies on the construction of a new sparse vector designed to approximate the normalized inner-product between underlying signal vectors. The resulting ANN search algorithm shows significant improvement compared to querying with the original sparse vectors. The system makes use of a proposed transform that succeeds in uniformly distributing the input dataset on the unit sphere while preserving relative angular distances. Joaquin Zepeda, Ewa Kijak, Christine Guillemot |
ICASSP | 2 |
| 2010 | Understanding the security and robustness of SIFTabstractMany content-based retrieval systems (CBIRS) describe images using the SIFT local features because of their very robust recognition capabilities. While SIFT features proved to cope with a wide spectrum of general purpose image distortions, its security has not fully been assessed yet. In one of their scenario, Hsu et al. in [2] show that very specific anti-SIFT attacks can jeopardize the keypoint detection. These attacks can delude systems using SIFT targeting application such as image authentication and (pirated) copy detection. Thanh-Toan Do, Ewa Kijak, Teddy Furon, Laurent Amsaleg |
ACM Multimedia | 2 |
| 2010 | Challenging the security of Content-Based Image Retrieval systemsabstractContent-Based Image Retrieval (CBIR) has been recently used as a filtering mechanism against the piracy of multimedia contents. Many publications in the last few years have proposed very robust schemes where pirated contents are detected despite severe modifications. As none of these systems have addressed the piracy problem from a security perspective, it is time to check whether they are secure: Can pirates mount violent attacks against CBIR systems by carefully studying the technology they use? This paper is an initial analysis of the security flaws of the typical technology blocks used in state-of-the-art CBIR systems. It is so far too early to draw any definitive conclusion about their inherent security, but it motivates and encourages further studies on this topic. Thanh-Toan Do, Ewa Kijak, Teddy Furon, Laurent Amsaleg |
MMSP | 2 |
| 2010 | The Iteration-Tuned Dictionary for sparse representationsabstractWe introduce a new dictionary structure for sparse representations better adapted to pursuit algorithms used in practical scenarios. The new structure, which we call an Iteration-Tuned Dictionary (ITD), consists of a set of dictionaries each associated to a single iteration index of a pursuit algorithm. In this work we first adapt pursuit decompositions to the case of ITD structures and then introduce a training algorithm used to construct ITDs. The training algorithm consists of applying a K-means to the (i -1)-th residuals of the training set to thus produce the i-th dictionary of the ITD structure. In the results section we compare our algorithm against the state-of-the-art dictionary training scheme and show that our method produces sparse representations yielding better signal approximations for the same sparsity level. Joaquin Zepeda, Christine Guillemot, Ewa Kijak |
MMSP | 3 |
| 2010 | Scalable object-based video retrieval in HD video databases
Claire Morand, Jenny Benois-Pineau, Jean-Philippe Domenger, Joaquin Zepeda, Ewa Kijak, Christine Guillemot |
Signal Process. Image Commun. | 5 |
| 2009 | SIFT-based local image description using sparse representationsabstractThis paper addresses the problem of efficient SIFT-based image description and searches in large databases within the framework of local querying. A descriptor called the bag-of-features has been introduced previously which first vector quantizes SIFT descriptors and then aggregates the set of resulting codeword indices (so-called visual words) into a histogram of occurrence of the different visual words in the image. The aim is to make the image search complexity tractable by transforming the set of local image descriptor vectors into a single sparse vector as sparsity particularly permits efficient inner product calculations. However, aggregating local descriptors into a single histogram decreases the discerning power of the system when performing local queries. In this paper, we propose a new approach that aims to enjoy the complexity benefits of sparsity while at the same time retaining the local quality of the input descriptor vectors. This is accomplished by searching for a sparse approximation of the input SIFT descriptors. The sparse approximation yields a sparse vector per local SIFT descriptor, and helps preserving local description properties by using each sparse-transformed descriptor independently in a voting system to retrieve indexed images. Our system is shown experimentally to perform better than histogram based systems under query locality, albeit at an increased complexity. Joaquin Zepeda, Ewa Kijak, Christine Guillemot |
MMSP | 2 |
| 2007 | The ARGOS campaign: Evaluation of video analysis and indexing tools
Philippe Joly, Jenny Benois-Pineau, Ewa Kijak, Georges Quénot |
Signal Process. Image Commun. | 3 |
| 2006 | Summarizing Video Information Using Self-Organizing MapsabstractFacing a huge amount of multimedia information available today, it becomes inevitably necessary to develop efficient methods for accessing, searching, structuring, and representing it. Multimedia retrieval systems especially in the case of video should support users in all of these tasks. Therefore, specialized systems that focus on each of these aspects have been developed. However, an open research perspective issue is the development of a retrieval tool that integrates all user interactions in a single interface. In this paper, we present a system that focuses on the summarization of one single video. We consider the structuring and visualization components including reasonable user interactions to have the most significant influence on the systems usability. Our prototype is based on growing self-organizing maps. The emphasis is on intuitive hierarchical interaction for content visualization exploiting the features of the maps and integrating additional potentially useful information. Thomas Bärecke, Ewa Kijak, Andreas Nürnberger, Marcin Detyniecki |
FUZZ-IEEE | 2 |
| 2006 | Using Self-Organizing Maps to Support Video Navigation
Thomas Bärecke, Ewa Kijak, Andreas Nürnberger, Marcin Detyniecki |
ICANN (1) | 2 |
| 2006 | Audiovisual integration for tennis broadcast structuring
Ewa Kijak, Guillaume Gravier, Lionel Oisel, Patrick Gros |
Multim. Tools Appl. | 1 |
| 2003 | Hierarchical structure analysis of sport videos using HMMSabstractThis paper focuses on the use of hidden Markov models (HMMs) for structure analysis of sport videos. The video structure parsing relies on the analysis of the temporal interleaving of video shots, with respect to a priori information about video content and editing rules. The basic temporal unit is the video shot and visual features are used to characterize its type of view. Our approach is validated in the particular domain of tennis videos. As a result, typical tennis scenes are identified. In addition, each shot is assigned to a level in the hierarchy described in terms of point, game and set. Ewa Kijak, Lionel Oisel, Patrick Gros |
ICIP (2) | 1 |
| 2003 | HMM based structuring of tennis videos using visual and audio cuesabstractThis paper focuses on the use of hidden Markov models (HMMs) for structure analysis of videos, and demonstrates how they can be efficiently applied to merge audio and visual cues. Our approach is validated in the particular domain of tennis videos. The basic temporal unit is the video shot. Visual features describe the audio events within a video shot. The video structure parsing relies on the analysis of the temporal interleaving of video shots, with respect to prior information about tennis content and editing rules. As a result, typical tennis scenes are identified. In addition, each shot is assigned to a level in the hierarchy described in terms of point, game and set. Ewa Kijak, Guillaume Gravier, Patrick Gros, Lionel Oisel, Frédéric Bimbot |
ICME | 1 |