EDBT 2026 Demo / reviewers in the wild / expert
Sabine Süsstrunk
dblp:s/SSusstrunk
· DBLP profile ↗
108ranked-venue papers
2as first author
32since 2021 · last 2025
0000-0002-0441-6068ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 91 · 2 first-author · 26 since 2021Artificial intelligence and machine learning · 42 · 21 since 2021Systems, architecture and hardware · 2Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | FDS: Frequency-Aware Denoising Score for Text-Guided Latent Diffusion Image EditingabstractText-guided image editing using Text-to-Image (T2I) models often fails to yield satisfactory results, frequently introducing unintended modifications, such as the loss of local detail and color changes. In this paper, we analyze these failure cases and attribute them to the indiscriminate optimization across all frequency bands, even though only specific frequencies may require adjustment. To address this, we introduce a simple yet effective approach that enables the selective optimization of specific frequency bands within localized spatial regions for precise edits. Our method leverages wavelets to decompose images into different spatial resolutions across multiple frequency bands, enabling precise modifications at various levels of detail. To extend the applicability of our approach, we provide a comparative analysis of different frequency-domain techniques. Additionally, we extend our method to 3D texture editing by performing frequency decomposition on the triplane representation, enabling frequency-aware adjustments for 3D textures. Quantitative evaluations and user studies demonstrate the effectiveness of our method in producing high-quality and precise edits. Further details are available on our project website: https://ivrl.github.io/fds-webpage/ Yufan Ren, Zicong Jiang, Tong Zhang 0023, Søren Forchhammer, Sabine Süsstrunk |
CVPR | 5 |
| 2024 | InNeRF360: Text-Guided 3D-Consistent Object Inpainting on 360° Neural Radiance FieldsabstractWe propose InNeRF360, an automatic system that accu-rately removes text-specified objectsfrom 360° Neural Radi-ance Fields (NeRF). The challenge is to effectively remove objects while inpainting perceptually consistent content for the missing regions, which is particularly demanding for existing NeRF models due to their implicit volumetric rep-resentation. Moreover, unbounded scenes are more prone to floater artifacts in the inpainted region than frontal-facing scenes, as the change of object appearance and background across views is more sensitive to inaccurate segmentations and inconsistent inpainting. With a trained NeRF and a text description, our method efficiently removes specified ob-jects and inpaints visually consistent content without arti-facts. We apply depth-space warping to enforce consistency across multiview text-encoded segmentations, and then re-fine the inpainted NeRF model using perceptual priors and 3D diffusion-based geometric priors to ensure visual plau-sibility. Through extensive experiments in segmentation and inpainting on 360° and frontal-facing NeRFs, we show that our approach is effective and enhances NeRF's ed-itability. Project page: https://ivr1.github.io/InNeRF360/. Dongqing Wang, Tong Zhang 0023, Alaa Abboud, Sabine Süsstrunk |
CVPR | 4 |
| 2024 | Mitigating Object Dependencies: Improving Point Cloud Self-Supervised Learning Through Object ExchangeabstractIn the realm of point cloud scene understanding, particularly in indoor scenes, objects are arranged following human habits, resulting in objects of certain semantics being closely positioned and displaying notable inter-object cor-relations. This can create a tendency for neural networks to exploit these strong dependencies, bypassing the individ-ual object patterns. To address this challenge, we introduce a novel self-supervised learning (SSL) strategy. Our approach leverages both object patterns and contextual cues to produce robust features. It begins with the formulation of an object-exchanging strategy, where pairs of objects with comparable sizes are exchanged across different scenes, effectively disentangling the strong contextual dependencies. Subsequently, we introduce a context-aware feature learning strategy, which encodes object patterns without relying on their specific context by aggregating object features across various scenes. Our extensive experiments demonstrate the superiority of our method over existing SSL techniques, further showing its better robustness to environmental changes. Moreover, we showcase the applicability of our approach by transferring pre-trained models to diverse point cloud datasets.11Our code is available at https:/lgithub.com/YanhaoWu/OESSL Yanhao Wu, Tong Zhang 0023, Wei Ke 0003, Congpei Qiu, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 5 |
| 2024 | Data Augmentation via Latent Diffusion for Saliency Prediction
Bahar Aydemir, Deblina Bhattacharjee, Tong Zhang 0023, Mathieu Salzmann, Sabine Süsstrunk |
ECCV (78) | 5 |
| 2024 | Message from the ChairsabstractWelcome to the 16th IEEE International Conference on Computational Photography (ICCP 2024), taking place at the Ecole Polytéchnique Fédérale (EPFL) in Lausanne, Switzerland! This is the first time that ICCP takes place in Continental Europe, specifically on the shore of beautiful Lake Geneva. This year's two-and-a-half day conference features 18 accepted papers, 3 keynote talks, 8 invited talks, and 50 posters and/or demos. Sabine Süsstrunk, Ashok Veeraraghavan, Roarke Horstmeyer, Wolfgang Heidrich |
ICCP | 1 |
| 2024 | Mind Your Augmentation: The Key to Decoupling Dense Self-Supervised LearningabstractDense Self-Supervised Learning (SSL) creates positive pairs by building positive paired regions or points, thereby aiming to preserve local features, for example of individual objects. However, existing approaches tend to couple objects by leaking information from the neighboring contextual regions when the pairs have a limited overlap. In this paper, we first quantitatively identify and confirm the existence of such a coupling phenomenon. We then address it by developing a remarkably simple yet highly effective solution comprising a novel augmentation method, Region Collaborative Cutout (RCC), and a corresponding decoupling branch. Importantly, our design is versatile and can be seamlessly integrated into existing SSL frameworks, whether based on Convolutional Neural Networks (CNNs) or Vision Transformers (ViTs). We conduct extensive experiments, incorporating our solution into two CNN-based and two ViT-based methods, with results confirming the effectiveness of our approach. Moreover, we provide empirical evidence that our method significantly contributes to the disentanglement of feature representations among objects, both in quantitative and qualitative terms. Congpei Qiu, Tong Zhang 0023, Yanhao Wu, Wei Ke 0003, Mathieu Salzmann, Sabine Süsstrunk |
ICLR | 6 |
| 2024 | AdanCA: Neural Cellular Automata As Adaptors For More Robust Vision TransformerabstractVision Transformers (ViTs) demonstrate remarkable performance in image classification through visual-token interaction learning, particularly when equipped with local information via region attention or convolutions. Although such architectures improve the feature aggregation from different granularities, they often fail to contribute to the robustness of the networks. Neural Cellular Automata (NCA) enables the modeling of global visual-token representations through local interactions, with its training strategies and architecture design conferring strong generalization ability and robustness against noisy input. In this paper, we propose Adaptor Neural Cellular Automata (AdaNCA) for Vision Transformers that uses NCA as plug-and-play adaptors between ViT layers, thus enhancing ViT's performance and robustness against adversarial samples as well as out-of-distribution inputs. To overcome the large computational overhead of standard NCAs, we propose Dynamic Interaction for more efficient interaction learning. Using our analysis of AdaNCA placement and robustness improvement, we also develop an algorithm for identifying the most effective insertion points for AdaNCA. With less than a 3% increase in parameters, AdaNCA contributes to more than 10% absolute improvement in accuracy under adversarial attacks on the ImageNet1K benchmark. Moreover, we demonstrate with extensive evaluations across eight robustness benchmarks and four ViT architectures that AdaNCA, as a plug-and-play module, consistently improves the robustness of ViTs. Yitao Xu 0002, Tong Zhang 0023, Sabine Süsstrunk |
NeurIPS | 3 |
| 2024 | Exploiting the Signal-Leak Bias in Diffusion ModelsabstractThere is a bias in the inference pipeline of most diffusion models. This bias arises from a signal leak whose distribution deviates from the noise distribution, creating a discrepancy between training and inference processes. We demonstrate that this signal-leak bias is particularly significant when models are tuned to a specific style, causing sub-optimal style matching. Recent research tries to avoid the signal leakage during training. We instead show how we can exploit this signal-leak bias in existing diffusion models to allow more control over the generated images. This enables us to generate images with more varied brightness, and images that better match a desired style or color. By modeling the distribution of the signal leak in the spatial frequency and pixel domains, and including a signal leak in the initial latent, we generate images that better match expected results without any additional training. Martin Nicolas Everaert, Athanasios Fitsios, Marco Bocchio, Sami Arpa, Sabine Süsstrunk, Radhakrishna Achanta |
WACV | 5 |
| 2024 | On the Impact of Hard Adversarial Instances on Overfitting in Adversarial TrainingabstractAdversarial training is a popular method to robustify models against adversarial attacks. However, it exhibits much more severe overfitting than training on clean inputs. In this work, we investigate this phenomenon from the perspective of training instances, i.e., training input-target pairs. Based on a quantitative metric measuring the relative difficulty of an instance in the training set, we analyze the model's behavior on training instances of different difficulty levels. This lets us demonstrate that the decay in generalization performance of adversarial training is a result of fitting hard adversarial instances. We theoretically verify our observations for both linear and general nonlinear models, proving that models trained on hard instances have worse generalization performance than ones trained on easy instances, and that this generalization gap increases with the size of the adversarial budget. Finally, we investigate solutions to mitigate adversarial overfitting in several scenarios, including fast adversarial training and fine-tuning a pretrained model with additional data. Our results demonstrate that using training data adaptively improves the model's robustness. Chen Liu 0027, Zhichao Huang 0002, Mathieu Salzmann, Tong Zhang 0023, Sabine Süsstrunk |
J. Mach. Learn. Res. | 5 |
| 2024 | Mesh Neural Cellular AutomataabstractTexture modeling and synthesis are essential for enhancing the realism of virtual environments. Methods that directly synthesize textures in 3D offer distinct advantages to the UV-mapping-based methods as they can create seamless textures and align more closely with the ways textures form in nature. We propose Mesh Neural Cellular Automata (MeshNCA), a method that directly synthesizes dynamic textures on 3D meshes without requiring any UV maps. MeshNCA is a generalized type of cellular automata that can operate on a set of cells arranged on non-grid structures such as the vertices of a 3D mesh. MeshNCA accommodates multi-modal supervision and can be trained using different targets such as images, text prompts, and motion vector fields. Only trained on an Icosphere mesh, MeshNCA shows remarkable test-time generalization and can synthesize textures on unseen meshes in real time. We conduct qualitative and quantitative comparisons to demonstrate that MeshNCA outperforms other 3D texture synthesis methods in terms of generalization and producing high-quality textures. Moreover, we introduce a way of grafting trained MeshNCA instances, enabling interpolation between textures. MeshNCA allows several user interactions including texture density/orientation controls, grafting/regenerate brushes, and motion speed/direction controls. Finally, we implement the forward pass of our MeshNCA model using the WebGL shading language and showcase our trained models in an online interactive demo, which is accessible on personal computers and smartphones and is available at https://meshnca.github.io/. Ehsan Pajouheshgar, Yitao Xu 0002, Alexander Mordvintsev, Eyvind Niklasson, Tong Zhang 0023, Sabine Süsstrunk |
ACM Trans. Graph. | 6 |
| 2023 | VETIM: Expanding the Vocabulary of Text-to-Image Models only with Text
Martin Nicolas Everaert, Marco Bocchio, Sami Arpa, Sabine Süsstrunk, Radhakrishna Achanta |
BMVC | 4 |
| 2023 | TempSAL - Uncovering Temporal Information for Deep Saliency PredictionabstractDeep saliency prediction algorithms complement the object recognition features, they typically rely on additional information such as scene context, semantic relationships, gaze direction, and object dissimilarity. However, none of these models consider the temporal nature of gaze shifts during image observation. We introduce a novel saliency prediction model that learns to output saliency maps in sequential time intervals by exploiting human temporal attention patterns. Our approach locally modulates the saliency predictions by combining the learned temporal maps. Our experiments show that our method outperforms the state-of-the-art models, including a multi-duration saliency model, on the SALICON benchmark and CodeCharts1k dataset. Our code is publicly available on GitHub11https://ivrl.github.io/Tempsal/. Bahar Aydemir, Ludo Hoffstetter, Tong Zhang 0023, Mathieu Salzmann, Sabine Süsstrunk |
CVPR | 5 |
| 2023 | DyNCA: Real-Time Dynamic Texture Synthesis Using Neural Cellular AutomataabstractCurrent Dynamic Texture Synthesis (DyTS) models can synthesize realistic videos. However, they require a slow iterative optimization process to synthesize a single fixed-size short video, and they do not offer any post-training control over the synthesis process. We propose Dynamic Neural Cellular Automata (DyNCA), a framework for real-time and controllable dynamic texture synthesis. Our method is built upon the recently introduced NCA models and can synthesize infinitely long and arbitrary-sized realistic video textures in real time. We quantitatively and qualitatively evaluate our model and show that our synthesized videos appear more realistic than the existing results. We improve the SOTA DyTS performance by 2 ~ 4 orders of magnitude. Moreover, our model offers several real-time video controls including motion speed, motion direction, and an editing brush tool. We exhibit our trained models in an online interactive demo that runs on local hardware and is accessible on personal computers and smartphones. Ehsan Pajouheshgar, Yitao Xu 0002, Tong Zhang 0023, Sabine Süsstrunk |
CVPR | 4 |
| 2023 | VolRecon: Volume Rendering of Signed Ray Distance Functions for Generalizable Multi-View ReconstructionabstractThe success of the Neural Radiance Fields (NeRF) in novel view synthesis has inspired researchers to propose neural implicit scene reconstruction. However, most existing neural implicit reconstruction methods optimize perscene parameters and therefore lack generalizability to new scenes. We introduce VolRecon, a novel generalizable implicit reconstruction method with Signed Ray Distance Function (SRDF). To reconstruct the scene with fine details and little noise, VolRecon combines projection features aggregated from multi-view features, and volume features interpolated from a coarse global feature volume. Using a ray transformer, we compute SRDF values of sampled points on a ray and then render color and depth. On DTU dataset, VolRecon outperforms SparseNeuS by about 30% in sparse view reconstruction and achieves comparable accuracy as MVSNet in full view reconstruction. Furthermore, our approach exhibits good generalization performance on the large-scale ETH3D benchmark. Code is available at https://github.com/IVRL/VolRecon/. Yufan Ren, Fangjinhua Wang, Tong Zhang 0023, Marc Pollefeys, Sabine Süsstrunk |
CVPR | 5 |
| 2023 | Spatiotemporal Self-Supervised Learning for Point Clouds in the WildabstractSelf-supervised learning (SSL) has the potential to benefit many applications, particularly those where manually annotating data is cumbersome. One such situation is the semantic segmentation of point clouds. In this context, existing methods employ contrastive learning strategies and define positive pairs by performing various augmentation of point clusters in a single frame. As such, these methods do not exploit the temporal nature of LiDAR data. In this paper, we introduce an SSL strategy that leverages positive pairs in both the spatial and temporal domain. To this end, we design (i) a point-to-cluster learning strategy that aggregates spatial information to distinguish objects; and (ii) a cluster-to-cluster learning strategy based on unsupervised object tracking that exploits temporal correspondences. We demonstrate the benefits of our approach via extensive experiments performed by self-supervised training on two large-scale LiDAR datasets and transferring the resulting models to other point cloud segmentation benchmarks. Our results evidence that our method outperforms the state-of-the-art point cloud SSL methods.11Our code and pretrained models will be found at https://github.com/YanhaoWu/STSSL. Correspondence to Ke Wei. Yanhao Wu, Tong Zhang 0023, Wei Ke 0003, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 4 |
| 2023 | Vision Transformer Adapters for Generalizable Multitask LearningabstractWe introduce the first multitasking vision transformer adapters that learn generalizable task affinities which can be applied to novel tasks and domains. Integrated into an off-the-shelf vision transformer backbone, our adapters can simultaneously solve multiple dense vision tasks in a parameter-efficient manner, unlike existing multitasking transformers that are parametrically expensive. In contrast to concurrent methods, we do not require retraining or fine-tuning whenever a new task or domain is added. We introduce a task-adapted attention mechanism within our adapter framework that combines gradient-based task similarities with attention-based ones. The learned task affinities generalize to the following settings: zero-shot task transfer, unsupervised domain adaptation, and generalization without fine-tuning to novel domains. We demonstrate that our approach outperforms not only the existing convolutional neural network-based multitasking methods but also the vision transformer-based ones. Our project page is at https://ivrl.github.io/VTAGML. Deblina Bhattacharjee, Sabine Süsstrunk, Mathieu Salzmann |
ICCV | 2 |
| 2023 | Diffusion in StyleabstractWe present Diffusion in Style, a simple method to adapt Stable Diffusion to any desired style, using only a small set of target images. It is based on the key observation that the style of the images generated by Stable Diffusion is tied to the initial latent tensor. Not adapting this initial latent tensor to the style makes fine-tuning slow, expensive, and impractical, especially when only a few target style images are available. In contrast, fine-tuning is much easier if this initial latent tensor is also adapted. Our Diffusion in Style is orders of magnitude more sample-efficient and faster. It also generates more pleasing images than existing approaches, as shown qualitatively and with quantitative comparisons. Martin Nicolas Everaert, Marco Bocchio, Sami Arpa, Sabine Süsstrunk, Radhakrishna Achanta |
ICCV | 4 |
| 2023 | NEMTO: Neural Environment Matting for Novel View and Relighting Synthesis of Transparent ObjectsabstractWe propose NEMTO, the first end-to-end neural rendering pipeline to model 3D transparent objects with complex geometry and unknown indices of refraction. Commonly used appearance modeling such as the Disney BSDF model cannot accurately address this challenging problem due to the complex light paths bending through refractions and the strong dependency of surface appearance on illumination. With 2D images of the transparent object as input, our method is capable of high-quality novel view and relighting synthesis. We leverage implicit Signed Distance Functions (SDF) to model the object geometry and propose a refraction-aware ray bending network to model the effects of light refraction within the object. Our ray bending network is more tolerant to geometric inaccuracies than traditional physically-based methods for rendering transparent objects. We provide extensive evaluations on both synthetic and real-world datasets to demonstrate our high-quality synthesis and the applicability of our method. Dongqing Wang, Tong Zhang 0023, Sabine Süsstrunk |
ICCV | 3 |
| 2023 | Towards Stable and Efficient Adversarial Training against l1 Bounded Adversarial AttacksabstractWe address the problem of stably and efficiently training a deep neural network robust to adversarial perturbations bounded by an $l_1$ norm. We demonstrate that achieving robustness against $l_1$-bounded perturbations is more challenging than in the $l_2$ or $l_\infty$ cases, because adversarial training against $l_1$-bounded perturbations is more likely to suffer from catastrophic overfitting and yield training instabilities. Our analysis links these issues to the coordinate descent strategy used in existing methods. We address this by introducing Fast-EG-$l_1$, an efficient adversarial training algorithm based on Euclidean geometry and free of coordinate descent. Fast-EG-$l_1$ comes with no additional memory costs and no extra hyper-parameters to tune. Our experimental results on various datasets demonstrate that Fast-EG-$l_1$ yields the best and most stable robustness against $l_1$-bounded adversarial attacks among the methods of comparable computational complexity. Code and the checkpoints are available at https://github.com/IVRL/FastAdvL. Yulun Jiang, Chen Liu 0027, Zhichao Huang 0002, Mathieu Salzmann, Sabine Süsstrunk |
ICML | 5 |
| 2023 | Fast Adversarial Training With Adaptive Step SizeabstractWhile adversarial training and its variants have shown to be the most effective algorithms to defend against adversarial attacks, their extremely slow training process makes it hard to scale to large datasets like ImageNet. The key idea of recent works to accelerate adversarial training is to substitute multi-step attacks (e.g., PGD) with single-step attacks (e.g., FGSM). However, these single-step methods suffer from catastrophic overfitting, where the accuracy against PGD attack suddenly drops to nearly 0% during training, and the network totally loses its robustness. In this work, we study the phenomenon from the perspective of training instances. We show that catastrophic overfitting is instance-dependent, and fitting instances with larger input gradient norm is more likely to cause catastrophic overfitting. Based on our findings, we propose a simple but effective method, Adversarial Training with Adaptive Step size (ATAS). ATAS learns an instance-wise adaptive step size that is inversely proportional to its gradient norm. Our theoretical analysis shows that ATAS converges faster than the commonly adopted non-adaptive counterparts. Empirically, ATAS consistently mitigates catastrophic overfitting and achieves higher robust accuracy on CIFAR10, CIFAR100, and ImageNet when evaluated on various adversarial budgets. Our code is released at https://github.com/HuangZhiChao95/ATAS. Zhichao Huang 0002, Yanbo Fan, Chen Liu 0027, Yong Zhang 0034, Mathieu Salzmann, Sabine Süsstrunk, Jue Wang 0001 |
IEEE Trans. Image Process. | 7 |
| 2023 | Training Provably Robust Models by Polyhedral Envelope RegularizationabstractTraining certifiable neural networks enables us to obtain models with robustness guarantees against adversarial attacks. In this work, we introduce a framework to obtain a provable adversarial-free region in the neighborhood of the input data by a polyhedral envelope, which yields more fine-grained certified robustness than existing methods. We further introduce polyhedral envelope regularization (PER) to encourage larger adversarial-free regions and thus improve the provable robustness of the models. We demonstrate the flexibility and effectiveness of our framework on standard benchmarks; it applies to networks of different architectures and with general activation functions. Compared with state of the art, PER has negligible computational overhead; it achieves better robustness guarantees and accuracy on the clean data in various settings. Chen Liu 0027, Mathieu Salzmann, Sabine Süsstrunk |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2022 | MuIT: An End-to-End Multitask Learning TransformerabstractWe propose an end-to-end Multitask Learning Transformer framework, named MulT, to simultaneously learn multiple high-level vision tasks, including depth estimation, semantic segmentation, reshading, surface normal estimation, 2D keypoint detection, and edge detection. Based on the Swin transformer model, our framework encodes the input image into a shared representation and makes predictions for each vision task using task-specific transformer-based decoder heads. At the heart of our approach is a shared attention mechanism modeling the dependencies across the tasks. We evaluate our model on several multitask benchmarks, showing that our MulT framework outperforms both the state-of-the art multitask convolutional neural network models and all the respective single task transformer models. Our experiments further highlight the benefits of sharing attention across all the tasks, and demonstrate that our MulT model is robust and generalizes well to new domains. Our project website is at https://ivrl.github.io/MulT/. Deblina Bhattacharjee, Tong Zhang 0023, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 3 |
| 2022 | Leverage Your Local and Global Representations: A New Self-Supervised Learning StrategyabstractSelf-supervised learning (SSL) methods aim to learn view-invariant representations by maximizing the similar-ity between the features extracted from different crops of the same image regardless of cropping size and content. In essence, this strategy ignores the fact that two crops may truly contain different image information, e.g., background and small objects, and thus tends to restrain the diversity of the learned representations. In this work, we address this issue by introducing a new self-supervised learning strat-egy, LoGo, that explicitly reasons about Local and Global crops. To achieve view invariance, LoGo encourages similarity between global crops from the same image, as well as between a global and a local crop. However, to correctly encode the fact that the content of smaller crops may differ entirely, LoGo promotes two local crops to have dissimi-lar representations, while being close to global crops. Our LoGo strategy can easily be applied to existing SSL meth-ods. Our extensive experiments on a variety of datasets and using different self-supervised learning frameworks vali-date its superiority over existing approaches. Noticeably, we achieve better results than supervised models on trans-fer learning when using only 1/10 of the data.11Our code and pretrained models can be found at https://github.com/ztt1024/LoGo-SSL. Tong Zhang 0023, Congpei Qiu, Wei Ke 0003, Sabine Süsstrunk, Mathieu Salzmann |
CVPR | 4 |
| 2022 | RC-MVSNet: Unsupervised Multi-View Stereo with Neural Rendering
Di Chang, Aljaz Bozic, Tong Zhang 0023, Qingsong Yan, Ying-Cong Chen, Sabine Süsstrunk, Matthias Nießner |
ECCV (31) | 6 |
| 2022 | Natural-Looking Adversarial Examples from Freehand SketchesabstractDeep neural networks (DNNs) have achieved great success in image classification and recognition compared to previous methods. However, recent works have reported that DNNs are very vulnerable to adversarial examples that are intentionally generated to mislead the predictions of the DNNs. Here, we present a novel freehand sketch-based natural-looking adversarial example generator that we call SketchAdv. To generate a natural-looking adversarial example from a sketch, we force the encoded edge information (i.e., the visual attributes) to be close to the latent random vector fed to the edge generator and adversarial example generator. This preserves the spatial consistency of the adversarial example generated from the random vector with the edge information. In addition, by employing a sketch-edge encoder with a novel sketch-edge matching loss, we reduce the gap between edges and sketches. We evaluate the proposed method on several dominant classes of SketchyCOCO, the benchmark dataset for sketch to image translation. Our experiments show that our SketchAdv produces visually plausible adversarial examples while remaining competitive with other adversarial attack methods. Hak Gu Kim, Davide Nanni, Sabine Süsstrunk |
ICASSP | 3 |
| 2022 | Optimizing Latent Space Directions for Gan-Based Local Image EditingabstractGenerative Adversarial Network (GAN) based localized image editing can suffer from ambiguity between semantic at-tributes. We thus present a novel objective function to evaluate the locality of an image edit. By introducing the super-vision from a pre-trained segmentation network and optimizing the objective function, our framework, called Locally Effective Latent Space Direction (LELSD), is applicable to any dataset and GAN architecture. Our method is also computationally fast and exhibits a high extent of disentanglement, which allows users to interactively perform a sequence of edits on an image. Our experiments on both GAN-generated and real images qualitatively demonstrate the high quality and advantages of our method. Ehsan Pajouheshgar, Tong Zhang 0023, Sabine Süsstrunk |
ICASSP | 3 |
| 2022 | Robust Binary Models by Pruning Randomly-initialized NetworksabstractRobustness to adversarial attacks was shown to require a larger model capacity, and thus a larger memory footprint. In this paper, we introduce an approach to obtain robust yet compact models by pruning randomly-initialized binary networks. Unlike adversarial training, which learns the model parameters, we initialize the model parameters as either +1 or −1, keep them fixed, and find a subnetwork structure that is robust to attacks. Our method confirms the Strong Lottery Ticket Hypothesis in the presence of adversarial attacks, and extends this to binary networks. Furthermore, it yields more compact networks with competitive performance than existing works by 1) adaptively pruning different network layers; 2) exploiting an effective binary initialization scheme; 3) incorporating a last batch normalization layer to improve training stability. Our experiments demonstrate that our approach not only always outperforms the state-of-the-art robust binary networks, but also can achieve accuracy better than full-precision ones on some datasets. Finally, we show the structured patterns of our pruned binary networks. Chen Liu 0027, Sabine Süsstrunk, Mathieu Salzmann |
NeurIPS | 3 |
| 2022 | Estimating Image Depth in the Comics DomainabstractEstimating the depth of comics images is challenging as such images a) are monocular; b) lack ground-truth depth annotations; c) differ across different artistic styles; d) are sparse and noisy. We thus, use an off-the-shelf unsupervised image to image translation method to translate the comics images to natural ones and then use an attention-guided monocular depth estimator to predict their depth. This lets us leverage the depth annotations of existing natural images to train the depth estimator. Furthermore, our model learns to distinguish between text and images in the comics panels to reduce text-based artefacts in the depth estimates. Our method consistently outperforms the existing state-of-the-art approaches across all metrics on both the DCM and eBDtheque images. Finally, we introduce a dataset to evaluate depth prediction on comics. Deblina Bhattacharjee, Martin Nicolas Everaert, Mathieu Salzmann, Sabine Süsstrunk |
WACV | 4 |
| 2022 | PoGaIN: Poisson-Gaussian Image Noise Modeling From Paired SamplesabstractImage noise can often be accurately fitted to a Poisson-Gaussian distribution. However, estimating the distribution parameters from a noisy image only is a challenging task. Here, we study the case when paired noisy and noise-free samples are accessible. No method is currently available to exploit the noise-free information, which may help to achieve more accurate estimations. To fill this gap, we derive a novel, cumulant-based, approach for Poisson-Gaussian noise modeling from paired image samples. We show its improved performance over different baselines, with special emphasis on MSE, effect of outliers, image dependence, and bias. We additionally derive the log-likelihood function for further insights and discuss real-world applicability. Nicolas Bähler, Majed El Helou, Étienne Objois, Kaan Okumus, Sabine Süsstrunk |
IEEE Signal Process. Lett. | 5 |
| 2022 | BIGPrior: Toward Decoupling Learned Prior Hallucination and Data Fidelity in Image RestorationabstractClassic image-restoration algorithms use a variety of priors, either implicitly or explicitly. Their priors are hand-designed and their corresponding weights are heuristically assigned. Hence, deep learning methods often produce superior image restoration quality. Deep networks are, however, capable of inducing strong and hardly predictable hallucinations. Networks implicitly learn to be jointly faithful to the observed data while learning an image prior; and the separation of original data and hallucinated data downstream is then not possible. This limits their wide-spread adoption in image restoration. Furthermore, it is often the hallucinated part that is victim to degradation-model overfitting. We present an approach with decoupled network-prior based hallucination and data fidelity terms. We refer to our framework as the Bayesian Integration of a Generative Prior (BIGPrior). Our method is rooted in a Bayesian framework and tightly connected to classic restoration methods. In fact, it can be viewed as a generalization of a large family of classic restoration algorithms. We use network inversion to extract image prior information from a generative network. We show that, on image colorization, inpainting and denoising, our framework consistently improves the inversion results. Our method, though partly reliant on the quality of the generative network inversion, is competitive with state-of-the-art supervised and task-specific restoration methods. It also provides an additional metric that sets forth the degree of prior reliance per pixel relative to data fidelity. Majed El Helou, Sabine Süsstrunk |
IEEE Trans. Image Process. | 2 |
| 2021 | Deep Gaussian Denoiser Epistemic Uncertainty and Decoupled Dual-Attention FusionabstractFollowing the performance breakthrough of denoising networks, improvements have come chiefly through novel architecture designs and increased depth. While novel denoising networks were designed for real images coming from different distributions, or for specific applications, comparatively small improvement was achieved on Gaussian denoising. The denoising solutions suffer from epistemic uncertainty that can limit further advancements. This uncertainty is traditionally mitigated through different ensemble approaches. However, such ensembles are prohibitively costly with deep networks, which are already large in size.Our work focuses on pushing the performance limits of state-of-the-art methods on Gaussian denoising. We propose a model-agnostic approach for reducing epistemic uncertainty while using only a single pretrained network. We achieve this by tapping into the epistemic uncertainty through augmented and frequency-manipulated images to obtain denoised images with varying error. We propose an ensemble method with two decoupled attention paths, over the pixel domain and over that of our different manipulations, to learn the final fusion. Our results significantly improve over the state-of-the-art baselines and across varying noise levels. Xiaoqi Ma, Majed El Helou, Sabine Süsstrunk |
ICIP | 4 |
| 2021 | Fidelity Estimation Improves Noisy-Image Classification With Pretrained NetworksabstractImage classification has significantly improved using deep learning. This is mainly due to convolutional neural networks (CNNs) that are capable of learning rich feature extractors from large datasets. However, most deep learning classification methods are trained on clean images and are not robust when handling noisy ones, even if a restoration preprocessing step is applied. While novel methods address this problem, they rely on modified feature extractors and thus necessitate retraining. We instead propose a method that can be applied on a $pretrained$ classifier. Our method exploits a fidelity map estimate that is fused into the internal representations of the feature extractor, thereby guiding the attention of the network and making it more robust to noisy data. We improve the noisy-image classification (NIC) results by significantly large margins, especially at high noise levels, and come close to the fully retrained approaches. Furthermore, as proof of concept, we show that when using our oracle fidelity map we even outperform the fully retrained methods, whether trained on noisy or restored images. Deblina Bhattacharjee, Majed El Helou, Sabine Süsstrunk |
IEEE Signal Process. Lett. | 4 |
| 2020 | Editing in Style: Uncovering the Local Semantics of GANsabstractWhile the quality of GAN image synthesis has improved tremendously in recent years, our ability to control and condition the output is still limited. Focusing on StyleGAN, we introduce a simple and effective method for making local, semantically-aware edits to a target output image. This is accomplished by borrowing elements from a source image, also a GAN output, via a novel manipulation of style vectors. Our method requires neither supervision from an external model, nor involves complex spatial morphing operations. Instead, it relies on the emergent disentanglement of semantic objects that is learned by StyleGAN during its training. Semantic editing is demonstrated on GANs producing human faces, indoor scenes, cats, and cars. We measure the locality and photorealism of the edits produced by our method, and find that it accomplishes both. Edo Collins, Raja Bala, Bob Price, Sabine Süsstrunk |
CVPR | 4 |
| 2020 | Stochastic Frequency Masking to Improve Super-Resolution and Denoising Networks
Majed El Helou, Ruofan Zhou, Sabine Süsstrunk |
ECCV (16) | 3 |
| 2020 | Volumetric Transformer Networks
Seungryong Kim, Sabine Süsstrunk, Mathieu Salzmann |
ECCV (28) | 2 |
| 2020 | AL2: Progressive Activation Loss for Learning General Representations in Classification Neural NetworksabstractThe large capacity of neural networks enables them to learn complex functions. To avoid overfitting, networks however require a lot of training data that can be expensive and time-consuming to collect. A common practical approach to attenuate overfitting is the use of network regularization techniques.We propose a novel regularization method that progressively penalizes the magnitude of activations during training. The combined activation signals produced by all neurons in a given layer form the representation of the input image in that feature space. We propose to regularize this representation in the last feature layer before classification layers. Our method's effect on generalization is analyzed with label randomization tests and cumulative ablations. Experimental results show the advantages of our approach in comparison with commonly-used regularizers on standard benchmark datasets. Majed El Helou, Frederike Dümbgen, Sabine Süsstrunk |
ICASSP | 3 |
| 2020 | Divergence-Based Adaptive Extreme Video CompletionabstractExtreme image or video completion, where, for instance, we only retain 1% of pixels in random locations, allows for very cheap sampling in terms of the required pre-processing. The consequence is, however, a reconstruction that is challenging for humans and inpainting algorithms alike. We propose an extension of a state-of-the-art extreme image completion algorithm to extreme video completion. We analyze a color-motion estimation approach based on color KL-divergence that is suitable for extremely sparse scenarios. Our algorithm leverages the estimate to adapt between its spatial and temporal filtering when reconstructing the sparse randomly-sampled video. We validate our results on 50 publicly-available videos using reconstruction PSNR and mean opinion scores. Majed El Helou, Ruofan Zhou, Frank Schmutz, Fabrice Guibert, Sabine Süsstrunk |
ICASSP | 5 |
| 2020 | On the Loss Landscape of Adversarial Training: Identifying Challenges and How to Overcome ThemabstractWe analyze the influence of adversarial training on the loss landscape of machine learning models. To this end, we first provide analytical studies of the properties of adversarial loss functions under different adversarial budgets. We then demonstrate that the adversarial loss landscape is less favorable to optimization, due to increased curvature and more scattered gradients. Our conclusions are validated by numerical analyses, which show that training under large adversarial budgets impede the escape from suboptimal random initialization, cause non-vanishing gradients and make the models' minima found sharper. Based on these observations, we show that a periodic adversarial scheduling (PAS) strategy can effectively overcome these challenges, yielding better results than vanilla adversarial training while being much less sensitive to the choice of learning rate. Chen Liu 0027, Mathieu Salzmann, Tao Lin 0004, Ryota Tomioka, Sabine Süsstrunk |
NeurIPS | 5 |
| 2020 | Evaluating salient object detection in natural images with multiple objects having multi-level saliencyabstractSalient object detection is evaluated using binary ground truth (GT) with the labels being salient object class and background. In this study, the authors corroborate based on three subjective experiments on a novel image dataset that objects in natural images are inherently perceived to have varying levels of importance. The authors' dataset, named SalMoN (saliency in multi‐object natural images), has 588 images containing multiple objects. The subjective experiments performed record spontaneous attention and perception through eye fixation duration, point clicking and rectangle drawing. As object saliency in a multi‐object image is inherently multi‐level, they propose that salient object detection must be evaluated for the capability to detect all multi‐level salient objects apart from the salient object class detection capability. For this purpose, they generate multi‐level maps as GT corresponding to all the dataset images using the results of the subjective experiments, with the labels being multi‐level salient objects and background. They then propose the use of mean absolute error, Kendall's rank correlation and average area under precision–recall curve to evaluate existing salient object detection methods on their multi‐level saliency GT dataset. Approaches that represent saliency detection on images as local‐global hierarchical processing of a graph perform well in their dataset. Gökhan Yildirim 0001, Debashis Sen, Mohan Kankanhalli, Sabine Süsstrunk |
IET Image Process. | 4 |
| 2020 | Blind Universal Bayesian Image Denoising With Gaussian Noise Level LearningabstractBlind and universal image denoising consists of using a unique model that denoises images with any level of noise. It is especially practical as noise levels do not need to be known when the model is developed or at test time. We propose a theoretically-grounded blind and universal deep learning image denoiser for additive Gaussian noise removal. Our network is based on an optimal denoising solution, which we call fusion denoising. It is derived theoretically with a Gaussian image prior assumption. Synthetic experiments show our network's generalization strength to unseen additive noise levels. We also adapt the fusion denoising network architecture for image denoising on real images. Our approach improves real-world grayscale additive image denoising PSNR results for training noise levels and further on noise levels not seen during training. It also improves state-of-the-art color image denoising performance on every single noise level, by an average of 0.1dB, whether trained on or not. Majed El Helou, Sabine Süsstrunk |
IEEE Trans. Image Process. | 2 |
| 2019 | Zero-Learning Fast Medical Image Fusion
Fayez Lahoud, Sabine Süsstrunk |
FUSION | 2 |
| 2019 | Kernel Modeling Super-Resolution on Real Low-Resolution ImagesabstractDeep convolutional neural networks (CNNs), trained on corresponding pairs of high- and low-resolution images, achieve state-of-the-art performance in single-image super-resolution and surpass previous signal-processing based approaches. However, their performance is limited when applied to real photographs. The reason lies in their training data: low-resolution (LR) images are obtained by bicubic interpolation of the corresponding high-resolution (HR) images. The applied convolution kernel significantly differs from real-world camera-blur. Consequently, while current CNNs well super-resolve bicubic-downsampled LR images, they often fail on camera-captured LR images. To improve generalization and robustness of deep super-resolution CNNs on real photographs, we present a kernel modeling super-resolution network (KMSR) that incorporates blur-kernel modeling in the training. Our proposed KMSR consists of two stages: we first build a pool of realistic blur-kernels with a generative adversarial network (GAN) and then we train a super-resolution network with HR and corresponding LR images constructed with the generated kernels. Our extensive experimental validations demonstrate the effectiveness of our single-image super-resolution approach on photographs with unknown blur-kernels. Ruofan Zhou, Sabine Süsstrunk |
ICCV | 2 |
| 2019 | Deep Semantic Segmentation Using Nir as Extra Physical InformationabstractDeep neural networks for semantic segmentation are most often trained with RGB color images, which encode the radiation visible to the human eyes. In this paper, we study if additional physical scene information, specifically Near-Infrared (NIR) images, improve the performance of neural networks. NIR information can be captured with conventional silicon-based cameras and provide complementary information to visible images regarding object boundaries and materials. In addition, extending the networks' input from a three to a four channel layer is trivial with respect to changes to the architecture and additional parameters. We perform experiments on several state-of-the-art neural networks trained both on RGB alone and on RGB plus NIR and show that the additional image channel consistently improves semantic segmentation accuracy over conventional RGB input even for powerful architectures. Siavash Arjomand Bigdeli, Sabine Süsstrunk |
ICIP | 2 |
| 2019 | Deep Feature Factorization for Content-Based Image Retrieval and LocalizationabstractState of the art content-based image retrieval algorithms owe their excellent performance to the rich semantics encoded in the deep activations of a convolutional neural network. The difference between these algorithms lies mostly in how activations are combined into a compact global image descriptor. In this paper, we propose to use deep feature factorization to achieve this goal. By factorizing CNN activations, we decompose an input image into semantic regions, represented by both spatial saliency heatmaps and basis vectors serving as descriptors for those regions. When combined to form a global image descriptor, our experiments show that DFF surpasses the state of the art in both image retrieval and localization of the region of interest within the set of retrieved images. Edo Collins, Sabine Süsstrunk |
ICIP | 2 |
| 2019 | Mirror, Mirror, on the Wall, Who's Got the Clearest Image of Them All? - A Tailored Approach to Single Image Reflection RemovalabstractRemoving reflection artefacts from a single image is a problem of both theoretical and practical interest, which still presents challenges because of the massively ill-posed nature of the problem. In this paper, we propose a technique based on a novel optimization problem. First, we introduce a simple user interaction scheme, which helps minimize information loss in the reflection-free regions. Second, we introduce an H2fidelity term, which preserves fine detail while enforcing the global color similarity. We show that this combination allows us to mitigate the shortcomings in structure and color preservation, which presents some of the most prominent drawbacks in the existing methods for reflection removal. We demonstrate, through numerical and visual experiments, that our method is able to outperform the state-of-the-art model-based methods and compete with recent deep-learning approaches. Daniel Heydecker, Georg Maierhofer, Angelica I. Avilés-Rivero, Qingnan Fan, Dongdong Chen 0001, Carola-Bibiane Schönlieb, Sabine Süsstrunk |
IEEE Trans. Image Process. | 7 |
| 2018 | Deep Feature Factorization for Concept Discovery
Edo Collins, Radhakrishna Achanta, Sabine Süsstrunk |
ECCV (14) | 3 |
| 2018 | Learning to see through reflectionsabstractPictures of objects behind a glass are difficult to interpret and understand due to the superposition of two real images: a reflection layer and a background layer. Separation of these two layers is challenging due to the ambiguities in assigning texture patterns and the average color in the input image to one of the two layers. In this paper, we propose a novel method to reconstruct these layers given a single input image by explicitly handling the ambiguities of the reconstruction. Our approach combines the ability of neural networks to build image priors on large image regions with an image model that accounts for the brightness ambiguity and saturation. We find that our solution generalizes to real images even in the presence of strong reflections. Extensive quantitative and qualitative experimental evaluations on both real and synthetic data show the benefits of our approach over prior work. Moreover, our proposed neural network is computationally and memory efficient. Meiguang Jin, Sabine Süsstrunk, Paolo Favaro |
ICCP | 2 |
| 2018 | Ar in VR: Simulating Infrared Augmented VisionabstractDeveloping an augmented reality (AR) system involves multiple algorithms such as image fusion, camera synchronization and calibration, and brightness control, each of them having diverse parameters. This abundance of settings, while allowing for many features, is detrimental to developers as they try to navigate between different combinations and pick the most suitable towards their application. Additionally, the temporally inconsistent nature of the real world makes it hard to build reproducible scenarios for testing and comparison. To help address these issues, we develop a virtual reality (VR) environment that allows simulating a variety of AR configurations”. We show the advantages of AR simulation in virtual reality, demonstrate an image fusion AR system and conduct an experiment to compare different fusion methods. Fayez Lahoud, Sabine Süsstrunk |
ICIP | 2 |
| 2018 | Deep Learning for Logic Optimization AlgorithmsabstractThe slowing down of Moore's law and the emergence of new technologies puts an increasing pressure on the field of EDA. There is a constant need to improve optimization algorithms. However, finding and implementing such algorithms is a difficult task, especially with the novel logic primitives and potentially unconventional requirements of emerging technologies. In this paper, we cast logic optimization as a deterministic Markov decision process (MDP). We then take advantage of recent advances in deep reinforcement learning to build a system that learns how to navigate this process. Our design has a number of desirable properties. It is autonomous because it learns automatically and does not require human intervention. It generalizes to large functions after training on small examples. Additionally, it intrinsically supports both single- and multi-output functions, without the need to handle special cases. Finally, it is generic because the same algorithm can be used to achieve different optimization objectives, e.g., size and depth. Winston Haaswijk, Edo Collins, Benoit Seguin, Mathias Soeken, Frédéric Kaplan, Sabine Süsstrunk, Giovanni De Micheli |
ISCAS | 6 |
| 2017 | Superpixels and Polygons Using Simple Non-iterative ClusteringabstractWe present an improved version of the Simple Linear Iterative Clustering (SLIC) superpixel segmentation. Unlike SLIC, our algorithm is non-iterative, enforces connectivity from the start, requires lesser memory, and is faster. Relying on the superpixel boundaries obtained using our algorithm, we also present a polygonal partitioning algorithm. We demonstrate that our superpixels as well as the polygonal partitioning are superior to the respective state-of-the-art algorithms on quantitative benchmarks. Radhakrishna Achanta, Sabine Süsstrunk |
CVPR | 2 |
| 2017 | Single Image Reflection SuppressionabstractReflections are a common artifact in images taken through glass windows. Automatically removing the reflection artifacts after the picture is taken is an ill-posed problem. Attempts to solve this problem using optimization schemes therefore rely on various prior assumptions from the physical world. Instead of removing reflections from a single image, which has met with limited success so far, we propose a novel approach to suppress reflections. It is based on a Laplacian data fidelity term and an l-zero gradient sparsity term imposed on the output. With experiments on artificial and real-world images we show that our reflection suppression method performs better than the state-of-the-art reflection removal techniques. Nikolaos Arvanitopoulos, Radhakrishna Achanta, Sabine Süsstrunk |
CVPR | 3 |
| 2017 | Webly Supervised Semantic SegmentationabstractWe propose a weakly supervised semantic segmentation algorithm that uses image tags for supervision. We apply the tags in queries to collect three sets of web images, which encode the clean foregrounds, the common backgrounds, and realistic scenes of the classes. We introduce a novel three-stage training pipeline to progressively learn semantic segmentation models. We first train and refine a class-specific shallow neural network to obtain segmentation masks for each class. The shallow neural networks of all classes are then assembled into one deep convolutional neural network for end-to-end training and testing. Experiments show that our method notably outperforms previous state-of-the-art weakly supervised semantic segmentation approaches on the PASCAL VOC 2012 segmentation benchmark. We further apply the class-specific shallow neural networks to object segmentation and obtain excellent results. Bin Jin, Maria V. Ortiz Segovia, Sabine Süsstrunk |
CVPR | 3 |
| 2017 | Simultaneous Geometric and Radiometric Calibration of a Projector-Camera PairabstractWe present a novel method that allows for simultaneous geometric and radiometric calibration of a projector-camera pair. It is simple and does not require specialized hardware. We prewarp and align a specially designed projection pattern onto a printed pattern of different colorimetric properties. After capturing the patterns in several orientations, we perform geometric calibration by estimating the corner locations of the two patterns in different color channels. We perform radiometric calibration of the projector by using the information contained inside the projected squares. We show that our method performs on par with current approaches that all require separate geometric and radiometric calibration, while being more efficient and user friendly. Marjan Shahpaski, Luis Ricardo Sapaico, Gaspard Chevassus, Sabine Süsstrunk |
CVPR | 4 |
| 2017 | Extreme image completionabstractIt is challenging to complete an image whose 99% pixels are randomly missing. We present a solution to this extreme image completion problem. As opposed to existing techniques, our solution has a computational complexity that is linear in the number of pixels of the full image and is realtime in practice. For comparable quality of reconstruction, our algorithm is thus almost 2 to 5 orders of magnitude faster than existing techniques. Radhakrishna Achanta, Nikolaos Arvanitopoulos, Sabine Süsstrunk |
ICASSP | 3 |
| 2017 | Face recognition in real-world imagesabstractFace recognition systems are designed to handle well-aligned images captured under controlled situations. However real-world images present varying orientations, expressions, and illumination conditions. Traditional face recognition algorithms perform poorly on such images. In this paper we present a method for face recognition adapted to real-world conditions that can be trained using very few training examples and is computationally efficient. Our method consists of performing a novel alignment process followed by classification using sparse representation techniques. We present our recognition rates on a difficult dataset that represents real-world faces where we significantly outperform state-of-the-art methods. Xavier Fontaine, Radhakrishna Achanta, Sabine Süsstrunk |
ICASSP | 3 |
| 2017 | Color representation in deep neural networksabstractConvolutional neural networks are top-performers on image classification tasks. Understanding how they make use of color information in images may be useful for various tasks. In this paper we analyze the representation learned by a popular CNN to detect and characterize color-related features. We confirm the existence of some object- and color-specific units, as well as the effect of layer-depth on color-sensitivity and class-invariance. Martin Engilberge, Edo Collins, Sabine Süsstrunk |
ICIP | 3 |
| 2017 | Correlation-based deblurring leveraging multispectral chromatic aberration in color and near-infrared joint acquisitionabstractJoint acquisition of color and near-infrared (NIR) images is of growing interest due to various applications that make use of the additional spectral information. An obstacle to this acquisition is the wavelength-dependent blurring caused by the chromatic aberration of optical lenses. When one of the spectral channels, for example the green channel, is in focus on the sensor plane, the images of the other channels, especially NIR, are blurred. This paper presents a study of spectral-spatial correlations between color and NIR channels and proposes a method to correct for chromatic aberrations. The algorithm we introduce leverages axial chromatic aberration to deblur the NIR image when the color image is in focus. The proposed technique improves image sharpness by 48.8% on average compared to state-of-the-art results. Moreover, our method generates an NIR image that has a larger depth-of-field compared to an NIR image originally captured in focus. Majed El Helou, Zahra Sadeghipoor, Sabine Süsstrunk |
ICIP | 3 |
| 2017 | Preserving perceptual contrast in decolorization with optimized color ordersabstractConverting a color image to a grayscale image, namely decolorization, is an important process for many real-world applications. Previous methods build contrast loss functions to minimize the contrast differences between the color images and the resultant grayscale images. In this paper, we improve upon a widely used decolorization method with two extensions. First, we relax the need for heuristics on color orders, which the baseline method relies on when computing the contrast differences. In our method, the color orders are incorporated into the loss function and are determined through optimization. Moreover, we apply a nonlinear function on the grayscale contrast to better model human perception of contrast. Both qualitative and quantitative results on the standard benchmark demonstrate the effectiveness of our two extensions. Bin Jin, Sabine Süsstrunk |
ICIP | 2 |
| 2017 | Keyword-based image color re-rendering with semantic segmentationabstractKeyword-based image color re-rendering is a convenient way to enhance the color of images. Most methods only focus on the color characteristics of the image while ignoring the semantic meaning of different regions. We propose to incorporate semantic information into the color re-rendering pipeline through semantic segmentation. Using semantic segmentation masks, we first generate more accurate correlations between keywords and color characteristics than the state-of-the-art approach. Such correlations are then adopted for rerendering the color of the input image, where the segmentation masks are used to indicate the regions for color rerendering. Qualitative comparisons show that our method generates visually better results than the state-of-the-art approach. We further validate with a psychophysical experiment, where the participants prefer the results of our method. Fayez Lahoud, Bin Jin, Maria V. Ortiz Segovia, Sabine Süsstrunk |
ICIP | 4 |
| 2016 | Image aesthetic predictors based on weighted CNNsabstractConvolutional Neural Networks (CNNs) have been widely adopted for many imaging applications. For image aesthetics prediction, state-of-the-art algorithms train CNNs on a recently-published large-scale dataset, AVA. However, the distribution of the aesthetic scores on this dataset is extremely unbalanced, which limits the prediction capability of existing methods. We overcome such limitation by using weighted CNNs. We train a regression model that improves the prediction accuracy of the aesthetic scores over state-of-the-art algorithms. In addition, we propose a novel histogram prediction model that not only predicts the aesthetic score, but also estimates the difficulty of performing aesthetics assessment for an input image. We further show an image enhancement application where we obtain an aesthetically pleasing crop of an input image using our regression model. Bin Jin, Maria V. Ortiz Segovia, Sabine Süsstrunk |
ICIP | 3 |
| 2016 | Non-Parametric Blur Map Regression for Depth of Field ExtensionabstractReal camera systems have a limited depth of field (DOF) which may cause an image to be degraded due to visible misfocus or too shallow DOF. In this paper, we present a blind deblurring pipeline able to restore such images by slightly extending their DOF and recovering sharpness in regions slightly out of focus. To address this severely ill-posed problem, our algorithm relies first on the estimation of the spatially varying defocus blur. Drawing on local frequency image features, a machine learning approach based on the recently introduced regression tree fields is used to train a model able to regress a coherent defocus blur map of the image, labeling each pixel by the scale of a defocus point spread function. A non-blind spatially varying deblurring algorithm is then used to properly extend the DOF of the image. The good performance of our algorithm is assessed both quantitatively, using realistic ground truth data obtained with a novel approach based on a plenoptic camera, and qualitatively with real images. Laurent D'Andres, Jordi Salvador, Axel Kochale, Sabine Süsstrunk |
IEEE Trans. Image Process. | 4 |
| 2015 | Image aesthetics depends on contextabstractWe investigate the influence of low-level image features for aesthetics prediction. We show that the aesthetic quality of a photography depends on its context. Image features learned from a specific image category are not necessarily the same as features learned from a generic image collection. Experiments conducted on specific image categories show that specific features obtain statistically significantly better results than generic ones. Florian Simond, Nikolaos Arvanitopoulos, Sabine Süsstrunk |
ICIP | 3 |
| 2015 | High Reliefs from 3D ScenesabstractAbstract We present a method for synthesizing high reliefs, a sculpting technique that attaches 3D objects onto a 2D surface within a limited depth range. The main challenges are the preservation of distinct scene parts by preserving depth discontinuities, the fine details of the shape, and the overall continuity of the scene. Bas relief depth compression methods such as gradient compression and depth range compression are not applicable for high relief production. Instead, our method is based on differential coordinates to bring scene elements to the relief plane while preserving depth discontinuities and surface details of the scene. We select a user‐defined number of attenuation points within the scene, attenuate these points towards the relief plane and recompute the positions of all scene elements by preserving the differential coordinates. Finally, if the desired depth range is not achieved we apply a range compression. High relief synthesis is semi‐automatic and can be controlled by user‐defined parameters to adjust the depth range, as well as the placement of the scene elements with respect to the relief plane. Sami Arpa, Sabine Süsstrunk, Roger D. Hersch |
Comput. Graph. Forum | 2 |
| 2015 | Semantic-Improved Color Imaging Applications: It Is All About ContextabstractMultimedia data with associated semantics is omnipresent in today's social online platforms in the form of keywords, user comments, and so forth. This article presents a statistical framework designed to infer knowledge in the imaging domain from the semantic domain. Note that this is the reverse direction of common computer vision applications. The framework relates keywords to image characteristics using a statistical significance test. It scales to millions of images and hundreds of thousands of keywords. We demonstrate the usefulness of the statistical framework with three color imaging applications: 1) semantic image enhancement: re-render an image in order to adapt it to its semantic context; 2) color naming: find the color triplet for a given color name; and 3) color palettes: find a palette of colors that best represents a given arbitrary semantic context and that satisfies established harmony constraints. Albrecht J. Lindner, Sabine Süsstrunk |
IEEE Trans. Multim. | 2 |
| 2014 | FASA: Fast, Accurate, and Size-Aware Salient Object Detection
Gökhan Yildirim 0001, Sabine Süsstrunk |
ACCV (3) | 2 |
| 2014 | Seam Carving for Text Line Extraction on Color and Grayscale Historical ManuscriptsabstractWe propose a novel algorithm for automatic text line extraction on color and gray scale manuscript pages without prior binarization. Our algorithm is based on seam carving to compute separating seams between text lines. Seam carving is likely to produce seams that move through gaps between neighboring lines, if no information about the text geometry is incorporated into the problem. By constraining the optimization procedure inside the region between two consecutive text lines, we can produce robust separating seams that do not cut through word and line components. Extensive experimental evaluations on diverse manuscript pages show that we improve upon the state-of-the-art for grayscale text line extraction. Nikolaos Arvanitopoulos, Sabine Süsstrunk |
ICFHR | 2 |
| 2014 | A practical method for measuring the spatial frequency response of light field camerasabstractThe spatial frequency response (SFR) is one of the most important and unbiased image quality measures of a digital camera. It evaluates to which extent a lens/sensor combination can resolve scene details. In this paper, we propose a simple and practical method to measure the SFR of microlens-based light field cameras. The particularity of such cameras resides in their ability to capture both spatial and angular information of the incoming light field thanks to an array of microlenses located in front of the sensor. Existing methods for measuring the SFR of conventional cameras are thus no longer applicable as the interaction between the main lens and the micro-lenses results in different resolving powers over the image plane that depend on the scene depths. By using a 3-dimensional target made of vertical lines printed on an inclined planar surface, we are able to measure the SFR across multiple depths in a single exposure. Our method allows SFR measurements from the raw light field itself as captured by the camera, and is thus independent of subsequent post-processing algorithms such as image reconstruction, digital refocusing or depth estimation. Our experimental results are consistent with theoretical bounds and reproducible. Damien Firmenich, Loïc Baboulaz, Sabine Süsstrunk |
ICIP | 3 |
| 2014 | Saliency Detection using regression trees on hierarchical image segmentsabstractThe currently best performing state-of-the-art saliency detection algorithms incorporate heuristic functions to evaluate saliency. They require parameter tuning, and the relationship between the parameter value and visual saliency is often not well understood. Instead of using parametric methods we follow a machine learning approach, which is parameter free, to estimate saliency. Our method learns data-driven saliency-estimation functions and exploits the contributions of visual properties on saliency. First, we over-segment the image into superpixels and iteratively connect them to form hierarchical image segments. Second, from these segments, we extract biologically-plausible visual features. Finally, we use regression trees to learn the relationship between the feature values and visual saliency. We show that our algorithm outperforms the most recent state-of-the-art methods on three public databases. Gökhan Yildirim 0001, Appu Shaji, Sabine Süsstrunk |
ICIP | 3 |
| 2014 | Automatic and Accurate Shadow Detection Using Near-Infrared InformationabstractWe present a method to automatically detect shadows in a fast and accurate manner by taking advantage of the inherent sensitivity of digital camera sensors to the near-infrared (NIR) part of the spectrum. Dark objects, which confound many shadow detection algorithms, often have much higher reflectance in the NIR. We can thus build an accurate shadow candidate map based on image pixels that are dark both in the visible and NIR representations. We further refine the shadow map by incorporating ratios of the visible to the NIR image, based on the observation that commonly encountered light sources have very distinct spectra in the NIR band. The results are validated on a new database, which contains visible/NIR images for a large variety of real-world shadow creating illuminant conditions, as well as manually labeled shadow ground truth. Both quantitative and qualitative evaluations show that our method outperforms current state-of-the-art shadow detection algorithms in terms of accuracy and computational efficiency. Dominic Rüfenacht, Clément Fredembach, Sabine Süsstrunk |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2013 | A novel compressive sensing approach to simultaneously acquire color and near-infrared images on a single sensorabstractSensors of most digital cameras are made of silicon that is inherently sensitive to both the visible and near-infrared parts of the electromagnetic spectrum. In this paper, we address the problem of color and NIR joint acquisition. We propose a framework for the joint acquisition that uses only a single silicon sensor and a slightly modified version of the Bayer color-filter array that is already mounted in most color cameras. Implementing such a design for an RGB and NIR joint acquisition system requires minor changes to the hardware of commercial color cameras. One of the important differences between this design and the conventional color camera is the post-processing applied to the captured values to reconstruct full resolution images. By using a CFA similar to Bayer, the sensor records a mixture of NIR and one color channel in each pixel. In this case, separating NIR and color channels in different pixels is equivalent to solving an under-determined system of linear equations. To solve this problem, we propose a novel algorithm that uses the tools developed in the field of compressive sensing. Our method results in high-quality RGB and NIR images (the average PSNR of more than 30 dB for the reconstructed images) and shows a promising path towards RGB and NIR cameras. Zahra Sadeghipoor, Yue M. Lu, Sabine Süsstrunk |
ICASSP | 3 |
| 2013 | Near-infrared guided color image dehazingabstractNear-infrared (NIR) light has stronger penetration capability than visible light due to its long wavelengths and is thus less scattered by particles in the air. This makes it desirable for image dehazing to unveil details of distant objects in landscape photographs. In this paper, we propose an improved image dehazing scheme using a pair of color and NIR images, which effectively estimates the airlight color and transfers details from the NIR. A two-stage dehazing method is proposed by exploiting the dissimilarity between RGB and NIR for airlight color estimation, followed by a dehazing procedure through an optimization framework. Experiments on captured haze images show that our method can achieve substantial improvements on the detail recovery and the color distribution over the existing image dehazing algorithms. Shaojie Zhuo, Xiaopeng Zhang 0001, Liang Shen 0007, Sabine Süsstrunk |
ICIP | 5 |
| 2013 | Automatic detection of dust and scratches in silver halide film using polarized dark-field illuminationabstractWe present a method to automatically detect dust and scratches on photographic material, in particular silver halide film, where traditional methods for detecting and removing defects fail. The film is digitized using a novel setup involving cross-polarization and dark-field illumination in a cardinal light configuration, which compresses the signal and highlights the parts that are due to defects in the film. Applying a principal component analysis (PCA) on the four cardinal images allows us to further separate the signal part of the film from the defects. Information from all four principal components is combined to produce a surface defect mask, which can be used as input to inpainting methods to remove the defects. Our method is able to detect most of the dust and scratches while keeping false-detections low. Dominic Rüfenacht, Giorgio Trumpy, Rudolf Gschwind, Sabine Süsstrunk |
ICIP | 4 |
| 2013 | Estimating beauty ratings of videos using supervoxelsabstractThe major low-level perceptual components that influence the beauty ratings of video are color, contrast, and motion. To estimate the beauty ratings of the NHK dataset, we propose to extract these features based on supervoxels, which are a group of pixels that share similar color and spatial information through the temporal domain. Recent beauty methods use frame-level processing for visual features and disregard the spatio-temporal aspect of beauty. In this paper, we explicitly model this property by introducing supervoxel-based visual and motion features. In order to create a beauty estimator, we first identify 60 videos (either beautiful or not beautiful) in the NHK dataset. We then train a neural network regressor using the supervoxel-based features and binary beauty ratings. We rate the 1000 videos in the NHK dataset and rank them according to their ratings. When comparing our rankings with the actual rankings of the NHK dataset, we obtain a Spearman correlation coefficient of 0.42. Gökhan Yildirim 0001, Appu Shaji, Sabine Süsstrunk |
ACM Multimedia | 3 |
| 2013 | What is the space of spectral sensitivity functions for digital color cameras?abstractCamera spectral sensitivity functions relate scene radiance with captured RGB triplets. They are important for many computer vision tasks that use color information, such as multispectral imaging, color rendering, and color constancy. In this paper, we aim to explore the space of spectral sensitivity functions for digital color cameras. After collecting a database of 28 cameras covering a variety of types, we find this space convex and two-dimensional. Based on this statistical model, we propose two methods to recover camera spectral sensitivities using regular reflective color targets (e.g., color checker) from a single image with and without knowing the illumination. We show the proposed model is more accurate and robust for estimating camera spectral sensitivities than other basis functions. We also show two applications for the recovery of camera spectral sensitivities - simulation of color rendering for cameras and computational color constancy. Dengyu Liu, Jinwei Gu, Sabine Süsstrunk |
WACV | 4 |
| 2012 | An efficient demosaicing technique using geometrical informationabstractColor image sensors use color filter arrays (CFA) to capture information at each sensor pixel position and require color demosaicing to reconstruct full color images. The quality of the demosaicked image is hindered by the sensor characteristics during the acquisition process. In this work, we propose a bandelet-based demosaicing method for color images. To this end, we have used a spatial multiplexing model of color in order to obtain the luminance and the chrominance components of the acquired image. Then, a luminance filter is used to reconstruct the luminance component. Thereafter, based on the concept of maximal gradient of multivalued images, we propose an extension of the bandelet representation for the case of multivalued images. Finally, demosaicing is performed by merging the luminance and each of the chrominance component in the multivalued bandelet transform domain. The experimental evaluation of the proposed scheme shows beneficial performance over existing demosaicing approaches. Aldo Maalouf, Mohamed-Chaker Larabi, Sabine Süsstrunk |
ICIP | 3 |
| 2012 | Joint statistical analysis of images and keywords with applications in semantic image enhancementabstractWith the advent of social image-sharing communities, millions of images with associated semantic tags are now available online for free and allow us to exploit this abundant data in new ways. We present a fast non-parametric statistical framework designed to analyze a large data corpus of images and semantic tag pairs and find correspondences between image characteristics and semantic concepts. We learn the relevance of different image characteristics for thousands of keywords from one million annotated images. We demonstrate the framework's effectiveness with three different examples of semantic image enhancement: we adapt the gray-level tone-mapping, emphasize semantically relevant colors, and perform a defocus magnification for an image based on its semantic context. The performance of our algorithms is validated with psychophysical experiments. Albrecht J. Lindner, Appu Shaji, Nicolas Bonnier, Sabine Süsstrunk |
ACM Multimedia | 4 |
| 2012 | Compression of multispectral images: Color (RGB) plus near-infrared (NIR)abstractWe propose a compression framework for four-channel images, composed of color (RGB) and near-infrared (NIR) channels, which exploits the correlation between the visible and the NIR information. The high-frequency components of both visible and NIR scene representations are strongly correlated. By encoding only the DCT components that differ above a chosen threshold, we significantly improve compression ratios for a given quality level. To evaluate our proposed method, we compare our results with standard JPEG compression, as well as PCA-based approaches that are often employed to compress multispectral images. Our experiments show that applying our proposed method yields the same quality at a lower bit-rate, compared to conventional JPEG and PCA-based algorithms. Neda Salamati, Zahra Sadeghipoor, Sabine Süsstrunk |
MMSP | 3 |
| 2012 | SLIC Superpixels Compared to State-of-the-Art Superpixel MethodsabstractComputer vision applications have come to rely increasingly on superpixels in recent years, but it is not always clear what constitutes a good superpixel algorithm. In an effort to understand the benefits and drawbacks of existing methods, we empirically compare five state-of-the-art superpixel algorithms for their ability to adhere to image boundaries, speed, memory efficiency, and their impact on segmentation performance. We then introduce a new superpixel algorithm, simple linear iterative clustering (SLIC), which adapts a k-means clustering approach to efficiently generate superpixels. Despite its simplicity, SLIC adheres to boundaries as well as or better than previous methods. At the same time, it is faster and more memory efficient, improves segmentation performance, and is straightforward to extend to supervoxel generation. Radhakrishna Achanta, Appu Shaji, Kevin Smith 0001, Aurélien Lucchi, Pascal Fua, Sabine Süsstrunk |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2012 | A New In-Camera Imaging Model for Color Computer Vision and Its ApplicationabstractWe present a study of in-camera image processing through an extensive analysis of more than 10,000 images from over 30 cameras. The goal of this work is to investigate if image values can be transformed to physically meaningful values, and if so, when and how this can be done. From our analysis, we found a major limitation of the imaging model employed in conventional radiometric calibration methods and propose a new in-camera imaging model that fits well with today's cameras. With the new model, we present associated calibration procedures that allow us to convert sRGB images back to their original CCD RAW responses in a manner that is significantly more accurate than any existing methods. Additionally, we show how this new imaging model can be used to build an image correction application that converts an sRGB input image captured with the wrong camera settings to an sRGB output image that would have been recorded under the correct settings of a specific camera. Seon Joo Kim, Hai Ting Lin, Zheng Lu 0002, Sabine Süsstrunk, Stephen Lin 0001, Michael S. Brown |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2011 | Multi-spectral SIFT for scene category recognitionabstractWe use a simple modification to a conventional SLR camera to capture images of several hundred scenes in colour (RGB) and near-infrared (NIR). We show that the addition of near-infrared information leads to significantly improved performance in a scene-recognition task, and that the improvements are greater still when an appropriate 4-dimensional colour representation is used. In particular we propose MSIFT - a multispectral SIFT descriptor that, when combined with a kernel based classifier, exceeds the performance of state-of-the-art scene recognition techniques (e.g., GIST) and their multispectral extensions. We extensively test our algorithms using a new dataset of several hundred RGB-NIR scene images, as well as benchmarking against Torralba's scene categorization dataset. Matthew A. Brown, Sabine Süsstrunk |
CVPR | 2 |
| 2011 | Revisiting radiometric calibration for color computer visionabstractWe present a study of radiometric calibration and the in-camera imaging process through an extensive analysis of more than 10,000 images from over 30 cameras. The goal is to investigate if image values can be transformed to physically meaningful values and if so, when and how this can be done. From our analysis, we show that the conventional radiometric model fits well for image pixels with low color saturation but begins to degrade as color saturation level increases. This is due to the color mapping process which includes gamut mapping in the in-camera processing that cannot be modeled with conventional methods. To this end, we introduce a new imaging model for radiometric calibration and present an effective calibration scheme that allows us to compensate for the nonlinear color correction to convert non-linear sRGB images to CCD RAW responses. Hai Ting Lin, Seon Joo Kim, Sabine Süsstrunk, Michael S. Brown |
ICCV | 3 |
| 2011 | Multispectral interest points for RGB-NIR image registrationabstractThis paper explores the use of joint colour and near-infrared (NIR) information for feature based matching and image registration. In particular, we investigate multispectral generalisations of two popular interest point detectors (Harris and difference of Gaussians), and show that these give a marked improvement in performance when the extra NIR channel is available. We also look at the problem of multimodal RGB to NIR registration, and propose a variant of the SIFT descriptor that gives improved performance. Damien Firmenich, Matthew A. Brown, Sabine Süsstrunk |
ICIP | 3 |
| 2011 | Correlation-based joint acquisition and demosaicing of visible and near-infrared imagesabstractJoint processing of visible (RGB) and near-infrared (NIR) images has recently found some appealing applications, which make joint capturing a pair of visible and NIR images an important problem. In this paper, we propose a new method to design color filter arrays (CFA) and demosaicing matrices for acquiring NIR and visible images using a single sensor. The proposed method modifies the optimum CFA algorithm proposed in [1] by taking advantage of the NIR/visible correlation in the design process. Simulation results show that by applying the proposed method, the quality of demo-saiced NIR and visible images is increased by about 1 dB in peak signal-to-noise ratio over the results of the optimum CFA algorithm. It is also shown that better visual quality can be obtained by using the proposed algorithm. Zahra Sadeghipoor, Yue M. Lu, Sabine Süsstrunk |
ICIP | 3 |
| 2011 | Removing shadows from images using color and near-infraredabstractShadows often introduce errors in the performance of computer vision algorithms, such as object detection and tracking. This paper proposes a method to remove shadows from real images based on a probability shadow map. The probability shadow map identifies how much light is impinging on a surface. The lightness of shadowed regions in an image is increased and then the color of that part of the surface is corrected so that it matches the lit part of the surface. The result is compared with two other shadow removal frameworks. The advantage of our method is that after removal, the texture and all the details in the shadowed regions remain intact. Neda Salamati, Arthur Germain, Sabine Süsstrunk |
ICIP | 3 |
| 2010 | Saliency detection using maximum symmetric surroundabstractDetection of visually salient image regions is useful for applications like object segmentation, adaptive compression, and object recognition. Recently, full-resolution salient maps that retain well-defined boundaries have attracted attention. In these maps, boundaries are preserved by retaining substantially more frequency content from the original image than older techniques. However, if the salient regions comprise more than half the pixels of the image, or if the background is complex, the background gets highlighted instead of the salient object. In this paper, we introduce a method for salient region detection that retains the advantages of such saliency maps while overcoming their shortcomings. Our method exploits features of color and luminance, is simple to implement and is computationally efficient. We compare our algorithm to six state-of-the-art salient region detection methods using publicly available ground truth. Our method outperforms the six algorithms by achieving both higher precision and better recall. We also show application of our saliency maps in an automatic salient object segmentation scheme using graph-cuts. Radhakrishna Achanta, Sabine Süsstrunk |
ICIP | 2 |
| 2010 | Automatic skin enhancement with visible and near-infrared image fusionabstractSkin tones, portraits in particular, are of critical importance in photography and video, but a number of factors, such as pigmentation irregularities (e.g., moles, freckles), irritation, roughness, or wrinkles can reduce their appeal. Moreover, such "defects" are oftentimes enhanced by scene lighting conditions. Sabine Süsstrunk, Clément Fredembach, Daniel Tamburrino |
ACM Multimedia | 1 |
| 2009 | Frequency-tuned salient region detectionabstractDetection of visually salient image regions is useful for applications like object segmentation, adaptive compression, and object recognition. In this paper, we introduce a method for salient region detection that outputs full resolution saliency maps with well-defined boundaries of salient objects. These boundaries are preserved by retaining substantially more frequency content from the original image than other existing techniques. Our method exploits features of color and luminance, is simple to implement, and is computationally efficient. We compare our algorithm to five state-of-the-art salient region detection methods with a frequency domain analysis, ground truth, and a salient object segmentation application. Our method outperforms the five algorithms both on the ground-truth evaluation and on the segmentation task by achieving both higher precision and better recall. Radhakrishna Achanta, Sheila S. Hemami, Francisco J. Estrada, Sabine Süsstrunk |
CVPR | 4 |
| 2009 | Appearance-based keypoint clusteringabstractWe present an algorithm for clustering sets of detected interest points into groups that correspond to visually distinct structure. Through the use of a suitable colour and texture representation, our clustering method is able to identify keypoints that belong to separate objects or background regions. These clusters are then used to constrain the matching of keypoints over pairs of images, resulting in greatly improved matching under difficult conditions. We present a thorough evaluation of each component of the algorithm, and show its usefulness on difficult matching problems. Francisco J. Estrada, Pascal Fua, Vincent Lepetit, Sabine Süsstrunk |
CVPR | 4 |
| 2009 | The gigavision cameraabstractWe propose a new image device called gigavision camera. The main differences between a conventional and a gigavision camera are that the pixels of the gigavision camera are binary and orders of magnitude smaller. A gigavision camera can be built using standard memory chip technology, where each memory bit is designed to be light sensitive. A conventional gray level image can be obtained from the binary gigavision image by low-pass filtering and sampling. The main advantage of the gigavision camera is that its response is non-linear and similar to a logarithmic function, which makes it suitable for acquiring high dynamic range scenes. The larger the number of binary pixels considered, the higher the dynamic range of the gigavision camera will be. In addition, the binary sensor of the gigavision camera can be combined with a lens array in order to realize an extremely thin camera. Due to the small size of the pixels, this design does not require deconvolution techniques typical of similar systems based on conventional sensors. Luciano Sbaiz, Edoardo Charbon, Sabine Süsstrunk, Martin Vetterli |
ICASSP | 4 |
| 2009 | Saliency detection for content-aware image resizingabstractContent aware image re-targeting methods aim to arbitrarily change image aspect ratios while preserving visually prominent features. To determine visual importance of pixels, existing re-targeting schemes mostly rely on grayscale intensity gradient maps. These maps show higher energy only at edges of objects, are sensitive to noise, and may result in deforming salient objects. In this paper, we present a computationally efficient, noise robust re-targeting scheme based on seam carving by using saliency maps that assign higher importance to visually prominent whole regions (and not just edges). This is achieved by computing global saliency of pixels using intensity as well as color features. Our saliency maps easily avoid artifacts that conventional seam carving generates and are more robust in the presence of noise. Also, unlike gradient maps, which may have to be recomputed several times during a seam carving based re-targeting operation, our saliency maps are computed only once independent of the number of seams added or removed. Radhakrishna Achanta, Sabine Süsstrunk |
ICIP | 2 |
| 2009 | Face image enhancement using 3D and spectral informationabstractThis paper presents a novel method of enhancing image quality of face pictures using 3D and spectral information. Most conventional techniques directly work on the image data, shifting the skin color to a predefined skin tone, and thus do not take into account the effects of shape and lighting. The proposed method first recovers the 3D shape of a face in an input image using a 3D morphable model. Then, using color constancy and inverse rendering techniques, specularities and the true skin color, i.e., its spectral reflectance, are recovered. The quality of the input image is improved by matching the skin reflectance to a predefined reference and reducing the amount of specularities. The method realizes the enhancement in a more physically accurate manner compared to previous ones. Subjective experiments on image quality demonstrate the validity of the proposed method. Charles Dubout, Masato Tsukada, Rui Ishiyama, Chisato Funayama, Sabine Süsstrunk |
ICIP | 5 |
| 2009 | Designing color filter arrays for the joint capture of visible and near-infrared imagesabstractDigital camera sensors are inherently sensitive to the near-infrared (NIR) part of the light spectrum. In this paper, we propose a general design for color filter arrays that allow the joint capture of visible/NIR images using a single sensor. We pose the CFA design as a novel spatial domain optimization problem, and provide an efficient iterative procedure that finds (locally) optimal solutions. Numerical experiments confirm the effectiveness of the proposed CFA design, which can simultaneously capture high quality visible and NIR image pairs. Yue M. Lu, Clément Fredembach, Martin Vetterli, Sabine Süsstrunk |
ICIP | 4 |
| 2009 | Color image dehazing using the near-infraredabstractIn landscape photography, distant objects often appear blurred with a blue color cast, a degradation caused by atmospheric haze. To enhance image contrast, pleasantness and information content, dehazing can be performed. Lex Schaul, Clément Fredembach, Sabine Süsstrunk |
ICIP | 3 |
| 2008 | Salient Region Detection and Segmentation
Radhakrishna Achanta, Francisco J. Estrada, Patricia Wils, Sabine Süsstrunk |
ICVS | 4 |
| 2008 | Color match: an imaging based mobile cosmetics advisory serviceabstractIn this paper we describe an exploratory study of a mobile cosmetic advisory system that enables women to select appropriate colors of cosmetics. This system is intended for commercial use to address the problem of foundation color selection. Although women are primarily responsible for making most purchasing decisions in the US, we found very few studies to assess the adoption of retail related mobile services by women. Based on surveys, semi-structured interviews, and focus groups, we have identified a number of design factors that should be considered when designing mobile services for women consumers. The results of our study indicate that while usefulness is an important factor, other design aspects such as mobile vs. kiosk, installed vs. existing software, technical comfort vs. social comfort, social vs. individual, privacy and trust should also be accounted for. Jhilmil Jain, Nina T. Bhatti, H. Harlyn Baker, Hui Chao, Mohamed Dekhil, Michael Harville, Nic Lyons, John Schettino, Sabine Süsstrunk |
Mobile HCI | 9 |
| 2008 | Higher Order SVD Analysis for Dynamic Texture SynthesisabstractVideos representing flames, water, smoke, etc., are often defined as dynamic textures: "textures" because they are characterized by the redundant repetition of a pattern and "dynamic" because this repetition is also in time and not only in space. Dynamic textures have been modeled as linear dynamic systems by unfolding the video frames into column vectors and describing their trajectory as time evolves. After the projection of the vectors onto a lower dimensional space by a singular value decomposition (SVD), the trajectory is modeled using system identification techniques. Synthesis is obtained by driving the system with random noise. In this paper, we show that the standard SVD can be replaced by a higher order SVD (HOSVD), originally known as Tucker decomposition. HOSVD decomposes the dynamic texture as a multidimensional signal (tensor) without unfolding the video frames on column vectors. This is a more natural and flexible decomposition, since it permits us to perform dimension reduction in the spatial, temporal, and chromatic domain, while standard SVD allows for temporal reduction only. We show that for a comparable synthesis quality, the HOSVD approach requires, on average, five times less parameters than the standard SVD approach. The analysis part is more expensive, but the synthesis has the same cost as existing algorithms. Our technique is, thus, well suited to dynamic texture synthesis on devices limited by memory and computational power, such as PDAs or mobile phones. Roberto Costantini, Luciano Sbaiz, Sabine Süsstrunk |
IEEE Trans. Image Process. | 3 |
| 2006 | Dynamic Texture Synthesis: Compact Models Based on Luminance-Chrominance Color RepresentationabstractDynamic textures are sequences of images showing temporal regularity. Examples can be found in videos representing smoke, flames, ocean waves, wind-shaken forests, etc. Dynamic texture modelling and synthesis has usually been done considering RGB color images. In this paper, we analyze the use of different color encodings, which permit to model luminance and chrominance information separately. We find that this separation is more appropriate, since it takes advantage of the spatial and temporal characteristics of the color channels and leads to more flexible and compact representations. We show that compared to RGB, similar synthesis performance can be achieved using YCbCror Lab color encodings, using half of the model coefficients and less computational power. Robert Constantini, Luciano Sbaiz, Sabine Süsstrunk |
ICIP | 3 |
| 2006 | High dynamic range image rendering with a retinex-based adaptive filterabstractWe propose a new method to render high dynamic range images that models global and local adaptation of the human visual system. Our method is based on the center-surround Retinex model. The novelties of our method is first to use an adaptive filter, whose shape follows the image high-contrast edges, thus reducing halo artifacts common to other methods. Second, only the luminance channel is processed, which is defined by the first component of a principal component analysis. Principal component analysis provides orthogonality between channels and thus reduces the chromatic changes caused by the modification of luminance. We show that our method efficiently renders high dynamic range images and we compare our results with the current state of the art. Laurence Meylan, Sabine Süsstrunk |
IEEE Trans. Image Process. | 2 |
| 2005 | Consistent image-based measurement and classification of skin colorabstractLittle prior image processing work has addressed estimation and classification of skin color in a manner that is independent of camera and illuminant. To this end, we first present new methods for 1) fast, easy-to-use image color correction, with specialization toward skin tones, and 2) fully automated estimation of facial skin color, with robustness to shadows, specularities, and blemishes. Each of these is validated independently against ground truth, and then combined with a classification method that successfully discriminates skin color across a population of people imaged with several different cameras. We also evaluate the effects of image quality and various algorithmic choices on our classification performance. We believe our methods are practical for relatively untrained operators, using inexpensive consumer equipment. Michael Harville, H. Harlyn Baker, Nina T. Bhatti, Sabine Süsstrunk |
ICIP (2) | 4 |
| 2005 | Super-resolution from highly undersampled imagesabstractAliasing artifacts in images are visually very disturbing. Therefore, most imaging devices apply a low-pass filter before sampling. This removes all aliasing from the image, but it also creates a blurred image. Actually, all the image information above half the sampling frequency is removed. In this paper, we present a new method for the reconstruction of a high resolution image from a set of highly undersampled and thus aliased images. We use the information in the entire frequency spectrum, including the aliased part, to create a sharp, high resolution image. The unknown relative shifts between the images are computed using a subspace projection approach. We show that the projection can be decomposed into multiple projections onto smaller subspaces. This allows for a considerable reduction of the overall computational complexity of the algorithm. A high resolution image can then be reconstructed from the registered low resolution images. Simulation results show the validity of our algorithm. Patrick Vandewalle, Luciano Sbaiz, Martin Vetterli, Sabine Süsstrunk |
ICIP (1) | 4 |
| 2005 | Linear demosaicing inspired by the human visual systemabstractThere is an analogy between single-chip color cameras and the human visual system in that these two systems acquire only one limited wavelength sensitivity band per spatial location. We have exploited this analogy, defining a model that characterizes a one-color per spatial position image as a coding into luminance and chrominance of the corresponding three colors per spatial position image. Luminance is defined with full spatial resolution while chrominance contains subsampled opponent colors. Moreover, luminance and chrominance follow a particular arrangement in the Fourier domain, allowing for demosaicing by spatial frequency filtering. This model shows that visual artifacts after demosaicing are due to aliasing between luminance and chrominance and could be solved using a preprocessing filter. This approach also gives new insights for the representation of single-color per spatial location images and enables formal and controllable procedures to design demosaicing algorithms that perform well compared to concurrent approaches, as demonstrated by experiments. David Alleysson, Sabine Süsstrunk, Jeanny Hérault |
IEEE Trans. Image Process. | 2 |
| 2004 | A measure for spatial dependence in natural stochastic textures
Roberto Costantini, Gloria Menegaz, Sabine Süsstrunk |
ICIP | 3 |
| 2004 | Mapping colour in image stitching applications
David Hasler, Sabine Süsstrunk |
J. Vis. Commun. Image Represent. | 2 |
| 2004 | Eigenregions for Image ClassificationabstractFor certain databases and classification tasks, analyzing images based region features instead of image features results in more accurate classifications. We introduce eigenregions, which are geometrical features that encompass area, location, and shape properties of an image region, even if the region is spatially incoherent. Eigenregions are calculated using principal component analysis (PCA). On a database of 77,000 different regions obtained through the segmentation of 13,500 real-scene photographic images taken by nonprofessionals, eigenregions improved the detection of localized image classes by a noticeable amount. Additionally, eigenregions allow us to prove that the largest variance in natural image region geometry is due to its area and not to shape or position. Clément Fredembach, Michael Schröder 0003, Sabine Süsstrunk |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Superresolution images reconstructed from aliased images
Patrick Vandewalle, Sabine Süsstrunk, Martin Vetterli |
VCIP | 2 |
| 2003 | Outlier Modeling in Image MatchingabstractWe address the question of how to characterize the outliers that may appear when matching two views of the same scene. The match is performed by comparing the difference of the two views at a pixel level aiming at a better registration of the images. When using digital photographs as input, we notice that an outlier is often a region that has been occluded, an object that suddenly appears in one of the images, or a region that undergoes an unexpected motion. By assuming that the error in pixel intensity generated by the outlier is similar to an error generated by comparing two random regions in the scene, we can build a model for the outliers based on the content of the two views. We illustrate our model by solving a pose estimation problem: the goal is to compute the camera motion between two views. The matching is expressed as a mixture of inliers versus outliers, and defines a function to minimize for improving the pose estimation. Our model has two benefits: First, it delivers a probability for each pixel to belong to the outliers. Second, our tests show that the method is substantially more robust than traditional robust estimators (M-estimators) used in image stitching applications, with only a slightly higher computational complexity. David Hasler, Luciano Sbaiz, Sabine Süsstrunk, Martin Vetterli |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2003 | Preface
Raimondo Schettini, Christine Fernandez, Sabine Süsstrunk |
Pattern Recognit. Lett. | 3 |
| 2000 | Digital photography-How long will it last?abstractPermanence issues for digital photographs arise in three different areas. First, the materials used for digital hard copies should preferably be as permanent as the conventional photographic materials. Longevity of digital hard copy materials is affected by the stability of the material and the storage conditions. Secondly, the digital files that are the counterpart of the conventional negatives need to be readable not only on various systems and platforms today but also in the future. Third, the encoding of digital images should be such that any improvement in processing algorithms and display/output technologies can be applied in future image workflows. The safe keeping of digital data requires an active and regular maintenance of the data. This paper discusses encoding issues for archival images and strategies to make image preservation happen for digital photographs. Franziska S. Frey, Sabine Süsstrunk |
ISCAS | 2 |