VLDB 2026 Research / reviewers in the wild / expert
Dimitris Samaras
dblp:s/DimitrisSamaras
· DBLP profile ↗
182ranked-venue papers
5as first author
74since 2021 · last 2026
0000-0002-1373-0294ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 135 · 5 first-author · 48 since 2021Graphics, computer vision, multimedia, augmented reality and games · 130 · 4 first-author · 56 since 2021Applied, interdisciplinary, general and emerging computing · 16 · 10 since 2021Computer networks · 3 · 2 since 2021Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LVLMs and Humans Ground Differently in Referential CommunicationabstractPeter Zeng, Weiling Li, Amie J. Paige, Zhengxiang Wang, Panagiotis Kaliosis, Dimitris Samaras, Gregory J. Zelinsky, Susan Brennan, Owen Rambow. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Peter Zeng, Weiling Li, Amie J. Paige, Zhengxiang Wang, Panagiotis Kaliosis, Dimitris Samaras, Gregory J. Zelinsky, Susan Brennan, Owen Rambow |
ACL (1) | 6 |
| 2026 | CORA: Consistency-Guided Semi-Supervised Framework for Reasoning SegmentationabstractReasoning segmentation seeks pixel-accurate masks for targets referenced by complex, often implicit instructions, requiring context-dependent reasoning over the scene. Recent multimodal language models have advanced instruction following segmentation, yet generalization remains limited. The key bottleneck is the high cost of curating diverse, high-quality pixel annotations paired with rich linguistic supervision leading to brittle performance under distribution shift. Therefore, we present CORA, a semi-supervised reasoning segmentation framework that jointly learns from limited labeled data and a large corpus of unlabeled images. CORA introduces three main components: 1) conditional visual instructions that encode spatial and contextual relationships between objects; 2) a noisy pseudo-label filter based on the consistency of Multimodal LLM’s outputs across semantically equivalent queries; and 3) a token-level contrastive alignment between labeled and pseudo-labeled samples to enhance feature consistency. These components enable CORA to perform robust reasoning segmentation with minimal supervision, outperforming existing baselines under constrained annotation settings. CORA achieves state-of-the-art results, requiring as few as 100 labeled images on Cityscapes, a benchmark dataset for urban scene understanding, surpassing the baseline by +2.3%. Similarly, CORA improves performance by +2.4% with only 180 labeled images on PanNuke, a histopathology dataset. Prantik Howlader, Hoang Nguyen-Canh, Srijan Das, Hieu Le 0001, Dimitris Samaras |
WACV | 6 |
| 2026 | Measuring and predicting where and when pathologists focus their visual attention while grading whole slide images of cancer
Souradeep Chakraborty, Ruoyu Xue, Rajarsi Gupta 0001, Oksana Yaskiv, Constantin Friedman, Natallia Sheuka, Dana Perez, Paul Friedman, Won-Tak Choi, Waqas Mahmud, Beatrice S. Knudsen, Gregory J. Zelinsky, Joel H. Saltz, Dimitris Samaras |
Medical Image Anal. | 14 |
| 2026 | Label-Efficient Deep Color Deconvolution of Brightfield Multiplex IHC ImagesabstractBrightfield Multiplex Immunohistochemistry (mIHC) provides simultaneous labeling of multiple protein biomarkers in the same tissue section. It enables the exploration of spatial relationships between the inflammatory microenvironment and tumor cells, and to uncover how tumor cell morphology relates to cancer biomarker expression. Color deconvolution is required to analyze and quantify the different cell phenotype populations present as indicated by the biomarkers. However, this becomes a challenging task as the number of multiplexed stains increase. In this work, we present self-supervised and semi-supervised approaches to mIHC color deconvolution. Our proposed methods are based on deep convolutional autoencoders and learn using innovative reconstruction losses inspired by physics. We show how we can integrate weak annotations and the abundant unlabeled data available to train a model to reliably unmix the multiplexed stains and generate stain segmentation maps. We demonstrate the effectiveness of our proposed methods through experiments on mIHC dataset of 7-plexed IHC images. Shahira Abousamra, Danielle Fassler, Rajarsi Gupta 0001, Tahsin M. Kurç, Luisa F. Escobar-Hoyos, Dimitris Samaras, Kenneth Shroyer, Joel H. Saltz, Chao Chen 0012 |
IEEE Trans. Medical Imaging | 6 |
| 2025 | Rig3DGS: Creating Controllable Portraits From Casual Monocular VideosabstractWe present Rig3DGS, a novel technique for creating reanimatable 3D portraits from short monocular smartphone videos. Rig3DGS learns to reconstruct a set of controllable 3D Gaussians from a monocular video of a dynamic subject captured with varying head poses and facial expressions in an in-the-wild scene. In contrast to synchronized multi-view studio captures, this in-the-wild, single camera setup brings fresh challenges to learning high quality 3D Gaussians. We address these challenges by learning to deform 3D Gaussians from a fixed canonical space to the deformed space that is consistent with the target facial expression and headpose. Our key contribution is a carefully designed deformation model that is guided by a 3D face morphable model. This deformation not only enables control over facial expression and head-poses but also allows our method to generates high-quality photorealistic renders of the whole scene. Once trained, Rig3DGS is able to generate photorealistic renders of a subject and their scene for novel facial expression, head-poses, and viewing directions. Through extensive experiments we demonstrate that Rig3DGS significantly outperforms prior art while being orders of magnitude faster. Alfredo Rivero, Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
3DV | 4 |
| 2025 | Direct and Explicit 3D Generation from a Single ImageabstractCurrent image-to-3D approaches suffer from high computational costs and lack scalability for high-resolution outputs. In contrast, we introduce a novel framework to directly generate explicit surface geometry and texture using multi-view 2D depth and RGB images along with 3D Gaussian features using a repurposed Stable Diffusion model. We introduce a depth branch into U-Net for efficient and high quality multi-view, cross-domain generation and incorporate epipolar attention into the latent-to-pixel decoder for pixel-level multi-view consistency. By back-projecting the generated depth pixels into 3D space, we create a structured 3D representation that can be either rendered via Gaussian splatting or extracted to high-quality meshes, thereby leveraging additional novel view synthesis loss to further improve our performance. Extensive experiments demonstrate that our method surpasses existing baselines in geometry and texture quality while achieving significantly faster generation time. Meher Gitika Karumuri, Chuhang Zou, Seungbae Bang, Dimitris Samaras, Sunil Hadap |
3DV | 6 |
| 2025 | mli-NeRF: Multi-Light Intrinsic-Aware Neural Radiance FieldsabstractCurrent methods for extracting intrinsic image components, such as reflectance and shading, primarily rely on statistical priors. These methods focus mainly on simple synthetic scenes and isolated objects and struggle to perform well on challenging real-world data. To address this issue, we propose MLI-NeRF, which integrates Multiple Light information in Intrinsic-aware Neural Radiance Fields. By leveraging scene information provided by different light source positions complementing the multi-view information, we generate pseudo-label images for reflectance and shading to guide intrinsic image decomposition without the need for ground truth data. Our method introduces straightforward supervision for intrinsic component separation and ensures robustness across diverse scene types. We validate our approach on both synthetic and real-world datasets, outperforming existing state-of-the-art methods. Additionally, we demonstrate its applicability to various image editing tasks. Code and data are available at https://github.com/liulisixin/MLI-NeRF. Yixiong Yang, Shilin Hu, Ramón Baldrich, Dimitris Samaras, María Vanrell 0001 |
3DV | 5 |
| 2025 | TopoCellGen: Generating Histopathology Cell Topology with a Diffusion ModelabstractAccurately modeling multi-class cell topology is crucial in digital pathology, as it provides critical insights into tissue structure and pathology. The synthetic generation of cell topology enables realistic simulations of complex tissue environments, enhances downstream tasks by augmenting training data, aligns more closely with pathologists' domain knowledge, and offers new opportunities for controlling and generalizing the tumor microenvironment. In this paper, we propose a novel approach that integrates topological constraints into a diffusion model to improve the generation of realistic, contextually accurate cell topologies. Our method refines the simulation of cell distributions and interactions, increasing the precision and interpretability of results in downstream tasks such as cell detection and classification. To assess the topological fidelity of generated layouts, we introduce a new metric, Topological Fréchet Distance (TopoFD), which overcomes the limitations of traditional metrics like FID in evaluating topological structure. Experimental results demonstrate the effectiveness of our approach in generating multi-class cell layouts that capture intricate topological relationships. Code is available at https://github.com/Melon-Xu/TopoCellGen. Meilong Xu, Saumya Gupta, Xiaoling Hu 0002, Chen Li 0045, Shahira Abousamra, Dimitris Samaras, Prateek Prasanna, Chao Chen 0012 |
CVPR | 6 |
| 2025 | Few-shot Personalized Scanpath PredictionabstractA personalized model for scanpath prediction provides insights into the visual preferences and attention patterns of individual subjects. However, existing methods for training scanpath prediction models are data-intensive and cannot be effectively personalized to new individuals with only a few available examples. In this paper, we propose few-shot personalized scanpath prediction task (FS-PSP) and a novel method to address it, which aims to predict scan-paths for an unseen subject using minimal support data of that subject’s scanpath behavior. The key to our method’s adaptability is the Subject-Embedding Network (SE-Net), specifically designed to capture unique, individualized representations for each subject’s scanpaths. SE-Net generates subject embeddings that effectively distinguish between subjects while minimizing variability among scanpaths from the same individual. The personalized scanpath prediction model is then conditioned on these subject embeddings to produce accurate, personalized results. Experiments on multiple eye-tracking datasets demonstrate that our method excels in FS-PSP settings and does not require any fine-tuning steps at test time. Code is available at: https://github.com/cvlab-stonybrook/few-shot-scanpath Ruoyu Xue, Sounak Mondal, Hieu Le 0001, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
CVPR | 7 |
| 2025 | ZoomLDM: Latent Diffusion Model for Multi-scale Image GenerationabstractDiffusion models have revolutionized image generation, yet several challenges restrict their application to large-image domains, such as digital pathology and satellite imagery. Given that it is infeasible to directly train a model on ’whole’ images from domains with potential gigapixel sizes, diffusion-based generative methods have focused on synthesizing small, fixed-size patches extracted from these images. However, generating small patches has limited applicability since patch-based models fail to capture the global structures and wider context of large images, which can be crucial for synthesizing (semantically) accurate samples. To overcome this limitation, we present ZoomLDM, a diffusion model tailored for generating images across multiple scales. Central to our approach is a novel magnification-aware conditioning mechanism that utilizes self-supervised learning (SSL) embeddings and allows the diffusion model to synthesize images at different ’zoom’ levels, i.e., fixed-size patches extracted from large images at varying scales. ZoomLDM synthesizes coherent histopathology images that remain contextually accurate and detailed at different zoom levels, achieving state-of-the-art image generation quality across all scales and excelling in the data-scarce setting of generating thumbnails of entire large images. The multi-scale nature of ZoomLDM unlocks additional capabilities in large image generation, enabling computationally tractable and globally coherent image synthesis up to 4096 × 4096 pixels and 4 × super-resolution. Additionally, multi-scale features extracted from ZoomLDM are highly effective in multiple instance learning experiments.1 Srikar Yellapragada, Alexandros Graikos, Kostas Triaridis, Prateek Prasanna, Rajarsi Gupta 0001, Joel H. Saltz, Dimitris Samaras |
CVPR | 7 |
| 2025 | 2DMamba: Efficient State Space Model for Image Representation with Applications on Giga-Pixel Whole Slide Image ClassificationabstractEfficiently modeling large 2D contexts is essential for various fields including Giga-Pixel Whole Slide Imaging (WSI) and remote sensing. Transformer-based models offer high parallelism but face challenges due to their quadratic complexity for handling long sequences. Recently, Mamba introduced a selective State Space Model (SSM) with linear complexity and high parallelism, enabling effective and efficient modeling of wide context in 1D sequences. However, extending Mamba to vision tasks, which inherently involve 2D structures, results in spatial discrepancies due to the limitations of 1D sequence processing. On the other hand, current 2D SSMs inherently model 2D structures but they suffer from prohibitively slow computation due to the lack of efficient parallel algorithms. In this work, we propose 2DMamba, a novel 2D selective SSM framework that incorporates the 2D spatial structure of images into Mamba, with a highly optimized hardware-aware operator, adopting both spatial continuity and computational efficiency. We validate the versatility of our approach on both WSIs and natural images. Extensive experiments on 10 public datasets for WSI classification and survival analysis show that 2DMamba improves up to 2.48% in AUC, 3.11% in F1 score, 2.47% in accuracy and 5.52% in C-index. Additionally, integrating our method with VMamba for natural imaging yields 0.5 to 0.7 improvements in mIoU on the ADE20k semantic segmentation dataset, and 0.2% accuracy improvement on ImageNet-1K classification dataset. Our code is available at https://github.com/AtlasAnalyticsLab/2DMamba. Anh Tien Nguyen, Xi Han 0002, Vincent Quoc-Huy Trinh, Hong Qin 0001, Dimitris Samaras, Mahdi S. Hosseini |
CVPR | 6 |
| 2025 | Leveraging Registers in Vision Transformers for Robust AdaptationabstractVision Transformers (ViTs) have shown success across a variety of tasks due to their ability to capture global image representations. Recent studies have identified the existence of high-norm tokens in ViTs, which can interfere with unsupervised object discovery. To address this, the use of "registers" which are additional tokens that isolate high norm patch tokens while capturing global image-level information has been proposed. While registers have been studied extensively for object discovery, their generalization properties particularly in out-of-distribution (OOD) scenarios, remains underexplored. In this paper, we examine the utility of register token embeddings in providing additional features for improving generalization and anomaly rejection. To that end, we propose a simple method that combines the special CLS token embedding commonly employed in ViTs with the average-pooled register embeddings to create feature representations which are subsequently used for training a downstream classifier. We find that this enhances OOD generalization and anomaly rejection, while maintaining in-distribution (ID) performance. Extensive experiments across multiple ViT backbones trained with and without registers reveal consistent improvements of 2-4% in top-1 OOD accuracy and a 2-3% reduction in false positive rates for anomaly detection. Importantly, these gains are achieved without additional computational overhead. Srikar Yellapragada, Kowshik Thopalli, Vivek Sivaraman Narayanaswamy, Wesam A. Sakla, Yamen Mubarka, Dimitris Samaras, Jayaraman J. Thiagarajan |
ICASSP | 7 |
| 2025 | AV-Flow: Transforming Text to Audio-Visual Human-Like InteractionsabstractWe introduce AV-Flow, an audio-visual generative model that animates photo-realistic 4D talking avatars given only text input. In contrast to prior work that assumes an existing speech signal, we synthesize speech and vision jointly. We demonstrate human-like speech synthesis, synchronized lip motion, lively facial expressions and head pose; all generated from just text characters. The core premise of our approach lies in the architecture of our two parallel diffusion transformers. Intermediate highway connections ensure communication between the audio and visual modalities, and thus, synchronized speech intonation and facial dynamics (e.g., eyebrow motion). Our model is trained with flow matching, leading to expressive results and fast inference. In case of dyadic conversations, AV-Flow produces an always-on avatar, that actively listens and reacts to the audio-visual input of a user. Through extensive experiments, we show that our method outperforms prior work, synthesizing natural-looking 4D talking avatars. Project page: https://aggelinacha.github.io/AV-Flow/ Aggelina Chatziagapi, Louis-Philippe Morency, Hongyu Gong, Michael Zollhöfer, Dimitris Samaras, Alexander Richard |
ICCV | 5 |
| 2025 | GECKO: Gigapixel Vision-Concept Contrastive Pretraining in HistopathologyabstractPretraining a Multiple Instance Learning (MIL) aggregator enables the derivation of Whole Slide Image (WSI)-level embeddings from patch-level representations without supervision. While recent multimodal MIL pretraining approaches leveraging auxiliary modalities have demonstrated performance gains over unimodal WSI pretraining, the acquisition of these additional modalities necessitates extensive clinical profiling. This requirement increases costs and limits scalability in existing WSI datasets lacking such paired modalities. To address this, we propose Gigapixel Vision-Concept Knowledge Contrastive pretraining (GECKO), which aligns WSIs with a Concept Prior derived from the available WSIs. First, we derive an inherently interpretable concept prior by computing the similarity between each WSI patch and textual descriptions of predefined pathology concepts. GECKO then employs a dual-branch MIL network: one branch aggregates patch embeddings into a WSI-level deep embedding, while the other aggregates the concept prior into a corresponding WSI-level concept embedding. Both aggregated embeddings are aligned using a contrastive objective, thereby pretraining the entire dual-branch MIL model. Moreover, when auxiliary modalities such as transcriptomics data are available, GECKO seamlessly integrates them. Across five diverse tasks, GECKO consistently outperforms prior unimodal and multimodal pretraining approaches while also delivering clinically meaningful interpretability that bridges the gap between computational models and pathology expertise. Code is made available at https://github.com/bmi-imaginelab/GECKO Saarthak Kapse, Pushpak Pati, Srikar Yellapragada, Srijan Das, Rajarsi Gupta 0001, Joel H. Saltz, Dimitris Samaras, Prateek Prasanna |
ICCV | 7 |
| 2025 | Multi-View Gaze Target EstimationabstractThis paper presents a method that utilizes multiple camera views for the gaze target estimation (GTE) task. The approach integrates information from different camera views to improve accuracy and expand applicability, addressing limitations in existing single-view methods that face challenges such as face occlusion, target ambiguity, and out-of-view targets. Our method processes a pair of camera views as input, incorporating a Head Information Aggregation (HIA) module for leveraging head information from both views for more accurate gaze estimation, an Uncertainty-based Gaze Selection (UGS) for identifying the most reliable gaze output, and an Epipolar-based Scene Attention (ESA) module for cross-view background information sharing. This approach significantly outperforms single-view baselines, especially when the second camera provides a clear view of the person's face. Additionally, our method can estimate the gaze target in the first view using the image of the person in the second view only, a capability not possessed by single-view GTE methods. Furthermore, the paper introduces a multi-view dataset for developing and evaluating multi-view GTE methods. Data and code are available at https://www3.cs.stonybrook.edu/~cvl/multiview_gte.html Qiaomu Miao, Vivek Raju Golani, Progga Paromita Dutta, Minh Hoai, Dimitris Samaras |
ICCV | 6 |
| 2025 | Importance-Based Token Merging for Efficient Image and Video Generation
Hieu Le 0001, Dimitris Samaras |
ICCV | 4 |
| 2025 | TopoDiffusionNet: A Topology-aware Diffusion ModelabstractDiffusion models excel at creating visually impressive images but often struggle to generate images with a specified topology. The Betti number, which represents the number of structures in an image, is a fundamental measure in topology. Yet, diffusion models fail to satisfy even this basic constraint. This limitation restricts their utility in applications requiring exact control, like robotics and environmental modeling. To address this, we propose TopoDiffusionNet (TDN), a novel approach that enforces diffusion models to maintain the desired topology. We leverage tools from topological data analysis, particularly persistent homology, to extract the topological structures within an image. We then design a topology-based objective function to guide the denoising process, preserving intended structures while suppressing noisy ones. Our experiments across four datasets demonstrate significant improvements in topological accuracy. TDN is the first to integrate topology with diffusion models, opening new avenues of research in this area. Saumya Gupta, Dimitris Samaras, Chao Chen 0012 |
ICLR | 2 |
| 2025 | Hummingbird: High Fidelity Image Generation via Multimodal Context AlignmentabstractWhile diffusion models are powerful in generating high-quality, diverse synthetic data for object-centric tasks, existing methods struggle with scene-aware tasks such as Visual Question Answering (VQA) and Human-Object Interaction (HOI) Reasoning, where it is critical to preserve scene attributes in generated images consistent with a multimodal context, i.e. a reference image with accompanying text guidance query. To address this, we introduce **Hummingbird**, the first diffusion-based image generator which, given a multimodal context, generates highly diverse images w.r.t. the reference image while ensuring high fidelity by accurately preserving scene attributes, such as object interactions and spatial relationships from the text guidance. Hummingbird employs a novel Multimodal Context Evaluator that simultaneously optimizes our formulated Global Semantic and Fine-grained Consistency Rewards to ensure generated images preserve the scene attributes of reference images in relation to the text guidance while maintaining diversity. As the first model to address the task of maintaining both diversity and fidelity given a multimodal context, we introduce a new benchmark formulation incorporating MME Perception and Bongard HOI datasets. Benchmark experiments show Hummingbird outperforms all existing methods by achieving superior fidelity while maintaining diversity, validating Hummingbird's potential as a robust multimodal context-aligned image generator in complex visual tasks. Project page: https://roar-ai.github.io/hummingbird Minh-Quan Le, Gaurav Mittal, Tianjian Meng, A S. M. Iftekhar, Vishwas Suryanarayanan, Barun Patra, Dimitris Samaras |
ICLR | 7 |
| 2025 | Pathology Image Compression with Pre-trained Autoencoders
Srikar Yellapragada, Alexandros Graikos, Kostas Triaridis, Zilinghan Li, Tarak Nath Nandi, Ravi K. Madduri, Prateek Prasanna, Joel H. Saltz, Dimitris Samaras |
MICCAI (2) | 9 |
| 2025 | Low-Rank Head Avatar Personalization with RegistersabstractWe introduce a novel method for low-rank personalization of a generic model for head avatar generation. Prior work proposes generic models that achieve high-quality face animation by leveraging large-scale datasets of multiple identities. However, such generic models usually fail to synthesize unique identity-specific details, since they learn a general domain prior. To adapt to specific subjects, we find that it is still challenging to capture high-frequency facial details via popular solutions like low-rank adaptation (LoRA). This motivates us to propose a specific architecture, a Register Module, that enhances the performance of LoRA, while requiring only a small number of parameters to adapt to an unseen identity. Our module is applied to intermediate features of a pre-trained model, storing and re-purposing information in a learnable 3D feature space. To demonstrate the efficacy of our personalization method, we collect a dataset of talking videos of individuals with distinctive facial details, such as wrinkles and tattoos. Our approach faithfully captures unseen faces, outperforming existing methods quantitatively and qualitatively. Sai Tanmay Reddy Chakkera, Aggelina Chatziagapi, Chen-Ping Yu, Yi-Hsuan Tsai, Dimitris Samaras |
NeurIPS | 6 |
| 2025 | Fast constrained sampling in pre-trained diffusion modelsabstractLarge denoising diffusion models, such as Stable Diffusion, have been trained on billions of image-caption pairs to perform text-conditioned image generation. As a byproduct of this training, these models have acquired general knowledge about image statistics, which can be useful for other inference tasks. However, when confronted with sampling an image under new constraints, e.g. generating the
missing parts of an image, using large pre-trained text-to-image diffusion models is inefficient and often unreliable. Previous approaches either utilized backpropagation through the denoiser network, making them significantly slower and more memory-demanding than simple text-to-image generation, or only enforced the constraint locally, failing to capture critical long-range correlations in the sampled
image. In this work, we propose an algorithm that enables fast, high-quality generation under arbitrary constraints. We show that in denoising diffusion models, we can employ an approximation to Newton’s optimization method that allows us to speed up inference and avoid the expensive backpropagation operations. Our approach produces results that rival or surpass the state-of-the-art training-free
inference methods while requiring a fraction of the time. We demonstrate the effectiveness of our algorithm under both linear (inpainting, super-resolution) and non-linear (style-guided generation) constraints. An implementation is provided at https://github.com/cvlab-stonybrook/fast-constrained-sampling. Alexandros Graikos, Nebojsa Jojic, Dimitris Samaras |
NeurIPS | 3 |
| 2025 | Shadow Removal Refinement via Material-Consistent Shadow EdgesabstractShadow boundaries can be confused with material boundaries as both exhibit sharp changes in luminance or contrast within a scene. However, shadows do not modify the intrinsic color or texture of surfaces. Therefore, on both sides of shadow edges traversing regions with the same material, the original color and texture should be the same if the shadow is removed properly. These shadow/shadow-free pairs are very useful but difficult-to-collect supervision signals. The crucial contribution of this paper is to learn how to identify those shadow edges that traverse material-consistent regions and how to use them as self-supervision for shadow removal refinement during test time. To achieve this, we fine-tune SAM, an image segmentation foundation model, to produce a shadow-invariant segmentation and then extract material-consistent shadow edges by comparing the SAM segmentation with the shadow mask. Utilizing these shadow edges, we in-troduce color- and texture-consistency losses to enhance the shadow removal process. We demonstrate the effectiveness of our method in improving shadow removal results on more challenging, in-the-wild images, outperforming the state-of-the-art shadow removal methods. Addition-ally, we propose a new metric and an annotated dataset for evaluating the performance of shadow removal methods without the need for paired shadow/shadow-free data. Our code and dataset are available at: https://github.com/cvlab-stonybrook/ShadowRemovalRefine Shilin Hu, Hieu Le 0001, Shahrukh Athar, Sagnik Das, Dimitris Samaras |
WACV | 5 |
| 2025 | Learning to Count from Pseudo-Labeled SegmentationabstractClass-agnostic counting (CAC) has numerous potential applications across various domains. The goal is to count objects of an arbitrary category during testing, based on only a few annotated exemplars. However, existing methods often count all objects in the image, including those from different categories than the exemplars. To address this issue, we propose localizing the area containing the objects of interest via an exemplar-based segmentation model before counting them. To train this model, we propose a novel method to obtain pseudo-labeled segmentation masks. Specifically, we use an unsupervised image clustering method to generate a set of candidate pseudo object masks, from which we select the optimal one using a pretrained CAC model. We show that the trained segmentation model can effectively localize objects of interest based on the exemplars and prevent the model from counting everything. To properly evaluate the performance of CAC methods in real-world scenarios, we introduce two new benchmarks: a synthetic test set and a new test set of real images containing countable objects from multiple classes. Our proposed method shows a significant advantage over previous CAC methods on these two benchmarks. Hieu Le 0001, Dimitris Samaras |
WACV | 3 |
| 2025 | Editorial: Introduction to the Special Section on Best of CVPR'2022
Kristin J. Dana, Gang Hua 0001, Stefan Roth 0001, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Controllable Dynamic Appearance for Neural 3D PortraitsabstractRecent advances in Neural Radiance Fields (NeRFs) have made it possible to reconstruct and reanimate dynamic portrait scenes with control over head-pose, facial expressions and viewing direction. However, training such models assumes photometric consistency over the deformed region e.g. the face must be evenly lit as it deforms with changing head-pose and facial expression. Such photometric consistency across frames of a video is hard to maintain, even in studio environments, thus making the created reanimatable neural portraits prone to artefacts during reanimation. In this work, we propose CoDyNeRF, a system that enables the creation of fully controllable 3D portraits in real-world capture conditions. CoDyNeRF learns to approximate illumination dependent effects via a dynamic appearance model in the canonical space that is conditioned on predicted surface normals and the facial expressions and head-pose deformations. The surface normals prediction is guided using 3DMM normals that act as a coarse prior for the normals of the human head, where direct prediction of normals is hard due to rigid and non-rigid deformations induced by head-pose and facial expression changes. Using only a smartphone-captured short video of a subject for training, we demonstrate the effectiveness of our method on free view synthesis of a portrait scene with explicit head pose and expression controls, and realistic lighting effects. Shahrukh Athar, Zhixin Shu, Zexiang Xu, Fujun Luan, Sai Bi, Kalyan Sunkavalli, Dimitris Samaras |
3DV | 7 |
| 2024 | JEAN: Joint Expression and Audio-guided NeRF-based Talking Face Generation
Sai Tanmay Reddy Chakkera, Aggelina Chatziagapi, Dimitris Samaras |
BMVC | 3 |
| 2024 | Unifying Top-Down and Bottom-Up Scanpath Prediction Using TransformersabstractMost models of visual attention aim at predicting either top-down or bottom-up control, as studied using different visual search and free-viewing tasks. In this paper we propose the Human Attention Transformer (HAT), a single model that predicts both forms of attention control. HAT uses a novel transformer-based architecture and a simplified foveated retina that collectively create a spatio-temporal awareness akin to the dynamic visual working memory of humans. HAT not only establishes a new state-of-the-art in predicting the scanpath of fixations made during target-present and target-absent visual search and "taskless" free viewing, but also makes human gaze behavior interpretable. Unlike previous methods that rely on a coarse grid of fixation cells and experience information loss due to fixation discretization, HAT features a sequential dense prediction architecture and outputs a dense heatmap for each fixation, thus avoiding discretizing fixations. HAT sets a new standard in computational attention, which emphasizes effectiveness, generality, and interpretability. HAT's demonstrated scope and applicability will likely inspire the development of new attention models that can better predict human behavior in various attention-demanding scenarios. Code is available at https://github.com/cvlab-stonybrook/HAT. Zhibo Yang 0002, Sounak Mondal, Seoyoung Ahn, Ruoyu Xue, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
CVPR | 7 |
| 2024 | Learned Representation-Guided Diffusion Models for Large-Image GenerationabstractTo synthesize high-fidelity samples, diffusion models typically require auxiliary data to guide the generation process. However, it is impractical to procure the painstaking patch-level annotation effort required in specialized domains like histopathology and satellite imagery; it is often performed by domain experts and involves hundreds of millions of patches. Modern-day self-supervised learning (SSL) representations encode rich semantic and visual information. In this paper, we posit that such representations are expressive enough to act as proxies to fine-grained human labels. We introduce a novel approach that trains diffusion models conditioned on embeddings from SSL. Our diffusion models successfully project these features back to high-quality histopathology and remote sensing images. In addition, we construct larger images by assembling spatially consistent patches inferred from SSL embeddings, preserving long-range dependencies. Augmenting real data by generating variations of real images improves downstream classifier accuracy for patch-level and larger, image-scale classification tasks. Our models are effective even on datasets not encountered during training, demonstrating their robustness and generalizability. Generating images from learned embeddings is agnostic to the source of the embeddings. The SSL embeddings used to generate a large image can either be extracted from a reference image, or sampled from an auxiliary model conditioned on any related modality (e.g. class labels, text, genomic data). As proof of concept, we introduce the text-to-large image synthesis paradigm where we successfully synthesize large pathology and satellite images out of text descriptions. Alexandros Graikos, Srikar Yellapragada, Minh-Quan Le, Saarthak Kapse, Prateek Prasanna, Joel H. Saltz, Dimitris Samaras |
CVPR | 7 |
| 2024 | SI-MIL: Taming Deep MIL for Self-Interpretability in Gigapixel HistopathologyabstractIntroducing interpretability and reasoning into Multiple Instance Learning (MIL) methods for Whole Slide Image (WSI) analysis is challenging, given the complexity of gigapixel slides. Traditionally, MIL interpretability is limited to identifying salient regions deemed pertinent for downstream tasks, offering little insight to the end-user (pathologist) regarding the rationale behind these selections. To address this, we propose Self-Interpretable MIL (SI-MIL), a method intrinsically designed for interpretability from the very outset. SI-MIL employs a deep MIL framework to guide an interpretable branch grounded on handcrafted pathological features, facilitating linear predictions. Beyond identifying salient regions, SI-MIL uniquely provides feature-level interpretations rooted in pathological insights for WSIs. Notably, SI-MIL, with its linear prediction constraints, challenges the prevalent myth of an inevitable trade-off between model interpretability and performance, demonstrating competitive results compared to state-of-the-art methods on WSI-level prediction tasks across three cancer types. In addition, we thoroughly benchmark the local-and global-interpretability of SI-MIL in terms of statistical analysis, a domain expert study, and desiderata of interpretability, namely, user-friendliness and faithfulness. Saarthak Kapse, Pushpak Pati, Srijan Das, Chao Chen 0012, Maria Vakalopoulou, Joel H. Saltz, Dimitris Samaras, Rajarsi Gupta 0001, Prateek Prasanna |
CVPR | 8 |
| 2024 | Self-supervised Co-salient Object Detection via Feature Correspondences at Multiple Scales
Souradeep Chakraborty, Dimitris Samaras |
ECCV (9) | 2 |
| 2024 | MIGS: Multi-Identity Gaussian Splatting via Tensor Decomposition
Aggelina Chatziagapi, Grigorios Chrysos 0002, Dimitris Samaras |
ECCV (24) | 3 |
| 2024 | Beyond Pixels: Semi-supervised Semantic Segmentation with a Multi-scale Patch-Based Multi-label Classifier
Prantik Howlader, Srijan Das, Hieu Le 0001, Dimitris Samaras |
ECCV (75) | 4 |
| 2024 | Weighting Pseudo-labels via High-Activation Feature Index Similarity and Object Detection for Semi-supervised Segmentation
Prantik Howlader, Hieu Le 0001, Dimitris Samaras |
ECCV (75) | 3 |
| 2024 | ∞-Brush: Controllable Large Image Synthesis with Diffusion Models in Infinite Dimensions
Minh-Quan Le, Alexandros Graikos, Srikar Yellapragada, Rajarsi Gupta 0001, Joel H. Saltz, Dimitris Samaras |
ECCV (32) | 6 |
| 2024 | Diffusion-Refined VQA Annotations for Semi-supervised Gaze Following
Qiaomu Miao, Alexandros Graikos, Sounak Mondal, Minh Hoai, Dimitris Samaras |
ECCV (39) | 6 |
| 2024 | Look Hear: Gaze Prediction for Speech-Directed Human Attention
Sounak Mondal, Seoyoung Ahn, Zhibo Yang 0002, Niranjan Balasubramanian, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai |
ECCV (42) | 5 |
| 2024 | Assessing Sample Quality via the Latent Space of Generative Models
Hieu Le 0001, Dimitris Samaras |
ECCV (59) | 3 |
| 2024 | Decoding the Visual Attention of Pathologists to Reveal Their Level of Expertise
Souradeep Chakraborty, Rajarsi Gupta 0001, Oksana Yaskiv, Constantin Friedman, Natallia Sheuka, Dana Perez, Paul Friedman, Gregory J. Zelinsky, Joel H. Saltz, Dimitris Samaras |
MICCAI (3) | 10 |
| 2024 | Unsupervised and semi-supervised co-salient object detection via segmentation frequency statisticsabstractIn this paper, we address the detection of co-occurring salient objects (CoSOD) in an image group using frequency statistics in an unsupervised manner, which further enable us to develop a semi-supervised method. While previous works have mostly focused on fully supervised CoSOD, less attention has been allocated to detecting co-salient objects when limited segmentation annotations are available for training. Our simple yet effective unsupervised method US-CoSOD combines the object co-occurrence frequency statistics of unsupervised single-image semantic segmentations with salient foreground detections using self-supervised feature learning. For the first time, we show that a large unlabeled dataset e.g. ImageNet-1k can be effectively leveraged to significantly improve unsupervised CoSOD performance. Our unsupervised model is a great pre-training initialization for our semi-supervised model SS-CoSOD, especially when very limited labeled data is available for training. To avoid propagating erroneous signals from predictions on unlabeled data, we propose a confidence estimation module to guide our semi-supervised training. Extensive experiments on three CoSOD benchmark datasets show that both of our unsupervised and semi-supervised models outperform the corresponding state-of-the-art models by a significant margin (e.g., on the Cosal2015 dataset, our US-CoSOD model has an 8.8% F-measure gain over a SOTA unsupervised co-segmentation model and our SS-CoSOD model has an 11.81% F-measure gain over a SOTA semi-supervised CoSOD model). Souradeep Chakraborty, Shujon Naha, Muhammet Bastan, Amit Kumar K. C, Dimitris Samaras |
WACV | 5 |
| 2024 | PathLDM: Text conditioned Latent Diffusion Model for HistopathologyabstractTo achieve high-quality results, diffusion models must be trained on large datasets. This can be notably prohibitive for models in specialized domains, such as computational pathology. Conditioning on labeled data is known to help in data-efficient model training. Therefore, histopathology reports, which are rich in valuable clinical information, are an ideal choice as guidance for a histopathology generative model. In this paper, we introduce PathLDM, the first text-conditioned Latent Diffusion Model tailored for generating high-quality histopathology images. Leveraging the rich contextual information provided by pathology text reports, our approach fuses image and textual data to enhance the generation process. By utilizing GPT's capabilities to distill and summarize complex text reports, we establish an effective conditioning mechanism. Through strategic conditioning and necessary architectural enhancements, we achieved a SoTA FID score of 7.64 for text-to-image generation on the TCGA-BRCA dataset, significantly outperforming the closest text-conditioned competitor with FID 30.1. Srikar Yellapragada, Alexandros Graikos, Prateek Prasanna, Tahsin M. Kurç, Joel H. Saltz, Dimitris Samaras |
WACV | 6 |
| 2024 | Attention De-sparsification Matters: Inducing diversity in digital pathology representation learning
Saarthak Kapse, Srijan Das, Rajarsi Gupta 0001, Joel H. Saltz, Dimitris Samaras, Prateek Prasanna |
Medical Image Anal. | 6 |
| 2023 | Conditional Generation from Pre-Trained Diffusion Models using Denoiser Representations
Alexandros Graikos, Srikar Yellapragada, Dimitris Samaras |
BMVC | 3 |
| 2023 | Topology-Guided Multi-Class Cell Context Generation for Digital PathologyabstractIn digital pathology, the spatial context of cells is important for cell classification, cancer diagnosis and prognosis. To model such complex cell context, however, is challenging. Cells form different mixtures, lineages, clusters and holes. To model such structural patterns in a learnable fashion, we introduce several mathematical tools from spatial statistics and topological data analysis. We incorporate such structural descriptors into a deep generative model as both conditional inputs and a differentiable loss. This way, we are able to generate high quality multi-class cell layouts for the first time. We show that the topology-rich cell layouts can be used for data augmentation and improve the performance of downstream tasks such as cell classification. Shahira Abousamra, Rajarsi Gupta 0001, Tahsin M. Kurç, Dimitris Samaras, Joel H. Saltz, Chao Chen 0012 |
CVPR | 4 |
| 2023 | AVFace: Towards Detailed Audio-Visual 4D Face ReconstructionabstractIn this work, we present a multimodal solution to the problem of 4D face reconstruction from monocular videos. 3D face reconstruction from 2D images is an under-constrained problem due to the ambiguity of depth. State-of-the-art methods try to solve this problem by leveraging visual information from a single image or video, whereas 3D mesh animation approaches rely more on audio. However, in most cases (e.g. AR/VR applications), videos include both visual and speech information. We propose AV-Face that incorporates both modalities and accurately re-constructs the 4D facial and lip motion of any speaker, without requiring any 3D ground truth for training. A coarse stage estimates the per-frame parameters of a 3D mor-phable model, followed by a lip refinement, and then a fine stage recovers facial geometric details. Due to the temporal audio and video information captured by transformer-based modules, our method is robust in cases when either modality is insufficient (e.g. face occlusions). Extensive qualitative and quantitative evaluation demonstrates the superiority of our method over the current state-of-the-art. Aggelina Chatziagapi, Dimitris Samaras |
CVPR | 2 |
| 2023 | Gazeformer: Scalable, Effective and Fast Prediction of Goal-Directed Human AttentionabstractPredicting human gaze is important in Human-Computer Interaction (HCI). However, to practically serve HCI applications, gaze prediction models must be scalable, fast, and accurate in their spatial and temporal gaze predictions. Recent scanpath prediction models focus on goaldirected attention (search). Such models are limited in their application due to a common approach relying on trained target detectors for all possible objects, and the availability of human gaze data for their training (both not scalable). In response, we pose a new task called ZeroGaze, a new variant of zero-shot learning where gaze is predicted for never-before-searched objects, and we develop a novel model, Gazeformer; to solve the ZeroGaze problem. In contrast to existing methods using object detector modules, Gazeformer encodes the target using a natural language model, thus leveraging semantic similarities in scanpath prediction. We use a transformer-based encoder-decoder architecture because transformers are particularly useful for generating contextual representations. Gazeformer surpasses other models by a large margin (19%-70%) on the ZeroGaze setting. It also outperforms existing target-detection models on standard gaze prediction for both target-present and target-absent search tasks. In addition to its improved performance, Gazeformer is more than five times faster than the state-of-the-art target-present visual search model. Code can be found at https://github.com/cvlab-stonybrook/Gazeformer/ Sounak Mondal, Zhibo Yang 0002, Seoyoung Ahn, Dimitris Samaras, Gregory J. Zelinsky, Minh Hoai |
CVPR | 4 |
| 2023 | Zero-Shot Object CountingabstractClass-agnostic object counting aims to count object instances of an arbitrary class at test time. Current methods for this challenging problem require human-annotated exemplars as inputs, which are often unavailable for novel categories, especially for autonomous systems. Thus, we propose zero-shot object counting (ZSC), a new setting where only the class name is available during test time. Such a counting system does not require human annotators in the loop and can operate automatically. Starting from a class name, we propose a method that can accurately identify the optimal patches which can then be used as counting exemplars. Specifically, we first construct a class prototype to select the patches that are likely to contain the objects of interest, namely class-relevant patches. Furthermore, we introduce a model that can quantitatively measure how suitable an arbitrary patch is as a counting exemplar. By applying this model to all the candidate patches, we can select the most suitable patches as exemplars for counting. Experimental results on a recent class-agnostic counting dataset, FSC-147, validate the effectiveness of our method. Code is available at https://github.com/cvlabstonybrook/zero-shot-counting. Hieu Le 0001, Vu Nguyen 0004, Viresh Ranjan, Dimitris Samaras |
CVPR | 5 |
| 2023 | Generating Features with Increased Crop-Related Diversity for Few-Shot Object DetectionabstractTwo-stage object detectors generate object proposals and classify them to detect objects in images. These proposals often do not contain the objects perfectly but overlap with them in many possible ways, exhibiting great variability in the difficulty levels of the proposals. Training a robust classifier against this crop-related variability requires abundant training data, which is not available in few-shot settings. To mitigate this issue, we propose a novel variational autoencoder (VAE) based data generation model, which is capable of generating data with increased croprelated diversity. The main idea is to transform the latent space such latent codes with different norms represent different crop-related variations. This allows us to generate features with increased crop-related diversity in difficulty levels by simply varying the latent norm. In particular, each latent code is rescaled such that its norm linearly correlates with the IoU score of the input crop w.r.t. the ground-truth box. Here the IoU score is a proxy that represents the difficulty level of the crop. We train this VAE model on base classes conditioned on the semantic code of each class and then use the trained model to generate features for novel classes. In our experiments our generated features consistently improve state-of-the-art few-shot object detection methods on the PASCAL VOC and MS COCO datasets. Hieu Le 0001, Dimitris Samaras |
CVPR | 3 |
| 2023 | FLAME-in-NeRF: Neural control of Radiance Fields for Free View Face AnimationabstractThis paper presents a neural rendering method for controllable portrait video synthesis. Recent advances in volumetric neural rendering, such as neural radiance fields (NeRF), have enabled the photorealistic novel view synthesis of static scenes with impressive results. However, modeling dynamic and controllable objects as part of a scene with such scene representations is still challenging. In this work, we design a system that enables 1) novel view synthesis for portrait video, of both the human subject and the scene they are in and 2) explicit control of the facial expressions through a low-dimensional expression representation. We represent the distribution of human facial expressions using the expression parameters of a 3D Morphable Model (3DMM) and condition the NeRF volumetric function on them. In order to guide the network to learn disentangled control for static scene appearance and dynamic facial actions, we impose a spatial prior via 3DMM fitting. We show the effectiveness of our method on free view synthesis of portrait videos with expression controls. To train a scene, our method only requires a short video of a subject captured by a mobile device. Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
FG | 3 |
| 2023 | LipNeRF: What is the right feature space to lip-sync a NeRF?abstractSynthesizing high-fidelity talking head videos of an arbitrary identity, lip-synced to a target speech segment, is a challenging problem. Recent GAN-based methods succeed by training a model on a large amount of videos, allowing the generator to learn a variety of audio-lip representations. However, they are unable to handle head pose changes. On the other hand, Neural Radiance Fields (NeRFs) model the 3D face geometry more accurately. Current audio-conditioned NeRFs are not as good in lip synchronization as GANs, since they are trained on limited video data of a single identity. In this work, we propose LipNeRF, a lip-syncing NeRF that bridges the gap between the accurate lip synchronization of GAN-based methods and the accurate 3D face modeling of NeRFs. LipNeRF is conditioned on the expression space of a 3DMM, instead of the audio feature space. We experimentally demonstrate that the expression space gives a better representation for the lip shape than the audio feature space. LipNeRF shows a significant improvement in lip-sync quality over the current state-of-the-art, especially in high-definition videos of cinematic content, with challenging pose, illumination and expression variations. Aggelina Chatziagapi, Shahrukh Athar, Rohith MV, Vimal Bhat, Dimitris Samaras |
FG | 6 |
| 2023 | S-VolSDF: Sparse Multi-View Stereo Regularization of Neural Implicit SurfacesabstractNeural rendering of implicit surfaces performs well in 3D vision applications. However, it requires dense input views as supervision. When only sparse input images are available, output quality drops significantly due to the shape-radiance ambiguity problem. We note that this ambiguity can be constrained when a 3D point is visible in multiple views, as is the case in multi-view stereo (MVS). We thus propose to regularize neural rendering optimization with an MVS solution. The use of an MVS probability volume and a generalized cross entropy loss leads to a noise-tolerant optimization process. In addition, neural rendering provides global consistency constraints that guide the MVS depth hypothesis sampling and thus improves MVS performance. Given only three sparse input views, experiments show that our method not only outperforms generic neural rendering models by a large margin but also significantly increases the reconstruction quality of MVS models. Alexandros Graikos, Dimitris Samaras |
ICCV | 3 |
| 2023 | Learning Probabilistic Topological Representations Using Discrete Morse Theory
Xiaoling Hu 0002, Dimitris Samaras, Chao Chen 0012 |
ICLR | 2 |
| 2023 | Prompt-MIL: Boosting Multi-instance Learning Schemes via Task-Specific Prompt Tuning
Saarthak Kapse, Ke Ma 0005, Prateek Prasanna, Joel H. Saltz, Maria Vakalopoulou, Dimitris Samaras |
MICCAI (8) | 7 |
| 2023 | Patch-level Gaze Distribution Prediction for Gaze FollowingabstractGaze following aims to predict where a person is looking in a scene, by predicting the target location, or indicating that the target is located outside the image. Recent works detect the gaze target by training a heatmap regression task with a pixel-wise mean-square error (MSE) loss, while formulating the in/out prediction task as a binary classification task. This training formulation puts a strict, pixel-level constraint in higher resolution on the single annotation available in training, and does not consider annotation variance and the correlation between the two subtasks. To address these issues, we introduce the patch distribution prediction (PDP) method. We replace the in/out prediction branch in previous models with the PDP branch, by predicting a patch-level gaze distribution that also considers the outside cases. Experiments show that our model regularizes the MSE loss by predicting better heatmap distributions on images with larger annotation variances, meanwhile bridging the gap between the target prediction and in/out prediction subtasks, showing a significant improvement in performance on both subtasks on public gaze following datasets. Qiaomu Miao, Minh Hoai, Dimitris Samaras |
WACV | 3 |
| 2023 | Predicting Visual Attention in Graphic Design DocumentsabstractWe present a model for predicting visual attention during the free viewing of graphic design documents. While existing works on this topic have aimed at predicting static saliency of graphic designs, our work is the first attempt to predict both spatial attention and dynamic temporal order in which the document regions are fixated by gaze using a deep learning based model. We propose a two-stage model for predicting dynamic attention on such documents, with webpages being our primary choice of document design for demonstration. In the first stage, we predict the saliency maps for each of the document components (e.g. logos, banners, texts, etc. for webpages) conditioned on the type of document layout. These component saliency maps are then jointly used to predict the overall document saliency. In the second stage, we use these layout-specific component saliency maps as the state representation for an inverse reinforcement learning model of fixation scanpath prediction during document viewing. To test our model, we collected a new dataset consisting of eye movements from 41 people freely viewing 450 webpages (the largest dataset of its kind). Experimental results show that our model outperforms existing models in both saliency and scanpath prediction for webpages, and also generalizes very well to other graphic design documents such as comics, posters, mobile UIs, etc. and natural images. Souradeep Chakraborty, Zijun Wei, Conor Kelton, Seoyoung Ahn, Aruna Balasubramanian, Gregory J. Zelinsky, Dimitris Samaras |
IEEE Trans. Multim. | 7 |
| 2022 | Learning an Isometric Surface Parameterization for Texture Unwrapping
Sagnik Das, Ke Ma 0001, Zhixin Shu, Dimitris Samaras |
ECCV (37) | 4 |
| 2022 | Target-Absent Human Attention
Zhibo Yang 0002, Sounak Mondal, Seoyoung Ahn, Gregory J. Zelinsky, Minh Hoai, Dimitris Samaras |
ECCV (4) | 6 |
| 2022 | Gigapixel Whole-Slide Images Classification Using Locally Supervised Learning
Ke Ma 0005, Rajarsi Gupta 0001, Joel H. Saltz, Maria Vakalopoulou, Dimitris Samaras |
MICCAI (2) | 7 |
| 2022 | Diffusion Models as Plug-and-Play PriorsabstractWe consider the problem of inferring high-dimensional data $x$ in a model that consists of a prior $p(x)$ and an auxiliary differentiable constraint $c(x,y)$ on $x$ given some additional information $y$. In this paper, the prior is an independently trained denoising diffusion generative model. The auxiliary constraint is expected to have a differentiable form, but can come from diverse sources. The possibility of such inference turns diffusion models into plug-and-play modules, thereby allowing a range of potential applications in adapting models to new domains and tasks, such as conditional generation or image segmentation. The structure of diffusion models allows us to perform approximate inference by iterating differentiation through the fixed denoising network enriched with different amounts of noise at each step. Considering many noised versions of $x$ in evaluation of its fitness is a novel search mechanism that may lead to new algorithms for solving combinatorial optimization problems. The code is available at https://github.com/AlexGraikos/diffusion_priors. Alexandros Graikos, Nikolay Malkin, Nebojsa Jojic, Dimitris Samaras |
NeurIPS | 4 |
| 2022 | Swift: Adaptive Video Streaming with Layered Neural Codecs
Mallesham Dasari, Kumara Kahatapitiya, Samir Ranjan Das, Aruna Balasubramanian, Dimitris Samaras |
NSDI | 5 |
| 2022 | Hierarchical Proxy-based Loss for Deep Metric LearningabstractProxy-based metric learning losses are superior to pair-based losses due to their fast convergence and low training complexity. However, existing proxy-based losses focus on learning class-discriminative features while overlooking the commonalities shared across classes which are potentially useful in describing and matching samples. Moreover, they ignore the implicit hierarchy of categories in real-world datasets, where similar subordinate classes can be grouped together. In this paper, we present a framework that leverages this implicit hierarchy by imposing a hierarchical structure on the proxies and can be used with any existing proxy-based loss. This allows our model to capture both class-discriminative features and class-shared characteristics without breaking the implicit data hierarchy. We evaluate our method on five established image retrieval datasets such as In-Shop and SOP. Results demonstrate that our hierarchical proxy-based loss framework improves the performance of existing proxy-based losses, especially on large datasets which exhibit strong hierarchical structure. Zhibo Yang 0002, Muhammet Bastan, Xinliang Zhu, Douglas Gray 0001, Dimitris Samaras |
WACV | 5 |
| 2022 | Artificial intelligence in drug discovery: applications and techniquesabstractArtificial intelligence (AI) has been transforming the practice of drug discovery in the past decade. Various AI techniques have been used in many drug discovery applications, such as virtual screening and drug design. In this survey, we first give an overview on drug discovery and discuss related applications, which can be reduced to two major tasks, i.e. molecular property prediction and molecule generation. We then present common data resources, molecule representations and benchmark platforms. As a major part of the survey, AI techniques are dissected into model architectures and learning paradigms. To reflect the technical development of AI in drug discovery over the years, the surveyed works are organized chronologically. We expect that this survey provides a comprehensive review on AI in drug discovery. We also provide a GitHub repository with a collection of papers (and codes, if applicable) as a learning resource, which is regularly updated. Jianyuan Deng, Zhibo Yang 0002, Iwao Ojima, Dimitris Samaras, Fusheng Wang 0001 |
Briefings Bioinform. | 4 |
| 2022 | Physics-Based Shadow Image Decomposition for Shadow RemovalabstractWe propose a novel deep learning method for shadow removal. Inspired by physical models of shadow formation, we use a linear illumination transformation to model the shadow effects in the image that allows the shadow image to be expressed as a combination of the shadow-free image, the shadow parameters, and a matte layer. We use two deep networks, namely SP-Net and M-Net, to predict the shadow parameters and the shadow matte respectively. This system allows us to remove the shadow effects from images. We then employ an inpainting network, I-Net, to further refine the results. We train and test our framework on the most challenging shadow removal dataset (ISTD). Our method improves the state-of-the-art in terms of mean absolute error (MAE) for the shadow area by 20%. Furthermore, this decomposition allows us to formulate a patch-based weakly-supervised shadow removal method. This model can be trained without any shadow- free images (that are cumbersome to acquire) and achieves competitive shadow removal results compared to state-of-the-art methods that are trained with fully paired shadow and shadow-free images. Last, we introduce SBU-Timelapse, a video shadow removal dataset for evaluating shadow removal methods. Hieu Le 0001, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | SIDER: Single-Image Neural Optimization for Facial Geometric Detail RecoveryabstractWe present SIDER (Single-Image neural optimization for facial geometric DEtail Recovery), a novel photometric optimization method that recovers detailed facial geometry from a single image in an unsupervised manner. Inspired by classical techniques of coarse-to-fine optimization and recent advances in implicit neural representations of 3D shape, SIDER combines a geometry prior based on statistical models and Signed Distance Functions (SDFs) to recover facial details from single images. First, it estimates a coarse geometry using a morphable model represented as an SDF. Next, it reconstructs facial geometry details by optimizing a photometric loss with respect to the ground truth image. In contrast to prior work, SIDER does not rely on any dataset priors and does not require additional supervision from multiple views, lighting changes or ground truth 3D shape. Extensive qualitative and quantitative evaluation demonstrates that our method achieves state-of-the-art on facial geometric detail recovery, using only a single in the-wild image. Aggelina Chatziagapi, Shahrukh Athar, Francesc Moreno-Noguer, Dimitris Samaras |
3DV | 4 |
| 2021 | Localization in the Crowd with Topological ConstraintsabstractWe address the problem of crowd localization, i.e., the prediction of dots corresponding to people in a crowded scene. Due to various challenges, a localization method is prone to spatial semantic errors, i.e., predicting multiple dots within a same person or collapsing multiple dots in a cluttered region. We propose a topological approach targeting these semantic errors. We introduce a topological constraint that teaches the model to reason about the spatial arrangement of dots. To enforce this constraint, we define a persistence loss based on the theory of persistent homology. The loss compares the topographic landscape of the likelihood map and the topology of the ground truth. Topological reasoning improves the quality of the localization algorithm especially near cluttered regions. On multiple public benchmarks, our method outperforms previous localization methods. Additionally, we demonstrate the potential of our method in improving the performance in the crowd counting task. Shahira Abousamra, Minh Hoai, Dimitris Samaras, Chao Chen 0012 |
AAAI | 3 |
| 2021 | Modeling Deep Learning Based Privacy Attacks on Physical Mail
Bingyao Huang, Ruyi Lian, Dimitris Samaras, Haibin Ling |
AAAI | 3 |
| 2021 | Multi-Class Cell Detection Using Spatial Context RepresentationabstractIn digital pathology, both detection and classification of cells are important for automatic diagnostic and prognostic tasks. Classifying cells into subtypes, such as tumor cells, lymphocytes or stromal cells is particularly challenging. Existing methods focus on morphological appearance of individual cells, whereas in practice pathologists often infer cell classes through their spatial context. In this paper, we propose a novel method for both detection and classification that explicitly incorporates spatial contextual information. We use the spatial statistical function to describe local density in both a multi-class and a multi-scale manner. Through representation learning and deep clustering techniques, we learn advanced cell representation with both appearance and spatial context. On various benchmarks, our method achieves better performance than state-of-the-arts, especially on the classification task. We also create a new dataset for multi-class cell detection and classification in breast cancer and we make both our code and data publicly available. Shahira Abousamra, David Belinsky, John S. Van Arnam, Felicia Allard, Eric Yee, Rajarsi Gupta 0001, Tahsin M. Kurç, Dimitris Samaras, Joel H. Saltz, Chao Chen 0012 |
ICCV | 8 |
| 2021 | End-to-end Piece-wise Unwarping of Document ImagesabstractDocument unwarping attempts to undo physical deformations of the paper and recover a ’flatbed’ scanned document-image for downstream tasks such as OCR. Current state-of-the-art relies on global unwarping of the document which is not robust to local deformation changes. Moreover, a global unwarping often produces spurious warping artifacts in less warped regions to compensate for severe warps present in other parts of the document. In this paper, we propose the first end-to-end trainable piece-wise unwarping1method that predicts local deformation fields and stitches them together with global information to obtain an improved unwarping. The proposed piece-wise formulation results in 4% improvement in terms of multi-scale structural similarity (MS-SSIM) and shows better performance in terms of OCR metrics, character error rate (CER) and word error rate (WER) compared to the state-of-the-art. Sagnik Das, Kunwar Yashraj Singh, Jon Wu, Erhan Bas, Vijay Mahadevan, Rahul Bhotika, Dimitris Samaras |
ICCV | 7 |
| 2021 | Variational Feature Disentangling for Fine-Grained Few-Shot ClassificationabstractData augmentation is an intuitive step towards solving the problem of few-shot classification. However, ensuring both discriminability and diversity in the augmented samples is challenging. To address this, we propose a feature disentanglement framework that allows us to augment features with randomly sampled intra-class variations while preserving their class-discriminative features. Specifically, we disentangle a feature representation into two components: one represents the intra-class variance and the other encodes the class-discriminative information. We assume that the intra-class variance induced by variations in poses, backgrounds, or illumination conditions is shared across all classes and can be modelled via a common distribution. Then we sample features repeatedly from the learned intra-class variability distribution and add them to the class-discriminative features to get the augmented features. Such a data augmentation scheme ensures that the augmented features inherit crucial class-discriminative features while exhibiting large intra-class variance. Our method significantly outperforms the state-of-the-art methods on multiple challenging fine-grained few-shot image classification benchmarks. Code is available at: https://github.com/cvlab-stonybrook/vfd-iccv21 Hieu Le 0001, Mingzhen Huang, Shahrukh Athar, Dimitris Samaras |
ICCV | 5 |
| 2021 | Topology-Aware Segmentation Using Discrete Morse Theory
Xiaoling Hu 0002, Yusu Wang 0001, Fuxin Li, Dimitris Samaras, Chao Chen 0012 |
ICLR | 4 |
| 2021 | Attention Based CNN-LSTM Network for Pulmonary Embolism Prediction on Chest Computed Tomography Pulmonary Angiograms
Sudhir Suman, Nicole Sakla, Rishabh Gattu, Jeremy Green, Tej Phatak, Dimitris Samaras, Prateek Prasanna |
MICCAI (7) | 7 |
| 2021 | Chest Radiograph Disentanglement for COVID-19 Outcome Prediction
Joseph Bae, Huidong Liu, Jeremy Green, Dimitris Samaras, Prateek Prasanna |
MICCAI (7) | 6 |
| 2021 | Large Scale Shadow Annotation and Detection Using Lazy Annotation and Stacked CNNsabstractRecent shadow detection algorithms have shown initial success on small datasets of images from specific domains. However, shadow detection on broader image domains is still challenging due to the lack of annotated training data, caused by the intense manual labor required for annotating shadow data. In this paper we propose "lazy annotation", an efficient annotation method where an annotator only needs to mark the important shadow areas and some non-shadow areas. This yields data with noisy labels that are not yet useful for training a shadow detector. We address the problem of label noise by jointly learning a shadow region classifier and recovering the labels in the training set. We consider the training labels as unknowns and formulate label recovery as the minimization of the sum of squared leave-one-out errors of a Least Squares SVM, which can be efficiently optimized. Experimental results show that a classifier trained with recovered labels achieves comparable performance to a classifier trained on the properly annotated data. These results motivated us to collect a new dataset that is 20 times larger than existing datasets and contains a large variety of scenes and image types. Naturally, such a large dataset is appropriate for training deep learning methods. Thus, we propose a stacked Convolutional Neural Network architecture that efficiently trains on patch level shadow examples while incorporating image level semantic information. This means that the detected shadow patches are refined based on image semantics. Our proposed pipeline, trained on recovered labels, performs at state-of-the art level. Furthermore, the proposed model performs exceptionally well on a cross dataset task, proving the generalization power of the proposed architecture and dataset. Le Hou, Tomás F. Yago Vicente, Minh Hoai, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2021 | Sequence-to-Segments Networks for Detecting Segments in VideosabstractDetecting segments of interest from videos is a common problem for many applications. And yet it is a challenging problem as it often requires not only knowledge of individual target segments, but also contextual understanding of the entire video and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segments Network (S2N), a novel and general end-to-end sequential encoder-decoder architecture. S2N first encodes the input video into a sequence of hidden states that capture information progressively, as it appears in the video. It then employs the Segment Detection Unit (SDU), a novel decoding architecture, that sequentially detects segments. At each decoding step, the SDU integrates the decoder state and encoder hidden states to detect a target segment. During training, we address the problem of finding the best assignment of predicted segments to ground truth using the Hungarian Matching Algorithm with Lexicographic Cost. Additionally we propose to use the squared Earth Mover's Distance to optimize the localization errors of the segments. We show the state-of-the-art performance of S2N across numerous tasks, including video highlighting, video summarization, and human action proposal generation. Zijun Wei, Boyu Wang 0001, Minh Hoai, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Radomír Mech, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 8 |
| 2021 | Mosaic: Advancing User Quality of Experience in 360-Degree Video Streaming With Machine LearningabstractConventional streaming solutions for streaming 360-degree panoramic videos are inefficient in that they download the entire 360-degree panoramic scene, while the user views only a small sub-part of the scene called the viewport. This can waste over 80% of the network bandwidth. We develop a comprehensive approach called Mosaic that combines a powerful neural network-based viewport prediction with a rate control mechanism that assigns rates to different tiles in the 360-degree frame such that the video quality of experience is optimized subject to a given network capacity. We model the optimization as a multi-choice knapsack problem and solve it using a greedy approach. We also develop an end-to-end testbed using standards-compliant components and provide a comprehensive performance evaluation of Mosaic along with five other streaming techniques - two for conventional adaptive video streaming and three for 360-degree tile-based video streaming. Mosaic outperforms the best of the competitions by as much as 47-191% in terms of average video quality of experience. Simulation-based evaluation as well as subjective user studies further confirm the superiority of the proposed approach. Sohee Kim Park, Arani Bhattacharya, Zhibo Yang 0002, Samir Ranjan Das, Dimitris Samaras |
IEEE Trans. Netw. Serv. Manag. | 5 |
| 2020 | Intrinsic Decomposition of Document Images In-the-Wild
Sagnik Das, Hassan Ahmed Sial, Ke Ma 0005, Ramón Baldrich, María Vanrell 0001, Dimitris Samaras |
BMVC | 6 |
| 2020 | Learning Visual Emotion Representations From Web DataabstractWe present a scalable approach for learning powerful visual features for emotion recognition. A critical bottleneck in emotion recognition is the lack of large scale datasets that can be used for learning visual emotion features. To this end, we curate a webly derived large scale dataset, StockEmotion, which has more than a million images. StockEmotion uses 690 emotion related tags as labels giving us a fine-grained and diverse set of emotion labels, circumventing the difficulty in manually obtaining emotion annotations. We use this dataset to train a feature extraction network, EmotionNet, which we further regularize using joint text and visual embedding and text distillation. Our experimental results establish that EmotionNet trained on the StockEmotion dataset outperforms SOTA models on four different visual emotion tasks. An aded benefit of our joint embedding training approach is that EmotionNet achieves competitive zero-shot recognition performance against fully supervised baselines on a challenging visual emotion dataset, EMOTIC, which further highlights the generalizability of the learned emotion features. Zijun Wei, Jianming Zhang 0001, Zhe Lin 0001, Joon-Young Lee, Niranjan Balasubramanian, Minh Hoai, Dimitris Samaras |
CVPR | 7 |
| 2020 | Predicting Goal-Directed Human Attention Using Inverse Reinforcement LearningabstractHuman gaze behavior prediction is important for behavioral vision and for computer vision applications. Most models mainly focus on predicting free-viewing behavior using saliency maps, but do not generalize to goal-directed behavior, such as when a person searches for a visual target object. We propose the first inverse reinforcement learning (IRL) model to learn the internal reward function and policy used by humans during visual search. We modeled the viewer's internal belief states as dynamic contextual belief maps of object locations. These maps were learned and then used to predict behavioral scanpaths for multiple target categories. To train and evaluate our IRL model we created COCO-Search18, which is now the largest dataset of high-quality search fixations in existence. COCO-Search18 has 10 participants searching for each of 18 target-object categories in 6202 images, making about 300,000 goal-directed fixations. When trained and evaluated on COCO-Search18, the IRL model outperformed baseline models in predicting search fixation scanpaths, both in terms of similarity to human search behavior and search efficiency. Finally, reward maps recovered by the IRL model reveal distinctive target-dependent patterns of object prioritization, which we interpret as a learned object context. Zhibo Yang 0002, Lihan Huang, Yupei Chen, Zijun Wei, Seoyoung Ahn, Gregory J. Zelinsky, Dimitris Samaras, Minh Hoai |
CVPR | 7 |
| 2020 | From Shadow Segmentation to Shadow Removal
Hieu Le 0001, Dimitris Samaras |
ECCV (11) | 2 |
| 2020 | TopoGAN: A Topology-Aware Generative Adversarial Network
Fan Wang 0010, Huidong Liu, Dimitris Samaras, Chao Chen 0012 |
ECCV (3) | 3 |
| 2020 | Self-supervised Deformation Modeling for Facial Expression EditingabstractDeep generative models have recently demonstrated impressive results in photo-realistic facial image synthesis and editing. Existing neural network-based approaches usually only rely on texture generation to edit expressions and largely neglect the motion information. However, facial expressions are inherently the result of muscle movement. In this work, we propose a novel end-to-end network that disentangles the task of facial editing into two steps: a “motionediting” step and a “texture-editing” step. In the “motion-editing” step, we explicitly model facial movement through an image deformation, warping the image into the desired expression. In the “texture-editing” step, we generate the necessary textures, such as teeth and shading effects, for a photorealistic result. Our physically-based task-disentanglement system design allows each step to learn a focused task, and thus need not generate texture to hallucinate motion. Our system is trained in a self-supervised manner, requiring no ground truth deformation annotation. Using Action Units [8] as the representation for facial expression, our method improves the state-of-the-art facial expression editing performance in both qualitative and quantitative evaluations. Shahrukh Athar, Zhixin Shu, Dimitris Samaras |
FG | 3 |
| 2020 | Nonverbal Behavioral Patterns Predict Social Rejection Elicited Aggression
Megan Quarmley, Zhibo Yang 0002, Shahrukh Athar, Gregory J. Zelinsky, Dimitris Samaras, Johanna M. Jarcho |
FG | 5 |
| 2020 | Learning Monocular Face Reconstruction using Multi-View SupervisionabstractWe present a method to reconstruct faces from a single portrait image. While traditional face reconstruction methods fit low-dimensional 3D morphable models to images, we train a deep network to regress depth from a single image directly. We do so by combining supervised losses on synthetic data with indirect supervision on real data using a novel multi-view photo-consistency loss. Furthermore, we regularize the depth estimation using a 3D morphable model (3DMM). We demonstrate that this leads to results that preserve facial features, capture facial geometry that goes beyond 3DMMs, and is also robust to viewpoint conditions. We evaluate our method on various datasets and via ablation studies, and demonstrate that it outperforms previous work significantly. Zhixin Shu, Duygu Ceylan, Kalyan Sunkavalli, Eli Shechtman, Sunil Hadap, Dimitris Samaras |
FG | 6 |
| 2020 | Distribution Matching for Crowd CountingabstractIn crowd counting, each training image contains multiple people, where each person is annotated by a dot. Existing crowd counting methods need to use a Gaussian to smooth each annotated dot or to estimate the likelihood of every pixel given the annotated point. In this paper, we show that imposing Gaussians to annotations hurts generalization performance. Instead, we propose to use Distribution Matching for crowd COUNTing (DM-Count). In DM-Count, we use Optimal Transport (OT) to measure the similarity between the normalized predicted density map and the normalized ground truth density map. To stabilize OT computation, we include a Total Variation loss in our model. We show that the generalization error bound of DM-Count is tighter than that of the Gaussian smoothed methods. In terms of Mean Absolute Error, DM-Count outperforms the previous state-of-the-art methods by a large margin on two large-scale counting datasets, UCF-QNRF and NWPU, and achieves the state-of-the-art results on the ShanghaiTech and UCF-CC50 datasets. DM-Count reduced the error of the state-of-the-art published result by approximately 16%. Code is available at https://github.com/cvlab-stonybrook/DM-Count. Boyu Wang 0001, Huidong Liu, Dimitris Samaras, Minh Hoai |
NeurIPS | 3 |
| 2019 | Exascale Deep Learning to Accelerate Cancer ResearchabstractDeep learning, through the use of neural networks, has demonstrated remarkable ability to automate many routine tasks when presented with sufficient data for training. The neural network architecture (e.g. number of layers, types of layers, connections between layers, etc.) plays a critical role in determining what, if anything, the neural network is able to learn from the training data. The trend for neural network architectures, especially those trained on ImageNet, has been to grow ever deeper and more complex. The result has been ever increasing accuracy on benchmark datasets with the cost of increased computational demands. In this paper we demonstrate that neural network architectures can be automatically generated, tailored for a specific application, with dual objectives: accuracy of prediction and speed of prediction. Using MENNDL- an HPC-enabled software stack for neural architecture search-we generate a neural network with comparable accuracy to state-of-the-art networks on a cancer pathology dataset that is also 16× faster at inference. The speedup in inference is necessary because of the volume and velocity of cancer pathology data; specifically, the previous state-of-the-art networks are too slow for individual researchers without access to HPC systems to keep pace with the rate of data generation. Our new model enables researchers with modest computational resources to analyze newly generated data faster than it is collected. Robert M. Patton, Shahira Abousamra, Dimitris Samaras, Joel H. Saltz, J. Travis Johnston, Steven R. Young, Catherine D. Schuman, Thomas E. Potok, Derek C. Rose, Seung-Hwan Lim, Junghoon Chae, Le Hou |
IEEE BigData | 3 |
| 2019 | SmartEye: Assisting Instant Photo Taking via Integrating User Preference with Deep View Proposal NetworkabstractInstant photo taking and sharing has become one of the most popular forms of social networking. However, taking high-quality photos is difficult as it requires knowledge and skill in photography that most non-expert users lack. In this paper we present SmartEye, a novel mobile system to help users take photos with good compositions in-situ. The back-end of SmartEye integrates the View Proposal Network (VPN), a deep learning based model that outputs composition suggestions in real time, and a novel, interactively updated module (P-Module) that adjusts the VPN outputs to account for personalized composition preferences. We also design a novel interface with functions at the front-end to enable real-time and informative interactions for photo taking. We conduct two user studies to investigate SmartEye qualitatively and quantitatively. Results show that SmartEye effectively models and predicts personalized composition preferences, provides instant high-quality compositions in-situ, and outperforms the non-personalized systems significantly. Shuai Ma 0005, Zijun Wei, Feng Tian 0001, Xiangmin Fan, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Jin Huang 0009, Radomír Mech, Dimitris Samaras, Hongan Wang |
CHI | 10 |
| 2019 | Robust Histopathology Image Analysis: To Label or to Synthesize?abstractDetection, segmentation and classification of nuclei are fundamental analysis operations in digital pathology. Existing state-of-the-art approaches demand extensive amount of supervised training data from pathologists and may still perform poorly in images from unseen tissue types. We propose an unsupervised approach for histopathology image segmentation that synthesizes heterogeneous sets of training image patches, of every tissue type. Although our synthetic patches are not always of high quality, we harness the motley crew of generated samples through a generally applicable importance sampling method. This proposed approach, for the first time, re-weighs the training loss over synthetic data so that the ideal (unbiased) generalization loss over the true data distribution is minimized. This enables us to use a random polygon generator to synthesize approximate cellular structures (i.e., nuclear masks) for which no real examples are given in many tissue types, and hence, GAN-based methods are not suited. In addition, we propose a hybrid synthesis pipeline that utilizes textures in real histopathology patches and GAN models, to tackle heterogeneity in tissue textures. Compared with existing state-of-the-art supervised models, our approach generalizes significantly better on cancer types without training data. Even in cancer types with training data, our approach achieves the same performance without supervision cost. We release code and segmentation results on over 5000 Whole Slide Images (WSI) in The Cancer Genome Atlas (TCGA) repository, a dataset that would be orders of magnitude larger than what is available today. Le Hou, Ayush Agarwal, Dimitris Samaras, Tahsin M. Kurç, Rajarsi Gupta 0001, Joel H. Saltz |
CVPR | 3 |
| 2019 | Reading detection in real-timeabstractObservable reading behavior, the act of moving the eyes over lines of text, is highly stereotyped among the users of a language, and this has led to the development of reading detectors-methods that input windows of sequential fixations and output predictions of the fixation behavior during those windows being reading or skimming. The present study introduces a new method for reading detection using Region Ranking SVM (RRSVM). An SVM-based classifier learns the local oculomotor features that are important for real-time reading detection while it is optimizing for the global reading/skimming classification, making it unnecessary to hand-label local fixation windows for model training. This RRSVM reading detector was trained and evaluated using eye movement data collected in a laboratory context, where participants viewed modified web news articles and had to either read them carefully for comprehension or skim them quickly for the selection of keywords (separate groups). Ground truth labels were known at the global level (the instructed reading or skimming task), and obtained at the local level in a separate rating task. The RRSVM reading detector accurately predicted 82.5% of the global (article-level) reading/skimming behavior, with accuracy in predicting local window labels ranging from 72-95%, depending on how tuned the RRSVM was for local and global weights. With this RRSVM reading detector, a method now exists for near real-time reading detection without the need for hand-labeling of local fixation windows. With real-time reading detection capability comes the potential for applications ranging from education and training to intelligent interfaces that learn what a user is likely to know based on previous detection of their reading behavior. Conor Kelton, Zijun Wei, Seoyoung Ahn, Aruna Balasubramanian, Samir Ranjan Das, Dimitris Samaras, Gregory J. Zelinsky |
ETRA | 6 |
| 2019 | DewarpNet: Single-Image Document Unwarping With Stacked 3D and 2D Regression NetworksabstractCapturing document images with hand-held devices in unstructured environments is a common practice nowadays. However, "casual" photos of documents are usually unsuitable for automatic information extraction, mainly due to physical distortion of the document paper, as well as various camera positions and illumination conditions. In this work, we propose DewarpNet, a deep-learning approach for document image unwarping from a single image. Our insight is that the 3D geometry of the document not only determines the warping of its texture but also causes the illumination effects. Therefore, our novelty resides on the explicit modeling of 3D shape for document paper in an end-to-end pipeline. Also, we contribute the largest and most comprehensive dataset for document image unwarping to date - Doc3D. This dataset features multiple ground-truth annotations, including 3D shape, surface normals, UV map, albedo image, etc. Training with Doc3D, we demonstrate state-of-the-art performance for DewarpNet with extensive qualitative and quantitative evaluations. Our network also significantly improves OCR performance on captured document images, decreasing character error rate by 42% on average. Both the code and the dataset are released. Sagnik Das, Ke Ma 0005, Zhixin Shu, Dimitris Samaras, Roy Shilkrot |
ICCV | 4 |
| 2019 | Shadow Removal via Shadow Image DecompositionabstractWe propose a novel deep learning method for shadow removal. Inspired by physical models of shadow formation, we use a linear illumination transformation to model the shadow effects in the image that allows the shadow image to be expressed as a combination of the shadow-free image, the shadow parameters, and a matte layer. We use two deep networks, namely SP-Net and M-Net, to predict the shadow parameters and the shadow matte respectively. This system allows us to remove the shadow effects on the images. We train and test our framework on the most challenging shadow removal dataset (ISTD). Compared to the state-of-the-art method, our model achieves a 40% error reduction in terms of root mean square error (RMSE) for the shadow area, reducing RMSE from 13.3 to 7.9. Moreover, we create an augmented ISTD dataset based on an image decomposition system by modifying the shadow parameters to generate new synthetic shadow images. Training our model on this new augmented ISTD dataset further lowers the RMSE on the shadow area to 7.4. Hieu Le 0001, Dimitris Samaras |
ICCV | 2 |
| 2019 | Wasserstein GAN With Quadratic Transport CostabstractWasserstein GANs are increasingly used in Computer Vision applications as they are easier to train. Previous WGAN variants mainly use the lι transport cost to compute the Wasserstein distance between the real and synthetic data distributions. The lι transport cost restricts the discriminator to be 1-Lipschitz. However, WGANs with lι transport cost were recently shown to not always converge. In this paper, we propose WGAN-QC, a WGAN with quadratic transport cost. Based on the quadratic transport cost, we propose an Optimal Transport Regularizer (OTR) to stabilize the training process of WGAN-QC. We prove that the objective of the discriminator during each generator update computes the exact quadratic Wasserstein distance between real and synthetic data distributions. We also prove that WGAN-QC converges to a local equilibrium point with finite discriminator updates per generator update. We show experimentally on a Dirac distribution that WGAN-QC converges, when many of the lι cost WGANs fail to [22]. Qualitative and quantitative results on the CelebA, CelebA-HQ, LSUN and the ImageNet dog datasets show that WGAN-QC is better than state-of-art GAN methods. WGAN-QC has much faster runtime than other WGAN variants. Huidong Liu, Xianfeng Gu, Dimitris Samaras |
ICCV | 3 |
| 2019 | Label super-resolution networks
Kolya Malkin, Caleb Robinson, Le Hou, Rachel Soobitsky, Jacob Czawlytko, Dimitris Samaras, Joel H. Saltz, Lucas Joppa, Nebojsa Jojic |
ICLR (Poster) | 6 |
| 2019 | Pancreatic Cancer Detection in Whole Slide Images Using Noisy Label Annotations
Han Le, Dimitris Samaras, Tahsin M. Kurç, Rajarsi Gupta 0001, Kenneth Shroyer, Joel H. Saltz |
MICCAI (1) | 2 |
| 2019 | Advancing User Quality of Experience in 360-degree Video StreamingabstractConventional streaming solutions for streaming 360-degree panoramic videos are inefficient in that they download the entire 360-degree panoramic scene, while the user views only a small sub-part of the scene called the viewport. This can waste over 80% of the network bandwidth. We develop a comprehensive approach called Mosaic that combines a powerful neural network-based viewport prediction with a rate control mechanism that assigns rates to different tiles in the 360-degree frame such that the video quality of experience is optimized subject to a given network capacity. We model the optimization as a multi-choice knapsack problem and solve it using a greedy approach. We also develop an end-to-end testbed using standards-compliant components and provide a comprehensive performance evaluation of Mosaic along with four other streaming techniques - two for conventional adaptive video streaming and two for 360-degree tile-based video streaming. Mosaic outperforms the best of the competition by as much as 50% in terms of median video quality. Sohee Kim Park, Arani Bhattacharya, Zhibo Yang 0002, Mallesham Dasari, Samir Ranjan Das, Dimitris Samaras |
Networking | 6 |
| 2019 | Topology-Preserving Deep Image SegmentationabstractSegmentation algorithms are prone to make topological errors on fine-scale struc- tures, e.g., broken connections. We propose a novel method that learns to segment with correct topology. In particular, we design a continuous-valued loss function that enforces a segmentation to have the same topology as the ground truth, i.e.,having the same Betti number. The proposed topology-preserving loss function is differentiable and can be incorporated into end-to-end training of a deep neural network. Our method achieves much better performance on the Betti number error, which directly accounts for the topological correctness. It also performs superior on other topology-relevant metrics, e.g., the Adjusted Rand Index and the Variation of Information, without sacrificing per-pixel accuracy. We illustrate the effectiveness of the proposed method on a broad spectrum of natural and biomedical datasets. Xiaoling Hu 0002, Fuxin Li, Dimitris Samaras, Chao Chen 0012 |
NeurIPS | 3 |
| 2019 | An Adversarial Neuro-Tensorial Approach for Learning Disentangled RepresentationsabstractSeveral factors contribute to the appearance of an object in a visual scene, including pose, illumination, and deformation, among others. Each factor accounts for a source of variability in the data, while the multiplicative interactions of these factors emulate the entangled variability, giving rise to the rich structure of visual object appearance. Disentangling such unobserved factors from visual data is a challenging task, especially when the data have been captured in uncontrolled recording conditions (also referred to as “in-the-wild”) and label information is not available. In this paper, we propose a pseudo-supervised deep learning method for disentangling multiple latent factors of variation in face images captured in-the-wild. To this end, we propose a deep latent variable model, where the multiplicative interactions of multiple latent factors of variation are explicitly modelled by means of multilinear (tensor) structure. We demonstrate that the proposed approach indeed learns disentangled representations of facial expressions and pose, which can be used in various applications, including face editing, as well as 3D face reconstruction and classification of facial expression, identity and pose. Mengjiao Wang 0002, Zhixin Shu, Shiyang Cheng 0001, Yannis Panagakis, Dimitris Samaras, Stefanos Zafeiriou |
Int. J. Comput. Vis. | 5 |
| 2019 | Sparse autoencoder for unsupervised nucleus detection and representation in histopathology images
Le Hou, Vu Nguyen 0004, Ariel B. Kanevsky, Dimitris Samaras, Tahsin M. Kurç, Rajarsi Gupta 0001, Yi Gao 0002, Wenjin Chen, David J. Foran, Joel H. Saltz |
Pattern Recognit. | 4 |
| 2018 | DocUNet: Document Image Unwarping via a Stacked U-NetabstractCapturing document images is a common way for digitizing and recording physical documents due to the ubiquitousness of mobile cameras. To make text recognition easier, it is often desirable to digitally flatten a document image when the physical document sheet is folded or curved. In this paper, we develop the first learning-based method to achieve this goal. We propose a stacked U-Net [25] with intermediate supervision to directly predict the forward mapping from a distorted image to its rectified version. Because large-scale real-world data with ground truth deformation is difficult to obtain, we create a synthetic dataset with approximately 100 thousand images by warping non-distorted document images. The network is trained on this dataset with various data augmentations to improve its generalization ability. We further create a comprehensive benchmark1 that covers various real-world conditions. We evaluate the proposed model quantitatively and qualitatively on the proposed benchmark, and compare it with previous non-learning-based methods. Ke Ma 0001, Zhixin Shu, Jue Wang 0001, Dimitris Samaras |
CVPR | 5 |
| 2018 | Good View Hunting: Learning Photo Composition From Dense View PairsabstractFinding views with good photo composition is a challenging task for machine learning methods. A key difficulty is the lack of well annotated large scale datasets. Most existing datasets only provide a limited number of annotations for good views, while ignoring the comparative nature of view selection. In this work, we present the first large scale Comparative Photo Composition dataset, which contains over one million comparative view pairs annotated using a cost-effective crowdsourcing workflow. We show that these comparative view annotations are essential for training a robust neural network model for composition. In addition, we propose a novel knowledge transfer framework to train a fast view proposal network, which runs at 75+ FPS and achieves state-of-the-art performance in image cropping and thumbnail generation tasks on three benchmark datasets. The superiority of our method is also demonstrated in a user study on a challenging experiment, where our method significantly outperforms the baseline methods in producing diversified well-composed views. Zijun Wei, Jianming Zhang 0001, Xiaohui Shen, Zhe Lin 0001, Radomír Mech, Minh Hoai, Dimitris Samaras |
CVPR | 7 |
| 2018 | A+D Net: Training a Shadow Detector with Adversarial Shadow Attenuation
Hieu Le 0001, Tomás F. Yago Vicente, Vu Nguyen 0004, Minh Hoai, Dimitris Samaras |
ECCV (2) | 5 |
| 2018 | Deforming Autoencoders: Unsupervised Disentangling of Shape and Appearance
Zhixin Shu, Mihir Sahasrabudhe, Riza Alp Güler, Dimitris Samaras, Nikos Paragios, Iasonas Kokkinos |
ECCV (10) | 4 |
| 2018 | A Two-Step Computation of the Exact GAN Wasserstein DistanceabstractIn this paper, we propose a two-step method to compute the Wasserstein distance in Wasserstein Generative Adversarial Networks (WGANs): 1) The convex part of our objective can be solved by linear programming; 2) The non-convex residual can be approximated by a deep neural network. We theoretically prove that the proposed formulation is equivalent to the discrete Monge-Kantorovich dual formulation. Furthermore, we give the approximation error bound of the Wasserstein distance and the error bound of generalizing the Wasserstein distance from discrete to continuous distributions. Our approach optimizes the exact Wasserstein distance, obviating the need for weight clipping previously used in WGANs. Results on synthetic data show that the our method computes the Wasserstein distance more accurately. Qualitative and quantitative results on MNIST, LSUN and CIFAR-10 datasets show that the proposed method is more efficient than state-of-the-art WGAN methods, and still produces images of comparable quality. Huidong Liu, Xianfeng Gu, Dimitris Samaras |
ICML | 3 |
| 2018 | Sequence-to-Segment Networks for Segment DetectionabstractDetecting segments of interest from an input sequence is a challenging problem which often requires not only good knowledge of individual target segments, but also contextual understanding of the entire input sequence and the relationships between the target segments. To address this problem, we propose the Sequence-to-Segment Network (S$^2$N), a novel end-to-end sequential encoder-decoder architecture. S$^2$N first encodes the input into a sequence of hidden states that progressively capture both local and holistic information. It then employs a novel decoding architecture, called Segment Detection Unit (SDU), that integrates the decoder state and encoder hidden states to detect segments sequentially. During training, we formulate the assignment of predicted segments to ground truth as bipartite matching and use the Earth Mover's Distance to calculate the localization errors. We experiment with S$^2$N on temporal action proposal generation and video summarization and show that S$^2$N achieves state-of-the-art performance on both tasks. Zijun Wei, Boyu Wang 0001, Minh Hoai, Jianming Zhang 0001, Zhe Lin 0001, Xiaohui Shen, Radomír Mech, Dimitris Samaras |
NeurIPS | 8 |
| 2018 | Leave-One-Out Kernel Optimization for Shadow Detection and RemovalabstractThe objective of this work is to detect shadows in images. We pose this as the problem of labeling image regions, where each region corresponds to a group of superpixels. To predict the label of each region, we train a kernel Least-Squares Support Vector Machine (LSSVM) for separating shadow and non-shadow regions. The parameters of the kernel and the classifier are jointly learned to minimize the leave-one-out cross validation error. Optimizing the leave-one-out cross validation error is typically difficult, but it can be done efficiently in our framework. Experiments on two challenging shadow datasets, UCF and UIUC, show that our region classifier outperforms more complex methods. We further enhance the performance of the region classifier by embedding it in a Markov Random Field (MRF) framework and adding pairwise contextual cues. This leads to a method that outperforms the state-of-the-art for shadow detection. In addition we propose a new method for shadow removal based on region relighting. For each shadow region we use a trained classifier to identify a neighboring lit region of the same material. Given a pair of lit-shadow regions we perform a region relighting transformation based on histogram matching of luminance values between the shadow region and the lit region. Once a shadow is detected, we demonstrate that our shadow removal approach produces results that outperform the state of the art by evaluating our method using a publicly available benchmark dataset. Tomás F. Yago Vicente, Minh Hoai, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2018 | Portrait Lighting Transfer Using a Mass Transport ApproachabstractLighting is a critical element of portrait photography. However, good lighting design typically requires complex equipment and significant time and expertise. Our work simplifies this task using a relighting technique that transfers the desired illumination of one portrait onto another. The novelty in our approach to this challenging problem is our formulation of relighting as a mass transport problem. We start from standard color histogram matching that only captures the overall tone of the illumination, and we show how to use the mass-transport formulation to make it dependent on facial geometry. We fit a three-dimensional (3D) morphable face model to the portrait, and for each pixel, we combine the color value with the corresponding 3D position and normal. We then solve a mass-transport problem in this augmented space to generate a color remapping that achieves localized, geometry-aware relighting. Our technique is robust to variations in facial appearance and small errors in face reconstruction. As we demonstrate, this allows our technique to handle a variety of portraits and illumination conditions, including scenarios that are challenging for previous methods. Zhixin Shu, Sunil Hadap, Eli Shechtman, Kalyan Sunkavalli, Sylvain Paris, Dimitris Samaras |
ACM Trans. Graph. | 6 |
| 2017 | ConvNets with Smooth Adaptive Activation Functions for RegressionabstractWithin Neural Networks (NN), the parameters of Adaptive Activation Functions (AAF) control the shapes of activation functions. These parameters are trained along with other parameters in the NN. AAFs have improved performance of Convolutional Neural Networks (CNN) in multiple classification tasks. In this paper, we propose and apply AAFs on CNNs for regression tasks. We argue that applying AAFs in the regression (second-to-last) layer of a NN can significantly decrease the bias of the regression NN. However, using existing AAFs may lead to overfitting. To address this problem, we propose a Smooth Adaptive Activation Function (SAAF) with a piecewise polynomial form which can approximate any continuous function to arbitrary degree of error, while having a bounded Lipschitz constant for given bounded model parameters. As a result, NNs with SAAF can avoid overfitting by simply regularizing model parameters. We empirically evaluated CNNs with SAAFs and achieved state-of-the-art results on age and pose estimation datasets. Le Hou, Dimitris Samaras, Tahsin M. Kurç, Yi Gao 0002, Joel H. Saltz |
AISTATS | 2 |
| 2017 | Large-scale Continual Road Inspection: Visual Infrastructure Assessment in the Wild
Ke Ma 0005, Minh Hoai, Dimitris Samaras |
BMVC | 3 |
| 2017 | Neural Face Editing with Intrinsic Image DisentanglingabstractTraditional face editing methods often require a number of sophisticated and task specific algorithms to be applied one after the other - a process that is tedious, fragile, and computationally intensive. In this paper, we propose an end-to-end generative adversarial network that infers a face-specific disentangled representation of intrinsic face properties, including shape (i.e. normals), albedo, and lighting, and an alpha matte. We show that this network can be trained on “in-the-wild” images by incorporating an in-network physically-based image formation module and appropriate loss functions. Our disentangling latent representation allows for semantically relevant edits, where one aspect offacial appearance can be manipulated while keeping orthogonal properties fixed, and we demonstrate its use for a number offacial editing applications. Zhixin Shu, Ersin Yumer, Sunil Hadap, Kalyan Sunkavalli, Eli Shechtman, Dimitris Samaras |
CVPR | 6 |
| 2017 | Shadow Detection with Conditional Generative Adversarial NetworksabstractWe introduce scGAN, a novel extension of conditional Generative Adversarial Networks (GAN) tailored for the challenging problem of shadow detection in images. Previous methods for shadow detection focus on learning the local appearance of shadow regions, while using limited local context reasoning in the form of pairwise potentials in a Conditional Random Field. In contrast, the proposed adversarial approach is able to model higher level relationships and global scene characteristics. We train a shadow detector that corresponds to the generator of a conditional GAN, and augment its shadow accuracy by combining the typical GAN loss with a data loss term. Due to the unbalanced distribution of the shadow labels, we use weighted cross entropy. With the standard GAN architecture, properly setting the weight for the cross entropy would require training multiple GANs, a computationally expensive grid procedure. In scGAN, we introduce an additional sensitivity parameter w to the generator. The proposed approach effectively parameterizes the loss of the trained detector. The resulting shadow detector is a single network that can generate shadow maps corresponding to different sensitivity levels, obviating the need for multiple models and a costly training procedure. We evaluate our method on the large-scale SBU and UCF shadow datasets, and observe up to 17% error reduction with respect to the previous state-of-the-art method. Vu Nguyen 0004, Tomás F. Yago Vicente, Maozheng Zhao, Minh Hoai, Dimitris Samaras |
ICCV | 5 |
| 2017 | Center-Focusing Multi-task CNN with Injected Features for Classification of Glioma Nuclear ImagesabstractClassifying the various shapes and attributes of a glioma cell nucleus is crucial for diagnosis and understanding of the disease. We investigate the automated classification of the nuclear shapes and visual attributes of glioma cells, using Convolutional Neural Networks (CNNs) on pathology images of automatically segmented nuclei. We propose three methods that improve the performance of a previously-developed semi-supervised CNN. First, we propose a method that allows the CNN to focus on the most important part of an image-the image's center containing the nucleus. Second, we inject (concatenate) pre-extracted VGG features into an intermediate layer of our Semi-Supervised CNN so that during training, the CNN can learn a set of additional features. Third, we separate the losses of the two groups of target classes (nuclear shapes and attributes) into a single-label loss and a multi-label loss in order to incorporate prior knowledge of inter-label exclusiveness. On a dataset of 2078 images, the combination of the proposed methods reduces the error rate of attribute and shape classification by 21.54% and 15.07% respectively compared to the existing state-of-the-art method on the same dataset. Veda Murthy, Le Hou, Dimitris Samaras, Tahsin M. Kurç, Joel H. Saltz |
WACV | 3 |
| 2017 | Portrait lighting transfer using a mass transport approachabstractLighting is a critical element of portrait photography. However, good lighting design typically requires complex equipment and significant time and expertise. Our work simplifies this task using a relighting technique that transfers the desired illumination of one portrait onto another. The novelty in our approach to this challenging problem is our formulation of relighting as a mass transport problem. We start from standard color histogram matching that only captures the overall tone of the illumination, and we show how to use the mass-transport formulation to make it dependent on facial geometry. We fit a three-dimensional (3D) morphable face model to the portrait, and for each pixel, we combine the color value with the corresponding 3D position and normal. We then solve a mass-transport problem in this augmented space to generate a color remapping that achieves localized, geometry-aware relighting. Our technique is robust to variations in facial appearance and small errors in face reconstruction. As we demonstrate, this allows our technique to handle a variety of portraits and illumination conditions, including scenarios that are challenging for previous methods. Zhixin Shu, Sunil Hadap, Eli Shechtman, Kalyan Sunkavalli, Sylvain Paris, Dimitris Samaras |
ACM Trans. Graph. | 6 |
| 2017 | EyeOpener: Editing Eyes in the WildabstractClosed eyes and look-aways can ruin precious moments captured in photographs. In this article, we present a new framework for automatically editing eyes in photographs. We leverage a user’s personal photo collection to find a “good” set of reference eyes and transfer them onto a target image. Our example-based editing approach is robust and effective for realistic image editing. A fully automatic pipeline for realistic eye editing is challenging due to the unconstrained conditions under which the face appears in a typical photo collection. We use crowd-sourced human evaluations to understand the aspects of the target-reference image pair that will produce the most realistic results. We subsequently train a model that automatically selects the top-ranked reference candidate(s) by narrowing the gap in terms of pose, local contrast, lighting conditions, and even expressions. Finally, we develop a comprehensive pipeline of three-dimensional face estimation, image warping, relighting, image harmonization, automatic segmentation, and image compositing in order to achieve highly believable results. We evaluate the performance of our method via quantitative and crowd-sourced experiments. Zhixin Shu, Eli Shechtman, Dimitris Samaras, Sunil Hadap |
ACM Trans. Graph. | 3 |
| 2016 | Geodesic Distance Histogram Feature for Video Segmentation
Hieu Le 0001, Vu Nguyen 0004, Chen-Ping Yu, Dimitris Samaras |
ACCV (1) | 4 |
| 2016 | Patch-Based Convolutional Neural Network for Whole Slide Tissue Image ClassificationabstractConvolutional Neural Networks (CNN) are state-of-the-art models for many image classification tasks. However, to recognize cancer subtypes automatically, training a CNN on gigapixel resolution Whole Slide Tissue Images (WSI) is currently computationally impossible. The differentiation of cancer subtypes is based on cellular-level visual features observed on image patch scale. Therefore, we argue that in this situation, training a patch-level classifier on image patches will perform better than or similar to an image-level classifier. The challenge becomes how to intelligently combine patch-level classification results and model the fact that not all patches will be discriminative. We propose to train a decision fusion model to aggregate patch-level predictions given by patch-level CNNs, which to the best of our knowledge has not been shown before. Furthermore, we formulate a novel Expectation-Maximization (EM) based method that automatically locates discriminative patches robustly by utilizing the spatial relationships of patches. We apply our method to the classification of glioma and non-small-cell lung carcinoma cases into subtypes. The classification accuracy of our method is similar to the inter-observer agreement between pathologists. Although it is impossible to train CNNs on WSIs, we experimentally demonstrate using a comparable non-cancer dataset of smaller images that a patch-based CNN can outperform an image-based CNN. Le Hou, Dimitris Samaras, Tahsin M. Kurç, Yi Gao 0002, James Davis 0001, Joel H. Saltz |
CVPR | 2 |
| 2016 | Noisy Label Recovery for Shadow Detection in Unfamiliar DomainsabstractRecent shadow detection algorithms have shown initial success on small datasets of images from specific domains. However, shadow detection on broader image domains is still challenging due to the lack of annotated training data. This is due to the intense manual labor in annotating shadow data. In this paper we propose "lazy annotation", an efficient annotation method where an annotator only needs to mark the important shadow areas and some non-shadow areas. This yields data with noisy labels that are not yet useful for training a shadow detector. We address the problem of label noise by jointly learning a shadow region classifier and recovering the labels in the training set. We consider the training labels as unknowns and formulate the label recovery problem as the minimization of the sum of squared leave-one-out errors of a Least Squares SVM, which can be efficiently optimized. Experimental results show that a classifier trained with recovered labels achieves comparable performance to a classifier trained on the properly annotated data. These results suggest a feasible approach to address the task of detecting shadows in an unfamiliar domain: collecting and lazily annotating some images from the new domain for training. As will be demonstrated, this approach outperforms methods that rely on precisely annotated but less relevant datasets. Initial results suggest more general applicability. Tomás F. Yago Vicente, Minh Hoai, Dimitris Samaras |
CVPR | 3 |
| 2016 | Large-Scale Training of Shadow Detectors with Noisily-Annotated Shadow Examples
Tomás F. Yago Vicente, Le Hou, Chen-Ping Yu, Minh Hoai, Dimitris Samaras |
ECCV (6) | 5 |
| 2016 | Design and evaluation of a foveated video streaming service for commodity client devicesabstractHumans see only a tiny region at the center of their visual field with the highest visual acuity, a behavior known as foveation. Visual acuity reduces drastically towards the visual periphery. 'Foveated' video coding/compression techniques exploit this non-uniformity to gain significant efficiency by compressing more in the periphery and less in the center. We propose a practical and scalable method to use such a technique for video streaming service over the Internet. The essential idea is to use a commodity webcam on the user side to provide real-time gaze feedback to the server with the server sending appropriately coded video to the client player. We develop a multi-resolution video coding approach that is scalable in that it is possible to pre-code the video in a small number of copies for a given set of resolutions. The coding approach is designed to match the error performance of an eye tracker built using commodity webcams. We demonstrate that the technique is energy efficient and thus usable in mobile devices. We develop a methodology for performance evaluation of such a system when network budgets may vary and video quality may fluctuate. Finally, we present a comprehensive user study that demonstrates a bandwidth reduction of a factor of 2 for the same user satisfaction. Jihoon Ryoo, Kiwon Yun, Dimitris Samaras, Samir Ranjan Das, Gregory J. Zelinsky |
MMSys | 3 |
| 2016 | Learned Region Sparsity and Diversity Also Predicts Visual AttentionabstractLearned region sparsity has achieved state-of-the-art performance in classification tasks by exploiting and integrating a sparse set of local information into global decisions. The underlying mechanism resembles how people sample information from an image with their eye movements when making similar decisions. In this paper we incorporate the biologically plausible mechanism of Inhibition of Return into the learned region sparsity model, thereby imposing diversity on the selected regions. We investigate how these mechanisms of sparsity and diversity relate to visual attention by testing our model on three different types of visual search tasks. We report state-of-the-art results in predicting the locations of human gaze fixations, even though our model is trained only on image-level labels without object location annotations. Notably, the classification performance of the extended model remains the same as the original. This work suggests a new computational perspective on visual attention mechanisms and shows how the inclusion of attention-based mechanisms can improve computer vision techniques. Zijun Wei, Hossein Adeli, Minh Hoai, Gregory J. Zelinsky, Dimitris Samaras |
NIPS | 5 |
| 2016 | Texture classification for rail surface condition evaluationabstractRail surface defects threaten train and passenger safety. Hence rail surfaces must be restored using different processes depending on measurement of the severity of the defects. In this paper, we propose a new method for automatic classification of rail surface defect severity from images collected by rail inspection vehicles. It contains 2 components: a rail surface segmentation module, which utilizes structured random forests to generate an edge map and a Generalized Hough Transform to locate the boundaries of the rail surface; and a defect severity classification module, which combines multiple classifiers through a stacked ensemble model. The first-level learners are trained using descriptors of the rail surface images extracted by texton forests and texton dictionaries, with x2-kernel SVM classifiers. The probability estimation output of the first-level learners is the input to a second level linear-kernel SVM. Our experiments on a dataset of 939 images categorized into 8 severity levels achieved 82% accuracy. Ke Ma 0005, Tomás F. Yago Vicente, Dimitris Samaras, Michael Petrucci, Daniel L. Magnus |
WACV | 3 |
| 2016 | Higher-Order Graph Principles towards Non-Rigid Surface RegistrationabstractThis paper casts surface registration as the problem of finding a set of discrete correspondences through the minimization of an energy function, which is composed of geometric and appearance matching costs, as well as higher-order deformation priors. Two higher-order graph-based formulations are proposed under different deformation assumptions. The first formulation encodes isometric deformations using conformal geometry in a higher-order graph matching problem, which is solved through dual-decomposition and is able to handle partial matching. Despite the isometry assumption, this approach is able to robustly match sparse feature point sets on surfaces undergoing highly anisometric deformations. Nevertheless, its performance degrades significantly when addressing anisometric registration for a set of densely sampled points. This issue is rigorously addressed subsequently through a novel deformation model that is able to handle arbitrary diffeomorphisms between two surfaces. Such a deformation model is introduced into a higher-order Markov Random Field for dense surface registration, and is inferred using a new parallel and memory efficient algorithm. To deal with the prohibitive search space, we also design an efficient way to select a number of matching candidates for each point of the source surface based on the matching results of a sparse set of points. A series of experiments demonstrate the accuracy and the efficiency of the proposed framework, notably in challenging cases of large and/or anisometric deformations, or surfaces that are partially occluded. Chaohui Wang, Xianfeng Gu, Dimitris Samaras, Nikos Paragios |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2015 | Leave-One-Out Kernel Optimization for Shadow DetectionabstractThe objective of this work is to detect shadows in images. We pose this as the problem of labeling image regions, where each region corresponds to a group of superpixels. To predict the label of each region, we train a kernel Least-Squares SVM for separating shadow and non-shadow regions. The parameters of the kernel and the classifier are jointly learned to minimize the leave-one-out cross validation error. Optimizing the leave-one-out cross validation error is typically difficult, but it can be done efficiently in our framework. Experiments on two challenging shadow datasets, UCF and UIUC, show that our region classifier outperforms more complex methods. We further enhance the performance of the region classifier by embedding it in an MRF framework and adding pairwise contextual cues. This leads to a method that significantly outperforms the state-of-the-art. Tomás F. Yago Vicente, Minh Hoai, Dimitris Samaras |
ICCV | 3 |
| 2015 | Efficient Video Segmentation Using Parametric Graph PartitioningabstractVideo segmentation is the task of grouping similar pixels in the spatio-temporal domain, and has become an important preprocessing step for subsequent video analysis. Most video segmentation and supervoxel methods output a hierarchy of segmentations, but while this provides useful multiscale information, it also adds difficulty in selecting the appropriate level for a task. In this work, we propose an efficient and robust video segmentation framework based on parametric graph partitioning (PGP), a fast, almost parameter free graph partitioning method that identifies and removes between-cluster edges to form node clusters. Apart from its computational efficiency, PGP performs clustering of the spatio-temporal volume without requiring a pre-specified cluster number or bandwidth parameters, thus making video segmentation more practical to use in applications. The PGP framework also allows processing sub-volumes, which further improves performance, contrary to other streaming video segmentation methods where sub-volume processing reduces performance. We evaluate the PGP method using the SegTrack v2 and Chen Xiph.org datasets, and show that it outperforms related state-of-the-art algorithms in 3D segmentation metrics and running time. Chen-Ping Yu, Hieu Le 0001, Gregory J. Zelinsky, Dimitris Samaras |
ICCV | 4 |
| 2014 | The Photometry of Intrinsic ImagesabstractIntrinsic characterization of scenes is often the best way to overcome the illumination variability artifacts that complicate most computer vision problems, from 3D reconstruction to object or material recognition. This paper examines the deficiency of existing intrinsic image models to accurately account for the effects of illuminant color and sensor characteristics in the estimation of intrinsic images and presents a generic framework which incorporates insights from color constancy research to the intrinsic image decomposition problem. The proposed mathematical formulation includes information about the color of the illuminant and the effects of the camera sensors, both of which modify the observed color of the reflectance of the objects in the scene during the acquisition process. By modeling these effects, we get a "truly intrinsic" reflectance image, which we call absolute reflectance, which is invariant to changes of illuminant or camera sensors. This model allows us to represent a wide range of intrinsic image decompositions depending on the specific assumptions on the geometric properties of the scene configuration and the spectral properties of the light source and the acquisition system, thus unifying previous models in a single general framework. We demonstrate that even partial information about sensors improves significantly the estimated reflectance images, thus making our method applicable for a wide range of sensors. We validate our general intrinsic image framework experimentally with both synthetic data and natural images. Marc Serra, Olivier Penacchio, Robert Benavente, María Vanrell 0001, Dimitris Samaras |
CVPR | 5 |
| 2014 | 3D assisted face recognition via progressive pose estimationabstractMost existing pose-independent Face Recognition (FR) techniques take advantage of 3D model to guarantee the naturalness while normalizing or simulating pose variations. Two nontrivial problems to be tackled are accurate measurement of pose parameters and computational efficiency. In this paper, we introduce an effective and efficient approach to estimate human head pose, which fundamentally ameliorates the performance of 3D aided FR systems. The proposed method works in a progressive way: firstly, a random forest (RF) is constructed utilizing synthesized images derived from 3D models; secondly, the classification result obtained by applying well-trained RF on a probe image is considered as the preliminary pose estimation; finally, this initial pose is transferred to shape-based 3D morphable model (3DMM) aiming at definitive pose normalization. Using such a method, similarity scores between frontal view gallery set and pose-normalized probe set can be computed to predict the identity. Experimental results achieved on the UHDB dataset outperform the ones so far reported. Additionally, it is much less time-consuming than prevailing 3DMM based approaches. Wuming Zhang, Di Huang 0001, Dimitris Samaras, Jean-Marie Morvan, Yunhong Wang 0001, Liming Chen 0002 |
ICIP | 3 |
| 2014 | Analysis and Control of Facial Expressions using Decomposable nonlinear Generative ModelsabstractFacial expressions convey personal characteristics and subtle emotional states. This paper presents a new framework for modeling subtle facial motions of different people with different types of expressions from high-resolution facial expression tracking data to synthesize new stylized subtle facial expressions. A conceptual facial motion manifold is used for a unified representation of facial motion dynamics from three-dimensional (3D) high-resolution facial motions as well as from two-dimensional (2D) low-resolution facial motions. Variant subtle facial motions in different people with different expressions are modeled by nonlinear mappings from the embedded conceptual manifold to input facial motions using empirical kernel maps. We represent facial expressions by a factorized nonlinear generative model, which decomposes expression style factors and expression type factors from different people with multiple expressions. We also provide a mechanism to control the high-resolution facial motion model from low-resolution facial video sequence tracking and analysis. Using the decomposable generative model with a common motion manifold embedding, we can estimate parameters to control 3D high resolution facial expressions from 2D tracking results, which allows performance-driven control of high-resolution facial expressions. Chan-Su Lee, Dimitris Samaras |
Int. J. Pattern Recognit. Artif. Intell. | 2 |
| 2013 | Single Image Shadow Detection Using Multiple Cues in a Supermodular MRFabstractWe propose a single region shadow classifier based on a multikernel SVM. Our multikernel model is a linear combination of χ2 and Earth Mover’s Distance(EMD)[5] kernels that operate on texture and color histograms disjointly. This single region classifier already outperforms the more complex state of art methods, without performing MRF/CRF optimization. The local appearance of a single region is often ambiguous. Even for a human observer it can be hard to discern if a region is in shadow or not, without considering its context. Hence, it is sensible to look beyond the boundaries of a single region to decide its shadow label [1] [6]. In contrast to previous work we strive to use such contextual information sparingly. For MRF optimization reasons we prefer that most of the work is handled by the single region classifier (unary MRF potentials), with sparse pairwise connections that smooth the label changes across regions. We build on the work of [1] to propose our own improved pairwise classifiers but constrained to adjacent regions: for pairs of regions sharing the same material and same illumination condition, and for same material pairs viewed under different illumination (first lit, second in shadow). We also propose a shadow boundary classifier. Since shadow boundaries often overlap with reflectance changes confounding the effects of the illumination change, our classifier focuses on boundaries of shadows cast over surfaces with the same underlying material. We integrate our single region classifier, our pairwise classifiers, and our boundary classifier using an MRF. Confident positive predictions of the pairwise and boundary classifiers are used to define the pairwise potentials and the graph topology of the MRF. The unary potentials are defined based on the single region classifier. We want to minimize the following functional: Tomás F. Yago Vicente, Chen-Ping Yu, Dimitris Samaras |
BMVC | 3 |
| 2013 | Studying Relationships between Human Gaze, Description, and Computer VisionabstractWe posit that user behavior during natural viewing of images contains an abundance of information about the content of images as well as information related to user intent and user defined content importance. In this paper, we conduct experiments to better understand the relationship between images, the eye movements people make while viewing images, and how people construct natural language to describe images. We explore these relationships in the context of two commonly used computer vision datasets. We then further relate human cues with outputs of current visual recognition systems and demonstrate prototype applications for gaze-enabled detection and annotation. Kiwon Yun, Dimitris Samaras, Gregory J. Zelinsky, Tamara L. Berg |
CVPR | 3 |
| 2013 | A Generic Deformation Model for Dense Non-rigid Surface Registration: A Higher-Order MRF-Based ApproachabstractWe propose a novel approach for dense non-rigid 3D surface registration, which brings together Riemannian geometry and graphical models. To this end, we first introduce a generic deformation model, called Canonical Distortion Coefficients (CDCs), by characterizing the deformation of every point on a surface using the distortions along its two principle directions. This model subsumes the deformation groups commonly used in surface registration such as isometry and conformality, and is able to handle more complex deformations. We also derive its discrete counterpart which can be computed very efficiently in a closed form. Based on these, we introduce a higher-order Markov Random Field (MRF) model which seamlessly integrates our deformation model and a geometry/texture similarity metric. Then we jointly establish the optimal correspondences for all the points via maximum a posteriori (MAP) inference. Moreover, we develop a parallel optimization algorithm to efficiently perform the inference for the proposed higher-order MRF model. The resulting registration algorithm outperforms state-of-the-art methods in both dense non-rigid 3D surface registration and tracking. Chaohui Wang, Xianfeng Gu, Dimitris Samaras, Nikos Paragios |
ICCV | 4 |
| 2013 | Intrinsic image evaluation on synthetic complex scenesabstractScene decomposition into its illuminant, shading, and reflectance intrinsic images is an essential step for scene understanding. Collecting intrinsic image groundtruth data is a laborious task. The assumptions on which the ground-truth procedures are based limit their application to simple scenes with a single object taken in the absence of indirect lighting and interreflections. We investigate synthetic data for intrinsic image research since the extraction of ground truth is straightforward, and it allows for scenes in more realistic situations (e.g, multiple illuminants and interreflections). With this dataset we aim to motivate researchers to further explore intrinsic image decomposition in complex scenes. Shida Kunz, Marc Serra, Joost van de Weijer 0001, Robert Benavente, María Vanrell 0001, Olivier Penacchio, Dimitris Samaras |
ICIP | 7 |
| 2013 | Extracting Brain Regions from Rest fMRI with Total-Variation Constrained Dictionary Learning
Alexandre Abraham, Elvis Dohmatob, Bertrand Thirion, Dimitris Samaras, Gaël Varoquaux |
MICCAI (2) | 4 |
| 2013 | Modeling Clutter Perception using Parametric Proto-object PartitioningabstractVisual clutter, the perception of an image as being crowded and disordered, affects aspects of our lives ranging from object detection to aesthetics, yet relatively little effort has been made to model this important and ubiquitous percept. Our approach models clutter as the number of proto-objects segmented from an image, with proto-objects defined as groupings of superpixels that are similar in intensity, color, and gradient orientation features. We introduce a novel parametric method of merging superpixels by modeling mixture of Weibull distributions on similarity distance statistics, then taking the normalized number of proto-objects following partitioning as our estimate of clutter perception. We validated this model using a new $\text{90}-$image dataset of realistic scenes rank ordered by human raters for clutter, and showed that our method not only predicted clutter extremely well (Spearman's $\rho = 0.81$, $p < 0.05$), but also outperformed all existing clutter perception models and even a behavioral object segmentation ground truth. We conclude that the number of proto-objects in an image affects clutter perception more than the number of objects or features. Chen-Ping Yu, Wen-Yu Hua, Dimitris Samaras, Gregory J. Zelinsky |
NIPS | 3 |
| 2013 | Simultaneous Cast Shadows, Illumination and Geometry Inference Using HypergraphsabstractThe cast shadows in an image provide important information about illumination and geometry. In this paper, we utilize this information in a novel framework in order to jointly recover the illumination environment, a set of geometry parameters, and an estimate of the cast shadows in the scene given a single image and coarse initial 3D geometry. We model the interaction of illumination and geometry in the scene and associate it with image evidence for cast shadows using a higher order Markov Random Field (MRF) illumination model, while we also introduce a method to obtain approximate image evidence for cast shadows. Capturing the interaction between light sources and geometry in the proposed graphical model necessitates higher order cliques and continuous-valued variables, which make inference challenging. Taking advantage of domain knowledge, we provide a two-stage minimization technique for the MRF energy of our model. We evaluate our method in different datasets, both synthetic and real. Our model is robust to rough knowledge of geometry and inaccurate initial shadow estimates, allowing a generic coarse 3D model to represent a whole class of objects for the task of illumination estimation, or the estimation of geometry parameters to refine our initial knowledge of scene geometry, simultaneously with illumination estimation. Alexandros Panagopoulos, Chaohui Wang, Dimitris Samaras, Nikos Paragios |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2012 | Reconstructing Shape from Dictionaries of Shading Primitives
Alexandros Panagopoulos, Sunil Hadap, Dimitris Samaras |
ACCV (4) | 3 |
| 2012 | Can a Single Brain Region Predict a Disorder?abstractWe perform prediction of diverse disorders (Cocaine Use, Schizophrenia and Alzheimers disease) in unseen subjects from brain fMRI. First, we show that for multi-subject prediction of simple cognitive states (e.g. motor vs. calculation and reading), voxels-as-features methods produce clusters that are similar for different leave-one-subject-out folds; while for group classification (e.g. cocaine addicted vs. control subjects), voxels are scattered and less stable. Therefore, we chose to use a single region per experimental condition and a majority vote classifier. Interestingly, our method outperforms state-of-the-art techniques. Our method can integrate multiple experimental conditions and successfully predict disorders in unseen subjects (leave-one-subjectout generalization accuracy: 89.3% and 90.9% for Cocaine Use, 96.4% for Schizophrenia and 81.5% for Alzheimers disease). Our experimental results not only span diverse disorders, but also different experimental designs (block design and event related tasks), facilities, magnetic fields (1.5Tesla, 3Tesla, 4Tesla) and speed of acquisition (interscan interval from 1600ms to 3500ms). We further argue that our method produces a meaningful low dimensional representation that retains discriminability. Jean Honorio, Dardo Tomasi, Rita Z. Goldstein, Hoi-Chung Leung, Dimitris Samaras |
IEEE Trans. Medical Imaging | 5 |
| 2011 | Illumination estimation and cast shadow detection through a higher-order graphical modelabstractIn this paper, we propose a novel framework to jointly recover the illumination environment and an estimate of the cast shadows in a scene from a single image, given coarse 3D geometry. We describe a higher-order Markov Random Field (MRF) illumination model, which combines low-level shadow evidence with high-level prior knowledge for the joint estimation of cast shadows and the illumination environment. First, a rough illumination estimate and the structure of the graphical model in the illumination space is determined through a voting procedure. Then, a higher order approach is considered where illumination sources are coupled with the observed image and the latent variables corresponding to the shadow detection. We examine two inference methods in order to effectively minimize the MRF energy of our model. Experimental evaluation shows that our approach is robust to rough knowledge of geometry and reflectance and inaccurate initial shadow estimates. We demonstrate the power of our MRF illumination model on various datasets and show that we can estimate the illumination in images of objects belonging to the same class using the same coarse 3D model to represent all instances of the class. Alexandros Panagopoulos, Chaohui Wang, Dimitris Samaras, Nikos Paragios |
CVPR | 3 |
| 2011 | Intrinsic dense 3D surface trackingabstractThis paper presents a novel intrinsic 3D surface distance and its use in a complete probabilistic tracking framework for dynamic 3D data. Registering two frames of a deforming 3D shape relies on accurate correspondences between all points across the two frames. In the general case such correspondence search is computationally intractable. Common prior assumptions on the nature of the deformation such as near-rigidity, isometry or learning from a training set, reduce the search space but often at the price of loss of accuracy when it comes to deformations not in the prior assumptions. If we consider the set of all possible 3D surface matchings defined by specifying triplets of correspondences in the uniformization domain, then we introduce a new matching cost between two 3D surfaces. The lowest feature differences across this set of matchings that cause two points to correspond, become the matching cost of that particular correspondence. We show that for surface tracking applications, the matching cost can be efficiently computed in the uniformization domain. This matching cost is then combined with regularization terms that enforce spatial and temporal motion consistencies, into a maximum a posteriori (MAP) problem which we approximate using a Markov Random Field (MRF). Compared to previous 3D surface tracking approaches that either assume isometric deformations or consistent features, our method achieves dense, accurate tracking results, which we demonstrate through a series of dense, anisometric 3D surface tracking experiments. Chaohui Wang, Yang Wang 0001, Xianfeng Gu, Dimitris Samaras, Nikos Paragios |
CVPR | 5 |
| 2011 | Analysis and synthesis of facial expressions using decomposable nonlinear generative modelsabstractThis paper presents a new framework that models facial expressions in multiple people with different expressions and synthesize new stylized subtle facial expressions using the generative models. As facial expressions pass through nonlinear shape deformations during facial expressions, we model the facial expression in nonlinear mapping space based on low dimensional embedding and kernel mapping. Characteristics of different type of expressions and variances in different people are decomposed by analyzing the nonlinear mapping between the Euclidean space of facial motion and a low dimensional embedding of these expressions. Using high resolution tracking of densely sampled 3D data, the generative model can control subtle facial expression characteristics of different person in different expression by low dimensional person dependent style factor and expression type dependent expression factor. The temporal characteristics of the motion can also be controlled by the trajectory sampling on the low dimensional embedding manifold which is independent of person style and expression type. Our experimental results are shown for subtle differences in different smile expressions in different people from dense 3D tracking. Chan-Su Lee, Dimitris Samaras |
FG | 2 |
| 2011 | Viewpoint invariant 3D landmark model inference from monocular 2D images using higher-order priorsabstractIn this paper, we propose a novel one-shot optimization approach to simultaneously determine both the optimal 3D landmark model and the corresponding 2D projections without explicit estimation of the camera viewpoint, which is also able to deal with misdetections as well as partial occlusions. To this end, a 3D shape manifold is built upon fourth-order interactions of landmarks from a training set where pose-invariant statistics are obtained in this space. The 3D-2D consistency is also encoded in such high-order interactions, which eliminate the necessity of viewpoint estimation. Furthermore, the modeling of visibility improves further the performance of the method by handling missing correspondences and occlusions. The inference is addressed through a MAP formulation which is naturally transformed into a higher-order MRF optimization problem and is solved using a dual-decomposition-based method. Promising results on standard face benchmarks demonstrate the potential of our approach. Chaohui Wang, Loïc Simon, Ioannis A. Kakadiaris, Dimitris Samaras, Nikos Paragios |
ICCV | 5 |
| 2010 | Dense non-rigid surface registration using high-order graph matchingabstractIn this paper, we propose a high-order graph matching formulation to address non-rigid surface matching. The singleton terms capture the geometric and appearance similarities (e.g., curvature and texture) while the high-order terms model the intrinsic embedding energy. The novelty of this paper includes: 1. casting 3D surface registration into a graph matching problem that combines both geometric and appearance similarities and intrinsic embedding information, 2. the first implementation of high-order graph matching algorithm that solves a non-convex optimization problem, and 3. an efficient two-stage optimization approach to constrain the search space for dense surface registration. Our method is validated through a series of experiments demonstrating its accuracy and efficiency, notably in challenging cases of large and/or non-isometric deformations, or meshes that are partially occluded. Chaohui Wang, Yang Wang 0001, Xianfeng Gu, Dimitris Samaras, Nikos Paragios |
CVPR | 5 |
| 2010 | Multi-Task Learning of Gaussian Graphical Models
Jean Honorio, Dimitris Samaras |
ICML | 2 |
| 2010 | Partial Face Biometry Using Shape Decomposition on 2D Conformal Maps of FacesabstractIn this paper, we introduce a new approach for partial 3D face recognition, which makes use of shape decomposition over the rigid part of a face. To explore the descriptiveness of shape dissimilarity over an isometric part of a face, which has lower probability to be influenced by expression, we transform a 3D shape to a 2D domain using conformal mapping and use shape decomposition as a similarity measurement. In our work we investigate several classifiers as well as several shape descriptors for recognition purposes. Recognition tests on a subset of the FRGC data set show approximately 80% rank-one recognition rate using only the eyes and nose part of the face. Przemyslaw Szeptycki, Mohsen Ardabilian, Liming Chen 0002, Wei Zeng 0002, Xianfeng Gu, Dimitris Samaras |
ICPR | 6 |
| 2010 | Ricci Flow for 3D Shape AnalysisabstractRicci flow is a powerful curvature flow method, which is invariant to rigid motion, scaling, isometric, and conformal deformations. We present the first application of surface Ricci flow in computer vision. Previous methods based on conformal geometry, which only handle 3D shapes with simple topology, are subsumed by the Ricci flow-based method, which handles surfaces with arbitrary topology. We present a general framework for the computation of Ricci flow, which can design any Riemannian metric by user-defined curvature. The solution to Ricci flow is unique and robust to noise. We provide implementation details for Ricci flow on discrete surfaces of either euclidean or hyperbolic background geometry. Our Ricci flow-based method can convert all 3D problems into 2D domains and offers a general framework for 3D shape analysis. We demonstrate the applicability of this intrinsic shape representation through standard shape analysis problems, such as 3D shape matching and registration, and shape indexing. Surfaces with large nonrigid anisotropic deformations can be registered using Ricci flow with constraints of feature points and curves. We show how conformal equivalence can be used to index shapes in a 3D surface shape space with the use of Teichmüller space coordinates. Experimental results are shown on 3D face data sets with large expression deformations and on dynamic heart data. Wei Zeng 0002, Dimitris Samaras, Xianfeng Gu |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Robust shadow and illumination estimation using a mixture modelabstractIlluminant estimation from shadows typically relies on accurate segmentation of the shadows and knowledge of exact 3D geometry, while shadow estimation is difficult in the presence of texture. These can be onerous requirements; in this paper we propose a graphical model to estimate the illumination environment and detect the shadows of a scene with textured surfaces from a single image and only coarse 3D information. We represent the illumination environment as a mixture of von Mises-Fisher distributions. Then, each shadow pixel becomes the combination of samples generated from this illumination environment. We integrate a number of low-level, illumination-invariant 2D cues in a graphical model to detect and estimate cast shadows on textured surfaces. Both 2D cues and approximate 3D reasoning are combined to infer a set of labels that identify the shadows in the image and estimate the positions, shapes and intensities of the light sources. Our results demonstrate that the probabilistic combination of multiple cues, unlike prior approaches, manages to differentiate both hard and soft shadows from the underlying surface texture even when we can only coarsely anticipate the effect of 3D geometry. We also experimentally demonstrate how correct estimation of the sharpness and shape of the light sources improves the augmented reality results. Alexandros Panagopoulos, Dimitris Samaras, Nikos Paragios |
CVPR | 2 |
| 2009 | Sparse and Locally Constant Gaussian Graphical ModelsabstractLocality information is crucial in datasets where each variable corresponds to a measurement in a manifold (silhouettes, motion trajectories, 2D and 3D images). Although these datasets are typically under-sampled and high-dimensional, they often need to be represented with low-complexity statistical models, which are comprised of only the important probabilistic dependencies in the datasets. Most methods attempt to reduce model complexity by enforcing structure sparseness. However, sparseness cannot describe inherent regularities in the structure. Hence, in this paper we first propose a new class of Gaussian graphical models which, together with sparseness, imposes local constancy through ${\ell}_1$-norm penalization. Second, we propose an efficient algorithm which decomposes the strictly convex maximum likelihood estimation into a sequence of problems with closed form solutions. Through synthetic experiments, we evaluate the closeness of the recovered models to the ground truth. We also test the generalization performance of our method in a wide range of complex real-world datasets and demonstrate that it can capture useful structures such as the rotation and shrinking of a beating heart, motion correlations between body parts during walking and functional interactions of brain regions. Our method outperforms the state-of-the-art structure learning techniques for Gaussian graphical models both for small and large datasets. Jean Honorio, Luis E. Ortiz, Dimitris Samaras, Nikos Paragios, Rita Z. Goldstein |
NIPS | 3 |
| 2009 | Face Relighting from a Single Image under Arbitrary Unknown Lighting ConditionsabstractIn this paper, we present a new method to modify the appearance of a face image by manipulating the illumination condition, when the face geometry and albedo information is unknown. This problem is particularly difficult when there is only a single image of the subject available. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using a spherical harmonic representation. Moreover, morphable models are statistical ensembles of facial properties such as shape and texture. In this paper, we integrate spherical harmonics into the morphable model framework by proposing a 3D spherical harmonic basis morphable model (SHBMM). The proposed method can represent a face under arbitrary unknown lighting and pose simply by three low-dimensional vectors, i.e., shape parameters, spherical harmonic basis parameters, and illumination coefficients, which are called the SHBMM parameters. However, when the image was taken under an extreme lighting condition, the approximation error can be large, thus making it difficult to recover albedo information. In order to address this problem, we propose a subregion-based framework that uses a Markov random field to model the statistical distribution and spatial coherence of face texture, which makes our approach not only robust to extreme lighting conditions, but also insensitive to partial occlusions. The performance of our framework is demonstrated through various experimental results, including the improved rates for face recognition under extreme lighting conditions. Yang Wang 0001, Lei Zhang 0002, Zicheng Liu 0001, Gang Hua 0001, Zhengyou Zhang, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2008 | 3D Non-rigid Surface Matching and Registration Based on Holomorphic Differentials
Wei Zeng 0002, Yang Wang 0001, Xiaotian Yin, Xianfeng Gu, Dimitris Samaras |
ECCV (3) | 6 |
| 2008 | Task-Specific Functional Brain Geometry from Model Maps
Georg Langs, Dimitris Samaras, Nikos Paragios, Jean Honorio, Nelly Alia-Klein, Dardo Tomasi, Nora D. Volkow, Rita Z. Goldstein |
MICCAI (1) | 2 |
| 2008 | Topology cuts: A novel min-cut/max-flow algorithm for topology preserving segmentation in N-D images
Dimitris Samaras, Wei Chen 0001, Qunsheng Peng 0001 |
Comput. Vis. Image Underst. | 2 |
| 2008 | High Resolution Tracking of Non-Rigid Motion of Densely Sampled 3D Data Using Harmonic Maps
Yang Wang 0001, Mohit Gupta 0001, Song Zhang 0002, Xianfeng Gu, Dimitris Samaras, Peisen Huang |
Int. J. Comput. Vis. | 6 |
| 2008 | Estimation of multiple directional illuminants from a single image
Yang Wang 0001, Dimitris Samaras |
Image Vis. Comput. | 2 |
| 2008 | Dependent Multiple Cue Integration for Robust TrackingabstractWe propose a new technique for fusing multiple cues to robustly segment an object from its background in video sequences that suffer from abrupt changes of both illumination and position of the target. Robustness is achieved by the integration of appearance and geometric object features and by their estimation using Bayesian filters, such as Kalman or particle filters. In particular, each filter estimates the state of a specific object feature, conditionally dependent on another feature estimated by a distinct filter. This dependence provides improved target representations, permitting to segment it out from the background even in non-stationary sequences. Considering that the procedure of the Bayesian filters may be described by a "hypotheses generation--hypotheses correction" strategy, the major novelty of our methodology compared to previous approaches is that the mutual dependence between filters is considered during the feature observation, i.e, into the "hypotheses correction" stage,instead of considering it when generating the hypotheses. This proves to be much more effective in terms of accuracy and reliability. The proposed method is analytically justified and applied to develop a robust tracking system that adapts online and simultaneously the color space where the image points are represented, the color distributions, the contour of the object and its bounding box. Results with synthetic data and real video sequences demonstrate the robustness and versatility of our method. Francesc Moreno-Noguer, Alberto Sanfeliu, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2007 | Capturing long-range correlations with patch modelsabstractThe use of image patches to capture local correlations between pixels has been growing in popularity for use in various low-level vision tasks. There is a trade-off between using larger patches to obtain additional high-order statistics and smaller patches to capture only the elemental features of the image. Previous work has leveraged short-range correlations between patches that share pixel values for use in patch matching. In this paper, long-range correlations between patches are introduced, where relations between patches that do not necessarily share pixels are learnt. Such correlations arise as an inherent property of the data itself. These long-range patch correlations are shown to be particularly important for video sequences where the patches have an additional time dimension, with correlation links in both space and time. We illustrate the power of our model on tasks such as multiple object registration and detection and missing data interpolation, including a difficult task of photograph relighting, where a single photograph is assumed to be the only observed part of a 3D volume whose two coordinates are the image 𝓍 and 𝒴 coordinates and the third coordinate is the illumination angle θ. We show that in some cases, the long-range correlations observed among the mappings of different volume patches in a small training set are sufficient to infer the possible complex intensity changes in a new photograph due to illumination angle variation. Vincent Cheung, Nebojsa Jojic, Dimitris Samaras |
CVPR | 3 |
| 2007 | Face Re-Lighting from a Single Image under Harsh Lighting ConditionsabstractIn this paper, we present a new method to change the illumination condition of a face image, with unknown face geometry and albedo information. This problem is particularly difficult when there is only one single image of the subject available and it was taken under a harsh lighting condition. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using spherical harmonic representation. However, the approximation error can be large under harsh lighting conditions thus making it difficult to recover albedo information. In order to address this problem, we propose a subregion based framework that uses a Markov Random Field to model the statistical distribution and spatial coherence of face texture, which makes our approach not only robust to harsh lighting conditions, but insensitive to partial occlusions as well. The performance of our framework is demonstrated through various experimental results, including the improvement to the face recognition rate under harsh lighting conditions. Yang Wang 0001, Zicheng Liu 0001, Gang Hua 0001, Zhengyou Zhang, Dimitris Samaras |
CVPR | 6 |
| 2007 | Ricci Flow for 3D Shape AnalysisabstractRicci flow is a powerful curvature flow method in geometric analysis. This work is the first application of surface Ricci flow in computer vision. We show that previous methods based on conformal geometries, such as harmonic maps and least-square conformal maps, which can only handle 3D shapes with simple topology are subsumed by our Ricci flow based method which can handle surfaces with arbitrary topology. Because the Ricci flow method is intrinsic and depends on the surface metric only, it is invariant to rigid motion, scaling, and isometric and conformal deformations. The solution to Ricci flow is unique and its computation is robust to noise. Our Ricci flow based method can convert all 3D problems into 2D domains and offers a general framework for 3D surface analysis. Large non-rigid deformations can be registered with feature constraints, hence we introduce a method that constrains Ricci flow computation using feature points and feature curves. Finally, we demonstrate the applicability of this intrinsic shape representation through standard shape analysis problems, such as 3D shape matching and registration. Xianfeng Gu, Yang Wang 0001, Hong Qin 0001, Dimitris Samaras |
ICCV | 7 |
| 2007 | Real-time Accurate Object Detection using Multiple ResolutionsabstractWe propose a multi-resolution framework inspired by human visual search for general object detection. Different resolutions are represented using a coarse-to-fine feature hierarchy. During detection, the lower resolution features are initially used to reject the majority of negative windows at relatively low cost, leaving a relatively small number of windows to be processed in higher resolutions. This enables the use of computationally more expensive higher resolution features to achieve high detection accuracy. We applied this framework on Histograms of Oriented Gradient (HOG) features for object detection. Our multi-resolution detector produced better performance for pedestrian detection than state-of-the-art methods (Dalal and Triggs, 2005), and was faster during both training and testing. Testing our method on motorbikes and cars from the VOC database revealed similar improvements in both speed and accuracy, suggesting that our approach is suitable for realtime general object detection applications. Wei Zhang 0002, Gregory J. Zelinsky, Dimitris Samaras |
ICCV | 3 |
| 2007 | Integration of deformable contours and a multiple hypotheses Fisher color model for robust tracking in varying illuminant environments
Francesc Moreno-Noguer, Alberto Sanfeliu, Dimitris Samaras |
Image Vis. Comput. | 3 |
| 2007 | Conformal Geometry and Its Applications on 3D Shape Matching, Recognition, and StitchingabstractThree-dimensional shape matching is a fundamental issue in computer vision with many applications such as shape registration, 3D object recognition, and classification. However, shape matching with noise, occlusion, and clutter is a challenging problem. In this paper, we analyze a family of quasi-conformal maps including harmonic maps, conformal maps, and least-squares conformal maps with regards to 3D shape matching. As a result, we propose a novel and computationally efficient shape matching framework by using least-squares conformal maps. According to conformal geometry theory, each 3D surface with disk topology can be mapped to a 2D domain through a global optimization and the resulting map is a diffeomorphism, i.e., one-to-one and onto. This allows us to simplify the 3D shape-matching problem to a 2D image-matching problem, by comparing the resulting 2D parametric maps, which are stable, insensitive to resolution changes and robust to occlusion, and noise. Therefore, highly accurate and efficient 3D shape matching algorithms can be achieved by using the above three parametric maps. Finally, the robustness of least-squares conformal maps is evaluated and analyzed comprehensively in 3D shape matching with occlusion, noise, and resolution variation. In order to further demonstrate the performance of our proposed method, we also conduct a series of experiments on two computer vision applications, i.e., 3D face recognition and 3D nonrigid surface alignment and stitching. Yang Wang 0001, Miao Jin, Xianfeng Gu, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2006 | 3D Surface Matching and Recognition Using Conformal Geometryabstract3D surface matching is a fundamental issue in computer vision with many applications such as shape registration, 3D object recognition and classification. However, surface matching with noise, occlusion and clutter is a challenging problem. In this paper, we analyze a family of conformal geometric maps including harmonic maps, conformal maps and least squares conformal maps with regards to 3D surface matching. As a result, we propose a novel and computationally efficient surface matching framework that uses least squares conformal maps. According to conformal geometry theory, each 3D surface with disk topology can be mapped to a 2D domain through a global optimization and the resulting map is a diffeomorphism, i.e., one-to-one and onto. This allows us to simplify the 3D surface-matching problem to a 2D image-matching problem, by comparing the resulting 2D conformal geometric maps, which are stable, insensitive to resolution changes and robust to occlusion and noise. Therefore, highly accurate and efficient 3D surface matching algorithms can be achieved by using conformal geometric maps. Finally, the performance of conformal geometric maps is evaluated and analyzed comprehensively in 3D surface matching with occlusion, noise and resolution variation. We also provide a series of experiments on real 3D face data that achieve high recognition rates. Yang Wang 0001, Miao Jin, Xianfeng Gu, Dimitris Samaras |
CVPR (2) | 5 |
| 2006 | Integration of Dependent Bayesian Filters for Robust TrackingabstractRobotics applications based on computer vision algorithms are highly constrained to indoor environments where conditions may be controlled. The development of robust visual algorithms is necessary for improving the capabilities of many autonomous systems in outdoor and dynamic environments. In particular, this paper proposes a tracking algorithm robust to several artifacts which may be found in real world applications, such as lighting changes, cluttered backgrounds and unexpected target movements. In order to deal with these difficulties the proposed tracking methodology integrates several Bayesian filters. Each filter estimates the state of a particular object feature which is conditionally dependent on another feature estimated by a distinct filter. This dependence provides improved representations of the target, allowing to segment it out from the background of the image. We describe the updating procedure of the Bayesian filters by a 'hypotheses generation and correction' scheme. The main difference with respect to previous approaches is that the dependence between filters is considered during the feature observation, i.e., into the 'hypotheses correction' stage, instead of considering it when generating the hypotheses. This proves to be much more effective in terms of accuracy and reliability Francesc Moreno-Noguer, Alberto Sanfeliu, Dimitris Samaras |
ICRA | 3 |
| 2006 | Face Recognition from a Single Training Image under Arbitrary Unknown Lighting Using Spherical HarmonicsabstractIn this paper, we propose two novel methods for face recognition under arbitrary unknown lighting by using spherical harmonics illumination representation, which require only one training image per subject and no 3D shape information. Our methods are based on the recent result which demonstrated that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace. We provide two methods to estimate the spherical harmonic basis images spanning this space from just one image. Our first method builds the statistical model based on a collection of 2D basis images. We demonstrate that, by using the learned statistics, we can estimate the spherical harmonic basis images from just one image taken under arbitrary illumination conditions if there is no pose variation. Compared to the first method, the second method builds the statistical models directly in 3D spaces by combining the spherical harmonic illumination representation and a 3D morphable model of human faces to recover basis images from images across both poses and illuminations. After estimating the basis images, we use the same recognition scheme for both methods: we recognize the face for which there exists a weighted combination of basis images that is the closest to the test face image. We provide a series of experiments that achieve high recognition rates, under a wide range of illumination conditions, including multiple sources of illumination. Our methods achieve comparable levels of accuracy with methods that have much more onerous training data requirements. Comparison of the two methods is also provided. Lei Zhang 0002, Dimitris Samaras |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2006 | Guest Editorial
Hanspeter Pfister, Dimitris Samaras |
Vis. Comput. | 2 |
| 2005 | Image-driven re-targeting and relighting of facial expressionsabstractSynthesis and re-targeting of facial expressions is central to facial animation and often involves significant manual work in order to achieve realistic expressions, due to the difficulty of capturing high quality expression data. Recent progress in dynamic 3-D scanning allows very accurate acquisition of dense point clouds of facial geometry and texture moving at video speeds. Often the new facial expressions need to be rendered in different environments where the illumination is different from the original capture conditions. In this paper we examine the problem of re-targeting captured facial motion under different illumination conditions when the information we have about the face we want to animate is minimal, a single input image. Given an input image of a face, a set of illumination example images (of other faces captured under different illumination) and a facial expression motion sequence, we aim to generate novel expression sequences of the input face under the lighting conditions in the illumination example images. The input image and illumination example images can be taken under arbitrary unknown lighting. In this paper, we propose two methods in which a 3D spherical harmonic morphable model (SHBMM) can generate images under new lighting conditions with remarkable quality even if only one single image under unknown lighting is available, not only for static poses but for dynamic sequences as well where the face is undergoing subtle high-detail motion. Lei Zhang 0002, Yang Wang 0001, Dimitris Samaras, Song Zhang 0002, Peisen Huang |
Computer Graphics International | 4 |
| 2005 | Face Modeling and Analysis in Stony Brook UniversityabstractIn this paper, we present our latest work on facial expression analysis, synthesis and face recognition. The advent of new technologies that allow the capture of massive amounts of high resolution, high frame rate face data, leads us to propose data-driven face models that accurately describe the appearance of faces under unknown pose and illumination conditions as well as to track subtle geometry changes that occur during expressions. In this paper, we also demonstrate our results for expression transfer among different subjects. We reduce the dimensionality of our data onto a lower dimensional space manifold and then decompose it into style and content parameters. This allows us to transfer subtle expression information (in the form of a style vector) between individuals to synthesize new expressions, as well as smoothly morph geometry and motion. Finally, we demonstrate the accuracy of our face modeling methods through an integrated example of image-driven re-targeting and relighting of facial expressions, where transfer of expression and illumination information between different individuals is possible. Dimitris Samaras, Yang Wang 0001, Lei Zhang 0002, Mohit Gupta 0001 |
CVPR (2) | 1 |
| 2005 | Machine Learning for Clinical Diagnosis from Functional Magnetic Resonance ImagingabstractFunctional magnetic resonance imaging (fMRI) has enabled scientists to look into the active human brain. FMRI provides a sequence of 3D brain images with intensities representing brain activations. Standard techniques for fMRI analysis traditionally focused on finding the area of most significant brain activation for different sensations or activities. In this paper, we explore a new application of machine learning methods to a more challenging problem: classifying subjects into groups based on the observed 3D brain images when the subjects are performing the same task. Here we address the separation of drug-addicted subjects from healthy non-drug-using controls. In this paper, we explore a number of classification approaches. We introduce a novel algorithm that integrates side information into the use of boosting. Our algorithm clearly outperformed well-established classifiers as documented in extensive experimental results. This is the first time that machine learning techniques based on 3D brain images are applied to a clinical diagnosis that currently is only performed through patient self-report. Our tools can therefore provide information not addressed by traditional analysis methods and substantially improve diagnosis. Lei Zhang 0002, Dimitris Samaras, Dardo Tomasi, Nora D. Volkow, Rita Z. Goldstein |
CVPR (1) | 2 |
| 2005 | Face Synthesis and Recognition from a Single Image under Arbitrary Unknown Lighting Using a Spherical Harmonic Basis Morphable ModelabstractUnderstanding and modifying the effects of arbitrary illumination on human faces in a realistic manner is a challenging problem both for face synthesis and recognition. Recent research demonstrates that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace using spherical harmonics representation. Morphable models are statistical ensembles of facial properties such as shape and texture. In this paper, we integrate spherical harmonics into the morphable model framework, by proposing a 3D spherical harmonic basis morphable model (SHBMM) and demonstrate that any face under arbitrary unknown lighting can be simply represented by three low-dimensional vectors: shape parameters, spherical harmonic basis parameters and illumination coefficients. We show that, with our SHBMM, given one single image under arbitrary unknown lighting, we can remove the illumination effects from the image (face "delighting") and synthesize new images under different illumination conditions (face "re-lighting"). Furthermore, we demonstrate that cast shadows can be detected and subsequently removed by using the image error between the input image and the corresponding rendered image. We also propose two illumination invariant face recognition methods based on the recovered SHBMM parameters and the de-lit images respectively. Experimental results show that using only a single image of a face under unknown lighting, we can achieve high recognition rates and generate photorealistic images of the face under a wide range of illumination conditions, including multiple sources of illumination. Lei Zhang 0002, Dimitris Samaras |
CVPR (2) | 3 |
| 2005 | Object Class Recognition Using Multiple Layer Boosting with Heterogeneous FeaturesabstractWe combine local texture features (PCA-SIFT), global features (shape context), and spatial features within a single multi-layer AdaBoost model of object class recognition. The first layer selects PCA-SIFT and shape context features and combines the two feature types to form a strong classifier. Although previous approaches have used either feature type to train an AdaBoost model, our approach is the first to combine these complementary sources of information into a single feature pool and to use Adaboost to select those features most important for class recognition. The second layer adds to these local and global descriptions information about the spatial relationships between features. Through comparisons to the training sample, we first find the most prominent local features in Layer I, then capture the spatial relationships between these features in Layer 2. Rather than discarding this spatial information, we therefore use it to improve the strength of our classifier. We compared our method to (R. Fergus et al., 2003, A. Opelt et al., 2004, J. Thureson et al., 2004) and in all cases our approach outperformed these previous methods using a popular benchmark for object class recognition (R. Fergus et al., 2003). ROC equal error rates approached 99%. We also tested our method using a dataset of images that better equates the complexity between object and non-object images, and again found that our approach outperforms previous methods. Wei Zhang 0002, Gregory J. Zelinsky, Dimitris Samaras |
CVPR (2) | 4 |
| 2005 | Integration of Conditionally Dependent Object Features for Robust Figure/Background SegmentationabstractWe propose a new technique for focusing multiple cues to robustly segment an object from its background in video sequences that suffer from abrupt changes of both illumination and position of the target. Robustness is achieved by tile integration of appearance and geometric object features and by their description using particle filters. Previous approaches assume independence of the object cues or apply the particle filter formulation to only one of the features, and assume a smooth change in the rest, which can prove is very limiting, especially when the state of some features needs to be updated using other cues or when their dynamics follow non-linear and unpredictable paths. Our technique offers a general framework to model the probabilistic relationship between features. The proposed method is analytically justified and applied to develop a robust tracking system that adapts online and simultaneously the color space where the image points are represented, the color distributions, and the contour of the object. Results with synthetic data and real video sequences demonstrate the robustness and versatility of our method Francesc Moreno-Noguer, Alberto Sanfeliu, Dimitris Samaras |
ICCV | 3 |
| 2005 | High Resolution Tracking of Non-Rigid 3D Motion of Densely Sampled Data Using Harmonic MapsabstractWe present a novel fully automatic method for high resolution, nonrigid dense 3D point tracking. High quality dense point clouds of nonrigid geometry moving at video speeds are acquired using a phase-shifting structured light ranging technique. To use such data for the temporal study of subtle motions such as those seen in facial expressions, an efficient nonrigid 3D motion tracking algorithm is needed to establish inter-frame correspondences. The novelty of this paper is the development of an algorithmic framework for 3D tracking that unifies tracking of intensity and geometric features, using harmonic maps with added feature correspondence constraints. While the previous uses of harmonic maps provided only global alignment, the proposed introduction of interior feature constraints guarantees that nonrigid deformations are accurately tracked as well. The harmonic map between two topological disks is a diffeomorphism with minimal stretching energy and bounded angle distortion. The map is stable, insensitive to resolution changes and is robust to noise. Due to the strong implicit and explicit smoothness constraints imposed by the algorithm and the high-resolution data, the resulting registration/deformation field is smooth, continuous and gives dense one-to-one inter-frame correspondences. Our method is validated through a series of experiments demonstrating its accuracy and efficiency. Yang Wang 0001, Mohit Gupta 0001, Song Zhang 0002, Xianfeng Gu, Dimitris Samaras, Peisen Huang |
ICCV | 6 |
| 2005 | Exploiting Temporal Information in Functional Magnetic Resonance Imaging Brain Data
Lei Zhang 0002, Dimitris Samaras, Dardo Tomasi, Nelly Alia-Klein, Lisa Cottone, Andreana Leskovjan, Nora D. Volkow, Rita Z. Goldstein |
MICCAI | 2 |
| 2005 | The Role of Top-down and Bottom-up Processes in Guiding Eye Movements during Visual SearchabstractTo investigate how top-down (TD) and bottom-up (BU) information is weighted in the guidance of human search behavior, we manipulated the proportions of BU and TD components in a saliency-based model. The model is biologically plausible and implements an artificial retina and a neuronal population code. The BU component is based on feature- contrast. The TD component is defined by a feature-template match to a stored target representation. We compared the model’s behavior at differ- ent mixtures of TD and BU components to the eye movement behavior of human observers performing the identical search task. We found that a purely TD model provides a much closer match to human behavior than any mixture model using BU information. Only when biological con- straints are removed (e.g., eliminating the retina) did a BU/TD mixture model begin to approximate human behavior. Gregory J. Zelinsky, Wei Zhang 0002, Dimitris Samaras |
NIPS | 5 |
| 2005 | Modeling Neuronal Interactivity using Dynamic Bayesian NetworksabstractFunctional Magnetic Resonance Imaging (fMRI) has enabled scientists to look into the active brain. However, interactivity between functional brain regions, is still little studied. In this paper, we contribute a novel framework for modeling the interactions between multiple active brain regions, using Dynamic Bayesian Networks (DBNs) as generative mod- els for brain activation patterns. This framework is applied to modeling of neuronal circuits associated with reward. The novelty of our frame- work from a Machine Learning perspective lies in the use of DBNs to reveal the brain connectivity and interactivity. Such interactivity mod- els which are derived from fMRI data are then validated through a group classification task. We employ and compare four different types of DBNs: Parallel Hidden Markov Models, Coupled Hidden Markov Models, Fully-linked Hidden Markov Models and Dynamically Multi- Linked HMMs (DML-HMM). Moreover, we propose and compare two schemes of learning DML-HMMs. Experimental results show that by using DBNs, group classification can be performed even if the DBNs are constructed from as few as 5 brain regions. We also demonstrate that, by using the proposed learning algorithms, different DBN structures charac- terize drug addicted subjects vs. control subjects. This finding provides an independent test for the effect of psychopathology on brain function. In general, we demonstrate that incorporation of computer science prin- ciples into functional neuroimaging clinical studies provides a novel ap- proach for probing human brain function. Lei Zhang 0002, Dimitris Samaras, Nelly Alia-Klein, Nora D. Volkow, Rita Z. Goldstein |
NIPS | 2 |
| 2005 | A Computational Model of Eye Movements during Object Class DetectionabstractWe present a computational model of human eye movements in an ob- ject class detection task. The model combines state-of-the-art computer vision object class detection methods (SIFT features trained using Ad- aBoost) with a biologically plausible model of human eye movement to produce a sequence of simulated fixations, culminating with the acqui- sition of a target. We validated the model by comparing its behavior to the behavior of human observers performing the identical object class detection task (looking for a teddy bear among visually complex non- target objects). We found considerable agreement between the model and human data in multiple eye movement measures, including number of fixations, cumulative probability of fixating the target, and scanpath distance. Wei Zhang 0002, Hyejin Yang, Dimitris Samaras, Gregory J. Zelinsky |
NIPS | 3 |
| 2004 | Shape Reconstruction from 3D and 2D Data Using PDE-Based Deformable Surfaces
Ye Duan, Hong Qin 0001, Dimitris Samaras |
ECCV (3) | 4 |
| 2004 | High Resolution Acquisition, Learning and Transfer of Dynamic 3D Facial ExpressionsabstractAbstract Synthesis and re‐targeting of facial expressions is central to facial animation and often involves significant manual work in order to achieve realistic expressions, due to the difficulty of capturing high quality dynamic expression data. In this paper we address fundamental issues regarding the use of high quality dense 3‐D data samples undergoing motions at video speeds, e.g. human facial expressions. In order to utilize such data for motion analysis and re‐targeting, correspondences must be established between data in different frames of the same faces as well as between different faces. We present a data driven approach that consists of four parts: 1) High speed, high accuracy capture of moving faces without the use of markers, 2) Very precise tracking of facial motion using a multi‐resolution deformable mesh, 3) A unified low dimensional mapping of dynamic facial motion that can separate expression style, and 4) Synthesis of novel expressions as a combination of expression styles. The accuracy and resolution of our method allows us to capture and track subtle expression details. The low dimensional representation of motion data in a unified embedding for all the subjects in the database allows for learning the most discriminating characteristics of each individual's expressions as that person's “expression style”. Thus new expressions can be synthesized, either as dynamic morphing between individuals, or as expression transfer from a source face to a target face, as demonstrated in a series of experiments. Categories and Subject Descriptors (according to ACM CCS): I.3.7 [Computer Graphics]: Animation; I.3.5 [Computer Graphics]: Curve, surface, solid, and object representations; I.3.3 [Computer Graphics]: Digitizing and scanning; I.2.10 [Artificial intelligence]: Motion ; I.2.10 [Artificial intelligence]: Representations, data structures, and transforms; I.2.10 [Artificial intelligence]: Shape; I.2.6 [Artificial intelligence]: Concept learning Yang Wang 0001, Sharon X. Huang, Chan-Su Lee, Song Zhang 0002, Dimitris Samaras, Dimitris N. Metaxas, Ahmed M. Elgammal, Peisen Huang |
Comput. Graph. Forum | 6 |
| 2003 | Using Multiple Cues for Hand Tracking and Model RefinementabstractWe present a model based approach to the integration of multiple cues for tracking high degree of freedom articulated motions and model refinement. We then apply it to the problem of hand tracking using a single camera sequence. Hand tracking is particularly challenging because of occlusions, shading variations, and the high dimensionality of the motion. The novelty of our approach is in the combination of multiple sources of information, which come from edges, optical flow, and shading information in order to refine the model during tracking. We first use a previously formulated generalized version of the gradient-based optical flow constraint, that includes shading flow i.e., the variation of the shading of the object as it rotates with respect to the light source. Using this model we track its complex articulated motion in the presence of shading changes. We use a forward recursive dynamic model to track the motion in response to data derived 3D forces applied to the model. However, due to inaccurate initial shape, the generalized optical flow constraint is violated. We use the error in the generalized optical flow equation to compute generalized forces that correct the model shape at each step. The effectiveness of our approach is demonstrated with experiments on a number of different hand motions with shading changes, rotations and occlusions of significant parts of the hand. Shan Lu 0010, Dimitris N. Metaxas, Dimitris Samaras, John Oliensis |
CVPR (2) | 3 |
| 2003 | Face Recognition Under Variable Lighting using Harmonic Image ExemplarsabstractWe propose a new approach for face recognition under arbitrary illumination conditions, which requires only one training image per subject (if there is no pose variation) and no 3D shape information. Our method is based on the result of Basri and Jacobs (2001), which demonstrated that the set of images of a convex Lambertian object obtained under a wide variety of lighting conditions can be approximated accurately by a low-dimensional linear subspace. In this paper, we show that we can recover basis images spanning this space from just one image taken under arbitrary illumination conditions. First, using a bootstrap set consisting of 3D face models, we compute a statistical model for each basis image. During training, given a novel face image under arbitrary illumination, we recover a set of images for this face. We prove that these images are the set of basis images with maximum probability. During testing, we recognize the face for which there exists a weighted combination of basis images that is the closest to the test face image. We provide a series of experiments that achieve high recognition rates, under a wide range of illumination conditions, including multiple sources of illumination. Our method achieves comparable levels of accuracy with methods that have much more onerous training data requirements. Lei Zhang 0002, Dimitris Samaras |
CVPR (1) | 2 |
| 2003 | Estimation of multiple directional light sources for synthesis of augmented reality images
Yang Wang 0001, Dimitris Samaras |
Graph. Model. | 2 |
| 2003 | Incorporating Illumination Constraints in Deformable Models for Shape from Shading and Light Direction EstimationabstractWe present a method for the integration of nonlinear holonomic constraints in deformable models and its application to the problems of shape and illuminant direction estimation from shading. Experimental results demonstrate that our method performs better than previous Shape from Shading algorithms applied to images of Lambertian objects under known illumination. It is also more general as it can be applied to non-Lambertian surfaces and it does not require knowledge of the illuminant direction. In this paper, (1) we first develop a theory for the numerically robust integration of nonlinear holonomic constraints within a deformable model framework. In this formulation, we use Lagrange multipliers and a Baumgarte stabilization approach (1972). (2) We also describe a fast new method for the computation of constraint based forces, in the case of high numbers of local parameters. (3) We demonstrate how any type of illumination constraint, from the simple Lambertian model to more complex highly nonlinear models can be incorporated in a deformable model framework. (4) We extend our method to work when the direction of the light source is not known. We couple our shape estimation method with a method for light estimation, in an iterative process, where improved shape estimation results in improved light estimation and vice versa. (5) We perform a series of experiments. Dimitris Samaras, Dimitris N. Metaxas |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2002 | Estimation of Multiple Illuminants from a Single Image of Arbitrary Known Geometry
Yang Wang 0001, Dimitris Samaras |
ECCV (3) | 2 |
| 2002 | Estimation of Multiple Directional Light Sources for Synthesis of Mixed Reality ImagesabstractWe present a new method for the detection and estimation of multiple directional illuminants, using a single image of any object with known geometry and Lambertian reflectance. We use the resulting highly accurate estimates to modify virtually the illumination and geometry of a real scene and produce correctly illuminated mixed reality images. Our method obviates the need to modify the imaged scene by inserting calibration objects of any particular geometry, relying instead on partial knowledge of the geometry of the scene. Thus, the recovered multiple illuminants can be used both for image-based rendering and for shape reconstruction. Our method combines information both from the shading of the object and from shadows cast on the scene by the object. Initially we use a method based on shadows and a method based on shading independently. The shadow based method utilizes brightness variation inside the shadows cast by the object, whereas the shading based method utilizes brightness variation on the directly illuminated portions of the object. We demonstrate how the two sources of information complement each other in a number of occasions. We then describe an approach that integrates the two methods, with results superior to those obtained if the two methods are used separately. The resulting illumination information can be used (i) to render synthetic objects in a real photograph with correct illumination effects, and (ii) to virtually re-light the scene. Yang Wang 0001, Dimitris Samaras |
PG | 2 |
| 2000 | Variable Albedo Surface Reconstruction from Stereo and Shape from ShadingabstractWe present a multiview method for the computation of object shape and reflectance characteristics based on the integration of shape from shading (SFS) and stereo, for nonconstant albedo and non-uniformly Lambertian surfaces. First we perform stereo fitting on the input stereo pairs or image sequences. When the images are uncalibrated, we recover the camera parameters using bundle adjustment. Eased on the stereo result, we can automatically segment the albedo map (which is taken to be piece-wise constant) using a minimum description length (MDL) based metric, to identify areas suitable for SFS (typically smooth textureless areas) and to derive illumination information. The shape and the illumination parameter estimates are refined using a deformable model SFS algorithm, which iterates between computing shape and illumination parameters. Our method takes into account the viewing angle dependent for shortening and specularity effects, and compensates as much as possible by utilizing information from more than one images. We demonstrate that we can extend the applicability of SFS algorithms to real world situations when some of its traditional assumptions are violated. We demonstrate our method by applying it to face shape reconstruction. Experimental results indicate a significant improvement over SFS-only or stereo-only based reconstruction. Model accuracy and detail are improved, especially in areas of low texture detail. Albedo information is retrieved and can be used to accurately re-render the model under different illumination conditions. Dimitris Samaras, Dimitris N. Metaxas, Pascal Fua, Yvan G. Leclerc |
CVPR | 1 |
| 1999 | Coupled Lighting Direction and Shape Estimation from Single ImagesabstractThis paper presents a new method for the simultaneous estimation of lighting direction and shape from shading. The method estimates the shape and the lighting direction using a two step iterative process. We assume an initial (possibly incorrect) estimate of the lighting position. A stiff deformable model is then fitted to the image, assuming this lighting position. Next, a least-squares estimate of the lighting position is derived from the model using the Levenberg-Marquart method. The two steps-model fitting and lighting-position estimation-are iterated. Once the light direction has converged to a stable solution the deformable model stiffness is lowered and the model fits accurately given the lighting model. In addition, we show how the method can be used with either orthographic or perspective projection assumptions. In a variety of experiments on real and synthetic data, the method is robust to errors both to the initial light position and shape estimates. Dimitris Samaras, Dimitris N. Metaxas |
ICCV | 1 |
| 1998 | Incorporating Illumination Constraints in Deformable ModelsabstractWe present a method for the integration of illumination constraints within a deformable model framework. These constraints are incorporated as nonlinear holonomic constraints in the Lagrange equations of motion governing the deformation of the model. For improved numerical performance we employ the Baumgarte stabilization method. Our methodology is general and can be used for a broad range of illumination constraints. This approach avoids commonly used approximations in shape from shading, such as linearization, and the use of partial differential equations, which require initial boundary conditions. Furthermore, global and local parameterizations of the deformable models allow an improved estimation of shape from shading. We demonstrate this improvement over previously used approaches through a series of experiments on standardized sets of real and synthetic data. Dimitris Samaras, Dimitris N. Metaxas |
CVPR | 1 |