Hossam Isack

dblp:93/8652 · also Hossam N. Isack · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
4since 2021 · last 2024
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 12 · 5 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 4 first-author · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
11 papers
3D vision · 41% Generative modeling · 31% Segmentation and scene understanding · 14%
Computer graphics and multimedia
5 papers
Rendering · 49% Geometric modeling and processing · 26% Image and video processing · 25%

Topics — the 30 heaviest of 31, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Generative modeling
diffusion model
1.422024
Unsupervised Keypoints from Pretrained Diffusion Models · CVPR 2024
Unsupervised Semantic Correspondence Using Stable Diffusion · NeurIPS 2023
Machine learning › Generative modeling › diffusion model › text-to-image generation
text-to-image diffusion model
1.422024
Unsupervised Keypoints from Pretrained Diffusion Models · CVPR 2024
Unsupervised Semantic Correspondence Using Stable Diffusion · NeurIPS 2023
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
importance sampling
0.812024
Accelerating Neural Field Training via Soft Mining · CVPR 2024
Computer vision › 3D vision › implicit neural representation
neural field
0.812024
Accelerating Neural Field Training via Soft Mining · CVPR 2024
Computer vision › 3D vision
neural field training
0.812024
Accelerating Neural Field Training via Soft Mining · CVPR 2024
Computer vision › 3D vision › low-level vision › feature detection
unsupervised keypoint learning
0.812024
Unsupervised Keypoints from Pretrained Diffusion Models · CVPR 2024
Rendering › gaussian splatting
3d gaussian splatting
0.812024
3D Gaussian Splatting as Markov Chain Monte Carlo · NeurIPS 2024
Rendering
neural rendering
0.812024
3D Gaussian Splatting as Markov Chain Monte Carlo · NeurIPS 2024
Computer vision › 3D vision › correspondence estimation
semantic correspondence
0.712023
Unsupervised Semantic Correspondence Using Stable Diffusion · NeurIPS 2023
Image and video processing
segmentation
0.312018
K-convexity Shape Priors for Segmentation · ECCV (11) 2018
Geometric modeling and processing
shape analysis
0.312018
K-convexity Shape Priors for Segmentation · ECCV (11) 2018
Geometric modeling and processing › shape analysis
shape prior
0.312018
K-convexity Shape Priors for Segmentation · ECCV (11) 2018
Computer vision › Segmentation and scene understanding › image segmentation
hierarchical segmentation
0.312017
Efficient Optimization for Hierarchically-Structured Interacting Segments (HINTS) · CVPR 2017
Computer vision › Segmentation and scene understanding › semantic segmentation
multi-label segmentation
0.312017
Efficient Optimization for Hierarchically-Structured Interacting Segments (HINTS) · CVPR 2017
Computer vision › Segmentation and scene understanding › object segmentation
multi-object segmentation
0.212016
Hedgehog Shape Priors for Multi-Object Segmentation · CVPR 2016
Computer vision › Segmentation and scene understanding
shape prior
0.212016
Hedgehog Shape Priors for Multi-Object Segmentation · CVPR 2016
Machine learning › Deep learning architectures and training › attention mechanism
cross-attention
0.212024
Unsupervised Keypoints from Pretrained Diffusion Models · CVPR 2024
Machine learning › Probabilistic and Bayesian machine learning › monte carlo methods
markov chain monte carlo
0.212024
3D Gaussian Splatting as Markov Chain Monte Carlo · NeurIPS 2024
Image and video processing
image reconstruction
0.212024
Accelerating Neural Field Training via Soft Mining · CVPR 2024
Computer vision › Segmentation and scene understanding
image segmentation
0.212015
Volumetric Bias in Segmentation and Reconstruction: Secrets and Solutions · ICCV 2015
Computer vision › 3D vision › 3d reconstruction
reconstruction
0.212015
Volumetric Bias in Segmentation and Reconstruction: Secrets and Solutions · ICCV 2015
Computer vision › 3D vision
3d reconstruction
0.212014
Energy Based Multi-model Fitting & Matching for 3D Reconstruction · CVPR 2014
Computer vision › 3D vision
feature matching
0.212014
Energy Based Multi-model Fitting & Matching for 3D Reconstruction · CVPR 2014
Computer vision › 3D vision › geometric estimation
geometric model fitting
0.212014
Energy Based Multi-model Fitting & Matching for 3D Reconstruction · CVPR 2014
Machine learning › Optimization for machine learning
energy minimization
0.112012
Fast Approximate Energy Minimization with Label Costs · Int. J. Comput. Vis. 2012
Geometric modeling and processing › model fitting
multi-model fitting
0.112012
Energy-Based Geometric Multi-model Fitting · Int. J. Comput. Vis. 2012
Computer vision › 3D vision › geometric estimation › geometric model fitting
multi-model fitting
0.112010
Fast approximate energy minimization with label costs · CVPR 2010
Image and video processing
energy minimization
0.112010
Fast approximate energy minimization with label costs · CVPR 2010
Image and video processing
image segmentation
0.112010
Fast approximate energy minimization with label costs · CVPR 2010
Mathematical optimization › discrete optimization
energy minimization
0.112015
Volumetric Bias in Segmentation and Reconstruction: Secrets and Solutions · ICCV 2015

Methods — techniques the papers use, named apart from their topics

stochastic gradient langevin dynamics · 1.5markov chain monte carlo · 1.5langevin monte carlo · 1.5l1 regularization · 1.5importance sampling · 1.5text embedding optimization · 0.8cross-attention map regularization · 0.8alpha-expansion · 0.8prompt embedding optimization · 0.7attention · 0.7convex optimization · 0.3probabilistic k-means · 0.2graph-cut optimization · 0.2KL divergence · 0.2geometric model fitting · 0.1energy minimization · 0.1label costs · 0.1
YearPublicationVenuePosition
2024 Unsupervised Keypoints from Pretrained Diffusion Models
abstract
Unsupervised learning of keypoints and landmarks has seen significant progress with the help of modern neural network architectures, but performance is yet to match the supervised counterpart, making their practicability questionable. We leverage the emergent knowledge within text-to-image diffusion models, towards more robust unsupervised keypoints. Our core idea is to find text embeddings that would cause the generative model to consistently attend to compact regions in images (i.e. keypoints). To do so, we simply optimize the text embedding such that the cross-attention maps within the denoising network are localized as Gaussians with small standard deviations. We validate our performance on multiple datasets: the CelebA, CUB-200-2011, Tai-Chi-HD, DeepFashion, and Human3.6m datasets. We achieve significantly improved accuracy, sometimes even outperforming supervised ones, particularly for data that is non-aligned and less curated. Our code is publicly available at the project page.
Eric Hedlin, Gopal Sharma, Shweta Mahajan, Xingzhe He, Hossam Isack, Abhishek Kar, Helge Rhodin, Andrea Tagliasacchi, Kwang Moo Yi
CVPR5
2024 Accelerating Neural Field Training via Soft Mining
abstract
We present an approach to accelerate Neural Field training by efficiently selecting sampling locations. While Neural Fields have recently become popular, it is often trained by uniformly sampling the training domain, or through handcrafted heuristics. We show that improved convergence and final training quality can be achieved by a soft mining technique based on importance sampling: rather than either considering or ignoring a pixel completely, we weigh the corresponding loss by a scalar. To implement our idea we use Langevin Monte-Carlo sampling. We show that by doing so, regions with higher error are being selected more frequently, leading to more than 2x improvement in convergence speed. The code and related resources for this study are publicly available at project page.
Shakiba Kheradmand, Daniel Rebain, Gopal Sharma, Hossam Isack, Abhishek Kar, Andrea Tagliasacchi, Kwang Moo Yi
CVPR4
2024 3D Gaussian Splatting as Markov Chain Monte Carlo
abstract
While 3D Gaussian Splatting has recently become popular for neural rendering, current methods rely on carefully engineered cloning and splitting strategies for placing Gaussians, which does not always generalize and may lead to poor-quality renderings. For many real-world scenes this leads to their heavy dependence on good initializations. In this work, we rethink the set of 3D Gaussians as a random sample drawn from an underlying probability distribution describing the physical representation of the scene—in other words, Markov Chain Monte Carlo (MCMC) samples. Under this view, we show that the 3D Gaussian updates can be converted as Stochastic Gradient Langevin Dynamics (SGLD) update by simply introducing noise. We then rewrite the densification and pruning strategies in 3D Gaussian Splatting as simply a deterministic state transition of MCMC samples, removing these heuristics from the framework. To do so, we revise the ‘cloning’ of Gaussians into a relocalization scheme that approximately preserves sample probability. To encourage efficient use of Gaussians, we introduce an L1-regularizer on the Gaussians. On various standard evaluation scenes, we show that our method provides improved rendering quality, easy control over the number of Gaussians, and robustness to initialization. The project website is available at https://3dgs-mcmc.github.io/.
Shakiba Kheradmand, Daniel Rebain, Gopal Sharma, Yang-Che Tseng, Hossam Isack, Abhishek Kar, Andrea Tagliasacchi, Kwang Moo Yi
NeurIPS6
2023 Unsupervised Semantic Correspondence Using Stable Diffusion
abstract
Text-to-image diffusion models are now capable of generating images that are often indistinguishable from real images. To generate such images, these models must understand the semantics of the objects they are asked to generate. In this work we show that, without any training, one can leverage this semantic knowledge within diffusion models to find semantic correspondences – locations in multiple images that have the same semantic meaning. Specifically, given an image, we optimize the prompt embeddings of these models for maximum attention on the regions of interest. These optimized embeddings capture semantic information about the location, which can then be transferred to another image. By doing so we obtain results on par with the strongly supervised state of the art on the PF-Willow dataset and significantly outperform (20.9% relative for the SPair-71k dataset) any existing weakly- or unsupervised method on PF-Willow, CUB-200 and SPair-71k datasets.
Eric Hedlin, Gopal Sharma, Shweta Mahajan, Hossam Isack, Abhishek Kar, Andrea Tagliasacchi, Kwang Moo Yi
NeurIPS4
2018 K-convexity Shape Priors for Segmentation
Hossam Isack, Lena Gorelick, Karin Ng, Olga Veksler, Yuri Boykov
ECCV (11)1
2017 Efficient Optimization for Hierarchically-Structured Interacting Segments (HINTS)
abstract
We propose an effective optimization algorithm for a general hierarchical segmentation model with geometric interactions between segments. Any given tree can specify a partial order over object labels defining a hierarchy. It is well-established that segment interactions, such as inclusion/exclusion and margin constraints, make the model significantly more discriminant. However, existing optimization methods do not allow full use of such models. Generic a-expansion results in weak local minima, while common binary multi-layered formulations lead to non-submodularity, complex high-order potentials, or polar domain unwrapping and shape biases. In practice, applying these methods to arbitrary trees does not work except for simple cases. Our main contribution is an optimization method for the Hierarchically-structured Interacting Segments (HINTS) model with arbitrary trees. Our Path-Moves algorithm is based on multi-label MRF formulation and can be seen as a combination of well-known a-expansion and Ishikawa techniques. We show state-of-the-art biomedical segmentation for many diverse examples of complex trees.
Hossam Isack, Olga Veksler, Ipek Oguz, Milan Sonka, Yuri Boykov
CVPR1
2016 Hedgehog Shape Priors for Multi-Object Segmentation
abstract
Star-convexity prior is popular for interactive single object segmentation due to its simplicity and amenability to binary graph cut optimization. We propose a more general multi-object segmentation approach. Moreover, each object can be constrained by a more descriptive shape prior, "hedgehog". Each hedgehog shape has its surface normals locally constrained by an arbitrary given vector field, e.g. gradient of the user-scribble distance transform. In contrast to star-convexity, the tightness of our normal constraint can be changed giving better control over allowed shapes. For example, looser constraints, i.e. wider cones of allowed normals, give more relaxed hedgehog shapes. On the other hand, the tightest constraint enforces skeleton consistency with the scribbles. In general, hedgehog shapes are more descriptive than a star, which is only a special case corresponding to a radial vector field and weakest tightness. Our approach has significantly more applications than standard single star-convex segmentation, e.g. in medical data we can separate multiple non-star organs with similar appearances and weak edges. Optimization is done by our modified -expansion moves shown to be submodular for multi-hedgehog shapes.
Hossam Isack, Olga Veksler, Milan Sonka, Yuri Boykov
CVPR1
2015 Volumetric Bias in Segmentation and Reconstruction: Secrets and Solutions
abstract
Many standard optimization methods for segmentation and reconstruction compute ML model estimates for appearance or geometry of segments, e.g. Zhu-Yuille [23], Torr [20], Chan-Vese [6], GrabCut [18], Delong et al. [8]. We observe that the standard likelihood term in these formu-lations corresponds to a generalized probabilistic K-means energy. In learning it is well known that this energy has a strong bias to clusters of equal size [11], which we express as a penalty for KL divergence from a uniform distribution of cardinalities. However, this volumetric bias has been mostly ignored in computer vision. We demonstrate signif- icant artifacts in standard segmentation and reconstruction methods due to this bias. Moreover, we propose binary and multi-label optimization techniques that either (a) remove this bias or (b) replace it by a KL divergence term for any given target volume distribution. Our general ideas apply to continuous or discrete energy formulations in segmenta- tion, stereo, and other reconstruction problems.
Yuri Boykov, Hossam Isack, Carl Olsson, Ismail Ben Ayed
ICCV2
2014 Energy Based Multi-model Fitting & Matching for 3D Reconstruction
abstract
Standard geometric model fitting methods take as an input a fixed set of feature pairs greedily matched based only on their appearances. Inadvertently, many valid matches are discarded due to repetitive texture or large baseline between view points. To address this problem, matching should consider both feature appearances and geometric fitting errors. We jointly solve feature matching and multi-model fitting problems by optimizing one energy. The formulation is based on our generalization of the assignment problem and its efficient min-cost-max-flow solver. Our approach significantly increases the number of correctly matched features, improves the accuracy of fitted models, and is robust to larger baselines.
Hossam Isack, Yuri Boykov
CVPR1
2012 Fast Approximate Energy Minimization with Label Costs
Andrew Delong, Anton Osokin, Hossam Isack, Yuri Boykov
Int. J. Comput. Vis.3
2012 Energy-Based Geometric Multi-model Fitting
Hossam Isack, Yuri Boykov
Int. J. Comput. Vis.1
2010 Fast approximate energy minimization with label costs
abstract
The α-expansion algorithm has had a significant impact in computer vision due to its generality, effectiveness, and speed. Thus far it can only minimize energies that involve unary, pairwise, and specialized higher-order terms. Our main contribution is to extend α-expansion so that it can simultaneously optimize “label costs” as well. An energy with label costs can penalize a solution based on the set of labels that appear in it. The simplest special case is to penalize the number of labels in the solution. Our energy is quite general, and we prove optimality bounds for our algorithm. A natural application of label costs is multi-model fitting, and we demonstrate several such applications in vision: homography detection, motion segmentation, and unsupervised image segmentation. Our C++/MATLAB implementation is publicly available.
Andrew Delong, Anton Osokin, Hossam Isack, Yuri Boykov
CVPR3