Mojtaba Seyedhosseini

dblp:93/8617 · also Seyed Mojtaba Seyedhosseini Tarzjani · DBLP profile ↗
← Back
15ranked-venue papers
7as first author
3since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 4 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Vision and language · 29% Segmentation and scene understanding · 26% Representation and self-supervised learning · 24%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 18 heaviest of 20, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Representation and self-supervised learning › representation learning › unsupervised representation learning › self-supervised representation learning › masked modeling
masked image modeling
0.912025
TIPS: Text-Image Pretraining with Spatial awareness · ICLR 2025
Machine learning › Representation and self-supervised learning
multimodal representation learning
0.912025
TIPS: Text-Image Pretraining with Spatial awareness · ICLR 2025
Computer vision › Vision and language
vision-language pretraining
0.912025
TIPS: Text-Image Pretraining with Spatial awareness · ICLR 2025
Computer vision › Vision and language › vision-language model › vision-language foundation model
multilingual vision-language models
0.812024
On Scaling Up a Multilingual Vision and Language Model · CVPR 2024
Information retrieval
image retrieval
0.712023
Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations · ICCV 2023
Information retrieval
retrieval evaluation
0.712023
Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations · ICCV 2023
Computer vision › Segmentation and scene understanding
image segmentation
0.632016
Image Segmentation Using Hierarchical Merge Tree · IEEE Trans. Image Process. 2016
Multi-Class Multi-Scale Series Contextual Model for Image Segmentation · IEEE Trans. Image Process. 2013
Image Segmentation with Cascaded Hierarchical Models and Logistic Disjunctive Normal Networks · ICCV 2013
Computer vision › Segmentation and scene understanding › image segmentation
contextual segmentation
0.322013
Multi-Class Multi-Scale Series Contextual Model for Image Segmentation · IEEE Trans. Image Process. 2013
Image Segmentation with Cascaded Hierarchical Models and Logistic Disjunctive Normal Networks · ICCV 2013
Computer vision › Video understanding and tracking
hierarchical context model
0.212016
Semantic Image Segmentation with Contextual Hierarchical Models · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Segmentation and scene understanding › image segmentation
hierarchical segmentation
0.212016
Image Segmentation Using Hierarchical Merge Tree · IEEE Trans. Image Process. 2016
Computer vision › Segmentation and scene understanding › image segmentation › region-based segmentation
region merging
0.212016
Image Segmentation Using Hierarchical Merge Tree · IEEE Trans. Image Process. 2016
Computer vision › Segmentation and scene understanding
semantic segmentation
0.212016
Semantic Image Segmentation with Contextual Hierarchical Models · IEEE Trans. Pattern Anal. Mach. Intell. 2016
Computer vision › Vision and language
image captioning
0.212024
On Scaling Up a Multilingual Vision and Language Model · CVPR 2024
Computer vision › Image recognition and object detection
object detection
0.212024
On Scaling Up a Multilingual Vision and Language Model · CVPR 2024
Computer vision › Vision and language
visual question answering
0.212024
On Scaling Up a Multilingual Vision and Language Model · CVPR 2024
Computer vision › Segmentation and scene understanding › semantic segmentation
multi-class segmentation
0.212013
Multi-Class Multi-Scale Series Contextual Model for Image Segmentation · IEEE Trans. Image Process. 2013
Machine learning › Learning theory › classification
neural network classifier
0.212013
Image Segmentation with Cascaded Hierarchical Models and Logistic Disjunctive Normal Networks · ICCV 2013
Computer vision › Segmentation and scene understanding
edge detection
0.112016
Semantic Image Segmentation with Contextual Hierarchical Models · IEEE Trans. Pattern Anal. Mach. Intell. 2016

Methods — techniques the papers use, named apart from their topics

synthetic captions · 0.9contrastive learning · 0.9scaling · 0.8multimodal pretraining · 0.8training and evaluation protocol · 0.7metric definition · 0.7multi-resolution contextual information · 0.2hierarchical merge tree · 0.2hierarchical classifier · 0.2ensemble boundary classifier · 0.2constrained conditional models · 0.2cascaded hierarchical model · 0.2
YearPublicationVenuePosition
2025 TIPS: Text-Image Pretraining with Spatial awareness
abstract
While image-text representation learning has become very popular in recent years, existing models tend to lack spatial awareness and have limited direct applicability for dense understanding tasks. For this reason, self-supervised image-only pretraining is still the go-to method for many dense vision applications (e.g. depth estimation, semantic segmentation), despite the lack of explicit supervisory signals. In this paper, we close this gap between image-text and self-supervised learning, by proposing a novel general-purpose image-text model, which can be effectively used off the shelf for dense and global vision tasks. Our method, which we refer to as Text-Image Pretraining with Spatial awareness (TIPS), leverages two simple and effective insights. First, on textual supervision: we reveal that replacing noisy web image captions by synthetically generated textual descriptions boosts dense understanding performance significantly, due to a much richer signal for learning spatially aware representations. We propose an adapted training method that combines noisy and synthetic captions, resulting in improvements across both dense and global understanding tasks. Second, on the learning technique: we propose to combine contrastive image-text learning with self-supervised masked image modeling, to encourage spatial coherence, unlocking substantial enhancements for downstream applications. Building on these two ideas, we scale our model using the transformer architecture, trained on a curated set of public images. Our experiments are conducted on $8$ tasks involving $16$ datasets in total, demonstrating strong off-the-shelf performance on both dense and global understanding, for several image-only and image-text tasks. Code and models are released at https://github.com/google-deepmind/tips .
Kevis-Kokitsi Maninis, Kaifeng Chen, Soham Ghosh 0001, Arjun Karpur, Koert Chen, Bingyi Cao, Daniel Salz, Guangxing Han, Jan Dlabal, Dan Gnanapragasam, Mojtaba Seyedhosseini, Howard Zhou, André Araújo 0001
ICLR12
2024 On Scaling Up a Multilingual Vision and Language Model
abstract
We explore the boundaries of scaling up a multilingual vision and language model, both in terms of size of the components and the breadth of its training task mixture. Our model achieves new levels of performance on a wide-range of varied and complex tasks, including multiple image-based captioning and question-answering tasks, image-based document understanding and few-shot (in-context) learning, as well as object detection, video question answering, and video captioning. Our model advances the state-of-the-art on most vision-and-language benchmarks considered (20+ of them). Finally, we observe emerging capabilities, such as complex counting and multilingual object detection, tasks that are not explicitly in the training mix.
Xi Chen 0071, Josip Djolonga, Piotr Padlewski, Basil Mustafa, Soravit Changpinyo, Carlos Riquelme, Sebastian Goodman, Xiao Wang 0038, Yi Tay, Siamak Shakeri, Mostafa Dehghani 0001, Daniel Salz, Mario Lucic, Michael Tschannen, Arsha Nagrani, Hexiang Hu, Mandar Joshi, Bo Pang 0001, Ceslee Montgomery, Paulina Pietrzyk, Marvin Ritter, A. J. Piergiovanni, Matthias Minderer, Filip Pavetic, Austin Waters, Gang Li 0021, Ibrahim Alabdulmohsin, Lucas Beyer, Julien Amelot, Kenton Lee, Andreas Steiner 0001, Yang Li 0058, Daniel Keysers, Anurag Arnab, Yuanzhong Xu, Keran Rong, Alexander Kolesnikov 0003, Mojtaba Seyedhosseini, Anelia Angelova, Xiaohua Zhai, Neil Houlsby, Radu Soricut
CVPR39
2023 Towards Universal Image Embeddings: A Large-Scale Dataset and Challenge for Generic Image Representations
abstract
Fine-grained and instance-level recognition methods are commonly trained and evaluated on specific domains, in a model per domain scenario. Such an approach, however, is impractical in real large-scale applications. In this work, we address the problem of universal image embedding, where a single universal model is trained and used in multiple domains. First, we leverage existing domain-specific datasets to carefully construct a new large-scale public benchmark for the evaluation of universal image embeddings, with 241k query images, 1.4M index images and 2.8M training images across 8 different domains and 349k classes. We define suitable metrics, training and evaluation protocols to foster future research in this area. Second, we provide a comprehensive experimental evaluation on the new dataset, demonstrating that existing approaches and simplistic extensions lead to worse performance than an assembly of models trained for each domain separately. Finally, we conducted a public research competition on this topic, leveraging industrial datasets, which attracted the participation of more than 1k teams world-wide. This exercise generated many interesting research ideas and findings which we present in detail. Project webpage: https://cmp.felk.cvut.cz/univ_emb/
Nikolaos-Antonios Ypsilantis, Kaifeng Chen, Bingyi Cao, Mário Lipovský, Pelin Dogan-Schönberger, Grzegorz Makosa, Boris Bluntschli, Mojtaba Seyedhosseini, Ondrej Chum, André Araújo 0001
ICCV8
2016 Disjunctive normal networks
Mehdi Sajjadi, Mojtaba Seyedhosseini, Tolga Tasdizen
Neurocomputing2
2016 Semantic Image Segmentation with Contextual Hierarchical Models
abstract
Semantic segmentation is the problem of assigning an object label to each pixel. It unifies the image segmentation and object recognition problems. The importance of using contextual information in semantic segmentation frameworks has been widely realized in the field. We propose a contextual framework, called contextual hierarchical model (CHM), which learns contextual information in a hierarchical framework for semantic segmentation. At each level of the hierarchy, a classifier is trained based on downsampled input images and outputs of previous levels. Our model then incorporates the resulting multi-resolution contextual information into a classifier to segment the input image at original resolution. This training strategy allows for optimization of a joint posterior probability at multiple resolutions through the hierarchy. Contextual hierarchical model is purely based on the input image patches and does not make use of any fragments or shape examples. Hence, it is applicable to a variety of problems such as object segmentation and edge detection. We demonstrate that CHM performs at par with state-of-the-art on Stanford background and Weizmann horse datasets. It also outperforms state-of-the-art edge detection methods on NYU depth dataset and achieves state-of-the-art on Berkeley segmentation dataset (BSDS 500).
Mojtaba Seyedhosseini, Tolga Tasdizen
IEEE Trans. Pattern Anal. Mach. Intell.1
2016 Image Segmentation Using Hierarchical Merge Tree
abstract
This paper investigates one of the most fundamental computer vision problems: image segmentation. We propose a supervised hierarchical approach to object-independent image segmentation. Starting with oversegmenting superpixels, we use a tree structure to represent the hierarchy of region merging, by which we reduce the problem of segmenting image regions to finding a set of label assignment to tree nodes. We formulate the tree structure as a constrained conditional model to associate region merging with likelihoods predicted using an ensemble boundary classifier. Final segmentations can then be inferred by finding globally optimal solutions to the model efficiently. We also present an iterative training and testing algorithm that generates various tree structures and combines them to emphasize accurate boundaries by segmentation accumulation. Experiment results and comparisons with other recent methods on six public data sets demonstrate that our approach achieves the state-of-the-art region accuracy and is competitive in image segmentation without semantic priors.
Ting Liu 0005, Mojtaba Seyedhosseini, Tolga Tasdizen
IEEE Trans. Image Process.2
2015 Disjunctive normal random forests
Mojtaba Seyedhosseini, Tolga Tasdizen
Pattern Recognit.1
2015 Nonlinear Regression with Logistic Product Basis Networks
abstract
We introduce a novel general regression model that is based on a linear combination of a new set of non-local basis functions that forms an effective feature space. We propose a training algorithm that learns all the model parameters simultaneously and offer an initialization scheme for parameters of the basis functions. We show through several experiments that the proposed method offers better coverage for high-dimensional space compared to local Gaussian basis functions and provides competitive performance in comparison to other state-of-the-art regression methods.
Mehdi Sajjadi, Mojtaba Seyedhosseini, Tolga Tasdizen
IEEE Signal Process. Lett.2
2013 Image Segmentation with Cascaded Hierarchical Models and Logistic Disjunctive Normal Networks
abstract
Contextual information plays an important role in solving vision problems such as image segmentation. However, extracting contextual information and using it in an effective way remains a difficult problem. To address this challenge, we propose a multi-resolution contextual framework, called cascaded hierarchical model (CHM), which learns contextual information in a hierarchical framework for image segmentation. At each level of the hierarchy, a classifier is trained based on downsampled input images and outputs of previous levels. Our model then incorporates the resulting multi-resolution contextual information into a classifier to segment the input image at original resolution. We repeat this procedure by cascading the hierarchical framework to improve the segmentation accuracy. Multiple classifiers are learned in the CHM; therefore, a fast and accurate classifier is required to make the training tractable. The classifier also needs to be robust against overfitting due to the large number of parameters learned during training. We introduce a novel classification scheme, called logistic disjunctive normal networks (LDNN), which consists of one adaptive layer of feature detectors implemented by logistic sigmoid functions followed by two fixed layers of logical units that compute conjunctions and disjunctions, respectively. We demonstrate that LDNN outperforms state-of-theart classifiers and can be used in the CHM to improve object segmentation performance.
Mojtaba Seyedhosseini, Mehdi Sajjadi, Tolga Tasdizen
ICCV1
2013 Watershed merge forest classification for electron microscopy image stack segmentation
abstract
Automated electron microscopy (EM) image analysis techniques can be tremendously helpful for connectomics research. In this paper, we extend our previous work [1] and propose a fully automatic method to utilize inter-section information for intra-section neuron segmentation of EM image stacks. A watershed merge forest is built via the watershed transform with each tree representing the region merging hierarchy of one 2D section in the stack. A section classifier is learned to identify the most likely region correspondence between adjacent sections. The inter-section information from such correspondence is incorporated to update the potentials of tree nodes. We resolve the merge forest using these potentials together with consistency constraints to acquire the final segmentation of the whole stack. We demonstrate that our method leads to notable segmentation accuracy improvement by experimenting with two types of EM image data sets.
Ting Liu 0005, Mojtaba Seyedhosseini, Mark H. Ellisman, Tolga Tasdizen
ICIP2
2013 Multi-Class Multi-Scale Series Contextual Model for Image Segmentation
abstract
Contextual information has been widely used as a rich source of information to segment multiple objects in an image. A contextual model uses the relationships between the objects in a scene to facilitate object detection and segmentation. Using contextual information from different objects in an effective way for object segmentation, however, remains a difficult problem. In this paper, we introduce a novel framework, called multiclass multiscale (MCMS) series contextual model, which uses contextual information from multiple objects and at different scales for learning discriminative models in a supervised setting. The MCMS model incorporates cross-object and inter-object information into one probabilistic framework and thus is able to capture geometrical relationships and dependencies among multiple objects in addition to local information from each single object present in an image. We demonstrate that our MCMS model improves object segmentation performance in electron microscopy images and provides a coherent segmentation of multiple objects. Through speeding up the segmentation process, the proposed method will allow neurobiologists to move beyond individual specimens and analyze populations paving the way for understanding neurodegenerative diseases at the microscopic level.
Mojtaba Seyedhosseini, Tolga Tasdizen
IEEE Trans. Image Process.1
2012 Watershed merge tree classification for electron microscopy image segmentation
Ting Liu 0005, Elizabeth Jurrus, Mojtaba Seyedhosseini, Mark H. Ellisman, Tolga Tasdizen
ICPR3
2011 Fast AdaBoost training using weighted novelty selection
abstract
In this paper, a new AdaBoost learning framework, called WNS-AdaBoost, is proposed for training discriminative models. The proposed approach significantly speeds up the learning process of adaptive boosting (AdaBoost) by reducing the number of data points. For this purpose, we introduce the weighted novelty selection (WNS) sampling strategy and combine it with AdaBoost to obtain an efficient and fast learning algorithm. WNS selects a representative subset of data thereby reducing the number of data points onto which AdaBoost is applied. In addition, WNS associates a weight with each selected data point such that the weighted subset approximates the distribution of all the training data. This ensures that AdaBoost can trained efficiently and with minimal loss of accuracy. The performance of WNS-AdaBoost is first demonstrated in a classification task. Then, WNS is employed in a probabilistic boosting-tree (PBT) structure for image segmentation. Results in these two applications show that the training time using WNS-AdaBoost is greatly reduced at the cost of only a few percent in accuracy.
Mojtaba Seyedhosseini, António R. C. Paiva, Tolga Tasdizen
IJCNN1
2011 Detection of Neuron Membranes in Electron Microscopy Images Using Multi-scale Context and Radon-Like Features
Mojtaba Seyedhosseini, Ritwik Kumar, Elizabeth Jurrus, Richard J. Giuly, Mark H. Ellisman, Hanspeter Pfister, Tolga Tasdizen
MICCAI (1)1
2010 Image Parsing with a Three-State Series Neural Network Classifier
abstract
We propose a three-state series neural network for effective propagation of context and uncertainty information for image parsing. The activation functions used in the proposed model have three states instead of the normal two states. This makes the neural network more flexible than the two-state neural network, and allows for uncertainty to be propagated through the stages. In other words, decisions about difficult pixels can be left for later stages which have access to more contextual information than earlier stages. We applied the proposed method to three different datasets and experimental results demonstrate higher performance of the three-state series neural network.
Mojtaba Seyedhosseini, António R. C. Paiva, Tolga Tasdizen
ICPR1