VLDB 2026 Research / reviewers in the wild / expert
Golnaz Ghiasi
dblp:17/8614
· DBLP profile ↗
22ranked-venue papers
12as first author
7since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 12 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 10 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
15 papers |
Segmentation and scene understanding · 23% Deep learning architectures and training · 16% Image recognition and object detection · 14% | |
| Computer graphics and multimedia
2 papers |
Visual content generation and editing · 92% Image and video processing · 8% |
Topics — the 30 heaviest of 33, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection
object detection |
1.8 | 5 | 2020 | Learning Data Augmentation Strategies for Object Detection · ECCV (27) 2020 SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization · CVPR 2020 MnasFPN: Learning Latency-Aware Pyramid Architecture for Object Detection on Mobile Devices · CVPR 2020 |
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search |
1.2 | 3 | 2020 | SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization · CVPR 2020 MnasFPN: Learning Latency-Aware Pyramid Architecture for Object Detection on Mobile Devices · CVPR 2020 NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection · CVPR 2019 |
Computer vision › Segmentation and scene understanding › semantic segmentation
open-vocabulary segmentation |
1.2 | 2 | 2023 | DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model · NeurIPS 2023 Scaling Open-Vocabulary Image Segmentation with Image-Level Labels · ECCV (36) 2022 |
Machine learning › Deep learning architectures and training
data augmentation |
0.9 | 2 | 2021 | Simple Copy-Paste Is a Strong Data Augmentation Method for Instance Segmentation · CVPR 2021 Learning Data Augmentation Strategies for Object Detection · ECCV (27) 2020 |
Machine learning › Transfer learning and domain adaptation › domain adaptation › unsupervised domain adaptation
self-training |
0.9 | 2 | 2021 | Multi-Task Self-Training for Learning General Representations · ICCV 2021 Rethinking Pre-training and Self-training · NeurIPS 2020 |
Natural language and speech › Language models and text generation
mathematical reasoning |
0.9 | 1 | 2025 | Towards Robust Mathematical Reasoning · EMNLP 2025 |
Natural language and speech › Language models and text generation › mathematical reasoning
robust mathematical reasoning |
0.9 | 1 | 2025 | Towards Robust Mathematical Reasoning · EMNLP 2025 |
Machine learning › Trustworthy machine learning › hallucination
hallucination evaluation |
0.8 | 1 | 2024 | HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning · ECCV (77) 2024 |
Computer vision › Vision and language
multimodal hallucination |
0.8 | 1 | 2024 | HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning · ECCV (77) 2024 |
Computer vision › Segmentation and scene understanding › image segmentation
multi-task segmentation |
0.7 | 1 | 2023 | DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model · NeurIPS 2023 |
Computer vision › Segmentation and scene understanding
panoptic segmentation |
0.7 | 1 | 2023 | DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model · NeurIPS 2023 |
Computer vision › Image recognition and object detection › object detection
feature pyramid network |
0.5 | 2 | 2020 | NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object Detection · CVPR 2019 SpineNet: Learning Scale-Permuted Backbone for Recognition and Localization · CVPR 2020 |
Machine learning › Deep learning architectures and training › data augmentation
copy-paste augmentation |
0.5 | 1 | 2021 | Simple Copy-Paste Is a Strong Data Augmentation Method for Instance Segmentation · CVPR 2021 |
Computer vision › Segmentation and scene understanding
instance segmentation |
0.5 | 1 | 2021 | Simple Copy-Paste Is a Strong Data Augmentation Method for Instance Segmentation · CVPR 2021 |
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
latency-aware architecture search |
0.4 | 1 | 2020 | MnasFPN: Learning Latency-Aware Pyramid Architecture for Object Detection on Mobile Devices · CVPR 2020 |
Machine learning › Representation and self-supervised learning
pre-training |
0.4 | 1 | 2020 | Rethinking Pre-training and Self-training · NeurIPS 2020 |
Visual content generation and editing › style transfer
real-time style transfer |
0.4 | 1 | 2020 | Adjustable Real-time Style Transfer · ICLR 2020 |
Visual content generation and editing
style transfer |
0.4 | 1 | 2020 | Adjustable Real-time Style Transfer · ICLR 2020 |
Computer vision › Segmentation and scene understanding
semantic segmentation |
0.4 | 2 | 2020 | Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation · ECCV (3) 2016 Rethinking Pre-training and Self-training · NeurIPS 2020 |
Machine learning › Deep learning architectures and training
convolutional neural network |
0.3 | 1 | 2018 | DropBlock: A regularization method for convolutional networks · NeurIPS 2018 |
Machine learning › Deep learning architectures and training
regularization |
0.3 | 1 | 2018 | DropBlock: A regularization method for convolutional networks · NeurIPS 2018 |
Machine learning › Deep learning architectures and training › regularization › dropout
structured dropout |
0.3 | 1 | 2018 | DropBlock: A regularization method for convolutional networks · NeurIPS 2018 |
Machine learning › Transfer learning and domain adaptation › multi-source learning
multi-dataset learning |
0.2 | 1 | 2023 | DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation Model · NeurIPS 2023 |
Computer vision › Segmentation and scene understanding › image segmentation › binary segmentation
foreground-background segmentation |
0.2 | 1 | 2014 | Parsing Occluded People · CVPR 2014 |
Computer vision › Face, body and person analysis
human pose estimation |
0.2 | 1 | 2014 | Parsing Occluded People · CVPR 2014 |
Computer vision › Face, body and person analysis › face detection
occluded face detection |
0.2 | 1 | 2014 | Occlusion Coherence: Localizing Occluded Faces with a Hierarchical Deformable Part Model · CVPR 2014 |
Computer vision › Segmentation and scene understanding › object segmentation
occlusion-aware segmentation |
0.2 | 1 | 2014 | Parsing Occluded People · CVPR 2014 |
Computer vision › Vision and language › visual grounding
vision-language segmentation |
0.2 | 1 | 2022 | Scaling Open-Vocabulary Image Segmentation with Image-Level Labels · ECCV (36) 2022 |
Machine learning › Learning paradigms
multi-task learning |
0.1 | 1 | 2021 | Multi-Task Self-Training for Learning General Representations · ICCV 2021 |
Image and video processing › multiscale analysis
multiscale image processing |
0.1 | 1 | 2016 | Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation · ECCV (3) 2016 |
Methods — techniques the papers use, named apart from their topics
pseudo-labeling · 1.0neural architecture search · 0.8visual question answering · 0.8multimodal reasoning · 0.8weak supervision · 0.7text embedding · 0.7mask proposal · 0.7image-level weak supervision · 0.6semi-supervised learning · 0.5self-training · 0.5laplacian pyramid · 0.2
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Towards Robust Mathematical ReasoningabstractThang Luong, Dawsen Hwang, Hoang H Nguyen, Golnaz Ghiasi, Yuri Chervonyi, Insuk Seo, Junsu Kim, Garrett Bingham, Jonathan Lee, Swaroop Mishra, Alex Zhai, Huiyi Hu, Henryk Michalewski, Jimin Kim, Jeonghyun Ahn, Junhwi Bae, Xingyou Song, Trieu Hoang Trinh, Quoc V Le, Junehyuk Jung. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Thang Luong, Dawsen Hwang, Golnaz Ghiasi, Yuri Chervonyi, Insuk Seo, Garrett Bingham, Jonathan Lee 0002, Swaroop Mishra, Alex Zhai, Clara Huiyi Hu, Henryk Michalewski, Jeonghyun Ahn, Junhwi Bae, Xingyou Song, Trieu H. Trinh, Quoc V. Le, Junehyuk Jung |
EMNLP | 4 |
| 2024 | HaloQuest: A Visual Hallucination Dataset for Advancing Multimodal Reasoning
Zhecan Wang, Garrett Bingham, Adams Wei Yu, Quoc V. Le, Thang Luong, Golnaz Ghiasi |
ECCV (77) | 6 |
| 2023 | DaTaSeg: Taming a Universal Multi-Dataset Multi-Task Segmentation ModelabstractObserving the close relationship among panoptic, semantic and instance segmentation tasks, we propose to train a universal multi-dataset multi-task segmentation model: DaTaSeg. We use a shared representation (mask proposals with class predictions) for all tasks. To tackle task discrepancy, we adopt different merge operations and post-processing for different tasks. We also leverage weak-supervision, allowing our segmentation model to benefit from cheaper bounding box annotations. To share knowledge across datasets, we use text embeddings from the same semantic embedding space as classifiers and share all network parameters among datasets. We train DaTaSeg on ADE semantic, COCO panoptic, and Objects365 detection datasets. DaTaSeg improves performance on all datasets, especially small-scale datasets, achieving 54.0 mIoU on ADE semantic and 53.5 PQ on COCO panoptic. DaTaSeg also enables weakly-supervised knowledge transfer on ADE panoptic and Objects365 instance segmentation. Experiments show DaTaSeg scales with the number of training datasets and enables open-vocabulary segmentation through direct transfer. In addition, we annotate an Objects365 instance segmentation set of 1,000 images and release it as a public evaluation benchmark on https://laoreja.github.io/dataseg. Xiuye Gu, Yin Cui, Jonathan Huang, Abdullah Rashwan, Xingyi Zhou, Golnaz Ghiasi, Weicheng Kuo, Huizhong Chen, Liang-Chieh Chen, David A. Ross |
NeurIPS | 7 |
| 2023 | Combined scaling for zero-shot transfer learningabstractRecent developments in multimodal training methodologies, including CLIP and ALIGN, obviate the necessity for individual data labeling. These approaches utilize pairs of data and corresponding textual information found online as a form of weak supervision signal. However, models employing this kind of weak supervision are not as competitive as their supervised and semi-supervised counterparts when sufficient labeled data is accessible. This performance gap constrains the applicability of weekly supervised models. In this paper, we narrow the gap by proposing a combined scaling method, named BASIC, that achieves 85.7% top-1 accuracy on the ImageNet ILSVRC-2012 validation set without learning from any labeled ImageNet example. This accuracy surpasses best-published similar models, CLIP and ALIGN, by 9.3%. Our BASIC model also shows significant improvements in robustness benchmarks. For instance, on 5 test sets with natural distribution shifts such as ImageNet-{A,R,V2,Sketch} and ObjectNet, our model achieves 84.3% top-1 average accuracy, only a small drop from its original ImageNet accuracy. To achieve these results, we first develop a theoretical framework which shows that larger contrastive batch sizes lead to smaller generalization gaps for image-text models such as CLIP and ALIGN. Based on this theoretical result, we scale up the contrastive learning framework of CLIP and ALIGN in three dimensions (data size, model size, and batch size) by proposing a new method using gradient checkpointing and model parallelism. As a result, our dataset has 6.6B noisy image-text pairs, which is 4x larger than ALIGN, and 16x larger than CLIP. Our largest model has 3B weights, which is 3.75x larger in parameters and 8x larger in FLOPs than ALIGN and CLIP. Finally, our batch size is 65536 which is 2x more than CLIP and 4x more than ALIGN. Hieu Pham 0001, Zihang Dai, Golnaz Ghiasi, Kenji Kawaguchi, Hanxiao Liu, Adams Wei Yu, Minh-Thang Luong, Mingxing Tan, Quoc V. Le |
Neurocomputing | 3 |
| 2022 | Scaling Open-Vocabulary Image Segmentation with Image-Level Labels
Golnaz Ghiasi, Xiuye Gu, Yin Cui, Tsung-Yi Lin |
ECCV (36) | 1 |
| 2021 | Simple Copy-Paste Is a Strong Data Augmentation Method for Instance SegmentationabstractBuilding instance segmentation models that are data-efficient and can handle rare object categories is an important challenge in computer vision. Leveraging data augmentations is a promising direction towards addressing this challenge. Here, we perform a systematic study of the Copy-Paste augmentation (e.g., [13], [12]) for instance segmentation where we randomly paste objects onto an image. Prior studies on Copy-Paste relied on modeling the surrounding visual context for pasting the objects. However, we find that the simple mechanism of pasting objects randomly is good enough and can provide solid gains on top of strong baselines. Furthermore, we show Copy-Paste is additive with semi-supervised methods that leverage extra data through pseudo labeling (e.g. self-training). On COCO instance segmentation, we achieve 49.1 mask AP and 57.3 box AP, an improvement of +0.6 mask AP and +1.5 box AP over the previous state-of-the-art. We further demonstrate that Copy-Paste can lead to significant improvements on the LVIS benchmark. Our baseline model outperforms the LVIS 2020 Challenge winning entry by +3.6 mask AP on rare categories.1 Golnaz Ghiasi, Yin Cui, Aravind Srinivas, Rui Qian 0003, Tsung-Yi Lin, Ekin Dogus Cubuk, Quoc V. Le, Barret Zoph |
CVPR | 1 |
| 2021 | Multi-Task Self-Training for Learning General RepresentationsabstractDespite the fast progress in training specialized models for various tasks, learning a single general model that works well for many tasks is still challenging for computer vision. Here we introduce multi-task self-training (MuST), which harnesses the knowledge in independent specialized teacher models (e.g., ImageNet model on classification) to train a single general student model. Our approach has three steps. First, we train specialized teachers independently on labeled datasets. We then use the specialized teachers to label an unlabeled dataset to create a multi-task pseudo labeled dataset. Finally, the dataset, which now contains pseudo labels from teacher models trained on different datasets/tasks, is then used to train a student model with multi-task learning. We evaluate the feature representations of the student model on 6 vision tasks including image recognition (classification, detection, segmentation) and 3D geometry estimation (depth and surface normal estimation). MuST is scalable with unlabeled or partially labeled datasets and outperforms both specialized supervised models and self-supervised models when training on large scale datasets. Lastly, we show MuST can improve upon already strong checkpoints [23] trained with billions of examples. The results suggest self-training is a promising direction to aggregate labeled and unlabeled training data for learning general feature representations. Golnaz Ghiasi, Barret Zoph, Ekin Dogus Cubuk, Quoc V. Le, Tsung-Yi Lin |
ICCV | 1 |
| 2020 | MnasFPN: Learning Latency-Aware Pyramid Architecture for Object Detection on Mobile DevicesabstractDespite the blooming success of architecture search for vision tasks in resource-constrained environments, the design of on-device object detection architectures have mostly been manual. The few automated search efforts are either centered around non-mobile-friendly search spaces or not guided by on-device latency. We propose MnasFPN, a mobile-friendly search space for the detection head, and combine it with latency-aware architecture search to produce efficient object detection models. The learned MnasFPN head, when paired with MobileNetV2 body, outperforms MobileNetV3+SSDLite by 1.8 mAP at similar latency on Pixel. It is both 1 mAP more accurate and 10\% faster than NAS-FPNLite. Ablation studies show that the majority of the performance gain comes from innovations in the search space. Further explorations reveal an interesting coupling between the search space design and the search algorithm, for which the complexity of MnasFPN search space is opportune. Bo Chen 0019, Golnaz Ghiasi, Hanxiao Liu, Tsung-Yi Lin, Dmitry Kalenichenko, Hartwig Adam, Quoc V. Le |
CVPR | 2 |
| 2020 | SpineNet: Learning Scale-Permuted Backbone for Recognition and LocalizationabstractConvolutional neural networks typically encode an input image into a series of intermediate features with decreasing resolutions. While this structure is suited to classification tasks, it does not perform well for tasks requiring simultaneous recognition and localization (e.g., object detection). The encoder-decoder architectures are proposed to resolve this by applying a decoder network onto a backbone model designed for classification tasks. In this paper, we argue encoder-decoder architecture is ineffective in generating strong multi-scale features because of the scale-decreased backbone. We propose SpineNet, a backbone with scale-permuted intermediate features and cross-scale connections that is learned on an object detection task by Neural Architecture Search. Using similar building blocks, SpineNet models outperform ResNet-FPN models by 3%+ AP at various scales while using 10-20% fewer FLOPs. In particular, SpineNet-190 achieves 52.1% AP on COCO, attaining the new state-of-the-art performance for single model object detection without test-time augmentation. SpineNet can transfer to classification tasks, achieving 5% top-1 accuracy improvement on a challenging iNaturalist fine-grained dataset. Code is at: https://github.com/tensorflow/tpu/tree/master/models/official/detection. Xianzhi Du, Tsung-Yi Lin, Pengchong Jin, Golnaz Ghiasi, Mingxing Tan, Yin Cui, Quoc V. Le, Xiaodan Song |
CVPR | 4 |
| 2020 | Learning Data Augmentation Strategies for Object Detection
Barret Zoph, Ekin Dogus Cubuk, Golnaz Ghiasi, Tsung-Yi Lin, Jonathon Shlens, Quoc V. Le |
ECCV (27) | 3 |
| 2020 | Adjustable Real-time Style Transfer
Mohammad Babaeizadeh, Golnaz Ghiasi |
ICLR | 2 |
| 2020 | Rethinking Pre-training and Self-trainingabstractPre-training is a dominant paradigm in computer vision. For example, supervised ImageNet pre-training is commonly used to initialize the backbones of object detection and segmentation models. He et al., however, show a striking result that ImageNet pre-training has limited impact on COCO object detection. Here we investigate self-training as another method to utilize additional data on the same setup and contrast it against ImageNet pre-training. Our study reveals the generality and flexibility of self-training with three additional insights: 1) stronger data augmentation and more labeled data further diminish the value of pre-training, 2) unlike pre-training, self-training is always helpful when using stronger data augmentation, in both low-data and high-data regimes, and 3) in the case that pre-training is helpful, self-training improves upon pre-training. For example, on the COCO object detection dataset, pre-training benefits when we use one fifth of the labeled data, and hurts accuracy when we use all labeled data. Self-training, on the other hand, shows positive improvements from +1.3 to +3.4AP across all dataset sizes. In other words, self-training works well exactly on the same setup that pre-training does not work (using ImageNet to help COCO). On the PASCAL segmentation dataset, which is a much smaller dataset than COCO, though pre-training does help significantly, self-training improves upon the pre-trained model. On COCO object detection, we achieve 53.8AP, an improvement of +1.7AP over the strongest SpineNet model. On PASCAL segmentation, we achieve 90.5mIOU, an improvement of +1.5mIOU over the previous state-of-the-art result by DeepLabv3+. Barret Zoph, Golnaz Ghiasi, Tsung-Yi Lin, Yin Cui, Hanxiao Liu, Ekin Dogus Cubuk, Quoc V. Le |
NeurIPS | 2 |
| 2020 | Writer identification with n-tuple direction feature from contourabstractThis study introduces an effective solution for text‐independent writer identification by generalising contour‐hinge feature, which is called n ‐tuple direction feature. For extracting n ‐tuple direction feature, the authors first obtain all contours from connected components, then n + 1 points are considered on the contour with a certain distance apart, and next, the directions of the fragments connecting two successive points are computed. The n + 1 points move on the contour and the n ‐dimensional histogram of directions is computed. The proposed method is evaluated on large Farsi and English databases. A correct writer identification rate of 92.2% for English handwritings from 900 persons and 97.7% for Farsi handwritings from 600 persons are achieved. Comparison between the proposed method and other studies shows the promising performance and superiority of the proposed method. Ali Reza Ghanbarian, Golnaz Ghiasi, Reza Safabakhsh, Narges Arastouie |
IET Image Process. | 2 |
| 2019 | NAS-FPN: Learning Scalable Feature Pyramid Architecture for Object DetectionabstractCurrent state-of-the-art convolutional architectures for object detection are manually designed. Here we aim to learn a better architecture of feature pyramid network for object detection. We adopt Neural Architecture Search and discover a new feature pyramid architecture in a novel scalable search space covering all cross-scale connections. The discovered architecture, named NAS-FPN, consists of a combination of top-down and bottom-up connections to fuse features across scales. NAS-FPN, combined with various backbone models in the RetinaNet framework, achieves better accuracy and latency tradeoff compared to state-of-the-art object detection models. NAS-FPN improves mobile detection accuracy by 2 AP compared to state-of-the-art SSDLite with MobileNetV2 model in [32] and achieves 48.3 AP which surpasses Mask R-CNN [10] detection accuracy with less computation time. Golnaz Ghiasi, Tsung-Yi Lin, Quoc V. Le |
CVPR | 1 |
| 2018 | DropBlock: A regularization method for convolutional networksabstractDeep neural networks often work well when they are over-parameterized and trained with a massive amount of noise and regularization, such as weight decay and dropout. Although dropout is widely used as a regularization technique for fully connected layers, it is often less effective for convolutional layers. This lack of success of dropout for convolutional layers is perhaps due to the fact that activation units in convolutional layers are spatially correlated so information can still flow through convolutional networks despite dropout. Thus a structured form of dropout is needed to regularize convolutional networks. In this paper, we introduce DropBlock, a form of structured dropout, where units in a contiguous region of a feature map are dropped together. We found that applying DropbBlock in skip connections in addition to the convolution layers increases the accuracy. Also, gradually increasing number of dropped units during training leads to better accuracy and more robust to hyperparameter choices. Extensive experiments show that DropBlock works better than dropout in regularizing convolutional networks. On ImageNet classification, ResNet-50 architecture with DropBlock achieves $78.13\%$ accuracy, which is more than $1.6\%$ improvement on the baseline. On COCO detection, DropBlock improves Average Precision of RetinaNet from $36.8\%$ to $38.4\%$. Golnaz Ghiasi, Tsung-Yi Lin, Quoc V. Le |
NeurIPS | 1 |
| 2017 | Exploring the structure of a real-time, arbitrary neural artistic stylization network
Golnaz Ghiasi, Honglak Lee, Manjunath Kudlur, Vincent Dumoulin, Jonathon Shlens |
BMVC | 1 |
| 2016 | Laplacian Pyramid Reconstruction and Refinement for Semantic Segmentation
Golnaz Ghiasi, Charless C. Fowlkes |
ECCV (3) | 1 |
| 2015 | Using Segmentation to Predict the Absence of Occluded Parts
Golnaz Ghiasi, Charless C. Fowlkes |
BMVC | 1 |
| 2014 | Occlusion Coherence: Localizing Occluded Faces with a Hierarchical Deformable Part ModelabstractThe presence of occluders significantly impacts performance of systems for object recognition. However, occlusion is typically treated as an unstructured source of noise and explicit models for occluders have lagged behind those for object appearance and shape. In this paper we describe a hierarchical deformable part model for face detection and keypoint localization that explicitly models occlusions of parts. The proposed model structure makes it possible to augment positive training data with large numbers of synthetically occluded instances. This allows us to easily incorporate the statistics of occlusion patterns in a discriminatively trained model. We test the model on several benchmarks for keypoint localization including challenging sets featuring significant occlusion. We find that the addition of an explicit model of occlusion yields a system that outperforms existing approaches in keypoint localization accuracy. Golnaz Ghiasi, Charless C. Fowlkes |
CVPR | 1 |
| 2014 | Parsing Occluded PeopleabstractOcclusion poses a significant difficulty for object recognition due to the combinatorial diversity of possible occlusion patterns. We take a strongly supervised, non-parametric approach to modeling occlusion by learning deformable models with many local part mixture templates using large quantities of synthetically generated training data. This allows the model to learn the appearance of different occlusion patterns including figure-ground cues such as the shapes of occluding contours as well as the co-occurrence statistics of occlusion between neighboring parts. The underlying part mixture-structure also allows the model to capture coherence of object support masks between neighboring parts and make compelling predictions of figure-ground-occluder segmentations. We test the resulting model on human pose estimation under heavy occlusion and find it produces improved localization accuracy. Golnaz Ghiasi, Yi Yang 0007, Deva Ramanan, Charless C. Fowlkes |
CVPR | 1 |
| 2013 | Offline text-independent writer identification using codebook and efficient code extraction methods
Golnaz Ghiasi, Reza Safabakhsh |
Image Vis. Comput. | 1 |
| 2010 | An Efficient Method for Offline Text Independent Writer IdentificationabstractThis paper proposes, an efficient method for text independent writer identification using a codebook. The occurrence histogram of the shapes in the codebook is used to create a feature vector for the handwriting. There is a wide variety of different shapes in the connected components obtained from handwriting. Small fragments of connected components should be used to avoid complex patterns. A new and more efficient method is introduced for this purpose. To evaluate the methods, writer identification is conducted on three varieties of a Farsi database. These varieties include texts of short, medium and large lengths. Experimental results show the efficiency of the method especially for short texts. Golnaz Ghiasi, Reza Safabakhsh |
ICPR | 1 |