Erhan Bas

dblp:33/9222 · DBLP profile ↗
← Back
12ranked-venue papers
5as first author
6since 2021 · last 2025
0000-0002-6719-0940ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021
YearPublicationVenuePosition
2025 Enhancing SAM with Efficient Prompting and Preference Optimization for Semi-supervised Medical Image Segmentation
abstract
Foundational models such as the Segment Anything Model (SAM) are gaining traction in medical imaging segmentation, supporting multiple downstream tasks. However, such models are supervised in nature, still relying on large annotated datasets or prompts supplied by experts. Conventional techniques such as active learning to alleviate such limitations are limited in scope and still necessitate continuous human involvement and complex domain knowledge for label refinement or establishing reward ground truth. To address these challenges, we propose an enhanced Segment Anything Model (SAM) framework that utilizes annotation-efficient prompts generated in a fully unsupervised fashion, while still capturing essential semantic, location, and shape information through contrastive language-image pretraining and visual question answering. We adopt the direct preference optimization technique to design an optimal policy that enables the model to generate high-fidelity segmentations with simple ratings or rankings provided by a virtual annotator simulating the human annotation process. State-of-the-art performance of our framework in tasks such as lung segmentation, breast tumor segmentation, and organ segmentation across various modalities, including X-ray, ultrasound, and abdominal CT, justifies its effectiveness in low-annotation data scenarios.
Aishik Konwer, Zhijian Yang, Erhan Bas, Cao Xiao, Prateek Prasanna, Parminder Bhatia, Taha A. Kass-Hout
CVPR3
2024 Detecting and Preventing Hallucinations in Large Vision Language Models
abstract
Instruction tuned Large Vision Language Models (LVLMs) have significantly advanced in generalizing across a diverse set of multi-modal tasks, especially for Visual Question Answering (VQA). However, generating detailed responses that are visually grounded is still a challenging task for these models. We find that even the current state-of-the-art LVLMs (InstructBLIP) still contain a staggering 30 percent of the hallucinatory text in the form of non-existent objects, unfaithful descriptions, and inaccurate relationships. To address this, we introduce M-HalDetect, a Multimodal Hallucination Detection Dataset that can be used to train and benchmark models for hallucination detection and prevention. M-HalDetect consists of 16k fine-grained annotations on VQA examples, making it the first comprehensive multi-modal hallucination detection dataset for detailed image descriptions. Unlike previous work that only consider object hallucination, we additionally annotate both entity descriptions and relationships that are unfaithful. To demonstrate the potential of this dataset for hallucination prevention, we optimize InstructBLIP through our novel Fine-grained Direct Preference Optimization (FDPO). We also train fine-grained multi-modal reward models from InstructBLIP and evaluate their effectiveness with best-of-n rejection sampling (RS). We perform human evaluation on both FDPO and rejection sampling, and find that they reduce hallucination rates in InstructBLIP by 41% and 55% respectively. We also find that our reward model generalizes to other multi-modal models, reducing hallucinations in LLaVA and mPLUG-OWL by 15% and 57% respectively, and has strong correlation with human evaluated accuracy scores. The dataset is available at https://github.com/hendryx-scale/mhal-detect.
Anisha Gunjal, Jihan Yin, Erhan Bas
AAAI3
2023 Masked Vision and Language Modeling for Multi-modal Representation Learning
Gukyeong Kwon, Zhaowei Cai, Avinash Ravichandran, Erhan Bas, Rahul Bhotika, Stefano Soatto
ICLR4
2023 Relaxing Contrastiveness in Multimodal Representation Learning
abstract
Multimodal representation learning for images with paired raw texts can improve the usability and generality of the learned semantic concepts while significantly reducing annotation costs. In this paper, we explore the design space of loss functions in visual-linguistic pretraining frameworks and propose a novel Relaxed Contrastive (ReCo) objective, which act as a drop-in replacement of the widely used InfoNCE loss. The key insight of ReCo is to allow a relaxed negative space by not penalizing unpaired multimodal samples (i.e., negative pairs) that are already orthogonal or negatively correlated. Unlike the widely-used InfoNCE, which keeps repelling negative pairs as long as they are not anti-correlated, ReCo by design embraces more diversity and flexibility of the learned embeddings. We conduct exten-sive experiments using ReCo with state-of-the-art models by pretraining on the MIMIC-CXR dataset that consists of chest radiographs and free-text radiology reports, and eval-uating on the CheXpert dataset for multimodal retrieval and disease classification. Our ReCo achieves an absolute improvement of 2.9% over the InfoNCE baseline on the CheXpert Retrieval dataset in average retrieval precision and re-ports better or comparable performance in the linear evaluation and finetuning for classification. We further show that ReCo outperforms InfoNCE on the Flickr30K dataset by 1.7% in retrieval Recall@1, demonstrating the generalizability of our approach to natural images.
Zudi Lin, Erhan Bas, Kunwar Yashraj Singh, Gurumurthy Swaminathan, Rahul Bhotika
WACV2
2022 X-DETR: A Versatile Architecture for Instance-wise Vision-Language Tasks
Zhaowei Cai, Gukyeong Kwon, Avinash Ravichandran, Erhan Bas, Zhuowen Tu, Rahul Bhotika, Stefano Soatto
ECCV (36)4
2021 End-to-end Piece-wise Unwarping of Document Images
abstract
Document unwarping attempts to undo physical deformations of the paper and recover a ’flatbed’ scanned document-image for downstream tasks such as OCR. Current state-of-the-art relies on global unwarping of the document which is not robust to local deformation changes. Moreover, a global unwarping often produces spurious warping artifacts in less warped regions to compensate for severe warps present in other parts of the document. In this paper, we propose the first end-to-end trainable piece-wise unwarping1method that predicts local deformation fields and stitches them together with global information to obtain an improved unwarping. The proposed piece-wise formulation results in 4% improvement in terms of multi-scale structural similarity (MS-SSIM) and shows better performance in terms of OCR metrics, character error rate (CER) and word error rate (WER) compared to the state-of-the-art.
Sagnik Das, Kunwar Yashraj Singh, Jon Wu, Erhan Bas, Vijay Mahadevan, Rahul Bhotika, Dimitris Samaras
ICCV4
2013 Contour-based shape representation using principal curves
Esra Ataer Cansizoglu, Erhan Bas, Jayashree Kalpathy-Cramer, Gregory C. Sharp, Deniz Erdogmus
Pattern Recognit.2
2012 Local tracing of curvilinear structures in volumetric color images: Application to the Brainbow analysis
Erhan Bas, Deniz Erdogmus, R. W. Draft, Jeff Lichtman
J. Vis. Commun. Image Represent.1
2011 Polytope kernel density estimates on Delaunay graphs
abstract
We present a polytope-kernel density estimation (PKDE) methodology that allows us to perform exact mean-shift up dates along the edges of the Delaunay graph of the data. We discuss explicit and implicit constructions of such a PKDE, where in the implicit construction one can exploit a smoother kernel such as the standard isotropic Gaussian. The resulting density estimate allows us to perform mean-shift clustering in a computationally efficient manner (similar to mediod shift), but in a manner that is exact and consistent with the underlying density assumption. The procedure also yields a hierarchical connectivity structure, a tree, that spans the dataset. We demonstrate how this tree, combined with density-weighted geodesic distance calculations between modal samples can be used to select number of clusters as well as a distance preserving dimension reduction technique.
Erhan Bas, Deniz Erdogmus
ICASSP1
2011 Sampling on locally defined principal manifolds
abstract
We start with a locally defined principal curve definition for a given probability density function (pdf) and define a pairwise manifold score based on local derivatives of the pdf. Proposed manifold score can be used to check if data pairs lie on the same manifold. We use this score to (i) cluster nonlinear manifolds having irregular shapes, and (ii) (down)sample a selected principal curve with sufficient accuracy sparsely. Our goal is to provide a heuristic-free formulation for principal graph generation and curve parametrization in order to form a basis for a principled principal manifold unwrapping method.
Erhan Bas, Deniz Erdogmus
ICASSP1
2011 Connectivity of projected high dimensional data charts on one-dimensional curves
Erhan Bas, Deniz Erdogmus
Signal Process.1
2010 Principal curve tracing
Erhan Bas, Deniz Erdogmus
ESANN1