VLDB 2026 Research / reviewers in the wild / expert
Iftekhar Naim
dblp:11/8759
· DBLP profile ↗
12ranked-venue papers
8as first author
2since 2021 · last 2023
0000-0002-2119-3273ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 7 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 first-authorHuman-computer interaction and ubiquitous computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Probabilistic and Bayesian machine learning · 37% Information extraction and text analysis · 33% Vision and language · 19% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Computer graphics and multimedia
1 paper |
Multimedia analysis and retrieval · 100% |
Topics — the 10 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Information retrieval › retrieval models › neural retrieval
late interaction retrieval |
0.7 | 1 | 2023 | Rethinking the Role of Token Retrieval in Multi-Vector Retrieval · NeurIPS 2023 |
Information retrieval › retrieval models › neural retrieval
multi-vector retrieval |
0.7 | 1 | 2023 | Rethinking the Role of Token Retrieval in Multi-Vector Retrieval · NeurIPS 2023 |
Information retrieval
retrieval models |
0.7 | 1 | 2023 | Rethinking the Role of Token Retrieval in Multi-Vector Retrieval · NeurIPS 2023 |
Natural language and speech › Information extraction and text analysis
sequence labeling |
0.6 | 1 | 2022 | Transforming Sequence Tagging Into A Seq2Seq Task · EMNLP 2022 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.6 | 1 | 2022 | Transforming Sequence Tagging Into A Seq2Seq Task · EMNLP 2022 |
Machine learning › Transfer learning and domain adaptation › domain alignment
unsupervised alignment |
0.2 | 1 | 2016 | Unsupervised Alignment of Actions in Video with Text Descriptions · IJCAI 2016 |
Computer vision › Vision and language › cross-modal alignment › visual-semantic alignment
video-text alignment |
0.2 | 1 | 2016 | Unsupervised Alignment of Actions in Video with Text Descriptions · IJCAI 2016 |
Computer vision › Vision and language › visual grounding
instruction grounding |
0.2 | 1 | 2014 | Unsupervised Alignment of Natural Language Instructions with Video Segments · AAAI 2014 |
Machine learning › Probabilistic and Bayesian machine learning › statistical inference › parameter estimation
expectation-maximization convergence |
0.1 | 1 | 2012 | Convergence of the EM Algorithm for Gaussian Mixtures with Unbalanced Mixing Coefficients · ICML 2012 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › mixture model
gaussian mixture model |
0.1 | 1 | 2012 | Convergence of the EM Algorithm for Gaussian Mixtures with Unbalanced Mixing Coefficients · ICML 2012 |
Methods — techniques the papers use, named apart from their topics
contrastive objective · 0.7contextualized token retrieval · 0.7seq2seq modeling · 0.6pre-trained language model · 0.6multilingual transfer learning · 0.6prosodic features · 0.5k-nearest neighbors · 0.5facial features · 0.5DBSCAN clustering · 0.5unsupervised learning · 0.2hidden markov model · 0.2generative model · 0.2IBM Model 1 · 0.2expectation-maximization · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Rethinking the Role of Token Retrieval in Multi-Vector RetrievalabstractMulti-vector retrieval models such as ColBERT [Khattab et al., 2020] allow token-level interactions between queries and documents, and hence achieve state of the art on many information retrieval benchmarks. However, their non-linear scoring function cannot be scaled to millions of documents, necessitating a three-stage process for inference: retrieving initial candidates via token retrieval, accessing all token vectors, and scoring the initial candidate documents. The non-linear scoring function is applied over all token vectors of each candidate document, making the inference process complicated and slow. In this paper, we aim to simplify the multi-vector retrieval by rethinking the role of token retrieval. We present XTR, ConteXtualized Token Retriever, which introduces a simple, yet novel, objective function that encourages the model to retrieve the most important document tokens first. The improvement to token retrieval allows XTR to rank candidates only using the retrieved tokens rather than all tokens in the document, and enables a newly designed scoring stage that is two-to-three orders of magnitude cheaper than that of ColBERT. On the popular BEIR benchmark, XTR advances the state-of-the-art by 2.8 nDCG@10 without any distillation. Detailed analysis confirms our decision to revisit the token retrieval stage, as XTR demonstrates much better recall of the token retrieval stage compared to ColBERT. Jinhyuk Lee, Zhuyun Dai, Sai Meher Karthik Duddu, Tao Lei 0001, Iftekhar Naim, Ming-Wei Chang, Vincent Y. Zhao |
NeurIPS | 5 |
| 2022 | Transforming Sequence Tagging Into A Seq2Seq TaskabstractPretrained, large, generative language models (LMs) have had great success in a wide range of sequence tagging and structured prediction tasks.Casting a sequence tagging task as a Seq2Seq one requires deciding the formats of the input and output sequences.However, we lack a principled understanding of the tradeoffs associated with these formats (such as the effect on model accuracy, sequence length, multilingual generalization, hallucination).In this paper, we rigorously study different formats one could use for casting input text sentences and their output labels into the input and target (i.e., output) of a Seq2Seq model.Along the way, we introduce a new format, which we show to to be both simpler and more effective.Additionally the new format demonstrates significant gains in the multilingual settings -both zero-shot transfer learning and joint training.Lastly, we find that the new format is more robust and almost completely devoid of hallucination -an issue we find common in existing formats.With well over a 1000 experiments studying 14 different formats, over 7 diverse public benchmarksincluding 3 multilingual datasets spanning 7 languages -we believe our findings provide a strong empirical basis in understanding how we should tackle sequence tagging tasks.Dataset Lang # Train # Valid # Test Token/ex Tagged span/ex % tokens tagged # Tag classes Tag Entropy mATIS en 4478 500 893 11.28 3.32 36.50 79 3. Karthik Raman 0001, Iftekhar Naim, Jiecao Chen, Kazuma Hashimoto, Kiran Yalasangi, Krishna Srinivasan |
EMNLP | 2 |
| 2018 | Feature-Based Decipherment for Machine TranslationabstractOrthographic similarities across languages provide a strong signal for unsupervised probabilistic transduction (decipherment) for closely related language pairs. The existing decipherment models, however, are not well suited for exploiting these orthographic similarities. We propose a log-linear model with latent variables that incorporates orthographic similarity features. Maximum likelihood training is computationally expensive for the proposed log-linear model. To address this challenge, we perform approximate inference via Markov chain Monte Carlo sampling and contrastive divergence. Our results show that the proposed log-linear model with contrastive divergence outperforms the existing generative decipherment models by exploiting the orthographic features. The model both scales to large vocabularies and preserves accuracy in low- and no-resource contexts. Iftekhar Naim, Parker Riley, Daniel Gildea |
Comput. Linguistics | 1 |
| 2018 | Automated Analysis and Prediction of Job Interview PerformanceabstractWe present a computational framework for automatically quantifying verbal and nonverbal behaviors in the context of job interviews. The proposed framework is trained by analyzing the videos of 138 interview sessions with 69 internship-seeking undergraduates at the Massachusetts Institute of Technology (MIT). Our automated analysis includes facial expressions (e.g., smiles, head gestures, facial tracking points), language (e.g., word counts, topic modeling), and prosodic information (e.g., pitch, intonation, and pauses) of the interviewees. The ground truth labels are derived by taking a weighted average over the ratings of nine independent judges. Our framework can automatically predict the ratings for interview traits such as excitement, friendliness, and engagement with correlation coefficients of 0.70 or higher, and can quantify the relative importance of prosody, language, and facial expressions. By analyzing the relative feature weights learned by the regression models, our framework recommends to speak more fluently, use fewer filler words, speak as “we” (versus “I”), use more unique words, and smile more. We also find that the students who were rated highly while answering the first interview question were also rated highly overall (i.e., first impression matters). Finally, our MIT Interview dataset is available to other researchers to further validate and expand our findings. Iftekhar Naim, Md. Iftekhar Tanveer, Daniel Gildea, Mohammed E. Hoque 0001 |
IEEE Trans. Affect. Comput. | 1 |
| 2016 | ROC comment: automated descriptive and subjective captioning of behavioral videosabstractWe present an automated interface, ROC Comment, for generating natural language comments on behavioral videos. We focus on the domain of public speaking, which many people consider their greatest fear. We collect a dataset of 196 public speaking videos from 49 individuals and gather 12,173 comments, generated by more than 500 independent human judges. We then train a k-Nearest-Neighbor (k-NN) based model by extracting prosodic (e.g., volume) and facial (e.g., smiles) features. Given a new video, we extract features and select the closest comments using k-NN model. We further filter the comments by clustering them using DBScan, and eliminating the outliers. Evaluation of our system with 30 participants conclude that while the generated comments are helpful, there is room for improvement in further personalizing them. Our model has been deployed online, allowing individuals to upload their videos and receive open-ended and interpretative comments. Our system is available at http://tinyurl.com/roccomment. Mohammad Rafayet Ali, Facundo Ciancio, Iftekhar Naim, Mohammed E. Hoque 0001 |
UbiComp | 4 |
| 2016 | Aligning movies with scripts by exploiting temporal ordering constraintsabstractScripts provide rich textual annotation of movies, including dialogs, character names, and other situational descriptions. Exploiting such rich annotations requires aligning the sentences in the scripts with the corresponding video frames. Previous work on aligning movies with scripts predominantly relies on time-aligned closed-captions or subtitles, which are not always available. In this paper, we focus on automatically aligning faces in movies with their corresponding character names in scripts without requiring closed-captions/subtitles. We utilize the intuition that faces in a movie generally appear in the same sequential order as their names are mentioned in the script. We first apply standard techniques for face detection and tracking, and cluster similar face tracks together. Next, we apply a generative Hidden Markov Model (HMM) and a discriminative Latent Conditional Random Field (LCRF) to align the clusters of face tracks with the corresponding character names. Our alignment models (especially LCRF) significantly outperform the previous state-of-the-art on two different movie datasets and for a wide range of face clustering algorithms. Iftekhar Naim, Abdullah Al Mamun 0002, Young Chol Song, Jiebo Luo 0001, Henry A. Kautz, Daniel Gildea |
ICPR | 1 |
| 2016 | Unsupervised Alignment of Actions in Video with Text Descriptions
Young Chol Song, Iftekhar Naim, Abdullah Al Mamun 0002, Kaustubh Kulkarni, Parag Singla, Jiebo Luo 0001, Daniel Gildea, Henry A. Kautz |
IJCAI | 2 |
| 2015 | Discriminative Unsupervised Alignment of Natural Language Instructions with Corresponding Video SegmentsabstractIftekhar Naim, Young C. Song, Qiguang Liu, Liang Huang, Henry Kautz, Jiebo Luo, Daniel Gildea. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2015. Iftekhar Naim, Young Chol Song, Qiguang Liu, Liang Huang 0001, Henry A. Kautz, Jiebo Luo 0001, Daniel Gildea |
HLT-NAACL | 1 |
| 2014 | Unsupervised Alignment of Natural Language Instructions with Video SegmentsabstractWe propose an unsupervised learning algorithm for automatically inferring the mappings between English nouns and corresponding video objects. Given a sequence of natural language instructions and an unaligned video recording, we simultaneously align each instruction to its corresponding video segment, and also align nouns in each instruction to their corresponding objects in video. While existing grounded language acquisition algorithms rely on pre-aligned supervised data (each sentence paired with corresponding image frame or video segment), our algorithm aims to automatically infer the alignment from the temporal structure of the video and parallel text instructions. We propose two generative models that are closely related to the HMM and IBM 1 word alignment models used in statistical machine translation. We evaluate our algorithm on videos of biological experiments performed in wetlabs, and demonstrate its capability of aligning video segments to text instructions and matching video objects to nouns in the absence of any direct supervision. Iftekhar Naim, Young Chol Song, Qiguang Liu, Henry A. Kautz, Jiebo Luo 0001, Daniel Gildea |
AAAI | 1 |
| 2013 | Text Alignment for Real-Time Crowd Captioning
Iftekhar Naim, Daniel Gildea, Walter S. Lasecki, Jeffrey P. Bigham |
HLT-NAACL | 1 |
| 2012 | Convergence of the EM Algorithm for Gaussian Mixtures with Unbalanced Mixing Coefficients
Iftekhar Naim, Daniel Gildea |
ICML | 1 |
| 2010 | Swift: Scalable weighted iterative sampling for flow cytometry clusteringabstractFlow cytometry (FC) is a powerful technology for rapid multivariate analysis and functional discrimination of cells. Current FC platforms generate large, high-dimensional datasets which pose a significant challenge for traditional manual bivariate analysis. Automated multivariate clustering, though highly desirable, is also stymied by the critical requirement of identifying rare populations that form rather small clusters, in addition to the computational challenges posed by the large size and dimensionality of the datasets. In this paper, we address these twin challenges by developing a two-stage scalable multivariate parametric clustering algorithm. In the first stage, we model the data as a mixture of Gaussians and use an iterative weighted sampling technique to estimate the mixture components successively in order of decreasing size. In the second stage, we apply a graph-based hierarchical merging technique to combine Gaussian components with significant overlaps into the final number of desired clusters. The resulting algorithm offers a reduction in complexity over conventional mixture modeling while simultaneously allowing for better detection of small populations. We demonstrate the effectiveness of our method both on simulated data and actual flow cytometry datasets. Iftekhar Naim, Suprakash Datta, Gaurav Sharma 0001, James S. Cavenaugh, Tim R. Mosmann |
ICASSP | 1 |