Kalmit Kulkarni

dblp:419/7612 · DBLP profile ↗
← Back
1ranked-venue papers
0as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 70% Machine translation · 30%

Topics — the 4 heaviest of 4, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Machine translation
machine translation evaluation
1.012026
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages · ACL (1) 2026
Natural language and speech › Language models and text generation › text summarization
summarization evaluation
1.012026
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages · ACL (1) 2026
Natural language and speech › Language models and text generation
text summarization
1.012026
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model evaluation
0.312026
Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

perturbation analysis · 1.0inter-metric correlation · 1.0human judgment alignment · 1.0
YearPublicationVenuePosition
2026 Revisiting Metric Reliability for Fine-grained Evaluation of Machine Translation and Summarization in Indian Languages
abstract
While automatic metrics drive progress in Machine Translation (MT) and Text Summarization (TS), existing metrics have been developed and validated almost exclusively for English and other high-resource languages. This narrow focus leaves Indian languages, spoken by over 1.5 billion people, largely overlooked, casting doubt on the universality of current evaluation practices. To address this gap, we introduce ITEM, a large-scale benchmark that systematically evaluates the alignment of 29 automatic metrics with human judgments across six major Indian languages, enriched with fine-grained annotations. Our extensive evaluation, covering agreement with human judgments, sensitivity to outliers, language-specific reliability, inter-metric correlations, and resilience to controlled perturbations reveals four central findings: (1) LLM-based evaluators show the strongest alignment with human judgments at both segment and system levels; (2) outliers exert a significant impact on metric-human agreement; (3) In TS, metrics are more effective at capturing content fidelity, whereas in MT, they better reflect fluency; and (4) Metrics differ in their robustness and sensitivity when subjected to diverse perturbations. Collectively, these findings offer critical guidance for advancing metric design and evaluation in Indian languages.
Amir Hossein Yari, Kalmit Kulkarni, Ahmad Raza Khan, Fajri Koto
ACL (1)2