Morteza Ziyadi

dblp:79/8297 · DBLP profile ↗
← Back
9ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0002-7022-0845ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Language models and text generation · 38% Trustworthy machine learning · 31% Video understanding and tracking · 13%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 87% Program synthesis and code generation · 13%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%

Topics — the 23 heaviest of 27, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding
1.012026
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
1.012026
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation · AAAI 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation
1.012026
JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation · AAAI 2026
Computer vision › Video understanding and tracking
video question answering
1.012026
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026
Machine learning › Trustworthy machine learning
fairness
0.912025
Certifying Counterfactual Bias in LLMs · ICLR 2025
Machine learning › Trustworthy machine learning › fairness
fairness evaluation
0.912025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025
Machine learning › Trustworthy machine learning › hallucination
hallucination evaluation
0.912025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025
Natural language and speech › Language models and text generation
large language model
0.912025
Certifying Counterfactual Bias in LLMs · ICLR 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025
Machine learning › Trustworthy machine learning › privacy
privacy evaluation
0.912025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025
Machine learning › Trustworthy machine learning
safety evaluation
0.912025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025
Information retrieval
agentic search
0.912025
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025
Information retrieval
retrieval evaluation
0.912025
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025
Information retrieval
search engines
0.912025
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025
Information retrieval
web search
0.912025
Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025
Empirical software engineering
developer studies
0.912025
Trust Dynamics in AI-Assisted Development: Definitions, Factors, and Implications · ICSE 2025
Empirical software engineering › software engineering research methodology
mixed-methods study
0.912025
Trust Dynamics in AI-Assisted Development: Definitions, Factors, and Implications · ICSE 2025
Natural language and speech › Language models and text generation
natural language reasoning
0.612022
Reasoning Like Program Executors · EMNLP 2022
Natural language and speech › Language models and text generation › pre-trained language model
table pre-training
0.612022
TAPEX: Table Pre-training via Learning a Neural SQL Executor · ICLR 2022
Natural language and speech › Information extraction and text analysis › document understanding
table understanding
0.612022
TAPEX: Table Pre-training via Learning a Neural SQL Executor · ICLR 2022
Natural language and speech › Language models and text generation
large language model evaluation
0.312026
MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026
Machine learning › Trustworthy machine learning › robustness › model robustness evaluation
adversarial robustness evaluation
0.312025
VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

tree-structured rubric design · 1.7survey · 1.7observation study · 1.7LLM-as-a-judge · 1.7multimodal benchmark · 1.0multi-agent judging · 1.0human baseline · 1.0elo rating · 1.0randomized smoothing · 0.9large language model evaluation · 0.9certification · 0.9benchmark construction · 0.9
YearPublicationVenuePosition
2026 JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
abstract
While small language models (SLMs) have shown promise on various reasoning tasks, their ability to judge the correctness of answers remains unclear compared to large language models (LLMs). Prior work on LLM-as-a-judge frameworks typically relies on comparing candidate answers against ground-truth labels or other candidate answers using predefined metrics like entailment. However, this approach is inherently indirect and difficult to fully automate, offering limited support for fine-grained and scalable evaluation of reasoning outputs. In this work, we propose JudgeBoard, a novel evaluation pipeline that directly queries models to assess the correctness of candidate answers without requiring extra answer comparisons. We focus on two core reasoning domains: mathematical reasoning and science/commonsense reasoning, and construct task-specific evaluation leaderboards using both accuracy-based ranking and an Elo-based rating system across five benchmark datasets, enabling consistent model comparison as judges rather than comparators. To improve judgment performance in lightweight models, we propose MAJ (Multi-Agent Judging), a novel multi-agent evaluation framework that leverages multiple interacting SLMs with distinct reasoning profiles to approximate LLM-level judgment accuracy through collaborative deliberation. Experimental results reveal a significant performance gap between SLMs and LLMs in isolated judging tasks. However, our MAJ framework substantially improves the reliability and consistency of SLMs. On the MATH dataset, MAJ using smaller-sized models as backbones performs comparatively well or even better than their larger-sized counterparts. Our findings highlight that multi-agent SLM systems can potentially match or exceed LLM performance in judgment tasks, with implications for scalable and efficient assessment.
Zhenyu Bi, Gaurav Srivastava 0012, Swastik Roy, Morteza Ziyadi, Xuan Wang 0008
AAAI6
2026 MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX
abstract
We introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with questions that closely mimic the multimodal understanding experiences available to humans during decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence, with the website publicly available below.
Liuyue Xie, Avik Kuthiala, George Z. Wei, Ananya Bal, Mosam Dabhi, Liting Wen, Taru Rustagi, Ethan Lai, Sushil Khyalia, Rohan Choudhury, Morteza Ziyadi, László A. Jeni
AAAI12
2025 Certifying Counterfactual Bias in LLMs
abstract
Large Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly evaluate biases across LLM responses for different demographic groups (a.k.a. counterfactual bias), as they do not scale to large number of inputs and do not provide guarantees. Therefore, we propose the first framework, LLMCert-B that certifies LLMs for counterfactual bias on distributions of prompts. A certificate consists of high-confidence bounds on the probability of unbiased LLM responses for any set of counterfactual prompts - prompts differing by demographic groups, sampled from a distribution. We illustrate counterfactual bias certification for distributions of counterfactual prompts created by applying prefixes sampled from prefix distributions, to a given set of prompts. We consider prefix distributions consisting random token sequences, mixtures of manual jailbreaks, and perturbations of jailbreaks in LLM’s embedding space. We generate non-trivial certificates for SOTA LLMs, exposing their vulnerabilities over distributions of prompts generated from computationally inexpensive prefix distributions.
Isha Chaudhary, Manoj Kumar 0007, Morteza Ziyadi, Rahul Gupta 0001, Gagandeep Singh 0001
ICLR4
2025 Trust Dynamics in AI-Assisted Development: Definitions, Factors, and Implications
abstract
Software developers increasingly rely on AI code generation utilities. To ensure that “good” code is accepted into the code base and “bad” code is rejected, developers must know when to trust an AI suggestion. Understanding how developers build this intuition is crucial to enhancing developer-AI collaborative programming. In this paper, we seek to understand how developers (1) define and (2) evaluate the trustworthiness of a code suggestion and (3) how trust evolves when using AI code assistants. To answer these questions, we conducted a mixed method study consisting of an in-depth exploratory survey with (n=29) developers followed by an observation study (n=10). We found that comprehensibility and perceived correctness were the most frequently used factors to evaluate code suggestion trustworthiness. However, the gap in developers' definition and evaluation of trust points to a lack of support for evaluating trustworthy code in real-time. We also found that developers often alter their trust decisions, keeping only 52% of original suggestions. Based on these findings, we extracted four guidelines to enhance developer-AI interactions. We validated the guidelines through a survey with (n=7) domain experts and survey members (n=8). We discuss the validated guidelines, how to apply them, and tools to help adopt them.
Sadra Sabouri, Philipp Eibl, Morteza Ziyadi, Nenad Medvidovic, Lars Lindemann, Souti Chattopadhyay
ICSE4
2025 Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge
abstract
Agentic search such as Deep Research systems-where agents autonomously browse the web, synthesize information, and return comprehensive citation-backed answers-represents a major shift in how users interact with web-scale information. While promising greater efficiency and cognitive offloading, the growing complexity and open-endedness of agentic search have outpaced existing evaluation benchmarks and methodologies, which largely assume short search horizons and static answers. In this paper, we introduce Mind2Web 2, a benchmark of 130 realistic, high-quality, and long-horizon tasks that require real-time web browsing and extensive information synthesis, constructed with over 1000 hours of human labor. To address the challenge of evaluating time-varying and complex answers, we propose a novel Agent-as-a-Judge framework. Our method constructs task-specific judge agents based on a tree-structured rubric design to automatically assess both answer correctness and source attribution. We conduct a comprehensive evaluation of ten frontier agentic search systems and human performance, along with a detailed error analysis to draw insights for future development. The best-performing system, OpenAI Deep Research, can already achieve 50-70% of human performance while spending half the time, highlighting its great potential. Altogether, Mind2Web 2 provides a rigorous foundation for developing and benchmarking the next generation of agentic search systems.
Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu 0016, Michael Lin, Weijian Qi, Andrei Kopanev, Botao Yu, Bernal Jimenez Gutierrez, Yiheng Shu, Chan Hee Song, Jiaman Wu, Hanane Nour Moussa, Tianshu Zhang 0001, Yifei Li 0005, Tianci Xue, Zeyi Liao, Kai Zhang 0033, Boyuan Zheng 0001, Zhaowei Cai, Viktor Rozgic, Morteza Ziyadi, Huan Sun 0001, Yu Su 0001
NeurIPS24
2025 VMDT: Decoding the Trustworthiness of Video Foundation Models
abstract
As foundation models become more sophisticated, ensuring their trustworthiness becomes increasingly critical; yet, unlike text and image, the video modality still lacks comprehensive trustworthiness benchmarks. We introduce VMDT (Video-Modal DecodingTrust), the first unified platform for evaluating text-to-video (T2V) and video-to-text (V2T) models across five key trustworthiness dimensions: safety, hallucination, fairness, privacy, and adversarial robustness. Through our extensive evaluation of 7 T2V models and 19 V2T models using VMDT, we uncover several significant insights. For instance, all open-source T2V models evaluated fail to recognize harmful queries and often generate harmful videos, while exhibiting higher levels of unfairness compared to image modality models. In V2T models, unfairness and privacy risks rise with scale, whereas hallucination and adversarial robustness improve---though overall performance remains low. Uniquely, safety shows no correlation with model size, implying that factors other than scale govern current safety levels. Our findings highlight the urgent need for developing more robust and trustworthy video foundation models, and VMDT provides a systematic framework for measuring and tracking progress toward this goal. The code is available at https://sunblaze-ucb.github.io/VMDT-page/.
Yujin Potter, Zhun Wang, Nicholas Crispino, Kyle Montgomery, Alexander Xiong, Ethan Y. Chang, Francesco Pinto, Rahul Gupta 0001, Morteza Ziyadi, Christos Christodoulopoulos 0001, Bo Li 0026, Chenguang Wang 0001, Dawn Song
NeurIPS10
2022 Reasoning Like Program Executors
abstract
Reasoning over natural language is a longstanding goal for the research community.However, studies have shown that existing language models are inadequate in reasoning.To address the issue, we present POET, a novel reasoning pre-training paradigm.Through pretraining language models with programs and their execution results, POET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach.POET is conceptually simple and can be instantiated by different kinds of program executors.In this paper, we showcase two simple instances POET-Math and POET-Logic, in addition to a complex instance, POET-SQL.Experimental results on six benchmarks demonstrate that POET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning.POET opens a new gate on reasoningenhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors.
Xinyu Pi, Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Qiang Fu 0015, Yan Gao 0002, Jian-Guang Lou, Weizhu Chen
EMNLP4
2022 TAPEX: Table Pre-training via Learning a Neural SQL Executor
Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, Jian-Guang Lou
ICLR4
2016 OFDM over mm-Wave OAM Channels in a Multipath Environment with Intersymbol Interference
abstract
This paper reports an experimental investigation of using orthogonal frequency division multiplexing (OFDM) on orbital-angular momentum (OAM) channels when there is intersymbol interference (ISI) caused by multipath effects. The impulse response measurement indicates that due to the divergence of OAM beams, channels of larger OAM number are more affected by the ISI caused the reflections from surrounding objects. Channel performance of three different OAM numbers l = 0, +1 and +3 are studied. When a single QPSK channel is used, the more channel performance degradation in term of BER and EVM is observed for higher OAM channels. OFDM with 4 sub-carriers is used on the OAM channels to help reduce the ISI and improve the channel performance to some extent.
Yan Yan 0010, Long Li 0001, Guodong Xie, Morteza Ziyadi, Amirhossein M. Ariaei, Yongxiong Ren, Olivier Renaudin, Zhe Zhao 0003, Zhe Wang 0020, Cong Liu 0010, Soji Sajuyigbe, Shilpa Talwar, Solyman Ashrafi, Andreas F. Molisch, Alan E. Willner
GLOBECOM4