EDBT 2026 Demo / reviewers in the wild / expert
Morteza Ziyadi
dblp:79/8297
· DBLP profile ↗
9ranked-venue papers
0as first author
8since 2021 · last 2026
0009-0002-7022-0845ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Computer networks · 1Software engineering, systems software and programming languages · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
7 papers |
Language models and text generation · 38% Trustworthy machine learning · 31% Video understanding and tracking · 13% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% | |
| Software engineering, system software, and programming languages
1 paper |
Empirical software engineering · 87% Program synthesis and code generation · 13% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 100% |
Topics — the 23 heaviest of 27, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge |
1.0 | 1 | 2026 | JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation |
1.0 | 1 | 2026 | JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation · AAAI 2026 |
Computer vision › Video understanding and tracking
video question answering |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Machine learning › Trustworthy machine learning
fairness |
0.9 | 1 | 2025 | Certifying Counterfactual Bias in LLMs · ICLR 2025 |
Machine learning › Trustworthy machine learning › fairness
fairness evaluation |
0.9 | 1 | 2025 | VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › hallucination
hallucination evaluation |
0.9 | 1 | 2025 | VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025 |
Natural language and speech › Language models and text generation
large language model |
0.9 | 1 | 2025 | Certifying Counterfactual Bias in LLMs · ICLR 2025 |
Natural language and speech › Language models and text generation
LLM agents |
0.9 | 1 | 2025 | Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025 |
Machine learning › Trustworthy machine learning › privacy
privacy evaluation |
0.9 | 1 | 2025 | VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025 |
Machine learning › Trustworthy machine learning
safety evaluation |
0.9 | 1 | 2025 | VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025 |
Information retrieval
agentic search |
0.9 | 1 | 2025 | Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025 |
Information retrieval
retrieval evaluation |
0.9 | 1 | 2025 | Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025 |
Information retrieval
search engines |
0.9 | 1 | 2025 | Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025 |
Information retrieval
web search |
0.9 | 1 | 2025 | Mind2Web 2: Evaluating Agentic Search with Agent-as-a-Judge · NeurIPS 2025 |
Empirical software engineering
developer studies |
0.9 | 1 | 2025 | Trust Dynamics in AI-Assisted Development: Definitions, Factors, and Implications · ICSE 2025 |
Empirical software engineering › software engineering research methodology
mixed-methods study |
0.9 | 1 | 2025 | Trust Dynamics in AI-Assisted Development: Definitions, Factors, and Implications · ICSE 2025 |
Natural language and speech › Language models and text generation
natural language reasoning |
0.6 | 1 | 2022 | Reasoning Like Program Executors · EMNLP 2022 |
Natural language and speech › Language models and text generation › pre-trained language model
table pre-training |
0.6 | 1 | 2022 | TAPEX: Table Pre-training via Learning a Neural SQL Executor · ICLR 2022 |
Natural language and speech › Information extraction and text analysis › document understanding
table understanding |
0.6 | 1 | 2022 | TAPEX: Table Pre-training via Learning a Neural SQL Executor · ICLR 2022 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Machine learning › Trustworthy machine learning › robustness › model robustness evaluation
adversarial robustness evaluation |
0.3 | 1 | 2025 | VMDT: Decoding the Trustworthiness of Video Foundation Models · NeurIPS 2025 |
Methods — techniques the papers use, named apart from their topics
tree-structured rubric design · 1.7survey · 1.7observation study · 1.7LLM-as-a-judge · 1.7multimodal benchmark · 1.0multi-agent judging · 1.0human baseline · 1.0elo rating · 1.0randomized smoothing · 0.9large language model evaluation · 0.9certification · 0.9benchmark construction · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning EvaluationabstractWhile small language models (SLMs) have shown promise on various reasoning tasks, their ability to judge the correctness of answers remains unclear compared to large language models (LLMs). Prior work on LLM-as-a-judge frameworks typically relies on comparing candidate answers against ground-truth labels or other candidate answers using predefined metrics like entailment. However, this approach is inherently indirect and difficult to fully automate, offering limited support for fine-grained and scalable evaluation of reasoning outputs. In this work, we propose JudgeBoard, a novel evaluation pipeline that directly queries models to assess the correctness of candidate answers without requiring extra answer comparisons. We focus on two core reasoning domains: mathematical reasoning and science/commonsense reasoning, and construct task-specific evaluation leaderboards using both accuracy-based ranking and an Elo-based rating system across five benchmark datasets, enabling consistent model comparison as judges rather than comparators. To improve judgment performance in lightweight models, we propose MAJ (Multi-Agent Judging), a novel multi-agent evaluation framework that leverages multiple interacting SLMs with distinct reasoning profiles to approximate LLM-level judgment accuracy through collaborative deliberation. Experimental results reveal a significant performance gap between SLMs and LLMs in isolated judging tasks. However, our MAJ framework substantially improves the reliability and consistency of SLMs. On the MATH dataset, MAJ using smaller-sized models as backbones performs comparatively well or even better than their larger-sized counterparts. Our findings highlight that multi-agent SLM systems can potentially match or exceed LLM performance in judgment tasks, with implications for scalable and efficient assessment. Zhenyu Bi, Gaurav Srivastava 0012, Swastik Roy, Morteza Ziyadi, Xuan Wang 0008 |
AAAI | 6 |
| 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeXabstractWe introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with questions that closely mimic the multimodal understanding experiences available to humans during decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence, with the website publicly available below. Liuyue Xie, Avik Kuthiala, George Z. Wei, Ananya Bal, Mosam Dabhi, Liting Wen, Taru Rustagi, Ethan Lai, Sushil Khyalia, Rohan Choudhury, Morteza Ziyadi, László A. Jeni |
AAAI | 12 |
| 2025 | Certifying Counterfactual Bias in LLMsabstractLarge Language Models (LLMs) can produce biased responses that can cause representational harms. However, conventional studies are insufficient to thoroughly
evaluate biases across LLM responses for different demographic groups (a.k.a.
counterfactual bias), as they do not scale to large number of inputs and do not
provide guarantees. Therefore, we propose the first framework, LLMCert-B that
certifies LLMs for counterfactual bias on distributions of prompts. A certificate
consists of high-confidence bounds on the probability of unbiased LLM responses
for any set of counterfactual prompts - prompts differing by demographic groups,
sampled from a distribution. We illustrate counterfactual bias certification for
distributions of counterfactual prompts created by applying prefixes sampled from
prefix distributions, to a given set of prompts. We consider prefix distributions consisting random token sequences, mixtures of manual jailbreaks, and perturbations
of jailbreaks in LLM’s embedding space. We generate non-trivial certificates for
SOTA LLMs, exposing their vulnerabilities over distributions of prompts generated
from computationally inexpensive prefix distributions. Isha Chaudhary, Manoj Kumar 0007, Morteza Ziyadi, Rahul Gupta 0001, Gagandeep Singh 0001 |
ICLR | 4 |
| 2025 | Trust Dynamics in AI-Assisted Development: Definitions, Factors, and ImplicationsabstractSoftware developers increasingly rely on AI code generation utilities. To ensure that “good” code is accepted into the code base and “bad” code is rejected, developers must know when to trust an AI suggestion. Understanding how developers build this intuition is crucial to enhancing developer-AI collaborative programming. In this paper, we seek to understand how developers (1) define and (2) evaluate the trustworthiness of a code suggestion and (3) how trust evolves when using AI code assistants. To answer these questions, we conducted a mixed method study consisting of an in-depth exploratory survey with (n=29) developers followed by an observation study (n=10). We found that comprehensibility and perceived correctness were the most frequently used factors to evaluate code suggestion trustworthiness. However, the gap in developers' definition and evaluation of trust points to a lack of support for evaluating trustworthy code in real-time. We also found that developers often alter their trust decisions, keeping only 52% of original suggestions. Based on these findings, we extracted four guidelines to enhance developer-AI interactions. We validated the guidelines through a survey with (n=7) domain experts and survey members (n=8). We discuss the validated guidelines, how to apply them, and tools to help adopt them. Sadra Sabouri, Philipp Eibl, Morteza Ziyadi, Nenad Medvidovic, Lars Lindemann, Souti Chattopadhyay |
ICSE | 4 |
| 2025 | Mind2Web 2: Evaluating Agentic Search with Agent-as-a-JudgeabstractAgentic search such as Deep Research systems-where agents autonomously browse the web, synthesize information, and return comprehensive citation-backed answers-represents a major shift in how users interact with web-scale information. While promising greater efficiency and cognitive offloading, the growing complexity and open-endedness of agentic search have outpaced existing evaluation benchmarks and methodologies, which largely assume short search horizons and static answers. In this paper, we introduce Mind2Web 2, a benchmark of 130 realistic, high-quality, and long-horizon tasks that require real-time web browsing and extensive information synthesis, constructed with over 1000 hours of human labor. To address the challenge of evaluating time-varying and complex answers, we propose a novel Agent-as-a-Judge framework. Our method constructs task-specific judge agents based on a tree-structured rubric design to automatically assess both answer correctness and source attribution. We conduct a comprehensive evaluation of ten frontier agentic search systems and human performance, along with a detailed error analysis to draw insights for future development. The best-performing system, OpenAI Deep Research, can already achieve 50-70% of human performance while spending half the time, highlighting its great potential. Altogether, Mind2Web 2 provides a rigorous foundation for developing and benchmarking the next generation of agentic search systems. Boyu Gou, Zanming Huang, Yuting Ning, Yu Gu 0016, Michael Lin, Weijian Qi, Andrei Kopanev, Botao Yu, Bernal Jimenez Gutierrez, Yiheng Shu, Chan Hee Song, Jiaman Wu, Hanane Nour Moussa, Tianshu Zhang 0001, Yifei Li 0005, Tianci Xue, Zeyi Liao, Kai Zhang 0033, Boyuan Zheng 0001, Zhaowei Cai, Viktor Rozgic, Morteza Ziyadi, Huan Sun 0001, Yu Su 0001 |
NeurIPS | 24 |
| 2025 | VMDT: Decoding the Trustworthiness of Video Foundation ModelsabstractAs foundation models become more sophisticated, ensuring their trustworthiness becomes increasingly critical; yet, unlike text and image, the video modality still lacks comprehensive trustworthiness benchmarks. We introduce VMDT (Video-Modal DecodingTrust), the first unified platform for evaluating text-to-video (T2V) and video-to-text (V2T) models across five key trustworthiness dimensions: safety, hallucination, fairness, privacy, and adversarial robustness. Through our extensive evaluation of 7 T2V models and 19 V2T models using VMDT, we uncover several significant insights. For instance, all open-source T2V models evaluated fail to recognize harmful queries and often generate harmful videos, while exhibiting higher levels of unfairness compared to image modality models. In V2T models, unfairness and privacy risks rise with scale, whereas hallucination and adversarial robustness improve---though overall performance remains low. Uniquely, safety shows no correlation with model size, implying that factors other than scale govern current safety levels. Our findings highlight the urgent need for developing more robust and trustworthy video foundation models, and VMDT provides a systematic framework for measuring and tracking progress toward this goal. The code is available at https://sunblaze-ucb.github.io/VMDT-page/. Yujin Potter, Zhun Wang, Nicholas Crispino, Kyle Montgomery, Alexander Xiong, Ethan Y. Chang, Francesco Pinto, Rahul Gupta 0001, Morteza Ziyadi, Christos Christodoulopoulos 0001, Bo Li 0026, Chenguang Wang 0001, Dawn Song |
NeurIPS | 10 |
| 2022 | Reasoning Like Program ExecutorsabstractReasoning over natural language is a longstanding goal for the research community.However, studies have shown that existing language models are inadequate in reasoning.To address the issue, we present POET, a novel reasoning pre-training paradigm.Through pretraining language models with programs and their execution results, POET empowers language models to harvest the reasoning knowledge possessed by program executors via a data-driven approach.POET is conceptually simple and can be instantiated by different kinds of program executors.In this paper, we showcase two simple instances POET-Math and POET-Logic, in addition to a complex instance, POET-SQL.Experimental results on six benchmarks demonstrate that POET can significantly boost model performance in natural language reasoning, such as numerical reasoning, logical reasoning, and multi-hop reasoning.POET opens a new gate on reasoningenhancement pre-training, and we hope our analysis would shed light on the future research of reasoning like program executors. Xinyu Pi, Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Qiang Fu 0015, Yan Gao 0002, Jian-Guang Lou, Weizhu Chen |
EMNLP | 4 |
| 2022 | TAPEX: Table Pre-training via Learning a Neural SQL Executor
Qian Liu 0033, Bei Chen 0008, Morteza Ziyadi, Zeqi Lin, Weizhu Chen, Jian-Guang Lou |
ICLR | 4 |
| 2016 | OFDM over mm-Wave OAM Channels in a Multipath Environment with Intersymbol InterferenceabstractThis paper reports an experimental investigation of using orthogonal frequency division multiplexing (OFDM) on orbital-angular momentum (OAM) channels when there is intersymbol interference (ISI) caused by multipath effects. The impulse response measurement indicates that due to the divergence of OAM beams, channels of larger OAM number are more affected by the ISI caused the reflections from surrounding objects. Channel performance of three different OAM numbers l = 0, +1 and +3 are studied. When a single QPSK channel is used, the more channel performance degradation in term of BER and EVM is observed for higher OAM channels. OFDM with 4 sub-carriers is used on the OAM channels to help reduce the ISI and improve the channel performance to some extent. Yan Yan 0010, Long Li 0001, Guodong Xie, Morteza Ziyadi, Amirhossein M. Ariaei, Yongxiong Ren, Olivier Renaudin, Zhe Zhao 0003, Zhe Wang 0020, Cong Liu 0010, Soji Sajuyigbe, Shilpa Talwar, Solyman Ashrafi, Andreas F. Molisch, Alan E. Willner |
GLOBECOM | 4 |