EDBT 2026 Demo / reviewers in the wild / expert
George Z. Wei
dblp:312/5276
· DBLP profile ↗
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Video understanding and tracking · 40% Trustworthy machine learning · 23% Vision and language · 20% |
Topics — the 7 heaviest of 7, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking › multimodal video understanding
audio-visual video understanding |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Computer vision › Video understanding and tracking
video question answering |
1.0 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Computer vision › Face, body and person analysis
face detection |
0.6 | 1 | 2022 | Robustness Disparities in Face Detection · NeurIPS 2022 |
Machine learning › Trustworthy machine learning › fairness › fairness evaluation
fairness benchmarking |
0.6 | 1 | 2022 | Robustness Disparities in Face Detection · NeurIPS 2022 |
Machine learning › Trustworthy machine learning
robustness |
0.6 | 1 | 2022 | Robustness Disparities in Face Detection · NeurIPS 2022 |
Natural language and speech › Language models and text generation
large language model evaluation |
0.3 | 1 | 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeX · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
multimodal benchmark · 1.0human baseline · 1.0robustness benchmarking · 0.6perturbation analysis · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MAVERIX: Multimodal Audio-Visual Evaluation and Recognition IndeXabstractWe introduce MAVERIX (Multimodal Audio-Visual Evaluation and Recognition IndeX), a unified benchmark to probe video understanding in multimodal LLMs, encompassing video, audio, and text inputs with human performance baselines. Although recent advancements in audiovisual models have shown substantial progress, the field lacks a standardized evaluation framework to thoroughly assess their cross-modality comprehension performance. MAVERIX curates 2,556 questions from 700 videos, in the form of both multiple-choice and open-ended formats, explicitly designed to evaluate multimodal models through questions that necessitate tight integration of video and audio information, spanning a broad spectrum of agentic scenarios. MAVERIX uniquely provides models with questions that closely mimic the multimodal understanding experiences available to humans during decision-making processes. To our knowledge, MAVERIX is the first benchmark aimed explicitly at assessing comprehensive audiovisual integration in such granularity. Experiments with state-of-the-art models, including Qwen 2.5 Omni and Gemini 2.5 Flash-Lite, show performance around 64% accuracy, while human experts reach near-ceiling performance of 92.8%, exposing a substantial gap to human-level comprehension. With standardized evaluation protocols, a rigorously annotated pipeline, and a public toolkit, MAVERIX establishes a challenging testbed for advancing audiovisual multimodal intelligence, with the website publicly available below. Liuyue Xie, Avik Kuthiala, George Z. Wei, Ananya Bal, Mosam Dabhi, Liting Wen, Taru Rustagi, Ethan Lai, Sushil Khyalia, Rohan Choudhury, Morteza Ziyadi, László A. Jeni |
AAAI | 3 |
| 2022 | Robustness Disparities in Face DetectionabstractFacial analysis systems have been deployed by large companies and critiqued by scholars and activists for the past decade. Many existing algorithmic audits examine the performance of these systems on later stage elements of facial analysis systems like facial recognition and age, emotion, or perceived gender prediction; however, a core component to these systems has been vastly understudied from a fairness perspective: face detection, sometimes called face localization. Since face detection is a pre-requisite step in facial analysis systems, the bias we observe in face detection will flow downstream to the other components like facial recognition and emotion prediction. Additionally, no prior work has focused on the robustness of these systems under various perturbations and corruptions, which leaves open the question of how various people are impacted by these phenomena. We present the first of its kind detailed benchmark of face detection systems, specifically examining the robustness to noise of commercial and academic models. We use both standard and recently released academic facial datasets to quantitatively analyze trends in face detection robustness. Across all the datasets and systems, we generally find that photos of individuals who are masculine presenting, older, of darker skin type, or have dim lighting are more susceptible to errors than their counterparts in other identities. Samuel Dooley, George Z. Wei, Tom Goldstein, John Dickerson 0001 |
NeurIPS | 2 |