VLDB 2026 Research / reviewers in the wild / expert
Yuheng Zha
dblp:295/8582
· DBLP profile ↗
7ranked-venue papers
3as first author
7since 2021 · last 2026
0009-0003-3489-8103ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
5 papers |
Vision and language · 28% Video understanding and tracking · 22% Reinforcement learning · 16% | |
| Databases, data mining, and information retrieval
1 paper |
Information retrieval · 100% |
Topics — the 13 heaviest of 15, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Reinforcement learning › reward design
reinforcement learning with verifiable rewards |
1.0 | 1 | 2026 | Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement Learning · AAAI 2026 |
Computer vision › Vision and language › multimodal reasoning
vision-language model reasoning |
1.0 | 1 | 2026 | Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement Learning · AAAI 2026 |
Computer vision › Vision and language › vision-language model
vision-language model training |
1.0 | 1 | 2026 | Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement Learning · AAAI 2026 |
Computer vision › Vision and language
visual reasoning |
1.0 | 1 | 2026 | Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement Learning · AAAI 2026 |
Computer vision › Video understanding and tracking
action detection |
0.9 | 1 | 2025 | CReLeRI: Explainable, Concept-centric, Representation, Learning, Reasoning, and Interaction Video Analysis System · ACM Multimedia 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | CReLeRI: Explainable, Concept-centric, Representation, Learning, Reasoning, and Interaction Video Analysis System · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation
large language model reasoning |
0.9 | 1 | 2025 | Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective · NeurIPS 2025 |
Machine learning › Reinforcement learning › reinforcement learning for NLP
reinforcement learning for language models |
0.9 | 1 | 2025 | Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain Perspective · NeurIPS 2025 |
Computer vision › Video understanding and tracking › action detection
temporal action localization |
0.9 | 1 | 2025 | CReLeRI: Explainable, Concept-centric, Representation, Learning, Reasoning, and Interaction Video Analysis System · ACM Multimedia 2025 |
Natural language and speech › Language models and text generation › large language model evaluation › truthfulness evaluation
factual consistency evaluation |
0.9 | 2 | 2023 | AlignScore: Evaluating Factual Consistency with A Unified Alignment Function · ACL (1) 2023 Text Alignment Is An Efficient Unified Model for Massive NLP Tasks · NeurIPS 2023 |
Natural language and speech › Information extraction and text analysis › text matching
text alignment |
0.7 | 1 | 2023 | Text Alignment Is An Efficient Unified Model for Massive NLP Tasks · NeurIPS 2023 |
Computer vision › Vision and language › vision-language model
multimodal large language model |
0.3 | 1 | 2026 | Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement Learning · AAAI 2026 |
Information retrieval › similarity measure
semantic similarity |
0.2 | 1 | 2023 | AlignScore: Evaluating Factual Consistency with A Unified Alignment Function · ACL (1) 2023 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.9question answering · 1.3natural language inference · 1.3multi-round data curriculum · 1.0influence function-based data filtering · 1.0scene segmentation · 0.9mixed-domain training · 0.9large language model · 0.9fine-tuning · 0.7RoBERTa · 0.7
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement LearningabstractRecent vision-language models (VLMs) show strong reasoning capabilities through training with reinforcement learning from verifiable rewards (RLVR). Despite their impressive capabilities, current VLMs focus on a limited range of reasoning tasks, such as mathematical and logical reasoning, due to the lack of readily available verifiable reward data in broader domains. As a result, these models struggle to generalize their reasoning abilities to the wide variety of challenges encountered in real-world environments. To address this limitation, we collect and assemble a comprehensive RL-ready visual reasoning training dataset encompassing 46 datasets across 13 dimensions of 5 domains, covering a wide range of realistic scenarios such as infographic reasoning, mathematical reasoning, spatial reasoning, and general science reasoning. Based on this dataset, we propose an influence function-based data filtering strategy and a multi-round data curriculum method to iteratively strengthen general visual reasoning abilities. Using this approach, we train a general reasoning VLM, namely Vision-G1. Our 7B model achieves state-of-the-art performance across nine visual reasoning benchmarks, surpassing previous similar-sized VLMs and even GPT-4o and Gemini-1.5 Flash. Yuheng Zha, Kun Zhou 0002, Yujia Wu, Yushu Wang, Shibo Hao, Zhengzhong Liu 0001, Eric P. Xing, Zhiting Hu |
AAAI | 1 |
| 2025 | CReLeRI: Explainable, Concept-centric, Representation, Learning, Reasoning, and Interaction Video Analysis SystemabstractExisting video analysis models often lack explainability, perform poorly on long videos, and frequently hallucinate. Commercial solutions are closed-source and costly. We introduce CReLeRI, an open-source system for action detection in untrimmed videos. CReLeRI segments videos using scene and action transitions, detects actions and their arguments and grounds them in 3D space to improve interpretability and reduce hallucinations. The system promotes transparency and trust in AI-driven analysis of complex, real-world videos. A demonstration video is also available. Michael Francis Perez, Yichi Yang, Yuheng Zha, Enze Ma, Danish Nisar Ahmed Tamboli, Haodi Ma, Reza Shahriari, Vyom Pathak, Dzmitry Kasinets, Rohith Venkatakrishnan, Daisy Zhe Wang, Jaime Ruiz 0002, Eric D. Ragan, Zhiting Hu, Eric P. Xing, Jun-Yan Zhu |
ACM Multimedia | 3 |
| 2025 | Revisiting Reinforcement Learning for LLM Reasoning from A Cross-Domain PerspectiveabstractReinforcement learning (RL) has shown promise in enhancing large language model (LLM) reasoning, yet progress towards broader capabilities is limited by the availability of high-quality, multi-domain datasets. This work introduces \ours, a 92K RL-for-reasoning dataset designed to address this gap, covering six reasoning domains: Math, Code, Science, Logic, Simulation, and Tabular, each with corresponding verifiers. We build \ours via a careful data-curation pipeline, including sourcing, deduplication, reward design, and domain-specific and difficulty-based filtering, to facilitate the systematic investigation of cross-domain RL generalization. Our study using \ours suggests the efficacy of a simple mixed-domain RL training approach and reveals several key aspects affecting cross-domain transferability. We further train two models {\ours}-7B and {\ours}-32B purely with RL on our curated data and observe largely improved performance over leading open RL reasoning model baselines, with gains of 7.3\% and 7.8\% respectively on an extensive 17-task, six-domain evaluation suite. We are releasing our dataset, code, and evaluation suite to the community, aiming to support further research and development of more general RL-enhanced reasoning models. Jorge (Zhoujun) Cheng, Shibo Hao, Tianyang Liu 0003, Yuexin Bian, Nilabjo Dey, Yonghao Zhuang 0001, Yuheng Zha, Yi Gu 0002, Kun Zhou 0002, Yuan Li 0032, Richard Fan, Jianshu She, Chengqian Gao, Abulhair Saparov, Taylor W. Killian, Haonan Li 0002, Mikhail Yurochkin, Eric P. Xing, Zhengzhong Liu 0001, Zhiting Hu |
NeurIPS | 10 |
| 2023 | AlignScore: Evaluating Factual Consistency with A Unified Alignment FunctionabstractMany text generation applications require the generated text to be factually consistent with input information.Automatic evaluation of factual consistency is challenging.Previous work has developed various metrics that often depend on specific functions, such as natural language inference (NLI) or question answering (QA), trained on limited data.Those metrics thus can hardly assess diverse factual inconsistencies (e.g., contradictions, hallucinations) that occur in varying inputs/outputs (e.g., sentences, documents) from different tasks.In this paper, we propose ALIGNSCORE, a new holistic metric that applies to a variety of factual inconsistency scenarios as above.ALIGN-SCORE is based on a general function of information alignment between two arbitrary text pieces.Crucially, we develop a unified training framework of the alignment function by integrating a large diversity of data sources, resulting in 4.7M training examples from 7 well-established tasks (NLI, QA, paraphrasing, fact verification, information retrieval, semantic similarity, and summarization).We conduct extensive experiments on large-scale benchmarks including 22 evaluation datasets, where 19 of the datasets were never seen in the alignment training.ALIGNSCORE achieves substantial improvement over a wide range of previous metrics.Moreover, ALIGNSCORE (355M parameters) matches or even outperforms metrics based on ChatGPT and GPT-4 that are orders of magnitude larger. 1 Yuheng Zha, Yichi Yang, Zhiting Hu |
ACL (1) | 1 |
| 2023 | Text Alignment Is An Efficient Unified Model for Massive NLP TasksabstractLarge language models (LLMs), typically designed as a function of next-word prediction, have excelled across extensive NLP tasks. Despite the generality, next-word prediction is often not an efficient formulation for many of the tasks, demanding an extreme scale of model parameters (10s or 100s of billions) and sometimes yielding suboptimal performance.
In practice, it is often desirable to build more efficient models---despite being less versatile, they still apply to a substantial subset of problems, delivering on par or even superior performance with much smaller model sizes.
In this paper, we propose text alignment as an efficient unified model for a wide range of crucial tasks involving text entailment, similarity, question answering (and answerability), factual consistency, and so forth. Given a pair of texts, the model measures the degree of alignment between their information. We instantiate an alignment model through lightweight finetuning of RoBERTa (355M parameters) using 5.9M examples from 28 datasets. Despite its compact size, extensive experiments show the model's efficiency and strong performance: (1) On over 20 datasets of aforementioned diverse tasks, the model matches or surpasses FLAN-T5 models that have around 2x or 10x more parameters; the single unified model also outperforms task-specific models finetuned on individual datasets; (2) When applied to evaluate factual consistency of language generation on 23 datasets, our model improves over various baselines, including the much larger GPT-3.5 (ChatGPT) and sometimes even GPT-4; (3) The lightweight model can also serve as an add-on component for LLMs such as GPT-3.5 in question answering tasks, improving the average exact match (EM) score by 17.94 and F1 score by 15.05 through identifying unanswerable questions. Yuheng Zha, Yichi Yang, Zhiting Hu |
NeurIPS | 1 |
| 2021 | KnowMeme: A Knowledge-enriched Graph Neural Network Solution to Offensive Meme DetectionabstractThis paper studies a critical problem of identifying offensive meme posts on online social media where images are superimposed with deliberately altered or crafted captions to deliver offensive information. Existing solutions often ignore the implicit relationship between visual and textual contents in the meme and their implied meanings, which are critical to accurately detect offensive memes. Two important challenges exist in solving the problem: i) how to effectively incorporate human commonsense knowledge to capture the symbolic or implied meaning of the meme contents that implicitly deliver offensive information? ii) How to accurately identify the cross-modal knowledge-based relations between entities in both the visual and textual content of the meme that jointly insinuate offensive messages? To address the above challenges, we develop KnowMeme, a knowledge-enriched graph neural network solution that leverages knowledge facts from human commonsense knowledge to effectively detect offensive meme posts on online social media. Evaluation results show that KnowMeme achieves significant performance gains compared to the state-of-the-art baseline methods in accurately detecting offensive memes. Lanyu Shang, Christina Youn, Yuheng Zha, Yang Zhang 0031, Dong Wang 0002 |
e-Science | 3 |
| 2021 | AOMD: An analogy-aware approach to offensive meme detection on social media
Lanyu Shang, Yang Zhang 0031, Yuheng Zha, Yingxi Chen, Christina Youn, Dong Wang 0002 |
Inf. Process. Manag. | 3 |