Yemin Wang

dblp:405/4209 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2026
0009-0009-6242-7361ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Trustworthy machine learning · 33% Question answering and dialogue systems · 25% Language models and text generation · 16%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 62% Knowledge graphs · 38%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
knowledge base question answering
1.012026
S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering · WWW 2026
Computer vision › Vision and language
medical report generation
1.012026
Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation · AAAI 2026
Natural language and speech › Question answering and dialogue systems › reasoning-based question answering
multi-hop question answering
1.012026
S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering · WWW 2026
Natural language and speech › Language models and text generation
retrieval-augmented generation
1.012026
S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering · WWW 2026
Machine learning › Reinforcement learning
reward learning
1.012026
Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation · AAAI 2026
Knowledge graphs › knowledge graph querying
knowledge graph question answering
1.012026
S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering · WWW 2026
Information retrieval › question answering
multi-hop question answering
1.012026
S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering · WWW 2026
Machine learning › Trustworthy machine learning
fairness
0.912025
F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations · EMNLP 2025
Machine learning › Trustworthy machine learning › fairness
fairness evaluation
0.912025
F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations · EMNLP 2025
Machine learning › Trustworthy machine learning › fairness
intersectional bias
0.912025
F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations · EMNLP 2025
Information retrieval › retrieval models
graph-based retrieval
0.312026
S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering · WWW 2026
Information retrieval
retrieval models
0.312026
S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering · WWW 2026
Natural language and speech › Language models and text generation
large language model evaluation
0.312025
F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

cross-attention · 2.0contrastive learning · 2.0constrained random walk · 2.0beam search · 2.0shortest-path search · 1.0shortest path search · 1.0reward model · 1.0reinforcement learning · 1.0large language model · 1.0open-ended evaluation · 0.9factuality grounding · 0.9
YearPublicationVenuePosition
2026 3D CoCa: Contrastive Learners are 3D Captioners
abstract
3D captioning, which aims to describe the content of 3D scenes in natural language, remains highly challenging due to the inherent sparsity of point clouds and weak cross-modal alignment in existing methods. To address these challenges, we propose 3D CoCa, a novel unified framework that seamlessly combines contrastive vision-language learning with 3D caption generation into a single architecture. We design a frozen CLIP vision-language backbone to provide rich semantic priors, a spatially-aware 3D scene encoder to capture geometric context, and a multi-modal decoder to generate descriptive captions. Unlike the prior two-stage methods that rely on explicit object proposals, 3D CoCa jointly optimizes contrastive and captioning objectives in a shared feature space, eliminating the need for external detectors or handcrafted proposals. This joint training paradigm yields stronger spatial reasoning and richer semantic grounding by aligning 3D and textual representations. Extensive experiments on the ScanRefer and Nr3D benchmarks demonstrate that 3D CoCa significantly outperforms current state-of-the-arts by 10.2% and 5.76% in [email protected], respectively. Code will be available at https://github.com/AIGeeksGroup/3DCoCa.
Zeyu Zhang 0006, Yemin Wang, Hao Tang 0005
3DV3
2026 Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation
abstract
Automatic medical report generation can greatly reduce the workload of doctors, but it is often unreliable for real-world deployment. Current methods can write formally fluent sentences but may be factually flawed, introducing serious medical errors known as clinical hallucinations, which make them untrustworthy for diagnosis. To bridge this gap, we introduce HiMed-RL, a Hierarchical Medical Reward Learning Framework designed to explicitly prioritize clinical quality. HiMed-RL moves beyond simple text matching by deconstructing reward learning into three synergistic levels: it first ensures linguistic fluency at the token-level, then enforces factual grounding at the concept-level by aligning key medical terms with expert knowledge, and finally assesses high-level diagnostic consistency at the semantic-level using a specialized LLM verifier. This hierarchical reward is implemented via a Human-inspired Dynamic Reward Adjustment, a strategy which first teaches the model to learn basic facts before progressing to more complex diagnostic reasoning. Experimentally, HiMed-3B achieves state-of-the-art performance on both in-domain and out-of-domain benchmarks, particularly on the latter, with an improvement of 10.8% over the second-best baseline. Our work provides a robust paradigm for generating reports that not only improve fluency but clinical fine-grained quality.
Shujian Gao, Songtao Jiang, Haoxiang Xia, Zhaolu Kang, Yemin Wang, Zuozhu Liu
AAAI8
2026 S-Path-RAG: Semantic-Aware Shortest-Path Retrieval Augmented Generation for Multi-Hop Knowledge Graph Question Answering
abstract
We present S-Path-RAG, a semantic-aware shortest-path Retrieval-Augmented Generation framework designed to improve multi-hop question answering over large knowledge graphs. S-Path-RAG departs from one-shot, text-heavy retrieval by enumerating bounded-length, semantically weighted candidate paths using a hybrid weighted $k$-shortest, beam, and constrained random-walk strategy, learning a differentiable path scorer together with a contrastive path encoder and lightweight verifier, and injecting a compact soft mixture of selected path latents into a language model via cross-attention. The system runs inside an iterative Neural-Socratic Graph Dialogue loop in which concise diagnostic messages produced by the language model are mapped to targeted graph edits or seed expansions, enabling adaptive retrieval when the model expresses uncertainty. This combination yields a retrieval mechanism that is both token-efficient and topology-aware while preserving interpretable path-level traces for diagnostics and intervention. We validate S-Path-RAG on standard multi-hop KGQA benchmarks and through ablations and diagnostic analyses. The results demonstrate consistent improvements in answer accuracy, evidence coverage, and end-to-end efficiency compared to strong graph- and LLM-based baselines. We further analyze trade-offs between semantic weighting, verifier filtering, and iterative updates, and report practical recommendations for deployment under constrained compute and token budgets.
Yemin Wang, Tianxiang Xu 0001, Yongtai Liu, Weizhi Tang, Wangyu Wu, Simon Fong 0001
WWW2
2025 F²Bench: An Open-ended Fairness Evaluation Benchmark for LLMs with Factuality Considerations
abstract
Warning: This paper contains content that may be offensive or harmful With the growing adoption of large language models (LLMs) in NLP tasks, concerns about their fairness have intensified.Yet, most existing fairness benchmarks rely on closed-ended evaluation formats, which diverge from realworld open-ended interactions.These formats are prone to position bias and introduce a "minimum score" effect, where models can earn partial credit simply by guessing.Moreover, such benchmarks often overlook factuality considerations rooted in historical, social, physiological, and cultural contexts, and rarely account for intersectional biases.To address these limitations, we propose F 2 Bench: an openended fairness evaluation benchmark for LLMs that explicitly incorporates factuality considerations.F 2 Bench comprises 2,568 instances across 10 demographic groups and two openended tasks.By integrating text generation, multi-turn reasoning, and factual grounding, F 2 Bench aims to more accurately reflect the complexities of real-world model usage.We conduct a comprehensive evaluation of several LLMs across different series and parameter sizes.Our results reveal that all models exhibit varying degrees of fairness issues.We further compare open-ended and closedended evaluations, analyze model-specific disparities, and provide actionable recommendations for future model development.Our code and dataset are publicly available at https: //github.com/VelikayaScarlet/F2Bench.
Jiang Li 0013, Yemin Wang, Xiangdong Su, Guanglai Gao
EMNLP3
2025 Robust Label Proportions Learning
abstract
Learning from Label Proportions (LLP) is a weakly-supervised paradigm that uses bag-level label proportions to train instance-level classifiers, offering a practical alternative to costly instance-level annotation. However, the weak supervision makes effective training challenging, and existing methods often rely on pseudo-labeling, which introduces noise. To address this, we propose RLPL, a two-stage framework. In the first stage, we use unsupervised contrastive learning to pretrain the encoder and train an auxiliary classifier with bag-level supervision. In the second stage, we introduce an LLP-OTD mechanism to refine pseudo labels and split them into high- and low-confidence sets. These sets are then used in LLPMix to train the final classifier. Extensive experiments and ablation studies on multiple benchmarks demonstrate that RLPL achieves comparable state-of-the-art performance and effectively mitigates pseudo-label noise.
Jueyu Chen, Wantao Wen, Yeqiang Wang, Erliang Lin, Yemin Wang, Yuheng Jia
NeurIPS5