Songtao Jiang

dblp:97/3729 · DBLP profile ↗
← Back
9ranked-venue papers
6as first author
8since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 3 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 44% Reinforcement learning · 28% Language models and text generation · 21%

Topics — the 15 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Vision and language
medical report generation
1.012026
Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation · AAAI 2026
Computer vision › Vision and language › visual question answering
medical visual question answering
1.012026
Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
reasoning consistency
1.012026
Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering · ACL (1) 2026
Machine learning › Reinforcement learning
reinforcement learning for reasoning
1.012026
Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering · ACL (1) 2026
Machine learning › Reinforcement learning
reward learning
1.012026
Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation · AAAI 2026
Computer vision › Vision and language
visual question answering
1.012026
Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering · ACL (1) 2026
Natural language and speech › Language models and text generation
alignment
0.912025
HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models · ACL (1) 2025
Natural language and speech › Language models and text generation
hallucination mitigation
0.912025
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment · IJCAI 2025
Natural language and speech › Language models and text generation
mathematical reasoning
0.912025
Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025
Computer vision › Vision and language › multimodal reasoning
multimodal mathematical reasoning
0.912025
Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025
Machine learning › Reinforcement learning › reinforcement learning from human feedback
process reward model
0.912025
Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025
Machine learning › Reinforcement learning
reinforcement learning from process rewards
0.912025
Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025
Computer vision › Vision and language › vision-language model
vision-language model alignment
0.912025
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment · IJCAI 2025
Natural language and speech › Language models and text generation › chain-of-thought reasoning
multimodal chain-of-thought
0.312025
Unlocking Multimodal Mathematical Reasoning via Process Reward Model · NeurIPS 2025
Computer vision › Vision and language › vision-language model
multimodal large language model
0.312025
Modality-Fair Preference Optimization for Trustworthy MLLM Alignment · IJCAI 2025

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 2.0reinforcement learning from human feedback · 1.7reward model · 1.0large language model · 1.0process reward model · 0.9preference optimization · 0.9hierarchical self-contrastive rewarding · 0.9group relative policy optimization · 0.9chain-of-thought · 0.9
YearPublicationVenuePosition
2026 Beyond N-grams: A Hierarchical Reward Learning Framework for Clinically-Aware Medical Report Generation
abstract
Automatic medical report generation can greatly reduce the workload of doctors, but it is often unreliable for real-world deployment. Current methods can write formally fluent sentences but may be factually flawed, introducing serious medical errors known as clinical hallucinations, which make them untrustworthy for diagnosis. To bridge this gap, we introduce HiMed-RL, a Hierarchical Medical Reward Learning Framework designed to explicitly prioritize clinical quality. HiMed-RL moves beyond simple text matching by deconstructing reward learning into three synergistic levels: it first ensures linguistic fluency at the token-level, then enforces factual grounding at the concept-level by aligning key medical terms with expert knowledge, and finally assesses high-level diagnostic consistency at the semantic-level using a specialized LLM verifier. This hierarchical reward is implemented via a Human-inspired Dynamic Reward Adjustment, a strategy which first teaches the model to learn basic facts before progressing to more complex diagnostic reasoning. Experimentally, HiMed-3B achieves state-of-the-art performance on both in-domain and out-of-domain benchmarks, particularly on the latter, with an improvement of 10.8% over the second-best baseline. Our work provides a robust paradigm for generating reports that not only improve fluency but clinical fine-grained quality.
Shujian Gao, Songtao Jiang, Haoxiang Xia, Zhaolu Kang, Yemin Wang, Zuozhu Liu
AAAI4
2026 Act as you think: Reinforcing Consistent Reasoning in Medical Visual Question Answering
abstract
Songtao Jiang, Yuan Wang, Ruizhe Chen, Yan Zhang, Ruilin Luo, Bohan Lei, Yeying Jin, Sibo Song, ZhiBo Yang, Jimeng Sun, Jian Wu, Zuozhu Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Songtao Jiang, Ruizhe Chen, Yan Zhang 0004, Ruilin Luo, Bohan Lei, Yeying Jin, Sibo Song, Jimeng Sun 0001, Jian Wu 0001, Zuozhu Liu
ACL (1)1
2025 HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models
abstract
Songtao Jiang, Yan Zhang, Yeying Jin, Zhihang Tang, Yangyang Wu, Yang Feng, Jian Wu, Zuozhu Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Songtao Jiang, Yan Zhang 0004, Yeying Jin, Zhihang Tang, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu
ACL (1)1
2025 Med-GLIP: Advancing Medical Language-Image Pre-Training with Large-Scale Grounded Dataset
abstract
Medical image grounding aims to align natural language phrases with specific regions in medical images, serving as a foundational task for intelligent diagnosis, visual question answering (VQA), and automated report generation(MRG). However, existing research is constrained by limited modality coverage, coarse-grained annotations, and the absence of a unified, generalizable grounding framework. To address these challenges, we construct a large-scale medical grounding dataset Med-GLIP-5M comprising over 5.3 million region-level annotations across seven imaging modalities, covering diverse anatomical structures and pathological findings. The dataset supports both segmentation and grounding tasks with hierarchical region labels, ranging from organ-level boundaries to fine-grained lesions. Based on this foundation, we propose Med-GLIP, a modality-aware grounding framework trained on Med-GLIP-5M. Rather than relying on explicitly designed expert modules, Med-GLIP implicitly acquires hierarchical semantic understanding from diverse training data-enabling it to recognize multi-granularity structures, such as distinguishing lungs from pneumonia lesions. Extensive experiments demonstrate that Med-GLIP consistently outperforms state-of-the-art baselines across multiple grounding benchmarks. Furthermore, integrating its spatial outputs into downstream tasks, including medical VQA and report generation, leads to substantial performance gains. Our dataset is available at Venn2025/Med-GLIP-5M.
Ziye Deng, Ruihan He, Zijie Meng, Songtao Jiang, Zuozhu Liu
BIBM6
2025 Syndrome-Molecule Graph Attention Networks (S-MoleGAT): Towards Interpretable TCM Prescription Generation
abstract
Traditional Chinese Medicine prescription recommendation faces significant challenges in integrating holistic diagnostic principles with molecular-level interpretability. To bridge this gap, we propose Syndrome-Molecule Graph Attention Networks (S-MoleGAT), a novel cross-modal framework that synergizes TCM syndrome differentiation with molecular interaction modeling. Our approach features three key innovations: (1) A temporal symptom perception module capturing dynamic syndrome progression through GRU networks, (2) A heterogeneous graph convolutional network integrating symptomherb relationships across medical paradigms, and (3) A dualchannel component propagation mechanism analyzing molecular synergies via substructure attention and message-passing networks. Evaluations demonstrate S-MoleGAT's superiority with 0.387 similarity ratio (17.3% above clinical benchmarks), identifying 5.02 similar pairs/prescription at optimal herb count (13.05). The model provides interpretable herb compatibility insights through cross-modal attention mechanisms, advancing scientifically-grounded TCM modernization.
Songtao Jiang, Zuxu Wang, Changqing Ji
BIBM1
2025 Modality-Fair Preference Optimization for Trustworthy MLLM Alignment
abstract
Multimodal large language models (MLLMs) have achieved remarkable success across various tasks. However, separate training of visual and textual encoders often results in a misalignment of the modality. Such misalignment may lead models to generate content that is absent from the input image, a phenomenon referred to as hallucination. These inaccuracies severely undermine the trustworthiness of MLLMs in real-world applications. Despite attempts to optimize text preferences to mitigate this issue, our initial investigation indicates that the trustworthiness of MLLMs remains inadequate. Specifically, these models tend to provide preferred answers even when the input image is heavily distorted. Analysis of visual token attention also indicates that the model focuses primarily on the surrounding context rather than the key object referenced in the question. These findings highlight a misalignment between the modalities, where answers inadequately leverage input images. Motivated by our findings, we propose Modality-Fair Preference Optimization (MFPO), which comprises three components: the construction of a multimodal preference dataset in which dispreferred images differ from originals solely in key regions; an image reward loss function encouraging the model to generate answers better aligned with the input images; and an easy-to-hard iterative alignment strategy to stabilize joint modality training. Extensive experiments on three trustworthiness benchmarks demonstrate that MFPO significantly enhances the trustworthiness of MLLMs. In particular, it enables the 7B models to attain trustworthiness levels on par with, or even surpass, those of the 13B, 34B, and larger models.
Songtao Jiang, Yan Zhang 0004, Ruizhe Chen, Tianxiang Hu, Yeying Jin, Qinglin He, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu
IJCAI1
2025 Knowing or Guessing? Robust Medical Visual Question Answering via Joint Consistency and Contrastive Learning
Songtao Jiang, Sibo Song, Yan Zhang 0004, Yeying Jin, Yang Feng 0011, Jian Wu 0001, Zuozhu Liu
MICCAI (11)1
2025 Unlocking Multimodal Mathematical Reasoning via Process Reward Model
abstract
Process Reward Models (PRMs) have shown promise in enhancing the mathematical reasoning capabilities of Large Language Models (LLMs) through Test-Time Scaling (TTS). However, their integration into multimodal reasoning remains largely unexplored. In this work, we take the first step toward unlocking the potential of PRMs in multimodal mathematical reasoning. We identify three key challenges: (i) the scarcity of high-quality reasoning data constrains the capabilities of foundation Multimodal Large Language Models (MLLMs), which imposes further limitations on the upper bounds of TTS and reinforcement learning (RL); (ii) a lack of automated methods for process labeling within multimodal contexts persists; (iii) the employment of process rewards in unimodal RL faces issues like reward hacking, which may extend to multimodal scenarios. To address these issues, we introduce URSA, a three-stage Unfolding multimodal pRocess-Supervision Aided training framework. We first construct MMathCoT-1M, a high-quality large-scale multimodal Chain-of-Thought (CoT) reasoning dataset, to build a stronger math reasoning foundation MLLM, URSA-8B. Subsequently, we go through an automatic process to synthesize process supervision data, which emphasizes both logical correctness and perceptual consistency. We introduce DualMath-1.1M to facilitate the training of URSA-8B-RM. Finally, we propose Process-Supervised Group-Relative-Policy-Optimization (PS-GRPO), pioneering a multimodal PRM-aided online RL method that outperforms vanilla GRPO. With PS-GRPO application, URSA-8B-PS-GRPO outperforms Gemma3-12B and GPT-4o by 8.4% and 2.7% on average across 6 benchmarks.
Ruilin Luo, Zhuofan Zheng, Xinzhe Ni, Zicheng Lin, Songtao Jiang, Yiyao Yu, Chufan Shi, Ruihang Chu, Yujiu Yang 0001
NeurIPS7
2007 Cost modeling of spatial operators using non-parametric regression
Songtao Jiang, Byung Suk Lee 0001, Zhen He 0002
Inf. Sci.1