VLDB 2026 Research / reviewers in the wild / expert
Jialong Tang
dblp:242/7975
· DBLP profile ↗
13ranked-venue papers
3as first author
10since 2021 · last 2026
0000-0003-1259-2931ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 3 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Systematic Assessment of Language Models with Linguistic Minimal Pairs in ChineseabstractAbstract We present ZhoBLiMP, the largest linguistic minimal pair benchmark for Chinese, with over 100 paradigms, ranging from topicalization to the Ba construction. We then train from scratch a suite of Chinese language models (LMs) with different tokenizers, parameter sizes, and token volumes, to study the learning curves of LMs on Chinese. To mitigate the biases introduced by unequal lengths of the sentences in a minimal pair, we propose a new metric named sub-linear length normalized log-probabilities (SLLN-LP). Using SLLN-LP as the metric, our results show that Anaphor, Quantifiers, and Ellipsis in Chinese are difficult for LMs even up to 32B parameters, and that SLLN-LP successfully mitigates biases in ZhoBLiMP, JBLiMP and BLiMP. We conclude that future evaluations should be more carefully designed to consider the intricate relations between linking functions, LMs, and targeted minimal pairs. Yikang Liu 0002, Yeting Shen, Hongao Zhu, Lilong Xu, Zhiheng Qian, Siyuan Song, Jialong Tang, Pei Zhang 0011, Baosong Yang, Rui Wang 0015, Hai Hu 0001 |
Trans. Assoc. Comput. Linguistics | 8 |
| 2025 | Locate-and-Focus: Enhancing Terminology Translation in Speech Language ModelsabstractDirect speech translation (ST) has garnered increasing attention nowadays, yet the accurate translation of terminology within utterances remains a great challenge. In this regard, current studies mainly concentrate on leveraging various translation knowledge into ST models. However, these methods often struggle with interference from irrelevant noise and can not fully utilize the translation knowledge. To address these issues, in this paper, we propose a novel Locate-and-Focus method for terminology translation. It first effectively locates the speech clips containing terminologies within the utterance to construct translation knowledge, minimizing irrelevant information for the ST model. Subsequently, it associates the translation knowledge with the utterance and hypothesis from both audio and textual modalities, allowing the ST model to better focus on translation knowledge during translation. Experimental results across various datasets demonstrate that our method effectively locates terminologies within utterances and enhances the success rate of terminology translation, while maintaining robust general translation performance. Suhang Wu, Jialong Tang, Pei Zhang 0011, Baosong Yang, Junhui Li 0001, Junfeng Yao, Min Zhang 0005, Jinsong Su |
ACL (1) | 2 |
| 2025 | AI Hospital: Benchmarking Large Language Models in a Multi-agent Medical Interaction SimulatorabstractArtificial intelligence has significantly revolutionized healthcare, particularly through large language models (LLMs) that demonstrate superior performance in static medical question answering benchmarks. However, evaluating the potential of LLMs for real-world clinical applications remains challenging due to the intricate nature of doctor-patient interactions. To address this, we introduce AI Hospital, a multi-agent framework emulating dynamic medical interactions between Doctor as player and NPCs including Patient and Examiner. This setup allows for more practical assessments of LLMs in simulated clinical scenarios. We develop the Multi-View Medical Evaluation (MVME) benchmark, utilizing high-quality Chinese medical records and multiple evaluation strategies to quantify the performance of LLM-driven Doctor agents on symptom collection, examination recommendations, and diagnoses. Additionally, a dispute resolution collaborative mechanism is proposed to enhance medical interaction capabilities through iterative discussions. Despite improvements, current LLMs (including GPT-4) still exhibit significant performance gaps in multi-turn interactive scenarios compared to non-interactive scenarios. Our findings highlight the need for further research to bridge these gaps and improve LLMs’ clinical decision-making capabilities. Our data, code, and experimental results are all open-sourced at https://github.com/LibertFan/AI_Hospital. Zhihao Fan, Jialong Tang, Wei Chen 0088, Siyuan Wang 0025, Zhongyu Wei, Fei Huang 0002 |
COLING | 3 |
| 2025 | Translationese-index: Using Likelihood Ratios for Graded and Generalizable Measurement of TranslationeseabstractYikang Liu, Wanyang Zhang, Yiming Wang, Jialong Tang, Pei Zhang, Baosong Yang, Fei Huang, Rui Wang, Hai Hu. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. 2025. Yikang Liu 0002, Wanyang Zhang, Yiming Wang 0011, Jialong Tang, Pei Zhang 0011, Baosong Yang, Fei Huang 0002, Rui Wang 0015, Hai Hu 0001 |
EMNLP | 4 |
| 2025 | PolyMath: Evaluating Mathematical Reasoning in Multilingual ContextsabstractIn this paper, we introduce PolyMath, a multilingual mathematical reasoning benchmark covering 18 languages and 4 easy-to-hard difficulty levels. Our benchmark ensures difficulty comprehensiveness, language diversity, and high-quality translation, making it a highly discriminative multilingual mathematical benchmark in the era of reasoning LLMs.We conduct a comprehensive evaluation for advanced LLMs and find that even Qwen-3-235B-A22B-Thinking and Gemini-2.5-pro, achieve only 54.6 and 52.2 benchmark scores, with about 40% accuracy under the highest level.From a language perspective, our benchmark reveals several key challenges of LLMs in multilingual reasoning:(1) Reasoning performance varies widely across languages for current LLMs;(2) Input-output language consistency is low in reasoning LLMs and may be correlated with performance;(3) The thinking length differs significantly by language for current LLMs.Additionally, we demonstrate that controlling the output language in the instructions has the potential to affect reasoning performance, especially for some low-resource languages, suggesting a promising direction for improving multilingual capabilities in LLMs. Yiming Wang 0011, Pei Zhang 0011, Jialong Tang, Baosong Yang, Rui Wang 0015, Chenshu Sun, Feitong Sun, Jiran Zhang, Junxuan Wu, Qiqian Cang, Yichang Zhang, Fei Huang 0002, Junyang Lin, Fei Huang 0005, Jingren Zhou 0001 |
NeurIPS | 3 |
| 2022 | Procedural Text Understanding via Scene-Wise EvolutionabstractProcedural text understanding requires machines to reason about entity states within the dynamical narratives. Current procedural text understanding approaches are commonly entity-wise, which separately track each entity and independently predict different states of each entity. Such an entity-wise paradigm does not consider the interaction between entities and their states. In this paper, we propose a new scene-wise paradigm for procedural text understanding, which jointly tracks states of all entities in a scene-by-scene manner. Based on this paradigm, we propose Scene Graph Reasoner (SGR), which introduces a series of dynamically evolving scene graphs to jointly formulate the evolution of entities, states and their associations throughout the narrative. In this way, the deep interactions between all entities and states can be jointly captured and simultaneously derived from scene graphs. Experiments show that SGR not only achieves the new state-of-the-art performance but also significantly accelerates the speed of reasoning. Jialong Tang, Meng Liao, Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Weijian Xie, Jin Xu 0014 |
AAAI | 1 |
| 2022 | End-to-end neural event coreference resolution
Yaojie Lu 0001, Jialong Tang, Xianpei Han, Le Sun 0001 |
Artif. Intell. | 3 |
| 2021 | Text2Event: Controllable Sequence-to-Structure Generation for End-to-end Event ExtractionabstractYaojie Lu, Hongyu Lin, Jin Xu, Xianpei Han, Jialong Tang, Annan Li, Le Sun, Meng Liao, Shaoyi Chen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Yaojie Lu 0001, Jin Xu 0014, Xianpei Han, Jialong Tang, Annan Li, Le Sun 0001, Meng Liao, Shaoyi Chen |
ACL/IJCNLP (1) | 5 |
| 2021 | From Discourse to Narrative: Knowledge Projection for Event Relation ExtractionabstractJialong Tang, Hongyu Lin, Meng Liao, Yaojie Lu, Xianpei Han, Le Sun, Weijian Xie, Jin Xu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jialong Tang, Meng Liao, Yaojie Lu 0001, Xianpei Han, Le Sun 0001, Weijian Xie, Jin Xu 0014 |
ACL/IJCNLP (1) | 1 |
| 2021 | Enhanced aspect-based sentiment analysis models with progressive self-supervised attention learning
Jinsong Su, Jialong Tang, Ziyao Lu, Yubin Ge, Linfeng Song, Deyi Xiong, Le Sun 0001, Jiebo Luo 0001 |
Artif. Intell. | 2 |
| 2020 | A Rigorous Study on Named Entity Recognition: Can Fine-tuning Pretrained Model Lead to the Promised Land?abstractFine-tuning pretrained model has achieved promising performance on standard NER benchmarks.Generally, these benchmarks are blessed with strong name regularity, high mention coverage and sufficient context diversity.Unfortunately, when scaling NER to open situations, these advantages may no longer exist.And therefore it raises a critical question of whether previous creditable approaches can still work well when facing these challenges.As there is no currently available dataset to investigate this problem, this paper proposes to conduct randomization test on standard benchmarks.Specifically, we erase name regularity, mention coverage and context diversity respectively from the benchmarks, in order to explore their impact on the generalization ability of models.To further verify our conclusions, we also construct a new open NER dataset that focuses on entity types with weaker name regularity and lower mention coverage to verify our conclusion.From both randomization test and empirical experiments, we draw the conclusions that 1) name regularity is critical for the models to generalize to unseen mentions; 2) high mention coverage may undermine the model generalization ability and 3) context patterns may not require enormous data to capture when using pretrained encoders. Yaojie Lu 0001, Jialong Tang, Xianpei Han, Le Sun 0001, Zhicheng Wei, Nicholas Jing Yuan |
EMNLP (1) | 3 |
| 2019 | Progressive Self-Supervised Attention Learning for Aspect-Level Sentiment AnalysisabstractIn aspect-level sentiment classification (ASC), it is prevalent to equip dominant neural models with attention mechanisms, for the sake of acquiring the importance of each context word on the given aspect. However, such a mechanism tends to excessively focus on a few frequent words with sentiment polarities, while ignoring infrequent ones. In this paper, we propose a progressive self-supervised attention learning approach for neural ASC models, which automatically mines useful attention supervision information from a training corpus to refine attention mechanisms. Specifically, we iteratively conduct sentiment predictions on all training instances. Particularly, at each iteration, the context word with the maximum attention weight is extracted as the one with active/misleading influence on the correct/incorrect prediction of every instance, and then the word itself is masked for subsequent iterations. Finally, we augment the conventional training objective with a regularization term, which enables ASC models to continue equally focusing on the extracted active context words while decreasing weights of those misleading ones. Experimental results on multiple datasets show that our proposed approach yields better attention mechanisms, leading to substantial improvements over the two state-of-the-art neural ASC models. Source code and trained models are available at https://github.com/DeepLearnXMU/PSSAttention. Jialong Tang, Ziyao Lu, Jinsong Su, Yubin Ge, Linfeng Song, Le Sun 0001, Jiebo Luo 0001 |
ACL (1) | 1 |
| 2019 | A neural image captioning model with caption-to-images semantic constructor
Jinsong Su, Jialong Tang, Ziyao Lu, Xianpei Han |
Neurocomputing | 2 |