VLDB 2026 Research / reviewers in the wild / expert
Juhao Liang
dblp:344/0709
· DBLP profile ↗
4ranked-venue papers
1as first author
4since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
3 papers |
Language models and text generation · 85% Deep learning architectures and training · 15% | |
| Interdisciplinary, comprehensive, and emerging computing
1 paper |
Medical and health informatics · 100% |
Topics — the 5 heaviest of 6, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Natural language and speech › Language models and text generation
multilingual language models |
2.5 | 3 | 2025 | Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts · ICLR 2025 Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary Expansion · ACL (1) 2025 Alignment at Pre-training! Towards Native Alignment for Arabic LLMs · NeurIPS 2024 |
Machine learning › Deep learning architectures and training
mixture of experts |
0.9 | 1 | 2025 | Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts · ICLR 2025 |
Natural language and speech › Language models and text generation
alignment |
0.8 | 1 | 2024 | Alignment at Pre-training! Towards Native Alignment for Arabic LLMs · NeurIPS 2024 |
Natural language and speech › Language models and text generation › multilingual language models
arabic language models |
0.8 | 1 | 2024 | Alignment at Pre-training! Towards Native Alignment for Arabic LLMs · NeurIPS 2024 |
Medical and health informatics › biomedical natural language processing
medical language model |
0.3 | 1 | 2025 | Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family Experts · ICLR 2025 |
Methods — techniques the papers use, named apart from their topics
sparse routing · 1.7language family experts · 1.7progressive vocabulary expansion · 0.9continued pretraining · 0.9reinforcement learning · 0.8pre-training · 0.8instruction tuning · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Second Language (Arabic) Acquisition of LLMs via Progressive Vocabulary ExpansionabstractJianqing Zhu, Huang Huang, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An, Juncai He, Xiangbo Wu, Fei Yu, Junying Chen, Ma Zhuoheng, Yuhao Du, He Zhang, Saied Alshahrani, Emad A. Alghamdi, Lian Zhang, Ruoyu Sun, Haizhou Li, Benyou Wang, Jinchao Xu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Jianqing Zhu, Zhihang Lin, Juhao Liang, Zhengyang Tang, Khalid Almubarak, Mosen Alharthi, Bang An 0004, Juncai He 0001, Xiangbo Wu, Fei Yu 0017, Zhuoheng Ma, Saied Alshahrani, Emad A. Alghamdi, Ruoyu Sun 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu |
ACL (1) | 4 |
| 2025 | Efficiently Democratizing Medical LLMs for 50 Languages via a Mixture of Language Family ExpertsabstractAdapting medical Large Language Models to local languages can reduce barriers to accessing healthcare services, but data scarcity remains a significant challenge, particularly for low-resource languages. To address this, we first construct a high-quality medical dataset and conduct analysis to ensure its quality. In order to leverage the generalization capability of multilingual LLMs to efficiently scale to more resource-constrained languages, we explore the internal information flow of LLMs from a multilingual perspective using Mixture of Experts (MoE) modularity. Technically, we propose a novel MoE routing method that employs language-specific experts and cross-lingual routing. Inspired by circuit theory, our routing analysis revealed a \textit{``Spread Out in the End``} information flow mechanism: while earlier layers concentrate cross-lingual information flow, the later layers exhibit language-specific divergence. This insight directly led to the development of the Post-MoE architecture, which applies sparse routing only in the later layers while maintaining dense others. Experimental results demonstrate that this approach enhances the generalization of multilingual models to other languages while preserving interpretability. Finally, to efficiently scale the model to 50 languages, we introduce the concept of \textit{language family} experts, drawing on linguistic priors, which enables scaling the number of languages without adding additional parameters. Guorui Zheng, Xidong Wang, Juhao Liang, Nuo Chen 0002, Yuping Zheng, Benyou Wang |
ICLR | 3 |
| 2025 | Smurfs: Multi-Agent System using Context-Efficient DFSDT for Tool PlanningabstractJunzhi Chen, Juhao Liang, Benyou Wang. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Junzhi Chen, Juhao Liang, Benyou Wang |
NAACL (Long Papers) | 2 |
| 2024 | Alignment at Pre-training! Towards Native Alignment for Arabic LLMsabstractThe alignment of large language models (LLMs) is critical for developing effective and safe language models. Traditional approaches focus on aligning models during the instruction tuning or reinforcement learning stages, referred to in this paper as `\textit{post alignment}'. We argue that alignment during the pre-training phase, which we term 'native alignment', warrants investigation. Native alignment aims to prevent unaligned content from the beginning, rather than relying on post-hoc processing. This approach leverages extensively aligned pre-training data to enhance the effectiveness and usability of pre-trained models. Our study specifically explores the application of native alignment in the context of Arabic LLMs. We conduct comprehensive experiments and ablation studies to evaluate the impact of native alignment on model performance and alignment stability. Additionally, we release open-source Arabic LLMs that demonstrate state-of-the-art performance on various benchmarks, providing significant benefits to the Arabic LLM community. Juhao Liang, Zhenyang Cai, Jianqing Zhu, Kewei Zong, Bang An 0004, Mosen Alharthi, Juncai He 0001, Haizhou Li 0001, Benyou Wang, Jinchao Xu |
NeurIPS | 1 |