Jihao Chen

dblp:377/0582 · DBLP profile ↗
← Back
4ranked-venue papers
0as first author
4since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Reinforcement learning · 77% Transfer learning and domain adaptation · 12% Multi-agent systems · 12%
Computer architecture, parallel and distributed computing, and storage systems
1 paper
High-performance computing · 56% Parallel and multicore computing · 44%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Smart cities and intelligent transportation · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Reinforcement learning › policy optimization
maximum a posteriori policy optimization
1.012026
Counterfactual baseline-based MAPPO for asymmetric UAV swarm confrontation game · Sci. China Inf. Sci. 2026
Machine learning › Reinforcement learning
multi-agent reinforcement learning
1.012026
Counterfactual baseline-based MAPPO for asymmetric UAV swarm confrontation game · Sci. China Inf. Sci. 2026
Smart cities and intelligent transportation › traffic prediction
spatio-temporal traffic prediction
1.012026
MoST: A Foundation Model for Multi-modality Spatio-temporal Traffic Prediction · KDD (1) 2026
Smart cities and intelligent transportation
traffic prediction
1.012026
MoST: A Foundation Model for Multi-modality Spatio-temporal Traffic Prediction · KDD (1) 2026
High-performance computing
large-scale training
0.912025
Hypertron: Efficiently Scaling Large Models by Exploring High-Dimensional Parallelization Space · SC 2025
Parallel and multicore computing
parallelization strategies
0.912025
Hypertron: Efficiently Scaling Large Models by Exploring High-Dimensional Parallelization Space · SC 2025
Machine learning › Transfer learning and domain adaptation › zero-shot learning
zero-shot forecasting
0.312026
MoST: A Foundation Model for Multi-modality Spatio-temporal Traffic Prediction · KDD (1) 2026

Methods — techniques the papers use, named apart from their topics

spatial expert selection · 2.0multi-modality refinement · 2.0counterfactual baseline · 1.0performance model · 0.9high-dimensional exploration · 0.9dimension fusion · 0.9
YearPublicationVenuePosition
2026 MoST: A Foundation Model for Multi-modality Spatio-temporal Traffic Prediction
abstract
Accurate spatio-temporal traffic prediction is essential for optimizing urban traffic management and resource allocation. To reduce the cost and complexity of cross-city deployment, recent studies have explored spatio-temporal foundation models capable of accurate zero-shot prediction. However, these models are limited to single-modal data, which restricts their capacity to capture the complexity of real-world traffic dynamics. The increasing availability of multi-modality data—such as satellite imagery and points of interest (POI)—offers a promising avenue for enhancing cross-city traffic prediction by providing richer background contexts. Despite this potential, developing foundational models for multi-modality spatio-temporal prediction presents two challenges: the availability and quality of multi-modality data vary significantly across cities, with some cities lacking certain modalities or containing noisy information; and spatial patterns are highly localized and specific to individual regions, which hinders generalization. To address these challenges, we propose MoST, a foundation model for multi-modality spatio-temporal traffic prediction. We introduce a Multi-modality Refinement Module that encodes available modality data and adaptively selects task-relevant modalities while suppressing noisy modalities. Furthermore, we design a Spatio-Temporal Prediction Module that incorporates a spatial expert selection mechanism guided by multi-modality cues. This mechanism dynamically identifies region-specific spatial patterns and assigns appropriate spatial experts to model local dependencies. Finally, we conduct extensive experiments on real-world datasets to validate the superior performance and strong generalization capability of MoST.
Ronghui Xu 0001, Jihao Chen, Jindong Tian, Chenjuan Guo, Bin Yang 0002
KDD (1)2
2026 Counterfactual baseline-based MAPPO for asymmetric UAV swarm confrontation game
Ershen Wang, Zeqi Tong, Xiaotong Wu, Mingming Xiao, Jihao Chen
Sci. China Inf. Sci.7
2025 CoT Reasoning-Based Content Adaptation and Image Generation for Chinese Poetry
Jihao Chen, Songtao Chen, Gaoqi He
ICIC (24)2
2025 Hypertron: Efficiently Scaling Large Models by Exploring High-Dimensional Parallelization Space
abstract
Large models are evolving towards massive scale, diverse model architectures (dense and sparse) and long-context processing, which makes it very challenging to efficiently scale large models on parallel machines. The current widely-used parallelization strategies are often sub-optimal due to their limited parallelization strategy space. To this end, we propose Hypertron, a scalable parallel large-model training framework which incorporates an unprecedented high-dimensional (up to 7D) parallelization space, a holistic scheme for efficient dimension fusion, and a comprehensive performance model to guide the high-dimensional exploration. By exploiting the high-dimensional space to discover the optimal strategy which is not supported by existing frameworks, Hypertron significantly reduces memory and communication cost while improving parallel scalability. Extensive evaluations demonstrate that Hypertron achieves up to 56.7% Model FLOPs Utilization (MFU) on 2,048 new-generation Ascend NPU accelerators (scaling with supernodes) for different large models (such as sparse 141B and dense 310B), with 1.33x speedup over the best configuration of the state-of-the-art frameworks.
Shigang Li 0002, Jingkun Dong, Jihao Chen, Zhongzhe Hu
SC3