VLDB 2026 Research / reviewers in the wild / expert
Tingjian Zhang
dblp:202/8402
· DBLP profile ↗
7ranked-venue papers
1as first author
1since 2021 · last 2023
0000-0002-2700-5683ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 1 first-authorArtificial intelligence and machine learning · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer architecture, parallel and distributed computing, and storage systems
6 papers |
High-performance computing · 83% Processor architecture and microarchitecture · 12% Parallel and multicore computing · 4% | |
| Artificial intelligence
1 paper |
Language models and text generation · 50% Question answering and dialogue systems · 50% |
Topics — the 13 heaviest of 14, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
High-performance computing › scientific computing systems
molecular dynamics simulation |
2.0 | 5 | 2020 | Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway Taihulight · IEEE Trans. Parallel Distributed Syst. 2020 Cell-list based molecular dynamics on many-core processors: a case study on sunway TaihuLight supercomputer · SC 2020 Neighbor-list-free molecular dynamics on sunway TaihuLight supercomputer · PPoPP 2020 |
High-performance computing
performance optimization at scale |
1.1 | 3 | 2020 | Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway Taihulight · IEEE Trans. Parallel Distributed Syst. 2020 SW_GROMACS: accelerate GROMACS on Sunway TaihuLight · SC 2019 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenarios · SC 2017 |
High-performance computing › supercomputing
sunway taihulight |
0.9 | 3 | 2020 | Cell-list based molecular dynamics on many-core processors: a case study on sunway TaihuLight supercomputer · SC 2020 Redesigning LAMMPS for peta-scale and hundred-billion-atom simulation on Sunway TaihuLight · SC 2018 Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway Taihulight · IEEE Trans. Parallel Distributed Syst. 2020 |
Processor architecture and microarchitecture
many-core architecture |
0.9 | 3 | 2020 | Cell-list based molecular dynamics on many-core processors: a case study on sunway TaihuLight supercomputer · SC 2020 Redesigning LAMMPS for peta-scale and hundred-billion-atom simulation on Sunway TaihuLight · SC 2018 SW_GROMACS: accelerate GROMACS on Sunway TaihuLight · SC 2019 |
High-performance computing
scientific computing systems |
0.7 | 2 | 2019 | SW_GROMACS: accelerate GROMACS on Sunway TaihuLight · SC 2019 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenarios · SC 2017 |
Natural language and speech › Language models and text generation › natural language understanding › question answering
explainable question answering |
0.7 | 1 | 2023 | Reasoning over Hierarchical Question Decomposition Tree for Explainable Question Answering · ACL (1) 2023 |
Natural language and speech › Question answering and dialogue systems › question understanding
question decomposition |
0.7 | 1 | 2023 | Reasoning over Hierarchical Question Decomposition Tree for Explainable Question Answering · ACL (1) 2023 |
High-performance computing › performance optimization
many-core processor optimization |
0.4 | 1 | 2020 | Neighbor-list-free molecular dynamics on sunway TaihuLight supercomputer · PPoPP 2020 |
High-performance computing › scientific computing systems › molecular dynamics simulation
reactive force field |
0.4 | 1 | 2020 | Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway Taihulight · IEEE Trans. Parallel Distributed Syst. 2020 |
High-performance computing › scientific computing systems
earthquake simulation |
0.3 | 1 | 2017 | 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenarios · SC 2017 |
Parallel and multicore computing › many-core systems
many-core parallelization |
0.3 | 1 | 2017 | 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenarios · SC 2017 |
High-performance computing
supercomputing |
0.1 | 1 | 2020 | Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway Taihulight · IEEE Trans. Parallel Distributed Syst. 2020 |
Memory systems › memory management
on-chip memory management |
0.1 | 1 | 2017 | 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenarios · SC 2017 |
Methods — techniques the papers use, named apart from their topics
vectorization · 0.8tree reasoning · 0.7replica summation · 0.4pipelined conjugate gradient · 0.4data reuse optimization · 0.4data layout reorganization · 0.4cutoff filtering · 0.4cell-list · 0.4deferred update · 0.4halo exchange · 0.3DMA coalescing · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2023 | Reasoning over Hierarchical Question Decomposition Tree for Explainable Question AnsweringabstractJiajie Zhang, Shulin Cao, Tingjian Zhang, Xin Lv, Juanzi Li, Lei Hou, Jiaxin Shi, Qi Tian. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Shulin Cao, Tingjian Zhang, Juan-Zi Li, Lei Hou 0001, Jiaxin Shi |
ACL (1) | 3 |
| 2020 | Neighbor-list-free molecular dynamics on sunway TaihuLight supercomputerabstractMolecular dynamics (MD) simulations are playing an increasingly important role in many research areas. Pair-wise potentials are widely used in MD simulations of bio-molecules, polymers, and nano-scale materials. Due to a low compute-to-memory-access ratio, their calculation is often bounded by memory transfer speeds. Sunway TaihuLight is one of the fastest supercomputers featuring a custom SW26010 many-core processor. Since the SW26010 has some critical limitations regarding main memory bandwidth and scratchpad memory size, it is considered as a good platform to investigate the optimization of pair-wise potentials especially in terms of data reusage. MD algorithms often use a neighbor-list data structure to reduce the computational workload. In this paper, we show that a cell-list-based approach is more suitable for the SW26010 processor. We apply a number of novel optimization methods including self-adaptable replica-summation for conflict-free parallelization, parameter profiles for flexible vectorization, and particle-cell cutoff checking filters for reducing the computational workload. We also established an open source standalone framework featuring the techniques above, ESMD1, which is at least 50% faster than the latest existing LAMMPS port on a single TaihuLight node. Furthermore, EMSD achieves a weak scaling efficiency of 88% on 4,096 nodes. Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Hongsong Meng, Bertil Schmidt, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002 |
PPoPP | 4 |
| 2020 | Cell-list based molecular dynamics on many-core processors: a case study on sunway TaihuLight supercomputerabstractMolecular dynamics (MD) simulations are playing an increasingly important role in several research areas. The most frequently used potentials in MD simulations are pair-wise potentials. Due to the memory wall, computing pair-wise potentials on many-core processors are usually memory bounded. In this paper, we take the SW26010 processor as an exemplary platform to explore the possibility to break the memory bottleneck by improving data reusage via cell-list-based methods. We use cell-lists instead of neighbor-lists in the potential computation, and apply a number of novel optimization methods. Theses methods include: an adaptive replica arrangement strategy, a parameter profile data structure, and a particle-cell cutoff checking filter. An incremental cell-list building method is also realized to accelerate the construction of cell-lists. Furthermore, we have established an open source standalone framework, ESMD, featuring the techniques above. Experiments show that ESMD is 50~170% faster than previous ports on a single node, and can scale to 1,024 nodes with a weak scalibility of 95%. Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Hongsong Meng, Bertil Schmidt, Haohuan Fu, Lin Gan 0001, Wei Xue 0003, Guangwen Yang 0002 |
SC | 4 |
| 2020 | Millimeter-Scale and Billion-Atom Reactive Force Field Simulation on Sunway TaihulightabstractLarge-scale molecular dynamics (MD) simulations on supercomputers play an increasingly important role in many research areas. With the capability of simulating charge equilibration (QEq), bonds and so on, Reactive force field (ReaxFF) enables the precise simulation of chemical reactions. Compared to the first principle molecular dynamics (FPMD), ReaxFF has far lower requirements on computational resources so that it can achieve higher efficiencies for large-scale simulations. In this article, we present our efforts on scaling ReaxFF on the Sunway TaihuLight Supercomputer (TaihuLight). We have carefully redesigned the force analysis and neighbor list building steps. By applying fine-grained optimizations we gain better single process performance. For the many-body interactions, we propose an isolated computation and update strategy and implement inverse trigonometric functions. For QEq, we implement a pipelined conjugate gradient (CG) approach to achieving better scalability. Furthermore, we reorganize the data layout and implement the update operation based on data locality in ReaxFF. Our experiments show that this approach can simulate chemical reactions with 1,358,954,496 atoms using 4,259,840 cores with a performance of 0.015 ns/day. To our best knowledge, this is the first realization of chemical reaction simulation with a millimeter-scale force field. Ping Gao 0005, Xiaohui Duan, Tingjian Zhang, Bertil Schmidt, Wusheng Zhang, Lin Gan 0001, Wei Xue 0003, Haohuan Fu, Guangwen Yang 0002 |
IEEE Trans. Parallel Distributed Syst. | 3 |
| 2019 | SW_GROMACS: accelerate GROMACS on Sunway TaihuLightabstractGROMACS is one of the most popular Molecular Dynamic (MD) applications and is widely used in the field of chemical and bimolecular system study. Similar to other MD applications, it needs long run-time for large-scale simulations. Therefore, many high performance platforms have been employed to accelerate it, such as Knights Landing (KNL), Cell Processor, Graphics Processing Unit (GPU) and so on. As the third fastest supercomputer in the world, Sunway TaihuLight contains 40960 SW26010 processors and SW26010 is a typical many-core processor. To make full use of the superior computation ability of TaihuLight, we port GROMACS to SW26010 with following new strategies: (1) a new deferred update strategy; (2) a new update mark strategy; (3) a full pipeline acceleration. Furthermore, we redesign GROMACS to enable all possible vectorization. Experiments show that our implementation achieves better performance than both Intel KNL and Nvidia P100 GPU when using appropriate number of SW26010 processors for a fair comparison. Tingjian Zhang, Ping Gao 0005, Mingshan Shao, Jinxiao Zhang, Xiaohui Duan, Lin Gan 0001, Haohuan Fu, Wei Xue 0003, Guangwen Yang 0002 |
SC | 1 |
| 2018 | Redesigning LAMMPS for peta-scale and hundred-billion-atom simulation on Sunway TaihuLight
Xiaohui Duan, Ping Gao 0005, Tingjian Zhang, Wusheng Zhang, Wei Xue 0003, Haohuan Fu, Lin Gan 0001, Dexun Chen, Xiangxu Meng, Guangwen Yang 0002 |
SC | 3 |
| 2017 | 18.9-Pflops nonlinear earthquake simulation on Sunway TaihuLight: enabling depiction of 18-Hz and 8-meter scenariosabstractThis paper reports our large-scale nonlinear earthquake simulation software on Sunway TaihuLight. Our innovations include: (1) a customized parallelization scheme that employs the 10 million cores efficiently at both the process and the thread levels; (2) an elaborate memory scheme that integrates on-chip halo exchange through register communcation, optimized blocking configuration guided by an analytic model, and coalesced DMA access with array fusion; (3) on-the-fly compression that doubles the maximum problem size and further improves the performance by 24%. With these innovations to remove the memory constraints of Sunway TaihuLight, our software achieves over 15% of the system's peak, better than the 11.8% efficiency achieved by a similar software running on Titan, whose byte to flop ratio is 5 times better than TaihuLight. The extreme cases demonstrate a sustained performance of over 18.9 Pflops, enabling the simulation of Tangshan earthquake as an 18-Hz scenario with an 8-meter resolution. Haohuan Fu, Conghui He, Bingwei Chen, Zekun Yin, Tingjian Zhang, Wei Xue 0003, Wanwang Yin, Guangwen Yang 0002 |
SC | 7 |