Dan Lu 0001

dblp:36/7513-1 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-5162-9843ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Deep learning architectures and training · 39% Efficient and distributed learning · 30% Trustworthy machine learning · 22%
Computer architecture, parallel and distributed computing, and storage systems
3 papers
High-performance computing · 100%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Environmental and earth informatics · 100%
Network and information security
1 paper
Security and privacy of machine learning · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention
0.912025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
large-scale distributed training
0.912025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
Machine learning › Deep learning architectures and training › transformer
vision transformer
0.912025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
High-performance computing › supercomputing
exascale computing
0.912025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training
0.812024
ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024
Machine learning › Time series and sequential data › time series analysis
time series forecasting
0.812024
ExoTST: Exogenous-Aware Temporal Sequence Transformer for Time Series Prediction · ICDM 2024
Machine learning › Deep learning architectures and training
transformer
0.812024
ExoTST: Exogenous-Aware Temporal Sequence Transformer for Time Series Prediction · ICDM 2024
Machine learning › Deep learning architectures and training › transformer › vision transformer
vision transformer scaling
0.812024
ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024
Environmental and earth informatics
climate modeling
0.812024
ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024
Machine learning › Trustworthy machine learning › robustness
adversarial robustness
0.612022
Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack · ECCV (5) 2022
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction intervals
0.612022
PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural Networks · ICLR 2022
Machine learning › Trustworthy machine learning
uncertainty estimation
0.612022
PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural Networks · ICLR 2022
Security and privacy of machine learning
adversarial attack
0.612022
Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack · ECCV (5) 2022
High-performance computing › large-scale training
distributed training on supercomputers
0.522025
Distributed Cross-Channel Hierarchical Aggregation for Foundation Models · SC 2025
ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024
Environmental and earth informatics › climate science
climate downscaling
0.312025
ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025
High-performance computing
performance optimization at scale
0.212024
ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.212022
PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural Networks · ICLR 2022

Methods — techniques the papers use, named apart from their topics

tile-wise sequence scaling · 2.6residual learning · 2.6bayesian regularization · 2.6hybrid tensor-data orthogonal parallelism · 2.3vision transformer · 1.7tensor parallelism · 1.7model sharding · 1.7black-box optimization · 1.1transformer · 0.8attention mechanism · 0.8
YearPublicationVenuePosition
2025 ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling
abstract
Sparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data.
Xiao Wang 0004, Jong-Youl Choi, Takuya Kurihana, Isaac Lyngaas, Hong-Jun Yoon, Xi Xiao 0003, David Pugmire, Nasik Muhammad Nafi, Aristeidis Tsaris, Ashwin M. Aji, Maliha Hossain, Mohamed Wahib, Dali Wang, Peter E. Thornton, Prasanna Balaprakash, Moetasim Ashfaq, Dan Lu 0001
SC18
2025 Distributed Cross-Channel Hierarchical Aggregation for Foundation Models
abstract
Vision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources—such as varying physical groundings or data acquisition systems—and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer.
Aristeidis Tsaris, Isaac Lyngaas, John H. Lagergren, Mohamed Wahib, Larry M. York, Prasanna Balaprakash, Dan Lu 0001, Feiyi Wang, Xiao Wang 0004
SC7
2024 ExoTST: Exogenous-Aware Temporal Sequence Transformer for Time Series Prediction
abstract
Accurate long-term predictions are the foundations for many machine learning applications and decision-making processes. Traditional time series approaches for prediction often focus on either autoregressive modeling, which relies solely on past observations of the target “endogenous variables”, or forward modeling, which considers only current covariate drivers “exogenous variables”. However, effectively integrating past endogenous and past exogenous with current exogenous variables remains a significant challenge. In this paper, we propose ExoTST, a novel transformer-based framework that effectively incorporates current exogenous variables alongside past context for improved time series prediction. To integrate exogenous information efficiently, ExoTST leverages the strengths of attention mechanisms and introduces a novel cross-temporal modality fusion module. This module enables the model to jointly learn from both past and current exogenous series, treating them as distinct modalities. By considering these series separately, ExoTST provides robustness and flexibility in handling data uncertainties that arise from the inherent distribution shift between historical and current exogenous variables. Extensive experiments on real-world carbon flux datasets and time series benchmarks demonstrate ExoTST's superior performance compared to state-of-the-art baselines, with improvements of up to 10% in prediction accuracy. Moreover, ExoTST exhibits strong robustness against missing values and noise in exogenous drivers, maintaining consistent performance in real-world situations where these imperfections are common.
Kshitij Tayal, Arvind Renganathan, Xiaowei Jia, Vipin Kumar 0001, Dan Lu 0001
ICDM5
2024 ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability
abstract
Earth system predictability is challenged by the complexity of environmental dynamics and the multitude of variables involved. Current AI foundation models, although advanced by leveraging large and heterogeneous data, are often constrained by their size and data integration, limiting their effectiveness in addressing the full range of Earth system prediction challenges. To overcome these limitations, we introduce the Oak Ridge Base Foundation Model for Earth System Predictability (ORBIT), an advanced vision transformer model that scales up to 113 billion parameters using a novel hybrid tensor-data orthogonal parallelism technique. As the largest model of its kind, ORBIT surpasses the current climate AI foundation model size by a thousandfold. Performance scaling tests conducted on the Frontier supercomputer have demonstrated that ORBIT achieves 684 petaFLOPS to 1.6 exaFLOPS sustained throughput, with scaling efficiency maintained at 41% to 85% across 49,152 AMD GPUs. These breakthroughs establish new advances in AIdriven climate modeling and demonstrate promise to significantly improve the Earth system predictability.
Xiao Wang 0004, Aristeidis Tsaris, Jong-Youl Choi, Ashwin M. Aji, Wei Zhang 0261, Junqi Yin, Moetasim Ashfaq, Dan Lu 0001, Prasanna Balaprakash
SC10
2023 Accelerating Scientific Simulations with Bi-Fidelity Weighted Transfer Learning
abstract
High-fidelity modeling is an essential design tool for many engineering applications. However, for complex systems, computational cost can be a limiting factor. Analyzing parameter sensitivity, uncertainty quantification, and design optimization require many model evaluations. Surrogate models are often used to develop the relationship between model parameters and quantities of interest. However, in the case of complex systems, surrogate models require several degrees of freedom and, thus, a large number of data points to determine the correct dependencies. For many applications, this may be prohibitively expensive. The reduction of computational requirements can be achieved by leveraging low-fidelity models. Low-fidelity models represent the system at a coarser resolution with the advantage of computational efficiency. Therefore, a bi-fidelity modeling paradigm, which augments the accuracy of a low-fidelity model in a computationally efficient manner by invoking limited runs of a high-fidelity model, can be leveraged to sufficiently balance the accuracy and computational requirements. In this work, a bi-fidelity weighted transfer learning method using neural networks was applied to a computational fluid dynamics heat transfer modeling problem. The transfer learning advantage was investigated as a function of hyperparameters. Our main finding is that the use of a bi-fidelity modeling paradigm achieves accuracy close to that of a high-fidelity Gaussian process model while significantly reducing computational cost. The bi-fidelity model achieves comparable performance with 90 high-fidelity samples-that is, 60% less than the samples needed to achieve similar accuracy without the use of bi-fidelity modeling,
Katarzyna Borowiec, Dan Lu 0001, Vikas Chandan, Samrat Chatterjee, Pradeep Ramuhalli, Ramakrishna Tipireddy, Mahantesh Halappanavar, Frank Liu 0001
ICMLA2
2022 Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack
Hoang Tran, Dan Lu 0001
ECCV (5)2
2022 PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural Networks
Dan Lu 0001
ICLR3
2021 Enabling long-range exploration in minimization of multimodal functions
abstract
We consider the problem of minimizing multi-modal loss functions with a large number of local optima. Since the local gradient points to the direction of the steepest slope in an infinitesimal neighborhood, an optimizer guided by the local gradient is often trapped in a local minimum. To address this issue, we develop a novel nonlocal gradient to skip small local minima by capturing major structures of the loss’s landscape in black-box optimization. The nonlocal gradient is defined by a directional Gaussian smoothing (DGS) approach. The key idea of DGS is to conducts 1D long-range exploration with a large smoothing radius along $d$ orthogonal directions in $R^d$, each of which defines a nonlocal directional derivative as a 1D integral. Such long-range exploration enables the nonlocal gradient to skip small local minima. The $d$ directional derivatives are then assembled to form the nonlocal gradient. We use the Gauss-Hermite quadrature rule to approximate the $d$ 1D integrals to obtain an accurate estimator. The superior performance of our method is demonstrated in three sets of examples, including benchmark functions for global optimization, and two real-world scientific problems.
Jiaxin Zhang 0005, Hoang Tran, Dan Lu 0001
UAI3