EDBT 2026 Demo / reviewers in the wild / expert
Dan Lu 0001
dblp:36/7513-1
· DBLP profile ↗
8ranked-venue papers
0as first author
8since 2021 · last 2025
0000-0001-5162-9843ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 5 since 2021Systems, architecture and hardware · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Deep learning architectures and training · 39% Efficient and distributed learning · 30% Trustworthy machine learning · 22% | |
| Computer architecture, parallel and distributed computing, and storage systems
3 papers |
High-performance computing · 100% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Environmental and earth informatics · 100% | |
| Network and information security
1 paper |
Security and privacy of machine learning · 100% |
Topics — the 17 heaviest of 18, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Deep learning architectures and training › attention mechanism
efficient attention |
0.9 | 1 | 2025 | ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025 |
Machine learning › Efficient and distributed learning › distributed training › distributed training systems
large-scale distributed training |
0.9 | 1 | 2025 | ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025 |
Machine learning › Deep learning architectures and training › transformer
vision transformer |
0.9 | 1 | 2025 | ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025 |
High-performance computing › supercomputing
exascale computing |
0.9 | 1 | 2025 | ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025 |
Machine learning › Efficient and distributed learning › large-scale learning
large-scale model training |
0.8 | 1 | 2024 | ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024 |
Machine learning › Time series and sequential data › time series analysis
time series forecasting |
0.8 | 1 | 2024 | ExoTST: Exogenous-Aware Temporal Sequence Transformer for Time Series Prediction · ICDM 2024 |
Machine learning › Deep learning architectures and training
transformer |
0.8 | 1 | 2024 | ExoTST: Exogenous-Aware Temporal Sequence Transformer for Time Series Prediction · ICDM 2024 |
Machine learning › Deep learning architectures and training › transformer › vision transformer
vision transformer scaling |
0.8 | 1 | 2024 | ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024 |
Environmental and earth informatics
climate modeling |
0.8 | 1 | 2024 | ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024 |
Machine learning › Trustworthy machine learning › robustness
adversarial robustness |
0.6 | 1 | 2022 | Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack · ECCV (5) 2022 |
Machine learning › Trustworthy machine learning › uncertainty estimation
prediction intervals |
0.6 | 1 | 2022 | PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural Networks · ICLR 2022 |
Machine learning › Trustworthy machine learning
uncertainty estimation |
0.6 | 1 | 2022 | PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural Networks · ICLR 2022 |
Security and privacy of machine learning
adversarial attack |
0.6 | 1 | 2022 | Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack · ECCV (5) 2022 |
High-performance computing › large-scale training
distributed training on supercomputers |
0.5 | 2 | 2025 | Distributed Cross-Channel Hierarchical Aggregation for Foundation Models · SC 2025 ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024 |
Environmental and earth informatics › climate science
climate downscaling |
0.3 | 1 | 2025 | ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate Downscaling · SC 2025 |
High-performance computing
performance optimization at scale |
0.2 | 1 | 2024 | ORBIT: Oak Ridge Base Foundation Model for Earth System Predictability · SC 2024 |
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection |
0.2 | 1 | 2022 | PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural Networks · ICLR 2022 |
Methods — techniques the papers use, named apart from their topics
tile-wise sequence scaling · 2.6residual learning · 2.6bayesian regularization · 2.6hybrid tensor-data orthogonal parallelism · 2.3vision transformer · 1.7tensor parallelism · 1.7model sharding · 1.7black-box optimization · 1.1transformer · 0.8attention mechanism · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ORBIT-2: Scaling Exascale Vision Foundation Models for Weather and Climate DownscalingabstractSparse observations and coarse-resolution climate models limit effective regional decision-making, underscoring the need for robust downscaling. However, existing AI methods struggle with generalization across variables and geographies and are constrained by the quadratic complexity of Vision Transformer (ViT) self-attention. We introduce ORBIT-2, a scalable foundation model for global, hyper-resolution climate downscaling. ORBIT-2 incorporates two key innovations: (1) Residual Slim ViT (Reslim), a lightweight architecture with residual learning and Bayesian regularization for efficient, robust prediction; and (2) TILES, a tile-wise sequence scaling algorithm that reduces self-attention complexity from quadratic to linear, enabling long-sequence processing and massive parallelism. ORBIT-2 scales to 10 billion parameters across 65,536 GPUs, achieving up to 4.1 ExaFLOPS sustained throughput and 74–98% strong scaling efficiency. It supports downscaling to 0.9 km global resolution and processes sequences up to 4.2 billion tokens. On 7 km resolution benchmarks, ORBIT-2 achieves high accuracy with R2 scores in range of 0.98–0.99 against observation data. Xiao Wang 0004, Jong-Youl Choi, Takuya Kurihana, Isaac Lyngaas, Hong-Jun Yoon, Xi Xiao 0003, David Pugmire, Nasik Muhammad Nafi, Aristeidis Tsaris, Ashwin M. Aji, Maliha Hossain, Mohamed Wahib, Dali Wang, Peter E. Thornton, Prasanna Balaprakash, Moetasim Ashfaq, Dan Lu 0001 |
SC | 18 |
| 2025 | Distributed Cross-Channel Hierarchical Aggregation for Foundation ModelsabstractVision-based scientific foundation models hold significant promise for advancing scientific discovery and innovation. This potential stems from their ability to aggregate images from diverse sources—such as varying physical groundings or data acquisition systems—and to learn spatio-temporal correlations using transformer architectures. However, tokenizing and aggregating images can be compute-intensive, a challenge not fully addressed by current distributed methods. In this work, we introduce the Distributed Cross-Channel Hierarchical Aggregation (D-CHAG) approach designed for datasets with a large number of channels across image modalities. Our method is compatible with any model-parallel strategy and any type of vision transformer architecture, significantly improving computational efficiency. We evaluated D-CHAG on hyperspectral imaging and weather forecasting tasks. When integrated with tensor parallelism and model sharding, our approach achieved up to a 75% reduction in memory usage and more than doubled sustained throughput on up to 1,024 AMD GPUs on the Frontier Supercomputer. Aristeidis Tsaris, Isaac Lyngaas, John H. Lagergren, Mohamed Wahib, Larry M. York, Prasanna Balaprakash, Dan Lu 0001, Feiyi Wang, Xiao Wang 0004 |
SC | 7 |
| 2024 | ExoTST: Exogenous-Aware Temporal Sequence Transformer for Time Series PredictionabstractAccurate long-term predictions are the foundations for many machine learning applications and decision-making processes. Traditional time series approaches for prediction often focus on either autoregressive modeling, which relies solely on past observations of the target “endogenous variables”, or forward modeling, which considers only current covariate drivers “exogenous variables”. However, effectively integrating past endogenous and past exogenous with current exogenous variables remains a significant challenge. In this paper, we propose ExoTST, a novel transformer-based framework that effectively incorporates current exogenous variables alongside past context for improved time series prediction. To integrate exogenous information efficiently, ExoTST leverages the strengths of attention mechanisms and introduces a novel cross-temporal modality fusion module. This module enables the model to jointly learn from both past and current exogenous series, treating them as distinct modalities. By considering these series separately, ExoTST provides robustness and flexibility in handling data uncertainties that arise from the inherent distribution shift between historical and current exogenous variables. Extensive experiments on real-world carbon flux datasets and time series benchmarks demonstrate ExoTST's superior performance compared to state-of-the-art baselines, with improvements of up to 10% in prediction accuracy. Moreover, ExoTST exhibits strong robustness against missing values and noise in exogenous drivers, maintaining consistent performance in real-world situations where these imperfections are common. Kshitij Tayal, Arvind Renganathan, Xiaowei Jia, Vipin Kumar 0001, Dan Lu 0001 |
ICDM | 5 |
| 2024 | ORBIT: Oak Ridge Base Foundation Model for Earth System PredictabilityabstractEarth system predictability is challenged by the complexity of environmental dynamics and the multitude of variables involved. Current AI foundation models, although advanced by leveraging large and heterogeneous data, are often constrained by their size and data integration, limiting their effectiveness in addressing the full range of Earth system prediction challenges. To overcome these limitations, we introduce the Oak Ridge Base Foundation Model for Earth System Predictability (ORBIT), an advanced vision transformer model that scales up to 113 billion parameters using a novel hybrid tensor-data orthogonal parallelism technique. As the largest model of its kind, ORBIT surpasses the current climate AI foundation model size by a thousandfold. Performance scaling tests conducted on the Frontier supercomputer have demonstrated that ORBIT achieves 684 petaFLOPS to 1.6 exaFLOPS sustained throughput, with scaling efficiency maintained at 41% to 85% across 49,152 AMD GPUs. These breakthroughs establish new advances in AIdriven climate modeling and demonstrate promise to significantly improve the Earth system predictability. Xiao Wang 0004, Aristeidis Tsaris, Jong-Youl Choi, Ashwin M. Aji, Wei Zhang 0261, Junqi Yin, Moetasim Ashfaq, Dan Lu 0001, Prasanna Balaprakash |
SC | 10 |
| 2023 | Accelerating Scientific Simulations with Bi-Fidelity Weighted Transfer LearningabstractHigh-fidelity modeling is an essential design tool for many engineering applications. However, for complex systems, computational cost can be a limiting factor. Analyzing parameter sensitivity, uncertainty quantification, and design optimization require many model evaluations. Surrogate models are often used to develop the relationship between model parameters and quantities of interest. However, in the case of complex systems, surrogate models require several degrees of freedom and, thus, a large number of data points to determine the correct dependencies. For many applications, this may be prohibitively expensive. The reduction of computational requirements can be achieved by leveraging low-fidelity models. Low-fidelity models represent the system at a coarser resolution with the advantage of computational efficiency. Therefore, a bi-fidelity modeling paradigm, which augments the accuracy of a low-fidelity model in a computationally efficient manner by invoking limited runs of a high-fidelity model, can be leveraged to sufficiently balance the accuracy and computational requirements. In this work, a bi-fidelity weighted transfer learning method using neural networks was applied to a computational fluid dynamics heat transfer modeling problem. The transfer learning advantage was investigated as a function of hyperparameters. Our main finding is that the use of a bi-fidelity modeling paradigm achieves accuracy close to that of a high-fidelity Gaussian process model while significantly reducing computational cost. The bi-fidelity model achieves comparable performance with 90 high-fidelity samples-that is, 60% less than the samples needed to achieve similar accuracy without the use of bi-fidelity modeling, Katarzyna Borowiec, Dan Lu 0001, Vikas Chandan, Samrat Chatterjee, Pradeep Ramuhalli, Ramakrishna Tipireddy, Mahantesh Halappanavar, Frank Liu 0001 |
ICMLA | 2 |
| 2022 | Exploiting the Local Parabolic Landscapes of Adversarial Losses to Accelerate Black-Box Adversarial Attack
Hoang Tran, Dan Lu 0001 |
ECCV (5) | 2 |
| 2022 | PI3NN: Out-of-distribution-aware Prediction Intervals from Three Neural Networks
Dan Lu 0001 |
ICLR | 3 |
| 2021 | Enabling long-range exploration in minimization of multimodal functionsabstractWe consider the problem of minimizing multi-modal loss functions with a large number of local optima. Since the local gradient points to the direction of the steepest slope in an infinitesimal neighborhood, an optimizer guided by the local gradient is often trapped in a local minimum. To address this issue, we develop a novel nonlocal gradient to skip small local minima by capturing major structures of the loss’s landscape in black-box optimization. The nonlocal gradient is defined by a directional Gaussian smoothing (DGS) approach. The key idea of DGS is to conducts 1D long-range exploration with a large smoothing radius along $d$ orthogonal directions in $R^d$, each of which defines a nonlocal directional derivative as a 1D integral. Such long-range exploration enables the nonlocal gradient to skip small local minima. The $d$ directional derivatives are then assembled to form the nonlocal gradient. We use the Gauss-Hermite quadrature rule to approximate the $d$ 1D integrals to obtain an accurate estimator. The superior performance of our method is demonstrated in three sets of examples, including benchmark functions for global optimization, and two real-world scientific problems. Jiaxin Zhang 0005, Hoang Tran, Dan Lu 0001 |
UAI | 3 |