Thomas George

dblp:69/641 · DBLP profile ↗
← Back
19ranked-venue papers
6as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-authorArtificial intelligence and machine learning · 6 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 first-author · 1 since 2021Software engineering, systems software and programming languages · 1Theory of computation · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer architecture, parallel and distributed computing, and storage systems
4 papers
High-performance computing · 43% Interconnection networks and networks-on-chip · 28% Parallel and multicore computing · 28%
Artificial intelligence
1 paper
Optimization for machine learning · 100%
Databases, data mining, and information retrieval
2 papers
Data mining · 53% Recommender systems · 35% Machine learning and data management · 12%

Topics — the 19 heaviest of 19, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Optimization for machine learning › second-order optimization
kronecker-factored approximate curvature
0.312018
Fast Approximate Natural Gradient Descent in a Kronecker Factored Eigenbasis · NeurIPS 2018
Machine learning › Optimization for machine learning › gradient-based optimization › gradient descent
natural gradient descent
0.312018
Fast Approximate Natural Gradient Descent in a Kronecker Factored Eigenbasis · NeurIPS 2018
High-performance computing
scientific computing systems
0.222012
A divide and conquer strategy for scaling weather simulations with multiple regions of interest · SC 2012
Sparse matrix factorization on massively parallel computers · SC 2009
High-performance computing
collective communication
0.112012
Collective algorithms for sub-communicators · PPoPP 2012
Interconnection networks and networks-on-chip
network topology
0.112012
A divide and conquer strategy for scaling weather simulations with multiple regions of interest · SC 2012
Parallel and multicore computing
parallel programming models and scheduling
0.112012
A divide and conquer strategy for scaling weather simulations with multiple regions of interest · SC 2012
Parallel and multicore computing
processor allocation
0.112012
A divide and conquer strategy for scaling weather simulations with multiple regions of interest · SC 2012
Interconnection networks and networks-on-chip › network topology
topology-aware mapping
0.112012
A divide and conquer strategy for scaling weather simulations with multiple regions of interest · SC 2012
Interconnection networks and networks-on-chip › network topology
torus network
0.112012
Collective algorithms for sub-communicators · PPoPP 2012
Parallel and multicore computing
parallel algorithms
0.112009
Sparse matrix factorization on massively parallel computers · SC 2009
High-performance computing
performance optimization at scale
0.112009
Sparse matrix factorization on massively parallel computers · SC 2009
High-performance computing › sparse linear algebra
sparse matrix factorization
0.112009
Sparse matrix factorization on massively parallel computers · SC 2009
High-performance computing
numerical linear algebra
0.112008
A Recommendation System for Preconditioned Iterative Solvers · ICDM 2008
Data mining
clustering
0.112005
A Scalable Collaborative Filtering Framework Based on Co-Clustering · ICDM 2005
Data mining › clustering
co-clustering
0.112005
A Scalable Collaborative Filtering Framework Based on Co-Clustering · ICDM 2005
Recommender systems
collaborative filtering
0.112005
A Scalable Collaborative Filtering Framework Based on Co-Clustering · ICDM 2005
Parallel and multicore computing › parallel programming models
message passing
0.012012
Collective algorithms for sub-communicators · PPoPP 2012
Machine learning and data management
machine learning for systems
0.012008
A Recommendation System for Preconditioned Iterative Solvers · ICDM 2008
Recommender systems › online recommendation
real-time recommendation
0.012005
A Scalable Collaborative Filtering Framework Based on Co-Clustering · ICDM 2005

Methods — techniques the papers use, named apart from their topics

kronecker factorization · 0.3eigenvalue decomposition · 0.3performance data regression · 0.2multi-stage learning · 0.2torus interconnect mapping · 0.1performance prediction · 0.1collective algorithm design · 0.1weighted co-clustering · 0.1matrix factorization · 0.1
YearPublicationVenuePosition
2026 Calibration Improves Detection of Mislabeled Examples
Ilies Chibane, Thomas George, Pierre Nodet, Vincent Lemaire 0001
DaWaK2
2022 Optimal tuning of FOPID controller for higher order process using hybrid approach
Thomas George, V. Ganesan
Appl. Intell.1
2021 Implicit Regularization via Neural Feature Alignment
abstract
We approach the problem of implicit regularization in deep learning from a geometrical viewpoint. We highlight a regularization effect induced by a dynamical alignment ofthe neural tangent features introduced by Jacot et al. (2018), along a small number of task-relevant directions. This can be interpreted as a combined mechanism of feature selection and compression. By extrapolating a new analysis of Rademacher complexity bounds for linear models, we motivate and study a heuristic complexity measure that captures this phenomenon, in terms of sequences of tangent kernel classes along optimization paths. The code for our experiments is available as https://github.com/tfjgeorge/ntk_alignment.
Aristide Baratin, Thomas George, César Laurent, R. Devon Hjelm, Guillaume Lajoie, Pascal Vincent, Simon Lacoste-Julien
AISTATS2
2019 HyTI: Thermal Hyperspectral Imaging from A Cubesat Platform
abstract
The HyTI (Hyperspectral Thermal Imager) mission, funded by NASA's Earth Science Technology Office InVEST (In-Space Validation of Earth Science Technologies) program, will demonstrate how high spectral and spatial long-wave infrared image data can be acquired from a 6U CubeSat platform. The mission will use a spatially modulated interferometric imaging technique to produce spectro-radiometrically calibrated image cubes, with 25 channels between 8-10.7 μm, at a ground sample distance of ~70 m. The HyTI performance model indicates narrow band NEΔTs of <; 0.3 K. The small form factor of HyTI is made possible via the use of a no-moving-parts Fabry-Perot interferometer, and JPL's cryogenically-cooled BIRD FPA technology. Launch is scheduled for no earlier than October 2020. The value of HyTI to Earth scientists will be demonstrated via on-board processing of the raw instrument data to generate L1 and L2 products, with a focus on rapid delivery of precision agriculture metrics.
Robert Wright, Chiara Ferrari-Wong, Abigail Flom, John Mecikalski, Prasad Thenkabail, Miguel Nunes, Paul Lucey, Luke Flynn, Thomas George, Sarath D. Gunapala, David Z. Ting, Sir Rafol, Alexander Soibel
IGARSS10
2018 Fast Approximate Natural Gradient Descent in a Kronecker Factored Eigenbasis
abstract
Optimization algorithms that leverage gradient covariance information, such as variants of natural gradient descent (Amari, 1998), offer the prospect of yielding more effective descent directions. For models with many parameters, the covari- ance matrix they are based on becomes gigantic, making them inapplicable in their original form. This has motivated research into both simple diagonal approxima- tions and more sophisticated factored approximations such as KFAC (Heskes, 2000; Martens & Grosse, 2015; Grosse & Martens, 2016). In the present work we draw inspiration from both to propose a novel approximation that is provably better than KFAC and amendable to cheap partial updates. It consists in tracking a diagonal variance, not in parameter coordinates, but in a Kronecker-factored eigenbasis, in which the diagonal approximation is likely to be more effective. Experiments show improvements over KFAC in optimization speed for several deep network architectures.
Thomas George, César Laurent, Xavier Bouthillier, Nicolas Ballas, Pascal Vincent
NeurIPS1
2015 Architecture Aware Resource Allocation for Structured Grid Applications: Flood Modelling Case
abstract
Numerous problems in science and engineering involve discretizing the problem domain as a regular structured grid and make use of domain decomposition techniques to obtain solutions faster using high performance computing. However, the load imbalance of the workloads among the various processing nodes can cause severe degradation in application performance. This problem is exacerbated for the case when the computational workload is non-uniform and the processing nodes have varying computational capabilities. In this paper, we present novel local search algorithms for regular partitioning of a structured mesh to heterogeneous compute nodes in a distributed setting. The algorithms seek to assign larger workloads to processing nodes having higher computation capabilities while maintaining the regular structure of the mesh in order to achieve a better load balance. We also propose a distributed memory (MPI) parallelization architecture that can be used to achieve a parallel implementation of scientific modelling software requiring structured grids on heterogeneous processing resources involving CPUs and GPUs. Our implementation can make use of the available CPU cores and multiple GPUs of the underlying platform simultaneously. Empirical evaluation on real world flood modelling domains on a heterogeneous architecture comprising of multicore CPUs and GPUs suggests that the proposed partitioning approach can provide a performance improvement of up to 8× over a naive uniform partitioning.
Vaibhav Saxena, Thomas George, Yogish Sabharwal, Lucas Correia Villa Real
CCGRID2
2014 IFM: A Scalable High Resolution Flood Modeling Framework
Swati Singhal, Sandhya Aneja, Frank Liu 0001, Lucas Correia Villa Real, Thomas George
Euro-Par5
2013 Evaluation and enhancement of weather application performance on Blue Gene/Q
abstract
Numerical weather prediction (NWP) models use mathematical models of the atmosphere to predict the weather. Ongoing efforts in the weather and climate community continuously try to improve the fidelity of weather models by employing higher order numerical methods suitable for solving model equations at high resolutions. In realistic weather forecasting scenario, simulating and tracking multiple regions of interest (nests) at fine resolutions is important in understanding the interplay between multiple weather phenomena and for comprehensive predictions. These multiple regions of interest in a simulation can be significantly different in resolution and other modeling parameters. Currently, the weather simulations involving these nested regions process them one after the other in a sequential fashion. There exists a lot of prior work in performance evaluation and optimization of weather models, however most of this work is either limited to simulations involving a single domain or multiple nests with same resolution and model parameters such as model physics options. In this paper, we evaluate and enhance the performance of popular WRF model on IBM Blue Gene/Q system. We consider nested simulations with multiple child domains and study how parameters such as physics options and simulation time steps for child domains affect the computational requirements. We also analyze how such configurations can benefit from parallel execution of the children domains rather than processing them sequentially. We demonstrate that it is important to allocate processors to nested child domains in proportion to the work load associated with them when executing them in parallel. This ensures that the time spent in the different nested simulations is nearly equal, and the nested domains reach the synchronization step with the parent simulation together. Our experimental evaluation using a simple heuristic for allocation of nodes shows that the performance of WRF simulations can be improved by up to 14% by parallel execution of sibling domains with different configuration of domain sizes, temporal resolutions and physics options.
Gurbinder Gill, Vaibhav Saxena, Rashmi Mittal, Thomas George, Yogish Sabharwal, Lalit Dagar
HiPC4
2013 A hybrid parallelization approach for high resolution operational flood forecasting
abstract
Accurate and timely flood forecasts are becoming highly essential due to the increased incidence of flood related disasters over the last few years. Such forecasts require a high resolution integrated flood modeling approach. In this paper, we present an integrated flood forecasting system with an automated workflow over the weather modeling, surface runoff estimation and water routing components. We primarily focus on the water routing process which is the most compute intensive phase and present two parallelization strategies to scale it up to large grid sizes. Specifically, we employ nature-inspired decomposition of a simulation domain into watershed basins and propose a master slave model of parallelization for distributed processing of the basins. We also propose an intra-basin shared memory parallelization approach using OpenMP. Empirical evaluation of the proposed parallelization strategies indicates a potential for high speedups for certain types of scenarios (e.g., speedup of 13× with 16 threads using OpenMP parallelization for the large Rio de Janeiro basin).
Swati Singhal, Lucas Correia Villa Real, Thomas George, Sandhya Aneja, Yogish Sabharwal
HiPC3
2012 Performance Evaluation and Optimization of Nested High Resolution Weather Simulations
Preeti Malakar, Vaibhav Saxena, Thomas George, Rashmi Mittal, Sameer Kumar 0001, Abdul Ghani Naim, Saiful Azmi bin Hj Husain
Euro-Par3
2012 Collective algorithms for sub-communicators
abstract
Collective communication over a group of processors is an integral and time consuming component in many high performance computing applications. Many modern day super- computers are based on torus interconnects and near optimal algorithms have been developed for collective communication over regular communicators on these systems. However, for an irregular communicator comprising of a subset of processors, the algorithms developed so far are not contention free in general and hence non-optimal. In this paper, we present a novel contention-free algorithm to perform collective operations over a subset of processors in a torus network. We also extend previous work on regular communicators to handle special cases of irregular communicators that occur frequently in parallel scientific applications. For the generic case where multiple node disjoint sub-communicators communicate simultaneously in a loosely synchronous fashion, we propose a novel cooperative approach to route the data for individual sub- communicators without contention. Empirical results demon- strate that our algorithms outperform the optimized MPI collective implementation on IBM's Blue Gene/P supercomputer for large data sizes and random node distributions.
Anshul Mittal, Thomas George, Yogish Sabharwal, Sameer Kumar 0001
ICS3
2012 Collective algorithms for sub-communicators
abstract
Collective communication over a group of processors is an integral and time consuming component in many HPC applications. Many modern day supercomputers are based on torus interconnects. On such systems, for an irregular communicator comprising of a subset of processors, the algorithms developed so far are not contention free in general and hence non-optimal.
Anshul Mittal, Thomas George, Yogish Sabharwal, Sameer Kumar 0001
PPoPP3
2012 A divide and conquer strategy for scaling weather simulations with multiple regions of interest
abstract
Accurate and timely prediction of weather phenomena, such as hurricanes and flash floods, require high-fidelity compute intensive simulations of multiple finer regions of interest within a coarse simulation domain. Current weather applications execute these nested simulations sequentially using all the available processors, which is sub-optimal due to their sub-linear scalability. In this work, we present a strategy for parallel execution of multiple nested domain simulations based on partitioning the 2-D processor grid into disjoint rectangular regions associated with each domain. We propose a novel combination of performance prediction, processor allocation methods and topology-aware mapping of the regions on torus interconnects. Experiments on IBM Blue Gene systems using WRF show that the proposed strategies result in performance improvement of up to 33% with topology-oblivious mapping and up to additional 7% with topology-aware mapping over the default sequential strategy.
Preeti Malakar, Thomas George, Sameer Kumar 0001, Rashmi Mittal, Vijay Natarajan, Yogish Sabharwal, Vaibhav Saxena, Sathish S. Vadhiyar
SC2
2012 An Empirical Analysis of the Performance of Preconditioners for SPD Systems
abstract
Preconditioned iterative solvers have the potential to solve very large sparse linear systems with a fraction of the memory used by direct methods. However, the effectiveness and performance of most preconditioners is not only problem dependent, but also fairly sensitive to the choice of their tunable parameters. As a result, a typical practitioner is faced with an overwhelming number of choices of solvers, preconditioners, and their parameters. The diversity of preconditioners makes it difficult to analyze them in a unified theoretical model. A systematic empirical evaluation of existing preconditioned iterative solvers can help in identifying the relative advantages of various implementations. We present the results of a comprehensive experimental study of the most popular preconditioner and iterative solver combinations for symmetric positive-definite systems. We introduce a methodology for a rigorous comparative evaluation of various preconditioners, including the use of some simple but powerful metrics. The detailed comparison of various preconditioner implementations and a state-of-the-art direct solver gives interesting insights into their relative strengths and weaknesses. We believe that these results would be useful to researchers developing preconditioners and iterative solvers as well as practitioners looking for appropriate sparse solvers for their applications.
Thomas George, Vivek Sarin
ACM Trans. Math. Softw.1
2011 Multifrontal Factorization of Sparse SPD Matrices on GPUs
abstract
Solving large sparse linear systems is often the most computationally intensive component of many scientific computing applications. In the past, sparse multifrontal direct factorization has been shown to scale to thousands of processors on dedicated supercomputers resulting in a substantial reduction in computational time. In recent years, an alternative computing paradigm based on GPUs has gained prominence, primarily due to its affordability, power-efficiency, and the potential to achieve significant speedup relative to desktop performance on regular and structured parallel applications. However, sparse matrix factorization on GPUs has not been explored sufficiently due to the complexity involved in an efficient implementation and concerns of low GPU utilization. In this paper, we present an adaptive hybrid approach for accelerating sparse multifrontal factorization based on a judicious exploitation of the processing power of the host CPU and GPU. We present four different policies for distributing and scheduling the workload between the host CPU and the GPU, and propose a mechanism for a runtime selection of the appropriate policy for each step of sparse Cholesky factorization. This mechanism relies on auto-tuning based on modeling the best policy predictor as a parametric classifier. We estimate the classifier parameters from the available empirical computation time data such that the expected computation time is minimized. This approach is readily adaptable for using the current or an extended set of policies for different CPU-GPU combinations as well as for different combinations of dense kernels for both the CPU and the GPU.
Thomas George, Vaibhav Saxena, Amik Singh, Anamitra R. Choudhury
IPDPS1
2009 Sparse matrix factorization on massively parallel computers
abstract
Direct methods for solving sparse systems of linear equations have a high asymptotic computational and memory requirements relative to iterative methods. However, systems arising in some applications, such as structural analysis, can often be too ill-conditioned for iterative solvers to be effective. We cite real applications where this is indeed the case, and using matrices extracted from these applications to conduct experiments on three different massively parallel architectures, show that a well designed sparse factorization algorithm can attain very high levels of performance and scalability. We present strong scalability results for test data from real applications on up to 8,192 cores, along with both analytical and experimental weak scalability results for a model problem on up to 16,384 cores---an unprecedented number for sparse factorization. For the model problem, we also compare experimental results with multiple analytical scaling metrics and distinguish between some commonly used weak scaling methods.
Seid Koric, Thomas George
SC3
2008 A Recommendation System for Preconditioned Iterative Solvers
abstract
Preconditioned iterative methods are often used to solve very large sparse systems of linear systems that arise in many scientific and engineering applications. The performance and robustness of these solvers is extremely sensitive to the choice of multiple preconditioner and solver parameters. Users of iterative methods often encounter an overwhelming number of combinations of choices for solvers, matrix preprocessing steps, preconditioners, and their parameters. The lack of a unified theoretical analysis of preconditioners coupled with limited knowledge of their interaction with linear systems makes it highly challenging for practitioners to choose good solver configurations. In this paper, we propose a novel, multi-stage learning based methodology for determining the best solver configurations to optimize the desired performance behavior for any given linear system. Empirical results over real performance data for the hyper iterative solver package demonstrate the efficacy and flexibility of the proposed approach.
Thomas George, Vivek Sarin
ICDM1
2005 A Scalable Collaborative Filtering Framework Based on Co-Clustering
abstract
Collaborative filtering-based recommender systems have become extremely popular due to the increase in Web-based activities such as e-commerce and online content distribution. Current collaborative filtering (CF) techniques such as correlation and SVD based methods provide good accuracy, but are computationally expensive and can be deployed only in static off-line settings. However, a number of practical scenarios require dynamic real-time collaborative filtering that can allow new users, items and ratings to enter the system at a rapid rate. In this paper, we consider a novel CF approach based on a proposed weighted co-clustering algorithm (Banerjee et al., 2004) that involves simultaneous clustering of users and items. We design incremental and parallel versions of the co-clustering algorithm and use it to build an efficient real-time CF framework. Empirical evaluation demonstrates that our approach provides an accuracy comparable to that of the correlation and matrix factorization based approaches at a much lower computational cost.
Thomas George, Srujana Merugu
ICDM1
2005 Loci: a rule-based framework for parallel multi-disciplinary simulation synthesis
abstract
We present a rule-based framework for the development of scalable parallel high performance simulations for a broad class of scientific applications (with particular emphasis on continuum mechanics). We take a pragmatic approach to our programming abstractions by implementing structures that are used frequently and have common high performance implementations on distributed memory architectures. The resulting framework borrows heavily from rule-based systems for relational database models, however limiting the scope to those parts that have obvious high performance implementation. Using our approach, we demonstrate predictable performance behavior and efficient utilization of large scale distributed memory architectures on problems of significant complexity involving multiple disciplines.
Edward A. Luke, Thomas George
J. Funct. Program.2