Stencil-Aware GPU Optimization of Iterative Solvers

Daniel Lowell,Jeswin Godwin,Justin Holewinski,Deepan Karthik,Chekuri Choudary,Azamat Mametjanov,Boyana Norris,Gerald Sabin,P. Sadayappan,Jason Sarich

Stencil-Aware GPU Optimization of Iterative Solvers

2013

Daniel Lowell
Jeswin Godwin
Justin Holewinski
Deepan Karthik
Chekuri Choudary
Azamat Mametjanov
Boyana Norris
Gerald Sabin
P. Sadayappan
Jason Sarich

Numerical solutions of nonlinear partial differential equations frequently rely on iterative Newton--Krylov methods, which linearize a finite-difference stencil-based discretization of a problem, producing a sparse matrix with regular structure. Knowledge of this structure can be used to exploit parallelism and locality of reference on modern cache-based multi and manycore architectures, achieving high performance for computations underlying commonly used iterative linear solvers. In this paper we describe our approach to sparse matrix data structure design and our implementation of the kernels underlying iterative linear solvers in PETSc. We also describe autotuning of CUDA implementations based on high-level descriptions of the stencil-based matrix and vector operations.

Keywords:

Locality of reference
Mathematical optimization
Data structure
Stencil
Discretization
Sparse matrix
Matrix (mathematics)
General-purpose computing on graphics processing units
Nonlinear system
Theoretical computer science
Computer science
CUDA
Parallel computing
Computational science

Correction
Source
Cite
Save
Machine Reading By IdeaReader

References

Citations