| marp | false |
|---|---|
| theme | cscs |
| paginate | true |
| backgroundColor | |
| backgroundImage | url('slides-support/common/4k-slide-bg-white.png') |
| size | 58140 |
- ICON grid is partly structured
- Investigate how expensive are indirect neighbor accesses compared to strided indexes of neighbors
- Investigate optimizations based on strided indexes
- Apply optimizations to current and future
GT4Pydevelopment
netcdffile with neighbor lists- Use of structured torus grid
- Generated using grid-generator from
DKRZby David Strassman - Read using icon4py GridManager
- Generated using grid-generator from
- Torus grid has periodic boundaries
- Adds complexity to compute strided indexes
- Not applicable in production scenario of CH
- Filter torus grid to end up with a cartesian grid with 2 level halo
- 1 would also be sufficient
- Discard periodic edges
- Calculate
XandYdimensions based on the distribution of vertices in spacex_dimequals the number of vertices withy == 0
- Fix
e2vordering of edges read by torus grid file to be by defaultper-vertexin GridManager- Some elements in
e2vdidn't follow this ordering by default
- Some elements in
- Select between
per-vertexandper-orientationordering in GridManager - Filter halo vertices and edges in benchmark python script
- Nabla4
- Neighbor tables
- e2c2v[4]
- e2ecv[4]
- Input fields
- u_vert[VertexDim, KDim]
- v_vert[VertexDim, KDim]
- primal_normal_vert_v1[EdgeDim * ECVDim]
- primal_normal_vert_v2[EdgeDim * ECVDim]
- z_nabla2_e[OutputEdgeDim, KDim]
- inv_vert_vert_length[OutputEdgeDim]
- inv_primal_edge_length[OutputEdgeDim]
- Output field
- z_nabla4_e2[OutputEdgeDim, KDim]
- Neighbor tables
- Interpolate
-
- Neighbor tables
- v2e[6]
- Input fields
- p_e_in[EdgeDim, KDim]
- ptr_coeff_1[OutputVertexDim, 6]
- ptr_coeff_2[OutputVertexDim, 6]
- Output fields
- p_u_out[OutputVertexDim, KDim]
- p_v_out[OutputVertexDim, KDim]
-
- verts2cells
- Artifical kernel
- Interpolates
p_u_outandp_v_outof vertices to cell
- Interpolates
- Neighbor tables
- c2v[3]
- Input fields
- p_u_in[VertexDim, KDim]
- p_v_in[VertexDim, KDim]
- ptr_c_coeff_1[OutputCellDim, 6]
- ptr_c_coeff_2[OutputCellDim, 6]
- Output fields
- p_cell_out[OutputCellDim, KDim]
- Artifical kernel
- Neighbor accesses
- Indirect
- Indirect accesses via neighbor tables
- Strided
- Neighbor accesses via strides
- Indirect
- Iteration strategies
gpu_naive- Default iteration strategy in
GridTools C++ - One GPU thread calculates 1 element in horizontal and vertical axis
- Default iteration strategy in
gpu_kloop- Optional iteration strategy in
GridTools C++ - One GPU thread calculates 1 element in horizontal axis but multiple vertical levels
- Save neighbors and vertically independent fields in registers and iterate over multiple vertical fields
- Optional iteration strategy in
- Separate
- Execute
nabla4kernel, theninterpolateand thenverts2cells
- Execute
- Inlined
- Compute
nabla4andinterpolateoutputs for every input ofverts2cellskernel - More computations
- Less writes to device memory
- Compute
- Inlined c2v (
indirectonly)- Compress
c2v[v2e[e2c2v]]neighbor accesses toc2v2e2c2v- Read fields for 12 vertices instead of (6*4*3=) 72 vertices
- Assumes certain order of cells in
c2v, vertices ine2c2vand edges inv2e
- Compress
- Inlined strided
- Computation of upward and downward cell corresponding to vertex
(i, j)is done by the same thread at the same time to save memory loads
- Computation of upward and downward cell corresponding to vertex
- Inlined_v2v_separate (
indirectonly)- Execute the
inlined_v2vimplementation fornabla4andinterpolatekernels and then execute separately theverts2cellskernel- Improves register pressure
- Execute the
- Inlined cached
- Save to shared memory the intermediate output of
nabla4andinterpolatekernels only for the necessary fields to calculate the output cells of each threadblock- Reduces overcomputations as much as possible
- Save to shared memory the intermediate output of
gtfn- Only
indirect - Based on
GridTools C++ - Improved
GridTools C++- Memory loads via
__ldg gpu_kloopoptionconstneighbor tables and input fields- Kudos to Felix Thaler
- Memory loads via
- Only
CUDA- Plain cuda kernels
- Launched by python script
- Random input in benchmarks
- Validated kernels with serialized data from
GT4Py
- Occupancy
- All kernels have been optimized for best occupancy
__launch_bounds__and__maxnreg__- Launch bounds are applied to all kernels except some that were performing better with register limitation to a certain number of registers
- Thread Block size
- Specifically for
gpu_naiveimplementations, increasing the vertical axis thread block size (ThreadBlockDim.y/z) was beneficial since there are more chances to find neighbor tables and vertically independent fields in cache - 4-8-9 most used number with 80 vertical levels in total. For some kernels 12 or 16
- Specifically for
gpu_kloop- For this implementation vertical
GridDimis set to 1. Iteration number is controlled byThreadBlockDim.y/zandKDim. Exceptions are theinline_cachedversions where the shared memory size is limiting the number of iterations possible - Optimized number of iterations per kernel. 5-40 iterations. 8-20 usually have the best performance
- For this implementation vertical
strided_gpu_{naive,kloop}_inlined_cached: Tried 2 different implementations, number of threads same as input but then deactivate for the output some of them and number of threads same as output where each thread caclulates multiple elements. Former is betterstrided_*_inlined: calculate only necessary indexes, similar toc2v2e2c2vindirect_*_inlined_c2v: passc2v2e2c2vas inputindirect:nabla4iterates on edges (per-orientation- 1 edge per thread)strided:nabla4iterates on vertices (per-vertex- 3 edges per vertex/thread)strided:verts2cellscalculates both upward and downward cells together- Both
stridedandindirectversions operate on data with same ordering in memory- No SFC. Vertices and cells are ordered per
iandjcoordinates andedgesperorientation/vertices
- No SFC. Vertices and cells are ordered per
strided:e2ecvis also computed
- Smaller grids benefit by more threads and less
k leveliterations - Loop in
kis done with strideblockDim.y/z * gridDim.y/z- It would be more beneficial for kernels that read data from adjacent
klevels to do the looping with stride1in each thread- Current implementation in
gtfn, not on theCUDAkernels since we don't evaluate such case
- Current implementation in
- It would be more beneficial for kernels that read data from adjacent
GH200GPU- Median runtime presented
- 10 dry runs (not taken into account)
- 101 runs to select median
- 229758 edges
- Close to the amount of edges that fit in a single GPU for ICON runs
- 915948 edges
- For exploration
- Commit 5272141
gpu_kloop~20-30% fasterindirectnabla4_interpolate_inlined_v2v/stridednabla4_interpolate_inlined&indirect/stridedverts2cellsseparately fastest variationstridedas fast or a bit faster thaninlined_v2v(up to 8%)- Inlining the 3 kernels is not beneficial
inlined_cachedversion is fastest- Bottleneck is register pressure
- Still ~70-90% off the maximum theoretical optimal performance
- Use cache hints for loads and non temporal stores
- Try
cachedapproach forindirectimplementation- Compute border coordinates for each Thread Block (should be the same as
stridedfor our grid)per-vertexordering should be better due to smaller range of vertices/edges that need to be saved to shared memory
- Load them to shared memory
- Compute border coordinates for each Thread Block (should be the same as
- Try
TMAimplementation- Tried
cuda::pipelineto useLDGSTS/LDGSTS.BYPASSinstructions for loading the input fields ofnabla4kernel combined with saving thenabla4output to memory- Didn't see an improvement
- Try with
TMAto see if it improves - More complex if possible due to memory alignment requirements
- Probably doesn't help because of extra synchronization in the kernel
- Tried