Repository navigation
Conversation
|
Nice. Make it the default, compute in geometry preprocessing when necessary, no opt-in |
|
The failed tests seem to be just due to floating-point ordering. Is the PR ready for review? |
Kind of. I wanted to run some more detailed measurements on larger runs and add the caching for periodic boundaries as well. I will have some spare time next week to do that. I also have to run the pre-commit hook |
|
I was finally able to implement the caching for period boundaries and run some benchmark cases. Periodic boundariesThe periodic contributions to the least-squares normal matrix are exchanged at the geometry level (a metric-only periodic communication added to the geometry, independent of any solver), so the cache is built once in Edge-coloring fixAn earlier version of this branch built the edge coloring eagerly in the cache setup, BenchmarkFull regression suites (parallel, hybrid, serial) run with both the develop base
Of the 147 cases with more than 1 s of compute time, 124 are faster, 19 are within the Note on regression referencesThe cached path accumulates the RHS in an edge loop instead of in a loop through point neighbours, which changes the roundoff of the gradients. This will make some regression cases fail. Those references will need to be updated once CI runs on this branch, unless we prefer to keep neighbour-order accumulation in the cached path at some cost in speed. IMO this is now ready for review. |
…the edge loop in the LSQ caching in\ dry run)
Proposed Changes
This PR
adds an opt-in option,caches the factorized metric terms of the (weighted or unweighted) least-squares gradients. The LSQ normal matrix A depends only on the node coordinates and on the weighting, yet it is currently re-assembled and re-factorized on every gradient evaluation, every nonlinear iteration. With the option enabled:CACHE_LSQ_METRICS(defaultNO),CGeometry.SetControlVolume(..., UPDATE)(the common point of all mesh motion/deformation paths, on fine and coarse MG levels) and rebuilt on the next evaluation, i.e. once per mesh update, so the savings scale with the number of inner iterations per time step.for periodic boundaries (the periodic LSQ communication fuses the matrix and RHS accumulations; an RHS-only exchange would be needed) andfor the discrete adjoint (the coordinate dependence of the metrics must remain on the tape, just a guess I am not an expert on the discrete adjoint implementation).Saved operations (per node per gradient evaluation, 3D, k edge-neighbors, N variables)
Eliminated from every evaluation after the first:
What remains is the RHS accumulation (~4N flops per node per edge) and one S*b product (18N flops per node). For the N=6 compressible primitives this roughly halves the gradient-kernel flops; for scalar solvers (N=1-2, e.g. SA) where the assembly dominates it is a ~4-5x reduction. Measured end-to-end (serial, linear-solver work held fixed): ~3.3% of total wall time on an inviscid ONERA M6 (57.5k nodes, tets, one gradient set per iteration) and ~6% on a 2D RANS-SA case (three LSQ gradient sets per iteration). Cases dominated by the linear solver will see proportionally less.
Memory footprint
The cache adds nDim*(nDim+1)/2 su2doubles per node per weighting used (48 B/node per weighting in 3D, 24 B/node in 2D; at most two weightings), allocated lazily per grid level and shared by all solvers. For reference, every solver already allocates an nDim*nDim
Rmatrixscratch (72 B/node in 3D), so for a typical RANS case the addition is small compared to the existing gradient machinery and negligible compared to overall solver memory.Validation
With the option enabled, results are identical to
develop(to output precision) for:NUM_METHOD_GRAD= WEIGHTED_LEAST_SQUARESandNUM_METHOD_GRAD_RECON= LEAST_SQUARES(all residuals including the turbulence variable),GRID_MOVEMENT= RIGID_MOTION(dual time stepping), exercising the invalidation/rebuild path,config_template.cfgdocuments the new option. Default behavior (CACHE_LSQ_METRICS= NO) is bit-identical to current develop.pre-commit run --allto format old commits.