The use of GPUs in large-scale technical and scientific computing is spreading, and several libraries for distributed parallel sparse linear equation solvers are GPU-compatible. While these libraries process a system consisting of a single coefficient matrix and right-hand-side (RHS) vector, in fields such as structural analysis based on the finite element method (FEM), the analysis mesh itself is divided into multiple subdomains and matrices are constructed independently in each subdomain. This represents a structural mismatch with the input format assumed by general-purpose direct solver libraries. In this study, a direct linear equation solver using NVIDIA CUDA libraries was newly implemented as an extension to FrontISTR, a parallel finite element structural analysis program. The objective is to enable solving linear equations based on domain decomposition by the direct method, which is a robust solution method applicable to a wide range of analyses. For sequential execution targeting models that fit within a single GPU’s memory, the implementation produced solutions with very small relative residuals on NVIDIA A100 GPUs, achieving speedups of approximately 7 to 25 times compared with CPU execution of the same routine. For larger models requiring distributed parallel execution across multiple GPUs, a parallel algorithm based on the multiplicative Schwarz procedure combined with CUDA-aware MPI was designed and implemented, and numerical experiments confirmed the algorithmic correctness of the iteration, although its wall-clock performance does not yet surpass that of FrontISTR’s CPU-based parallel direct solver. The presented framework is therefore positioned as an alternative GPU-compatible solution pathway that operates directly on the domain-decomposed layout, using a single-GPU sparse-direct routine as a per-subdomain local solver, for models that exceed the memory of a single GPU.
抄録全体を表示