Recommended Citation
1. N. Choi, K. M. Kim, H. G. Joo, “Initial Development of PRAGMA – A GPU-Based Continuous Energy Monte Carlo Code for Practical Applications” Proceedings of the Korean Nuclear Society Autumn Meeting, Goyang, Korea, Oct. 24-25, 2019.
2. N. Choi and H. G. Joo, "Domain Decomposition for GPU-Based Continuous Energy Monte Carlo Power Reactor Calculation," Nuclear Engineering and Technology, 2020, In Press, DOI:
https://doi.org/10.1016/j.net.2020.04.024.
GPU-based Continuous-Energy Monte Carlo Code
PRAGMA is a continuous-energy Monte Carlo (MC) criticality calculation code dedicated for commercial reactor applications. By employing cutting-edge GPU acceleration techniques, PRAGMA realizes massive particle simulations of full-scale PWRs in practical time span with limited computing resources.
Neutron Tracking Algorithm
For the vectorization of the random walk simulation, which is necessary to exploit GPUs, PRAGMA employs a special neutron tracking algorithm named event-based algorithm. In contrast to the history-based algorithm wherein the basic unit of work is a neutron history from birth to death, the event-based algorithm treats a particular event in a neutron history as the unit of work. That is, the events are processed one at a time, and a collection of neutrons that are about to undergo the event are traced together. Special sorting schemes are introduced to exploit the benefit of the event-based algorithm; at each migration, neutrons are sorted by their statuses (dead / alive), events (collision / surface), energies, and positions to maximize memory throughput and minimize branch divergences. This is a unique feature of the GPU-based MC simulation.
Reduction of Inter-Cycle Correlations
Various approaches to reduce inter-cycle correlations in MC k-eigenvalue calculations have been suggested, but the most favorable approach in terms of statistics is to simply use massive amount of particles. In this regard, PRAGMA employs a significant quantity of particles – at least over 108 per cycle–in actual simulations, powered by the exclusive parallel computing power of GPUs. At this rate of particles per cycle, the number of active cycles required to obtain a statistically meaningful solution often drops to several dozens. However, the number of required inactive cycles is independent to the number of particles, as it is only dependent to the dominance ratio of the problem.
Therefore, PRAGMA employs CMFD and ramp-up acceleration schemes together during the inactive cycles to reduce the cost of the inactive cycles under massive particle condition. Combining the two acceleration schemes together gives ~ 97% reduction to the inactive cycle computing cost compared to the standard MC algorithm, which makes the massive particle simulations practical.
Domain Decomposition
To carry out large-scale power reactor simulations with limited GPU memory, PRAGMA employs a domain decomposition algorithm. This algorithm provides an automated way of domain generation which optimizes the load balance. And through the inner-outer tracking algorithm, communications and computations are proceeded asynchronously and the communication overheads are effectively hidden behind the computations.
Validation of PRAGMA
PRAGMA provides full-featured continuous-energy calculation capability and achieves considerable speedup over well-known conventional production-grade MC codes. With a single consumer-grade GPU, PRAGMA is capable of delivering the performance equivalent to that of the production-grade codes with hundreds of CPU cores.
Also, T/H feedback, boron search, and control rod insertion capabilities have been prepared for the application to operating power reactor calculations and nuclear design. Currently, a depletion module is under development and PRAGMA will be capable of carrying out a complete core follow simulation in the upcoming release.
| # of GPUs |
24(NVIDIA GeForce RTX 2080 Ti) |
| # of Particles per Cycle |
240,000,000(Total 11.4 Billion) |
| # of Inactive Cycles |
25 |
| # of Active Cycles |
40 |
| Computing Time |
13m 54s |
| CMFD |
On(Assembly-wise) |
| Ramp-up |
On(Exponential, 20) |
| Core Power |
3983MW |
| Core Flow Rate |
20553.7kg/s |
| Inlet Temperature |
564.45K |
| Pressure |
15.514MPa |
| Boron |
1085ppm |
| Library Temperatures (K) |
500/600/700/ … /1700/1800 |