From Performance Assessment to Faster Simulations: Optimising Alya through POP3

Tuesday, October 6, 2026

Alya, developed at the Barcelona Supercomputing Center (BSC), is a high-performance computational mechanics code capable of solving complex, coupled-physics problems and is one of the codes included in the EXCELLERAT CoE.

In collaboration with EXCELLERAT we have conducted a performance assessment of an Alya use case to analyse a combustion simulation. The simulation uses the external ISAT-CK7 library to accelerate the computation of chemical reactions by retrieving results from a previously computed table.

The assessment identified two performance bottlenecks affecting the application's scalability: a very high load imbalance and an abnormally low CPU frequency in the region responsible for handling chemical reactions.

While the load imbalance is a well known issue in this kind of simulation because chemical reactions are computed only in a small part of the simulation domain, the low CPU frequency was an unexpected finding. Moreover, we also identified the CPU frequency to be worsening the load imbalance problem.

Identifying the source of the bottleneck

Figure 1: Timeline showing 4 timesteps of the Alya use case analyzed. Code color represents the different code regions.

In Figure 1 we can see the structure of the use case analyzed, it shows 4 timesteps and the main computational regions of the application. The Chemic begin time step region, which includes the processing associated with the chemical reactions, accounts for nearly 89% of the total useful execution time per timestep, making it a key target for optimisation.

A closer look at the average CPU frequency provided a further clue. As shown in Table 1, the Chemic begin time step region runs at just 0.63 GHz, compared with 2.49 GHz in Chemic do iterations and 2.72 GHz in Nastin do iteration.

Region   Average CPU frequency (GHz)
Chemic begin time step     0.63
Chemic do iterations     2.49
Nastin do iteration     2.72

Table 1. Average CPU frequency in the main computational regions.

The assessment identified the calls to the external ISAT-CK7 library as the source for the low-frequency. This finding narrowed the investigation from the application as a whole to a specific component, providing a clear direction for the proof of concept (PoC): determine the cause of the unexpectedly low CPU frequency and implement a targeted solution.

A targeted Proof-of-Concept: Investigating ISAT-CK7

To isolate the issue, the team analysed ISAT-CK7 independently and compared its performance when compiled with different compilers. This uncovered that using a different compiler (NVFortran instead of ifort) reduced the execution time by half. Further investigation showed that the number of system calls was also drastically reduced with this change. This pointed us to an issue in the way the library interacted with the operating system through the Fortran runtime.

The investigation identified calls to uname and getrusage, functions provided by the operating system and used by the Fortran runtime implementation of cputime(), as the source of the unnecessary overhead. The proposed solution was to replace the original cputime() subroutine with an equivalent C implementation that avoids this overhead. The workaround solution can be seen in Listing 1.

Listing 1. Implementation of a C wrapper to replace the original cputime() subroutine.

Validating the Optimisation in Alya

The team then linked the new C implementation into ISAT-CK7 and integrated the modified library into Alya. The results confirmed that this targeted change improved performance across the full range of scalability tests.

As shown in Figure 2, the optimised implementation consistently outperformed the original version, demonstrating that a relatively small change within an external library could benefit the performance of the complete application.

Figure 2. Alya scalability results with the original and optimised ISAT-CK7 implementations.

The detailed trace analysis provides further evidence of the improvement. Figure 3 compares two execution traces from the 3D case using 1,120 MPI ranks. In the selected timestep, the optimised implementation achieves a 5× speedup over the whole time step duration. While the average CPU frequency in the Chemic begin timestep region increases from 0.63 GHz to 2.97 GHz. The optimisation also improves load balance by 10% globally.

Figure 3. Execution traces for the 3D case using 1,120 MPI ranks, comparing the original and optimised implementations.

From Performance Insight to Measurable Impact

This use case demonstrates the value of the POP3 performance assessment methodology: combining performance metrics and trace analysis to identify a bottleneck, investigate its underlying cause, and turn the resulting insight into a concrete optimisation.

By isolating the overhead in an external library and validating a targeted solution in the full application, the assessment and subsequent proof of concept delivered measurable performance improvements in Alya. The results show how a focused intervention in a specific software component can have a substantial impact on a complex HPC application running at scale.

– Álvaro LLorente and Marta Garcia-Gasulla (BSC)