0
votes

Say that I have a construct like this:

for(int i=0;i<5000;i++){
  const int upper_bound = f(i);
  #pragma acc parallel loop
  for(int j=0;j<upper_bound;j++){
    //Do work...
  }
}

Where f is a monotonically-decreasing function of i.

Since num_gangs, num_workers, and vector_length are not set, OpenACC chooses what it thinks is an appropriate scheduling.

But does it choose such a scheduling afresh each time it encounters the pragma, or only once the first time the pragma is encountered?

Looking at the output of PGI_ACC_TIME suggests that scheduling is only performed once.

2

2 Answers

1
votes

The PGI compiler will choose how to decompose the work at compile-time, but will generally determine the number of gangs at runtime. Gangs are inherently scalable parallelism, so the decision on how many can be deferred until runtime. The vector length and number of workers affects how the underlying kernel gets generated, so they're generally selected at compile-time to maximize optimization opportunities. With loops like these, where the bounds aren't really known at compile-time, the compiler has to generate some extra code in the kernel to ensure exactly the correct number of iterations are performed.

0
votes

According to OpenAcc 2.6 specification[1] Line 1357 and 1358:

A loop associated with a loop construct that does not have a seq clause must be written such that the loop iteration count is computable when entering the loop construct.

Which seems to be the case, so your code is valid.

However, note it is implementation defined how to distribute the work among the gangs and workers, and it may be that the PGI compiler is simply doing some simple partitioning of the iterations. You could manually define values of gang/workers using num_gangs and num_workers, and the integer expression passed to those clauses can depend on the value of your function (See 2.5.7 and 2.5.8 on OpenACC specification).

[1] https://www.openacc.org/sites/default/files/inline-files/OpenACC.2.6.final.pdf