I am launching a given problem that parallelizates by means of OpenMP. It runs a given number of iterations of the same piece of code that processes a volume of data. Is in that level where OpenMP is applied, making each thread process a subvolume. Every iteration should have the same workload, as well as every subvolume.
When compiled with ICC, iterations last always the same amount of time, as expected. But there comes the weird thing: when compiled with GCC, the time per iteration starts to increase, reaches a maximum and then decreases once again until it reaches a given value where it stabilises. The same program compiled without OpenMP makes no difference when using ICC or GCC.
Does anyone observed that behaviour in OpenMP in those compilers?
[EDIT 1]: guided and static scheduling policies have been tested.
[EDIT 2]: The code looks somewhat like this:
#pragma omp parallel for schedule(static) private(i,j,k)
for(i = 0; i < N; i++)
for(j = 0; j < N; j++)
for(k = 0; k < N; k++){
a[ k+j*N+i*NN] = 0.f;
b[ k+j*N+i*NN] = 0.f;
c[ k+j*N+i*NN] = 0.f;
d[ k+j*N+i*NN] = 0.f;
}
for( t = 0; t < T; t+=dt){
/* ... change some discrete values in a,b,c .... */
/* and propagate changes */
#pragma omp parallel for schedule(static) private(i,j,k)
for(i = 0; i < N; i++)
for(j = 0; j < N; j++)
for(k = 0; k < N; k++){
d[ k+j*N+i*NN ] = COMP( a,b,c,k+j*N+i*NN );
}
}
Where COMP performs some kind of linear application of values in a,b,c in the position k+j*N+i*NN (and some of their neighbours). The point is that this code in GCC and ICC caused the problem I described. The point is that I found out that I change the initialisation of a,b,c,d to some value other than 0.0f (f.ex, 0.5f) that thing that the time spent per time step increases doesn't occur.
[EDIT 3] : It seems is not GOMP's fault. The same happens with OpenMP disabled. Once again, with ICC (without or with openmp) doesn't occur at all. Is there any way I can close this thread?
GOMP_CPU_AFFINITY=0-31where 31 is the number of cpu cores -1; andOMP_WAIT_POLICY=activeto get more predictable results. - osgx