I need to execute a for loop that acts on many data points, each one independently of any other data point.
for x in range(1,10000000) some_procedure(x)
and i need to make it run faster, by exploiting data parallelism, but I have to work with MPI.
I am using mpi4py but am completely new to this. Which one is supposed to work faster in theory between option a and b, and why:
a. in foo.py:
from mpi4py import MPI
sendbuf=[]
root=0
comm = MPI.COMM_WORLD
if comm.rank==0:
sendbuf= range(0,10000000)
v=comm.scatter(sendbuf,root)
some_procedure(v)
and later do the following:
mpiexec -np 10000000 python foo.py
b. or
from mpi4py import MPI
import sys
client_script = 'some_procedure.py'
comm = MPI.COMM_SELF.Spawn(sys.executable, args=[client_script], maxprocs=10000000)
Many thanks beforehand!!