2
votes

I am building a Spark App, in which I submit several jobs (pyspark). I am using threads to run them in parallel, and also I am setting: conf.set("spark.scheduler.mode", "FAIR")

Still, I see the jobs run serially in FIFO way. Am I missing something?

EDIT: After writing to the Spark Mailing List, I got a couple of things more:

I was totally missing this last point

2
I read the documentation, but I can assure you that they are still running in FIFO way, without any round robin - Enrico D' Urso
What master url did you pass? - Ajeet Shah

2 Answers

0
votes

Apparently, spark doesn't support threading for all kinds of jobs. You can only submit a job in parallel if it is a spark action.

Inside a given Spark application (SparkContext instance), multiple parallel jobs can run simultaneously if they were submitted from separate threads. By “job”, in this section, we mean a Spark action (e.g. save, collect) and any tasks that need to run to evaluate that action. Spark’s scheduler is fully thread-safe and supports this use case to enable applications that serve multiple requests (e.g. queries for multiple users).

please follow this link

0
votes

How many executors do you have?

If you have 1 executor then FIFO is same as FAIR.

I am saying so because by default standalone mode creates 2 executors and in "cluster" mode you job would take 1 for driver and 1 for executor.

So you need 4 executors to run 2 jobs in cluster mode.