Is it possible to run Hive on Spark with YARN capacity scheduler?

Question

I use Apache Hive 2.1.1-cdh6.2.1 (Cloudera distribution) with MR as execution engine and YARN's Resource Manager using Capacity scheduler.

I'd like to try Spark as an execution engine for Hive. While going through the docs, I found a strange limitation:

Instead of the capacity scheduler, the fair scheduler is required. This fairly distributes an equal share of resources for jobs in the YARN cluster.

Having all the queues set up properly, that's very undesirable for me.

Is it possible to run Hive on Spark with YARN capacity scheduler? If not, why?

Kenry Sanchez Kenry Sanchez · Accepted Answer · 2020-06-27T01:12:39

I'm not sure you can execute Hive using spark Engines. I highly recommend you configure Hive to use Tez https://cwiki.apache.org/confluence/display/Hive/Hive+on+Tez which is faster than MR and it's pretty similar to Spark due to it uses DAG as the task execution engine.

Is it possible to run Hive on Spark with YARN capacity scheduler?

2 Answers