How to indicate the database in SparkSQL over Hive in Spark 1.3

Question

I have a simple Scala code that retrieves data from the Hive database and creates an RDD out of the result set. It works fine with HiveContext. The code is similar to this:

val hc = new HiveContext(sc)
val mySql = "select PRODUCT_CODE, DATA_UNIT from account"
hc.sql("use myDatabase")
val rdd = hc.sql(mySql).rdd

The version of Spark that I'm using is 1.3. The problem is that the default setting for hive.execution.engine is 'mr' that makes Hive to use MapReduce which is slow. Unfortunately I can't force it to use "spark". I tried to use SQLContext by replacing hc = new SQLContext(sc) to see if performance will improve. With this change the line

hc.sql("use myDatabase")

is throwing the following exception:

Exception in thread "main" java.lang.RuntimeException: [1.1] failure: ``insert'' expected but identifier use found

use myDatabase
^

The Spark 1.3 documentation says that SparkSQL can work with Hive tables. My question is how to indicate that I want to use a certain database instead of the default one.

Did you try the regular Hive syntax i.e. select * from mydb.mytable? — Samson Scharfrichter
Yes - getting another error: java.lang.RuntimeException: Table Not Found: myDatabase.account — Michael D

WestCoastProjects WestCoastProjects · Accepted Answer · 2017-11-16T19:03:09

use database

is supported in later Spark versions

https://docs.databricks.com/spark/latest/spark-sql/language-manual/use-database.html

You need to put the statement in two separate spark.sql calls like this:

spark.sql("use mydb")
spark.sql("select * from mytab_in_mydb").show

How to indicate the database in SparkSQL over Hive in Spark 1.3

3 Answers

use database