Run SparkR in Hive Context to reach the Hive UDFs and window functions the SQL Context lacks, compared against dplyr, plus building Spark with Hive.
Run SparkR in local and cluster mode from an R project, setting environment variables and the library path, then reading JSON and CSV files.
Set up a two node Spark standalone cluster on VirtualBox Ubuntu guests, covering machine preparation, copying the VDI image and password-less SSH.