AWS Redshift Data Processing

Question

I'm working with a small company currently that stores all of their app data in an AWS Redshift cluster. I have been tasked with doing some data processing and machine learning on the data in that Redshift cluster.

The first task I need to do requires some basic transforming of existing data in that cluster into some new tables based on some fairly simple SQL logic. In an MSSQL environment, I would simply put all the logic into a parameterized stored procedure and schedule it via SQL Server Agent Jobs. However, sprocs don't appear to be a thing in Redshift. How would I go about creating a SQL job and scheduling it to run nightly (for example) in an AWS environment?

The other task I have involves developing a machine learning model (in Python) and scoring records in that Redshift database. What's the best way to host my python logic and do the data processing if the plan is to pull data from that Redshift cluster, score it, and then insert it into a new table on the same cluster? It seems like I could spin up an EC2 instance, host my python scripts on there, do the processing on there as well, and schedule the scripts to run via cron?

I see tons of AWS (and non-AWS) products that look like they might be relevant (AWS Glue/Data Pipeline/EMR), but there's so many that I'm a little overwhelmed. Thanks in advance for the assistance!

This is a really broad question and there are many ways to implement what you're talking about. You are generally asking about ETL (Extract, Transform, Load), so I would advise searching books and docs on that. — Dan Kowalczyk
Also, since you're new to SO, you may not know that it's more likely to get answers if you keep your questions focused. I see it more that a focused question will have a general answer than the other way around where a general question gets a focused answer. — Dan Kowalczyk
please accept my answer or whichever is the best answer below — Jon Scott
Stored Procedures are now supported in Amazon Redshift from version 1.0.7287 (late April 2019). Please review the document "Creating Stored Procedures in Amazon Redshift" for more information on getting started with stored procedures. — Joe Harris

John Rotenstein John Rotenstein · Accepted Answer · 2017-10-07T22:40:05

ETL

Amazon Redshift does not support stored procedures. Also, I should point out that stored procedures are generally a bad thing because you are putting logic into a storage layer, which makes it very hard to migrate to other solutions in the future. (I know of many Oracle customers who have locked themselves into never being able to change technologies!)

You should run your ETL logic external to Redshift, simply using Redshift as a database. This could be as simple as running a script that uses psql to call Redshift, such as:

`psql <authentication stuff> -c 'insert into z select a, b, from x'`

(Use psql v8, upon which Redshift was based.)

Alternatively, you could use more sophisticated ETL tools such as AWS Glue (not currently in every Region) or 3rd-party tools such as Bryte.

Machine Learning

Yes, you could run code on an EC2 instance. If it is small, you could use AWS Lambda (maximum 5 minutes run-time). Many ML users like using Spark on Amazon EMR. It depends upon the technology stack you require.

Amazon CloudWatch Events can schedule Lambda functions, which could then launch EC2 instances that could do your processing and then self-Terminate.

Lots of options, indeed!

AWS Redshift Data Processing

3 Answers