AWS Glue can read data, transform it with Spark, and load the result into Amazon Redshift on a schedule or in response to an event. This guide shows the parts that make that pipeline work: a source in Amazon S3, a Glue job, a Redshift connection, and an S3 staging location used during the transfer. It assumes your source data and Redshift database already exist.

Pipeline: S3 source → Glue ETL job → S3 temporary staging → Redshift table. Glue uses Redshift COPY and UNLOAD commands for data movement through S3, so the job, Redshift, and staging location each need the right permissions and network path. See AWS’s Redshift connection guide for the supported connection patterns.

How Glue and Redshift Work Together

AWS Glue runs the transformation; Redshift stores and queries the prepared data. S3 can hold the source files and the temporary files exchanged with Redshift. Keep the source, temporary staging, and any curated output in clearly named prefixes so their permissions and retention can be managed separately.

The Data Catalog Holds Metadata

The AWS Glue Data Catalog stores table definitions and other metadata. A crawler can inspect a source and create or update catalog tables; the table points to data elsewhere, such as files in S3. A crawler is optional when your job can read a known path and schema, and it does not validate the business meaning of the records.

If you want a wider introduction to Glue’s catalog, crawlers, and jobs, read our AWS Glue core concepts tutorial. For current service behavior, use the AWS Glue Data Catalog documentation.

Redshift Uses S3 Staging

When a Glue job reads from or writes to Redshift, the connector stages data in S3 and asks Redshift to run UNLOAD or COPY. The temporary directory is a working location, not the permanent destination table. Give the Glue job access to that S3 prefix, and give Redshift access to the same prefix through an IAM role when the connector uses role-based access. AWS lists the required Redshift-side S3 permissions in its guide to permissions for COPY and UNLOAD. For the default Redshift Spark connector path, keep the staging bucket in the same Region as Redshift unless you configure a documented cross-Region option.

Plan IAM, S3, and Network Access

Separate Glue and Redshift Permissions

The Glue job runs with an IAM execution role. Grant that role the Glue permissions it needs, read access to source objects, access to its script and logs, and the S3 permissions required for the temporary staging prefix. Add Glue Data Catalog, Secrets Manager, and KMS permissions only when the job uses those resources. Scope S3 actions to the required bucket and prefixes rather than attaching broad bucket access. AWS documents the role and S3 permissions in Create an IAM role for AWS Glue.

For the recommended role-based access path, associate a dedicated IAM role with the Redshift cluster or Serverless namespace and limit its S3 permissions to the staging prefix. AWS Glue can also pass temporary credentials from the job role, but AWS does not recommend that option because those credentials expire after one hour. The database user in the Glue connection also needs SQL privileges for the target schema and table; those database grants are separate from IAM permissions. See AWS’s Redshift connection and role guidance.

Configure VPC Connectivity

Create an AWS Glue Redshift connection with the VPC, subnet, security groups, endpoint, database, and credentials for your environment. Add that connection to the job’s network configuration. Glue creates private network interfaces in the selected subnet; Redshift must be reachable from that subnet through the VPC or an established network route. AWS describes this setup in its guide to network access for Glue data stores.

On the Redshift security group, allow inbound traffic to the database port from the Glue job’s security group. On a Glue security group, AWS requires a self-referencing inbound rule for all TCP ports so Spark executors can communicate with one another. Also make sure Glue’s outbound rules and the subnet routes allow connections to Redshift and S3. Keep these rules scoped to the relevant security groups; do not open the Redshift port to all IP addresses. For S3 access from a private subnet, configure an S3 VPC endpoint and its route and policy. A same-Region S3 endpoint does not require a NAT gateway; add NAT only if the job needs other internet destinations or network paths that require it. See AWS’s S3 VPC endpoint guidance for Glue.

Protect Credentials and Staged Data

Prefer IAM-based Redshift authentication when your chosen Glue connector supports it. Otherwise, use a Glue connection backed by AWS Secrets Manager rather than putting credentials in the script or job arguments; restrict who can read the secret and allow only the Glue role that needs it. Encrypt source and staging objects as required by your data policy, and grant the Glue and Redshift roles the KMS permissions for the keys they use. Treat staged files as sensitive data: restrict their prefix, avoid logging record contents, and apply an S3 lifecycle policy. The Redshift Spark connector does not automatically clean its temporary directory, so lifecycle cleanup matters for both data retention and cost. See AWS’s guidance on Redshift Spark connector options and considerations.

Build and Run the Glue Job

Create and Test a Redshift Connection

  1. Confirm the Redshift endpoint, port, database, target schema, and table. Create the destination table with deliberate column types and a documented key strategy.
  2. In AWS Glue, create a Redshift connection that uses the intended VPC, subnet, security groups, and credential source.
  3. Test the connection, then attach it to the job’s network connections. Confirm that the Glue role can reach the source and staging S3 prefixes and that Redshift can reach its staging prefix.
  4. Choose a source from the Data Catalog or read a known S3 path in the job. Use a crawler when schema discovery or catalog updates are useful; define the schema directly when you need tighter control over schema changes.

Write Through the Redshift Connector

This script shows the shape of a catalog-source job that writes a DynamicFrame to Redshift through a Glue connection. It uses AWS Glue’s documented from_jdbc_conf API. Glue 4.0 and later also include the Redshift integration for Apache Spark, with its own connection options; follow the documentation for the path you choose. Replace the example names and IAM role ARN, pass a valid TempDir job argument, and add the transformations and validation your data requires. The Redshift role in aws_iam_role is used for Redshift’s access to the S3 staging location.

import sys
from awsglue.context import GlueContext
from awsglue.job import Job
from awsglue.utils import getResolvedOptions
from pyspark.context import SparkContext

args = getResolvedOptions(sys.argv, ["JOB_NAME", "TempDir"])
glue_context = GlueContext(SparkContext.getOrCreate())
job = Job(glue_context)
job.init(args["JOB_NAME"], args)

source = glue_context.create_dynamic_frame.from_catalog(
    database="raw_data",
    table_name="daily_sales",
    transformation_ctx="sales_source",
)

# Add explicit mappings, filters, and data-quality checks here.
glue_context.write_dynamic_frame.from_jdbc_conf(
    frame=source,
    catalog_connection="redshift-analytics",
    connection_options={
        "database": "warehouse",
        "dbtable": "staging.daily_sales",
        "aws_iam_role": "arn:aws:iam::123456789012:role/RedshiftS3StageAccess",
    },
    redshift_tmp_dir=args["TempDir"],
)

# Validate staging rows and complete the final table write here.
# Raise an error before committing if validation or promotion fails.
job.commit()

The example writes to a staging table to make a controlled final load possible; the comments show where that final step belongs. When job bookmarks are enabled, validate the staged rows and complete the final write to the reporting table before calling job.commit(). This keeps the source bookmark from advancing past a batch whose final load did not succeed. The Glue Studio Redshift target’s documented write options include append and configured upsert or merge behavior. With a script-based job, choose and test the write semantics explicitly.

Choose the Target Write Behavior

Decide how a run should affect rows already in Redshift before you schedule it. The right choice depends on whether the source contains only new records, a complete snapshot, or changed records keyed to existing rows. Redshift does not enforce primary-key or unique constraints, so validate key uniqueness in your load logic instead of relying on a table declaration to reject duplicates; see AWS’s table-constraint guidance.

Write patternUse it whenDesign check
AppendEach input batch contains only records that should be added.Retries or replayed files must not create unwanted duplicates.
Full refreshThe source is a complete snapshot and replacing the target is acceptable.Plan what readers see if the load fails partway through.
Upsert or mergeIncoming records can update existing rows using stable keys.Define key uniqueness, update columns, and delete behavior explicitly.

AWS Glue job bookmarks can track input already processed for supported sources. Enable them in the job and preserve the source’s transformation_ctx; initialize and commit the job state in the script. Bookmarks manage source progress. They do not decide whether Redshift appends, replaces, or updates target rows, so make that target behavior explicit. See Using job bookmarks for the supported patterns and requirements.

Schedule, Monitor, and Tune the Pipeline

Schedule Jobs and Manage Dependencies

Use an AWS Glue schedule trigger for recurring loads. Use a Glue workflow when the pipeline has multiple jobs or crawlers with dependencies. For event-driven runs, a Glue workflow can start from an EventBridge event. In the Glue workflow pattern for S3 object data events, enable the documented CloudTrail delivery to EventBridge. AWS Glue does not guarantee EventBridge message delivery and does not deduplicate duplicate events, so design the target load to handle retries and repeated starts safely. See starting a Glue workflow from EventBridge.

Monitor Job Runs and Data Results

Review each run’s status, duration, driver and executor logs, and available Glue or CloudWatch metrics. Alert on failed runs and on duration or output changes that matter to your pipeline. Compare source and target row counts or business totals after a load, and keep enough context to identify the input batch and job run. AWS explains the available Glue logs, metrics, and monitoring.

Control Runtime and Cost

Start with a modest worker configuration, then measure representative runs before changing worker type, capacity, or concurrency. Filter source data early and read only the S3 partitions the job needs. Avoid generating excessive tiny files in intermediate S3 output. Redshift’s parallel load path depends on the staged data and target table, so compare complete job runtime and Redshift load behavior instead of tuning Glue workers in isolation.

FAQs

Does an AWS Glue Job Need a Crawler?

No. A crawler is useful when you want Glue to discover or refresh table metadata. If the source schema and S3 path are known, you can define the schema in your job or create a Data Catalog table manually, then read the source without crawling it. Crawlers do not replace data validation in the ETL job.

Do Job Bookmarks Prevent Duplicate Redshift Rows?

Not by themselves. A bookmark tracks eligible source progress. Select append, full-refresh, or upsert behavior separately, and make retries and replayed events safe for that target behavior.

Does Glue Send Data Directly to Redshift?

The Redshift connector stages data in S3 and uses Redshift COPY and UNLOAD operations. Configure the S3 staging path, the permissions on both sides, and the network routes as part of the connection setup.