AWS Glue Tutorial for Beginners: Core Concepts and ETL Jobs
Published on · Updated on
Rewrote the tutorial with a practical S3-to-Parquet job and current guidance on catalogs, roles, bookmarks, triggers, and costs.
AWS Glue helps you discover, prepare, and move data for analytics. This beginner tutorial explains how its Data Catalog, crawlers, and ETL jobs fit together, then walks through a small batch job that reads CSV files from Amazon S3 and writes Parquet files back to S3.
The key distinction is that the Data Catalog stores descriptions of data; your files and database rows remain in their data stores. A crawler can infer and record metadata, while an ETL job reads and changes the actual records.
Introduction to AWS Glue
What AWS Glue does
AWS Glue is a managed data integration service. You define a job that reads from one or more sources, applies transformations, and writes to a destination. AWS manages the job runtime, so you do not create and operate a Spark cluster yourself. You still choose the job type, supported Glue version, worker configuration, permissions, and run behavior; auto scaling is an optional job setting where supported, not an assumption for every job.
Glue is useful when you need repeatable data preparation, such as cleaning records, joining datasets, changing file formats, or loading results into another data store. A Glue job does not run merely because you uploaded data. You start it yourself or configure a trigger or workflow.
The Data Catalog, crawlers, and jobs
- Data Catalog: A regional metadata repository with databases, table definitions, schemas, partitions, and other information that helps AWS services locate and interpret data. A catalog table describes data stored elsewhere; it does not hold the rows.
- Crawler: An optional discovery tool that connects to a supported source, infers its format and schema, and creates or updates table metadata in the Data Catalog. It does not transform or move the source data.
- ETL job: The code or visual workflow that reads records, transforms them, and writes output. Jobs can use Apache Spark, Ray, or Python shell runtimes. AWS Glue Studio provides visual job authoring as well as script editing.
- Trigger: A way to start jobs or crawlers on demand, on a schedule, or after configured job or crawler events.
For the current list of job types and runtime versions, see the AWS Glue Studio job guide and AWS Glue version notes. The Glue version determines the Spark and Python runtime for Spark jobs, so check the available versions before relying on a specific library or language version.
Prepare a small S3 dataset
Create input and output locations
You need an AWS account with permission to use AWS Glue and Amazon S3. Create or choose an S3 bucket, and keep the raw input and processed output in separate prefixes. If you need a refresher on buckets, objects, and basic access controls, see Getting Started with AWS S3: Essential Concepts.
For this example, upload a UTF-8 CSV file to a prefix such as s3://YOUR_BUCKET/raw/orders/. Give it a header row with these columns:
order_id,customer_id,amount,order_date
101,c-7,19.95,2026-10-01
102,c-9,5.00,2026-10-02
Create a different output prefix, such as s3://YOUR_BUCKET/processed/orders/. Do not place the output beneath the input prefix, or later runs could read the job's own results.
Use a crawler only when discovery helps
You can run an ETL job directly against a known S3 path; the example below does not require a crawler or a catalog table. To explore the Data Catalog, create a crawler for the raw prefix, choose a Data Catalog database, and run it on demand or on a schedule. Review the resulting table name, columns, and types before using them in a catalog-source job. Inferred schemas are useful starting points, but they can be wrong when files have mixed layouts or ambiguous values. The example below reads the S3 path directly whether or not a crawler ran. See AWS's guide to populating the Data Catalog with crawlers.
A crawler records metadata and can group directory layouts into catalog partitions. It does not create, move, or rewrite the underlying files. Schema discovery also does not tell you whether records meet your business rules. AWS Glue has separate data quality features for that work; our AWS Glue Data Quality guide explains how catalog evaluations and in-job checks fit different needs.
Give each role the access it needs
The person or role using the console needs permission to create and start Glue resources. The Glue job runs with its own IAM execution role, whose trust policy must allow glue.amazonaws.com to assume it. The person or role creating the job also needs iam:PassRole scoped to that execution role. For this example, the job role needs permission to read the input objects and job script, write and replace objects in the output prefix, and send logs to CloudWatch. AWS lists s3:ListBucket and s3:GetObject for S3 sources, plus s3:ListBucket, s3:PutObject, and s3:DeleteObject for targets; this overwrite example needs the delete permission. Add access to the Data Catalog, KMS keys, connections, or other services only if the job uses them.
If you run a crawler, its configured role needs to read the source and create or update the catalog metadata. Keep permissions scoped to the required bucket prefixes and Glue resources. A bucket path can be used directly in this job; a Glue connection is needed only when the selected source or target requires one, such as some database connections. See AWS's minimum permissions guide for Glue jobs and IAM role guide.
Create an S3-to-Parquet Glue job
Choose a Spark job for this example
This walkthrough uses a batch Apache Spark job because it reads a collection of CSV files, filters rows, and writes a columnar format. In AWS Glue Studio, create a Spark job and use the script editor. Select a supported Glue version and a modest worker configuration suitable for a small test dataset. The visual editor is another way to build Spark ETL flows; it is not required for this script.
Before running the job, replace YOUR_BUCKET in both paths below. Confirm that the job role can read raw/orders/ and write processed/orders/. The CSV header must contain the four column names used in the script.
Read CSV, filter it, and write Parquet
import sys
from awsglue.context import GlueContext
from awsglue.job import Job
from awsglue.utils import getResolvedOptions
from pyspark.context import SparkContext
from pyspark.sql import functions as F
args = getResolvedOptions(sys.argv, ["JOB_NAME"])
glue_context = GlueContext(SparkContext.getOrCreate())
job = Job(glue_context)
job.init(args["JOB_NAME"], args)
raw_orders = glue_context.create_dynamic_frame.from_options(
connection_type="s3",
connection_options={"paths": ["s3://YOUR_BUCKET/raw/orders/"]},
format="csv",
format_options={"withHeader": True},
transformation_ctx="raw_orders",
).toDF()
clean_orders = (
raw_orders
.filter(F.col("order_id").isNotNull() & (F.trim(F.col("order_id")) != ""))
.select("order_id", "customer_id", "amount", "order_date")
)
clean_orders.write.mode("overwrite").parquet(
"s3://YOUR_BUCKET/processed/orders/"
)
job.commit()
The job reads the CSV header, removes rows with a missing order ID, keeps the four selected columns, and writes Parquet output. Spark writes one or more part files under the output prefix; it does not produce one file named orders.parquet. This demonstration uses overwrite, so each full-input run replaces the contents of the output prefix. Test with a disposable destination and choose a deliberate write strategy for production.
Run the job and inspect the result
- Save the job and start a run in AWS Glue Studio.
- Wait for the run status to become successful, then review its duration and logs if it fails.
- Open the output prefix in Amazon S3 and confirm that Parquet part files were written.
Job logs and run metrics help you find permission, parsing, and transformation errors. A successful job status confirms that the script completed; it does not verify that the output satisfies your business rules.
Crawlers, bookmarks, and scheduling
What job bookmarks track
For supported AWS Glue Spark sources, job bookmarks can track input that the job has already processed. They require the bookmark setting to be enabled and source state to be configured as documented; custom scripts use job.init, job.commit, and a stable transformation_ctx for the source. This example includes the initialization and source context, but leave bookmarks disabled for its full-refresh, overwrite behavior.
Bookmarks track source progress; they do not make a destination write idempotent or prevent duplicate output rows. Their behavior depends on the source: for example, S3 bookmarks use object modification times and may process a changed object again, while JDBC bookmarks rely on configured key columns and do not detect changes to previously read rows. Python shell jobs do not support job bookmarks. Read AWS's job bookmark documentation before using them, and design your target writes for retries or replayed input.
Start jobs on demand or with triggers
Start a job manually while learning. For repeat runs, a Glue trigger can start a job on demand, on a cron schedule, or after configured job or crawler events. Glue cron schedules use UTC. A new S3 object does not automatically start this job; for object-arrival processing, configure an EventBridge event and a Glue workflow. EventBridge messages can be duplicated, so make the workflow safe to retry. AWS documents Glue trigger types, schedule timing, and the EventBridge workflow pattern.
Understand cost and choose the right tool
AWS Glue charges for job runtime and configured processing capacity; crawlers also incur charges, and Data Catalog pricing depends on usage. Storage, requests, and other services such as S3 or Redshift are billed separately. Pricing varies by Region and usage, so check the current AWS Glue pricing page before running larger or frequent jobs.
Use Glue when you need a managed, repeatable job to transform or move data. If your main task is to query data in S3 with SQL, Amazon Athena can query it in place using the Data Catalog. If your destination is an Amazon Redshift data warehouse, Glue can prepare and load the data as one part of that pipeline; see our practical guide to automating Redshift ETL with AWS Glue. These services have different roles: Glue processes data, S3 stores objects, Athena queries data, and Redshift is a data warehouse.
Conclusion: Start with one small Glue job
Start with a small input prefix, a job role limited to that input and a separate output prefix, and a manual Spark job run. Add a crawler when metadata discovery helps, and add schedules, bookmarks, or event-driven workflows only after the basic read-transform-write path behaves as you expect. Check the output data as well as the job status before using it downstream.