Amazon Redshift Tutorial for Beginners: Serverless First Steps
Published on · Updated on
Rebuilt the tutorial around Redshift Serverless, added SQL and S3 loading examples, and corrected pricing, permissions, table constraints, and cleanup guidance.
Amazon Redshift is a data warehouse for SQL analytics: combining records, calculating totals, and preparing reports across large datasets. This tutorial uses Redshift Serverless to create a small learning environment, query four sample sales, and optionally load the same data from Amazon S3.
You need an AWS account, permission to create Redshift resources, and basic SQL familiarity. The examples use disposable tutorial data. Creating resources and running queries can incur charges; review the cost controls before starting.
Understand what Redshift does
Redshift stores table data by column and distributes analytical work across compute resources. That suits queries such as total revenue by region. It serves a different purpose from a transactional database handling individual orders. Although its SQL has PostgreSQL roots, it is not a drop-in PostgreSQL replacement. AWS explains these differences in Redshift and PostgreSQL.
Redshift can also transform data with SQL, including joins, aggregations, and table-building queries. You can transform before loading with an ETL tool or load first and transform inside the warehouse, an approach called ELT. For a separate transformation pipeline, see our practical guide to Redshift ETL with AWS Glue.
Serverless or a provisioned cluster?
| Option | What you configure | Why choose it? |
|---|---|---|
| Redshift Serverless | A namespace, a workgroup, networking, permissions, and capacity limits. | You want to start without selecting a cluster node type and count. |
| Provisioned Redshift | A cluster with supported node types and counts, plus workload and network settings. | You want explicit cluster capacity and can evaluate its cost against your workload. |
Both run Redshift SQL. Serverless is the path used here; it is not limited to querying S3 files. A namespace holds database objects, users, and associated storage settings. Its associated workgroup holds compute and network settings. AWS documents their relationship in workgroups and namespaces.
Prepare permissions and cost controls
Sign in through an IAM role or your organization's identity provider. Ask your administrator for access to create the learning environment, use query editor v2, and obtain the chosen database credentials. IAM permissions to operate AWS resources and SQL privileges inside the database are separate. For this exercise, your database identity must be able to create a schema and tables and query them. See AWS's query editor v2 account setup.
On-demand Serverless compute is charged for capacity used over time, in RPU-hours, metered per second with a 60-second minimum. An RPU is a Redshift Processing Unit. This is not a fixed charge per query. Managed storage, retained manual snapshots, and applicable transfer charges are separate. Check the Serverless billing guide and current account eligibility before relying on trial credits.
Choose a base capacity appropriate to a small exercise and set a maximum compute capacity. Also configure a daily RPU-hour usage limit; select Turn off user queries if you want that limit to disable further user queries. Logging and alert actions only notify you. These controls do not cap storage or every AWS charge. The usage-limit documentation explains the available actions.
Step 1: Create the Serverless environment
- Open the Amazon Redshift console and select a Region where Serverless is available.
- Open the Serverless creation workflow. Create a workgroup named
tutorial-workgroupand a namespace namedtutorial-namespace, or use clearly identified disposable resources your administrator provides. - Use the
devdatabase for this exercise. Review database authentication and encryption settings. - Choose an approved VPC, subnets, and security groups. Check the subnet, Availability Zone, and available IP requirements for your selected configuration. Keep public access disabled for this console-based tutorial.
- Review capacity and usage limits, then create the environment. Wait until the workgroup is available.
Console labels and defaults can change. Follow AWS's workgroup creation instructions for the current screens. There is no need to expose a database port to the internet for this exercise. If you later connect an external JDBC or ODBC client, plan its network path and scoped security-group rules separately.
Step 2: Connect and run SQL
Choose Query data to open query editor v2. Select your workgroup and connect using the authentication method your administrator configured, such as federated access or database credentials. Select the dev database before running statements. AWS's Serverless getting-started guide also shows this connection flow and built-in sample notebooks.
Create a dedicated schema and table. Run this setup once in your disposable environment:
CREATE SCHEMA tutorial;
CREATE TABLE tutorial.sales (
order_id INTEGER NOT NULL,
sold_on DATE NOT NULL,
region VARCHAR(30) NOT NULL,
amount DECIMAL(12,2) NOT NULL
)
DISTSTYLE AUTO
SORTKEY AUTO
ENCODE AUTO;
INSERT INTO tutorial.sales VALUES
(1, '2026-10-01', 'East', 120.50),
(2, '2026-10-01', 'West', 75.00),
(3, '2026-10-02', 'East', 45.25),
(4, '2026-10-02', 'West', 200.00);
The AUTO options let Redshift manage distribution, sorting, and column compression rather than requiring you to choose them for four rows. See CREATE TABLE for their syntax. Inserting a few rows is convenient for learning; bulk loads should use COPY.
Summarize the sample:
SELECT region,
COUNT(*) AS order_count,
SUM(amount) AS revenue
FROM tutorial.sales
GROUP BY region
ORDER BY region;
| region | order_count | revenue |
|---|---|---|
| East | 2 | 165.75 |
| West | 2 | 275.00 |
These are the expected results from the four rows above. Rerunning the INSERT adds those rows again. Redshift enforces NOT NULL, but primary-key, foreign-key, and unique constraints are informational rather than enforced. Validate uniqueness in your ingestion process; do not declare an untrue key, because the query optimizer can rely on it. See AWS's table-constraint guidance.
Step 3: Load CSV data from S3
This optional step demonstrates the usual bulk-loading path. Save the following as sales.csv and upload it to a private S3 bucket in the same Region as your workgroup, under redshift-tutorial/input/:
order_id,sold_on,region,amount
1,2026-10-01,East,120.50
2,2026-10-01,West,75.00
3,2026-10-02,East,45.25
4,2026-10-02,West,200.00
Associate an IAM role with your Serverless namespace and make it the default role. Use the Redshift role-creation workflow or an administrator-managed role with the correct service trust policy. Grant it s3:ListBucket for the source bucket, limited to the input prefix, and s3:GetObject for that prefix's objects. SSE-KMS objects also require appropriate KMS decrypt authorization. Your permission to run SQL does not automatically give Redshift permission to read S3. Follow AWS's IAM role guidance for COPY; our S3 access management guide explains bucket and object permission scopes.
Replace YOUR-BUCKET below. Create a separate table so the S3 exercise does not duplicate the earlier rows:
CREATE TABLE tutorial.sales_from_s3 (LIKE tutorial.sales);
COPY tutorial.sales_from_s3
FROM 's3://YOUR-BUCKET/redshift-tutorial/input/sales.csv'
IAM_ROLE default
FORMAT AS CSV
IGNOREHEADER 1;
SELECT COUNT(*) AS rows_loaded, SUM(amount) AS revenue
FROM tutorial.sales_from_s3;
Expect four rows and revenue of 440.75. COPY appends data, so rerunning it is not an upsert. An S3 path can match a prefix: keep this input location free of other files whose keys begin with sales.csv, or use a manifest to identify exact objects. Consult the COPY reference before adapting the example to other formats or Regions.
If enhanced VPC routing is enabled, also configure a route to S3, such as a same-Region S3 gateway endpoint, and check its endpoint policy. AWS explains the network requirements in enhanced VPC routing.
If a load fails, check the namespace's role association, S3 and KMS permissions, network path, input location, delimiter, and column types. For parsing failures, inspect the SYS_LOAD_ERROR_DETAIL view. Correct the source data or schema rather than skipping errors without understanding them.
Monitor before tuning
Use the Serverless dashboard and query history to inspect query duration, queueing, compute usage, and failures. The Serverless monitoring guide lists the relevant views and metrics. Four rows cannot tell you whether a warehouse configuration will suit a production workload; measure representative data and query concurrency before changing capacity.
Start with automatic table optimization and inspect actual query behavior before assigning manual sort or distribution keys. Redshift does not use PostgreSQL-style secondary indexes. When performance becomes a concern, examine scans, joins, data distribution, and statistics rather than trying to add indexes.
Step 4: Clean up the tutorial
- Save any SQL or results you want to keep. If you used a shared warehouse, remove only the tutorial objects you created, with the owner's approval.
- For a disposable environment you own, delete the workgroup, then delete the namespace. Decide whether a final snapshot is needed before deleting data.
- Remove unneeded manual snapshots, uploaded S3 objects, and any bucket or IAM role created solely for this tutorial. Retained snapshots and S3 objects can continue to incur charges.
- Confirm the resources are gone and review usage after billing data updates. Closing query editor v2 does not remove your warehouse or its stored data.
Build on the first exercise
You now have a path from a table definition to validated query results, with a separate S3 loading exercise. Next, try filtering by date and joining sales to a small lookup table. Keep row counts and business totals as checks whenever you change a load.
For queries over files kept in S3, external tables describe supported data formats and locations; they do not make arbitrary unstructured files directly queryable. AWS explains the setup in external tables for Redshift Spectrum. Federated queries are a different feature for supported RDS and Aurora PostgreSQL or MySQL databases; see the federated query overview.
If another team needs live warehouse data without building a second copy, continue with our Redshift data sharing setup and permissions guide.