Working With Aws From Python A Boto3 Crash Course
You’ve been manually downloading Parquet files from the AWS web console for every training run. Click. Wait. Download. Move to S3 bucket. Click again. It’s Tuesday morning, your model just failed because you grabbed yesterday’s CSV instead of today’s, and you’re staring at a browser tab full of folders you didn’t create. You know there has to be a better way.
There is. AWS S3 is the de facto data lake for the modern data stack, and boto3 is your Python passport. By the end of this article, you’ll upload, download, list, and load data into pandas — all from Python, no browser required.
Let’s start with the basics: what is S3? Think of a bucket as a folder that lives on Amazon’s computers instead of yours. Each bucket has a globally unique name — no two AWS accounts in the world can have a bucket with the same name. Inside a bucket, you have objects (files) organized by keys (paths). That’s it. One bucket, many objects, all accessible via Python.
Throughout this article, we’ll use a single running example: a small CSV of NYC taxi trip data. You’ll upload it, download it, list it, and load it into pandas. By the end, you’ll have a complete data pipeline that never touches a web browser.
The Two Faces of boto3: Resource vs. Client
Most frustration with boto3 comes from not understanding its two personalities. The Resource API feels like regular Python objects — buckets have objects, objects have methods. The Client API feels like command-line calls — every S3 API action has an exact method.
Think of it this way: Resource is like a rental car — turn the key and go. Client is the engine block — you can tweak everything. One is easier to read, one gives you full control.
Here’s the catch: resources aren’t available for every AWS service. S3 has a resource, but newer services like SQS or Step Functions don’t. When in doubt, use the client — it always works.
Let’s see both in action:
import boto3
# Resource API (object-oriented)
s3_resource = boto3.resource('s3')
# List buckets using the resource
print("Resource API:")
for bucket in s3_resource.buckets.all():
print(f" - {bucket.name}")
# Client API (low-level, maps to REST API)
s3_client = boto3.client('s3')
# List buckets using the client
print("\nClient API:")
response = s3_client.list_buckets()
for bucket in response['Buckets']:
print(f" - {bucket['Name']}")
# They're connected: you can reach the client from the resource
print("\nThey share the same underlying connection:")
print(type(s3_resource.meta.client)) # <class 'botocore.client.S3'>
When you run this, you’ll see your bucket names printed twice — once from the resource, once from the client. The resource is a thin wrapper around the client. If you need fine-grained control (like setting request headers), use the client. If you want readable code, use the resource.
Credentials, Regions, and the First ‘Hello, S3’ Call
Before you can write to S3, boto3 needs to know who you are and where to put things. This is where beginners panic — the credentials dance. But the logic is simple: a structured file that boto3 reads automatically.
Boto3 looks for credentials in this order (from Boto3 Configuration docs):
- Explicit parameters when creating a client/resource
- Environment variables (
AWS_ACCESS_KEY_ID,AWS_SECRET_ACCESS_KEY,AWS_DEFAULT_REGION) - Shared credential file (
~/.aws/credentials) - Config file (
~/.aws/config) - IAM role (if running on EC2 or Lambda)
For local development, the credential file is simplest. Here’s what it looks like:
# ~/.aws/credentials
[default]
aws_access_key_id = YOUR_ACCESS_KEY
aws_secret_access_key = YOUR_SECRET_KEY
And the config file:
# ~/.aws/config
[default]
region = us-east-1
You can create these files with the AWS CLI (aws configure) or manually with a text editor. The region matters — your bucket lives in a specific region, and your code must match it.
Now, the first real call:
import boto3
# Create the S3 resource (reads credentials from ~/.aws/credentials)
s3 = boto3.resource('s3')
# List all buckets
print("Your buckets:")
for bucket in s3.buckets.all():
print(f" - {bucket.name}")
If this code returns nothing, come back and check your region matches your bucket’s region. The most common beginner mistake is having credentials in one region and buckets in another.
Uploading Your First File — and Choosing the Right Method
Uploading a file to S3 is the data scientist’s equivalent of copy-pasting into a cloud folder. Three methods exist, and choosing the right one matters:
upload_file()— for files on disk. Handles large files with multipart upload automatically.put_object()— for small objects (under 5GB). Maps directly to the S3 API, no multipart.upload_fileobj()— for data still in memory (like a DataFrame). We’ll use this later.
Let’s start with upload_file():
import boto3
from botocore.exceptions import ClientError
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket' # Replace with your bucket
local_file = 'nyc_taxi_sample.csv'
s3_key = 'raw/nyc_taxi_sample.csv'
try:
s3.upload_file(local_file, bucket_name, s3_key)
print(f"Successfully uploaded {local_file} to s3://{bucket_name}/{s3_key}")
except ClientError as e:
print(f"Upload failed: {e}")
Now compare with put_object() for a string (like a JSON config):
import boto3
import json
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket'
s3_key = 'config/model_params.json'
config_data = {
'learning_rate': 0.01,
'max_depth': 5,
'n_estimators': 100
}
# Convert dict to JSON string, then to bytes
config_bytes = json.dumps(config_data).encode('utf-8')
s3.put_object(
Bucket=bucket_name,
Key=s3_key,
Body=config_bytes,
Metadata={'data': 'science'} # Custom metadata
)
print(f"Uploaded config to s3://{bucket_name}/{s3_key}")
The key difference: upload_file() handles large files automatically (multipart upload for files over 5GB). put_object() gives you fine-grained control but maxes out at 5GB. For most data science work, upload_file() is your go-to.
Downloading Files — and Reading Them Straight Into pandas
Downloading a file is the mirror operation. But here’s the magic: you can stream the data straight into pandas without saving to disk first.
First, the simple disk-based approach:
import boto3
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket'
s3_key = 'raw/nyc_taxi_sample.csv'
local_file = 'downloaded_nyc_taxi.csv'
s3.download_file(bucket_name, s3_key, local_file)
print(f"Downloaded to {local_file}")
But what if you’re on a server with no disk space? Or you want to avoid writing temp files? Here’s the streaming approach:
import boto3
import pandas as pd
import io
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket'
s3_key = 'raw/nyc_taxi_sample.csv'
# Get the object as a stream
response = s3.get_object(Bucket=bucket_name, Key=s3_key)
content = response['Body'].read()
# Read into pandas using BytesIO
buffer = io.BytesIO(content)
df = pd.read_csv(buffer)
print(df.head())
print(f"\nLoaded {len(df)} rows into memory")
Let’s break down what’s happening:
s3.get_object()returns a dictionary with aBodykey — that’s the file content as a stream..read()reads the entire stream into bytes.io.BytesIO(content)wraps those bytes in a file-like object that pandas can read.pd.read_csv(buffer)reads the CSV directly from memory.
Now your data is in a DataFrame, never written to disk. This pattern is essential for serverless environments where /tmp is limited.
Listing More Than 1000 Objects — the Paginator Pattern
S3 returns at most 1000 keys per listing call. If you have 10,000 CSV files, you’ll think boto3 is broken. It isn’t — you need a paginator.
Think of a paginator as a for loop over the API responses, fetching the next page each time. Here’s how it works:
import boto3
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket'
prefix = 'processed/' # Only list objects under this prefix
# Create a paginator
paginator = s3.get_paginator('list_objects_v2')
# Paginate through all objects
print("All objects under 'processed/':")
for page in paginator.paginate(Bucket=bucket_name, Prefix=prefix):
if 'Contents' in page:
for obj in page['Contents']:
print(f" - {obj['Key']} ({obj['Size']} bytes)")
But here’s where it gets really useful — you can filter with JMESPath, a query language for JSON. Let’s find all objects larger than 5MB:
import boto3
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket'
paginator = s3.get_paginator('list_objects_v2')
# Use JMESPath to filter objects larger than 5MB
large_files = paginator.paginate(
Bucket=bucket_name
).search('Contents[?Size > `5000000`]')
print("Objects larger than 5MB:")
for obj in large_files:
if obj: # search can return None
print(f" - {obj['Key']} ({obj['Size'] / 1024 / 1024:.2f} MB)")
The .search() method applies a JMESPath expression to each page. The expression Contents[?Size > 5000000] filters the list to only objects where Size is greater than 5,000,000 bytes. This returns only the large files, skipping the rest — much more efficient than filtering in Python after fetching everything.
Error Handling: Don’t Just Catch Exception
New boto3 users throw broad except Exception and don’t distinguish between ‘file not found’ (no such key) and ‘you’re not allowed to see this’ (bad permissions). Let’s parse the error response correctly.
Boto3 has two types of errors:
- Static errors: caught before the request is sent (e.g.,
ParamValidationError,NoCredentialsError) - Dynamic errors: returned by AWS (always a
ClientError)
Here’s how to handle them properly:
import boto3
from botocore.exceptions import ClientError, NoCredentialsError, ParamValidationError
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket'
s3_key = 'nonexistent-file.csv'
try:
response = s3.get_object(Bucket=bucket_name, Key=s3_key)
print("File found!")
except ClientError as e:
error_code = e.response['Error']['Code']
error_message = e.response['Error']['Message']
if error_code == 'NoSuchKey':
print(f"The key '{s3_key}' does not exist in this bucket.")
elif error_code == 'AccessDenied':
print(f"You don't have permission to access this object.")
elif error_code == 'NoSuchBucket':
print(f"The bucket '{bucket_name}' does not exist.")
else:
print(f"AWS error {error_code}: {error_message}")
except NoCredentialsError:
print("No AWS credentials found. Check ~/.aws/credentials")
except ParamValidationError as e:
print(f"Invalid parameters: {e}")
A generic catch hides the difference between a missing file and a permission problem — and wastes hours of debugging. Always check the specific error code.
For newer boto3 versions (>= 1.26), you can use the even cleaner pattern:
import boto3
s3 = boto3.client('s3')
try:
s3.get_object(Bucket='my-bucket', Key='missing.csv')
except s3.exceptions.NoSuchKey:
print("Key not found — specific exception!")
except s3.exceptions.NoSuchBucket:
print("Bucket not found!")
This is cleaner but not available for all services yet. Stick with the ClientError pattern for maximum compatibility.
The Data Scientist’s Workflow: Uploading a DataFrame Directly
This is the most common data science task: write a DataFrame to S3 without an intermediate file. The BytesIO dance feels weird at first, but it’s the only way to do it with plain boto3.
Here’s the pattern:
import boto3
import pandas as pd
import io
# Create a sample DataFrame
df = pd.DataFrame({
'trip_id': [1, 2, 3],
'pickup_datetime': ['2024-01-01', '2024-01-02', '2024-01-03'],
'passenger_count': [1, 2, 3],
'trip_distance': [2.5, 5.1, 3.8],
'fare_amount': [12.50, 25.30, 18.90]
})
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket'
s3_key = 'processed/taxi_trips.csv'
# Step 1: Create a buffer
buffer = io.BytesIO()
# Step 2: Write DataFrame to the buffer
df.to_csv(buffer, index=False)
# Step 3: Reset the buffer pointer to the start
buffer.seek(0)
# Step 4: Upload the buffer
s3.upload_fileobj(buffer, bucket_name, s3_key)
print(f"DataFrame uploaded to s3://{bucket_name}/{s3_key}")
The critical line is buffer.seek(0). After df.to_csv(buffer), the buffer’s internal pointer is at the end of the data. When boto3 tries to read from the buffer, it finds nothing — the pointer is past the end. seek(0) resets it to the beginning, so boto3 reads the entire CSV.
Now let’s do the same with Parquet — the format data scientists love for its speed and compression:
import boto3
import pandas as pd
import io
# Same sample DataFrame
df = pd.DataFrame({
'trip_id': [1, 2, 3],
'pickup_datetime': ['2024-01-01', '2024-01-02', '2024-01-03'],
'passenger_count': [1, 2, 3],
'trip_distance': [2.5, 5.1, 3.8],
'fare_amount': [12.50, 25.30, 18.90]
})
s3 = boto3.client('s3')
bucket_name = 'my-data-science-bucket'
s3_key = 'processed/taxi_trips.parquet'
buffer = io.BytesIO()
df.to_parquet(buffer, index=False)
buffer.seek(0)
s3.upload_fileobj(buffer, bucket_name, s3_key)
print(f"Parquet DataFrame uploaded to s3://{bucket_name}/{s3_key}")
You’ve never written a byte to disk. Your DataFrame went from memory directly to S3. This pattern is the foundation of every production data pipeline.
Beyond Raw boto3: When You Outgrow the Boilerplate
The BytesIO pattern works, but you won’t want to write it 50 times. Once you master the raw boto3 boilerplate, you’ll wish it were simpler. Enter awswrangler (AWS SDK for pandas) — it wraps all this into one-liners.
import awswrangler as wr
# Read a Parquet file from S3 — one line, no buffer, no seek
df = wr.s3.read_parquet('s3://my-data-science-bucket/processed/taxi_trips.parquet')
print(df.head())
# Write a DataFrame to S3 — one line
wr.s3.to_csv(
df,
's3://my-data-science-bucket/processed/taxi_trips.csv',
index=False
)
print("Written with awswrangler!")
Under the hood, awswrangler calls boto3 — you just don’t see the boilerplate. It handles the BytesIO buffer, the seek(0), and the upload for you. The trade-off? An extra dependency, but it cuts your code by 80%.
Another alternative is s3fs, which lets you use pandas’ built-in S3 support:
import pandas as pd
# Requires s3fs to be installed: pip install s3fs
df = pd.read_csv('s3://my-data-science-bucket/processed/taxi_trips.csv')
print(df.head())
Both awswrangler and s3fs are built on boto3. They don’t replace it — they simplify it. Start with raw boto3 to understand what’s happening, then graduate to these wrappers when you’re tired of the boilerplate.
You Now Speak boto3
Let’s recap the five key moves you can now make from Python:
- Configure credentials: Create a
~/.aws/credentialsfile or set environment variables. - Upload a file: Use
upload_file()for disk files,upload_fileobj()for in-memory data. - Download into pandas: Use
get_object()+BytesIOto stream directly into a DataFrame. - Paginate a list: Use
get_paginator('list_objects_v2')to list more than 1000 objects. - Catch specific errors: Check
error['Error']['Code']instead of catching generic exceptions.
Here’s a clean summary block of all five patterns:
import boto3
import pandas as pd
import io
# 1. Create client
s3 = boto3.client('s3')
# 2. Upload file
s3.upload_file('local.csv', 'my-bucket', 'data/local.csv')
# 3. Download into pandas
response = s3.get_object(Bucket='my-bucket', Key='data/local.csv')
df = pd.read_csv(io.BytesIO(response['Body'].read()))
# 4. List all objects
paginator = s3.get_paginator('list_objects_v2')
for page in paginator.paginate(Bucket='my-bucket'):
for obj in page.get('Contents', []):
print(obj['Key'])
# 5. Handle errors
from botocore.exceptions import ClientError
try:
s3.get_object(Bucket='my-bucket', Key='missing.csv')
except ClientError as e:
if e.response['Error']['Code'] == 'NoSuchKey':
print("File not found")
You didn’t need to read the AWS API docs — you just needed a running example and clear intuition. Your data pipeline just got a whole lot faster.
In the next part of this series, we’ll automate model artifact storage — saving trained models to S3 and loading them back with a single function call. See you there.
Check Your Understanding
Remember: What are the two faces of boto3, and when would you use each?
Understand: Explain why buffer.seek(0) is necessary after writing a DataFrame to a BytesIO buffer.
Apply: Write a function that takes a pandas DataFrame and an S3 path (like s3://bucket/key.parquet) and uploads it using the BytesIO pattern.
Analyze: Compare the error handling approach using ClientError vs. using s3.exceptions.NoSuchKey. What are the trade-offs?
Evaluate: When would you choose raw boto3 over awswrangler for a data pipeline? Consider factors like dependencies, code readability, and control.
Create: Design a Python class S3DataLake that wraps the five patterns above into methods: upload_df(), download_df(), list_files(), delete_file(), and exists(). The class should handle credentials and error cases gracefully.
Related articles
- Part 1: SQLAlchemy for Data Scientists — Query databases without raw SQL strings, then load results into pandas and upload to S3 using the patterns from this article.
- Python Data Science Toolbox series — The complete collection of practical data engineering patterns for data scientists.
Apply What You Learned is for Supporter and Insider subscribers.
Subscribe to unlock the exercises on this post.
See plansRelated articles
- Python Engineering Under review
SQLAlchemy for Data Scientists: Querying Databases Without Raw SQL Strings
Picture this: You're a data scientist, and you've just spent three hours building a beautiful data pipeline in a Jupyter notebook. You've got your SQL queries working perfectly in your local SQLite database.
- Python Engineering Under review
Refactoring A God Class Into Four Single Responsib
Before we touch anything, let’s look at what we’re dealing with. The `DataProcessor` class has four groups of methods, each doing a completely different job:
- Python Engineering Under review
Data Flow Decomposition Why Every Pandas Sklearn P
Here's a realistic pandas pipeline. It loads customer transaction data, cleans it, groups it, and merges it with customer info. You've written something like this before:
- Python Engineering Under review
Test-First Design: Writing the Assertion Before the Function Exists
You've just finished writing a complex function. Maybe it's a causal inference estimator. Maybe it's a simulation that generates synthetic data. You run it. It doesn't crash.
Looking for something else?
Search every article by title, summary or topic.