S3 Multipart Uploads and Object Consistency

Updated on
9 min read

Large objects are difficult to transfer reliably: a single interrupted request may force an application to start again, while parallel connections can saturate a network or overwhelm a client. S3 multipart uploads address that transfer problem, but they also introduce upload IDs, part tracking, completion and cleanup steps. Knowing how those pieces interact with object consistency helps storage operators build reliable media pipelines, backups, and data ingestion jobs.

What Are S3 Multipart Uploads?

An S3 multipart upload is a protocol for creating one object from several separately transferred parts. A client starts an upload, sends numbered parts to the service, and asks the service to assemble them when it has all the required part identifiers and entity tags. The resulting object has one key and is read like any other object; its parts are not separate files in the bucket.

The Amazon S3 multipart upload documentation describes the initiate, upload, and complete steps, along with aborting an upload that will not finish. The Amazon S3 service overview explains the broader object-storage service that exposes this API. Other systems that implement the S3 API may differ in supported features, limits, consistency guarantees, and cleanup behavior, so check the selected provider’s documentation rather than assuming every implementation behaves identically.

Why Does Multipart Upload Exist?

A one-request upload is easy to reason about, but its failure boundary is the entire object. If a connection drops near the end of a multi-gigabyte transfer, an application may have to resend data that already crossed the network. A single request can also be constrained by one stream, network path, or client process.

Multipart upload changes the unit of work from one huge request to independently tracked pieces. Clients can send several parts concurrently, retry a failed part, and resume work while the upload remains active. This is especially useful when the object is larger than the cost of coordinating multiple requests.

The trade-off is more state. The client must keep the upload ID, part numbers, and each successful part’s entity tag (ETag). It must complete the upload with the right ordered part list or explicitly abort it. Until completion succeeds, the uploaded parts remain an unfinished upload, not a normal object available at the destination key.

Concern Single-request upload Multipart upload
Transfer unit One request carries the object Numbered parts are separate requests
Retry scope Usually the entire request A failed part can be retried
Parallel transfer Not typically part of the basic request Parts can be uploaded concurrently
Client state Object key and request status Upload ID, part numbers, ETags, and completion state
Visibility at the object key Object appears after a successful write Parts stay hidden until successful completion
Best fit Small objects and simple clients Large objects, unreliable links, and resumable workflows

How Multipart Uploads and Object Consistency Work

An upload has three main phases:

  1. Initiate: The client requests a new multipart upload for a bucket and key. The service returns an upload ID that identifies this in-progress operation.
  2. Upload parts: The client sends numbered parts associated with that upload ID. Parts can be sent in parallel and retried. Reusing a part number replaces that part within the same upload, so a retry must contain the intended bytes.
  3. Complete or abort: To complete, the client submits the part numbers and ETags it recorded. The service assembles the object and makes the completed object available. To abandon the work, the client aborts the upload; this removes its stored parts.

For AWS S3, part numbers range from 1 to 10,000. Each part must be at least 5 MiB except the final part, and a part cannot exceed 5 GiB. AWS lists the maximum object size as 48.8 TiB in its multipart upload quotas. These are service-specific limits, not rules guaranteed by the S3-compatible API as a whole. Select a part size that stays within the part-count limit while keeping retries manageable: tiny parts multiply request overhead, and very large parts make each retry expensive.

Consistency concerns what a reader sees at the object key. AWS documents strong read-after-write consistency for object PUT and DELETE operations, including overwrites, in its Amazon S3 consistency guidance. Before a multipart upload is complete, a GET for a new key does not expose a partially assembled object. After a successful completion, reads and listings reflect the completed write according to the service’s consistency contract. For an overwrite, readers see an old or new complete object, not an intermediate sequence of uploaded parts.

That behavior is not a transaction across several object keys. A workflow that replaces a manifest and multiple data objects still needs its own ordering, versioning, or commit-marker design. Nor should an AWS guarantee be generalized to every S3-compatible server, gateway, cache, or downstream replication target.

Retry behavior also needs care. The IETF’s HTTP Semantics specification defines idempotent methods as having the same intended effect when repeated. That is a useful lens for designing clients, not a blanket statement that every S3 API call is safe to repeat. Retrying a part with the same upload ID, part number, and bytes is different from blindly initiating a second upload or repeating an ambiguous completion. When a response is lost, inspect the upload or destination state before starting over.

Key Components and Integrity Concepts

  • Upload ID: The identifier returned at initiation. It scopes part uploads, part listings, completion, and abort operations. Persist it durably if the client must resume after restarting.
  • Part number: An integer that marks a part’s position. A later upload to the same upload ID and number replaces the earlier part.
  • ETag: A response value the client supplies for each part when completing. For multipart objects, the final ETag is not generally the MD5 digest of the full object. Do not use it alone as proof that the downloaded bytes match the source.
  • Checksum: A checksum verifies data integrity. Services and SDKs can support checksums for individual parts or the complete object, and may report a full-object or composite checksum. Confirm which algorithm and checksum type your provider supports; compare like with like.
  • Incomplete-upload storage: Parts belonging to an unfinished upload consume storage until completion or abort. A lifecycle rule that aborts old incomplete multipart uploads is a useful safety net, but it should allow legitimate long-running transfers enough time.

Completion errors can be ambiguous if the client loses its connection after the service has committed the object. Before initiating another upload, check whether the key exists and compare its expected size and checksum, then inspect the upload state if the service still reports it. A missing upload ID may mean it was completed or aborted; it is not sufficient evidence by itself that the data is correct.

Real-World Use Cases

Media platforms upload large recordings and video files from browsers or mobile clients, where connections can drop and individual parts can be retried. Backup systems use multipart transfers to move large archives to object storage without restarting from byte zero after a transient network failure. Data pipelines can transfer model checkpoints, analytics exports, and large scientific files with parallel workers.

Multipart upload improves the transfer path, not the whole storage design. It does not make an upload a backup, protect against deletion, verify that an application produced the right source file, or create atomic updates across a collection of keys. For broader object-storage concepts and architecture, see the object storage implementation guide. For deployment and client configuration on an S3-compatible server, see the S3-compatible storage server guide.

Getting Started and Verifying an Upload

For a first transfer, use the AWS CLI high-level s3 cp command. It manages multipart initiation, parallel part transfers, retries, and completion when the object exceeds the configured multipart threshold. The example assumes credentials and a region are already configured, and that the bucket exists:

aws s3 cp ./large-archive.tar s3://archive-prod/backups/large-archive.tar

The AWS CLI’s transfer settings can be adjusted in the selected profile’s configuration file when a workload needs different part sizes or concurrency. For example:

[profile archive]
s3 =
  multipart_threshold = 64MB
  multipart_chunksize = 32MB

Use a part size large enough to avoid exceeding 10,000 parts for the largest expected object. Test settings with realistic network, disk, and memory limits: concurrency can increase throughput, but each active transfer also uses connections and buffers.

After the command succeeds, inspect object metadata and request checksum information where supported:

aws s3api head-object \
  --bucket archive-prod \
  --key backups/large-archive.tar \
  --checksum-mode ENABLED \
  --query '{Size:ContentLength,ETag:ETag,ChecksumSHA256:ChecksumSHA256,ChecksumType:ChecksumType}'

Compare the returned size with the local file and compare checksums only when the source and destination use the same algorithm and checksum type. An ETag that looks like a hash should not be assumed to be a full-object checksum.

If you use a low-level SDK or API instead of aws s3 cp, persist the upload ID and each completed part’s number and ETag. On failure, list the parts for that upload and retry only missing or damaged parts. Once it has been verified that an upload is abandoned, abort it:

aws s3api abort-multipart-upload \
  --bucket archive-prod \
  --key backups/large-archive.tar \
  --upload-id 'UPLOAD_ID_FROM_INITIATE'

Use the actual upload ID returned during initiation; do not abort an upload another worker owns. At the bucket level, configure an incomplete-multipart-upload lifecycle rule for abandoned transfers and monitor it alongside failed upload counts. To connect multipart recovery to a complete retention and restore plan, read the backup strategy best practices guide.

Common Misconceptions

“A multipart upload is visible as a partially written object.” No. Uploaded parts are associated with an upload ID. The destination object becomes available when the multipart operation completes successfully.

“An ETag is always an MD5 checksum.” No. A multipart ETag is generally not the MD5 of the full object. Use an explicitly supported checksum and validate its algorithm and scope.

“Multipart makes every upload faster and is portable across providers.” No. Extra requests and concurrency can slow small transfers or strain a client. Compatible services may have different limits, checksum support, completion rules, or consistency guarantees; test the exact provider and SDK combination you plan to operate.

Changelog

  • 2026-10-08: Published the canonical explainer for S3 multipart uploads and object consistency.
TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.