Container Image Lazy Loading and Remote Snapshotters
Container image lazy loading lets a container start before every byte of its image has been downloaded and unpacked. A remote snapshotter presents an image as a usable filesystem while fetching missing file data from a registry when the workload needs it. This can reduce the time between scheduling a container and starting its process, especially on newly provisioned nodes. It does not make the image smaller or remove the cost of reading data that the application eventually uses.
What Is Container Image Lazy Loading?
A container image is a set of metadata and filesystem layers. The Open Container Initiative (OCI) Image Specification defines the standard image manifest, configuration, and layer descriptors. At runtime, a snapshotter turns those layers into the filesystem view used by a container. In containerd, for example, snapshotters are plugins that manage container filesystem snapshots; the containerd snapshotter documentation describes the built-in and external options.
With a conventional local snapshotter such as OverlayFS, the runtime generally downloads image layers and prepares them before the container can use the root filesystem. A lazy-loading snapshotter can instead mount a filesystem backed partly by remote data. It reads metadata needed to locate files, then fetches their data as files are opened or read. A local cache keeps fetched data available for later access.
The term remote snapshotter describes this runtime-side filesystem integration, not a new image registry protocol by itself. Implementations differ in their image format, index metadata, cache behavior, and runtime setup. Lazy loading must be supported by both the image or its associated metadata and the snapshotter configured on the host.
Why Lazy Loading Exists
Large images can make startup slow even when an application initially needs only a small fraction of their contents. A node may spend time downloading compressed layers, verifying them, and unpacking files before a process can start. That delay compounds during a rollout, a burst of autoscaled workloads, or a job that runs on a fresh worker and exits soon afterward.
Lazy loading changes when image bytes are transferred. The runtime obtains enough metadata to mount the filesystem and start the process, while the snapshotter fetches file data on demand. A workload that reads only a small portion of a large image may avoid transferring and unpacking unused files altogether. A workload that eventually reads everything still needs that data, so the total transfer may be similar or even slower if remote reads add latency.
This is particularly useful when reducing time to first process matters more than maximizing steady-state local I/O. It complements node provisioning rather than replacing it: even after a node joins a Kubernetes cluster, pulling a large image can delay Pod readiness. For the separate capacity side of this startup path, see Kubernetes cluster autoscaling and node provisioning.
How Remote Snapshotters Work
An image registry stores manifests and content-addressed blobs. The OCI format gives runtimes consistent ways to identify image content, but it does not require all runtimes to fetch a layer lazily. A remote snapshotter adds the indexes, filesystem implementation, and network reads needed to provide that behavior.
The flow is generally:
- The runtime resolves the image manifest and platform-specific layers from the registry.
- The snapshotter obtains metadata or an index that maps filesystem paths to compressed data ranges or chunks.
- It mounts a filesystem view and makes the container root filesystem available without first materializing every file locally.
- When a process reads a file, the snapshotter checks its local cache. On a miss, it requests the relevant data from the registry and verifies it before exposing it to the container.
- The fetched data remains cached according to the snapshotter’s policy, so later reads or container starts can reuse it.
An index is essential because compressed image layers are not generally random-access filesystems on their own. The snapshotter needs to locate the bytes for a requested path without downloading and unpacking the entire layer first. Some approaches make a prepared, seekable image layer; others attach or publish index metadata associated with an otherwise standard OCI image.
The runtime, registry, and network all remain part of the critical path. Remote reads can depend on registry support for the required requests, node credentials, DNS, TLS, and outbound network access. Registry compatibility is implementation-specific, so a deployment should test the exact registry and authentication flow rather than assuming every OCI-compatible endpoint behaves identically.
Components and Approaches
The default local snapshotter and remote snapshotters solve related but different problems. OverlayFS optimizes local filesystem composition; a remote snapshotter trades some remote I/O for earlier access to image contents.
| Feature | OverlayFS (local) | Stargz Snapshotter | SOCI Snapshotter |
|---|---|---|---|
| Data available at start | Image layers are downloaded and prepared locally | Metadata and prioritized files can be available before the full image | Indexed layers can be mounted before all their data is downloaded |
| Image metadata | Standard OCI layers | eStargz layers include data arranged for seeking and an index | SOCI index metadata describes locations within image layers |
| Preparation | No lazy-loading index step | Images are commonly converted or built as eStargz | A SOCI index is generated and published with the image |
| Reads after start | Local reads after preparation | Missing chunks are fetched remotely and cached | Missing indexed data is fetched remotely and cached |
| Main operational requirement | Local disk capacity and unpack time | Stargz plugin, compatible image, registry access | SOCI plugin, compatible index, registry support and access |
Stargz Snapshotter uses eStargz, a seekable image format with metadata for locating files and prioritizing important content. Its nerdctl lazy-pulling guide shows how to register the plugin and select it for a command. Existing OCI-compatible runtimes can still run eStargz images without lazy pulling, but only a runtime configured with the snapshotter gets the lazy behavior.
The SOCI Snapshotter project uses SOCI indexes and per-layer tables of contents (zTOCs) to locate data in compressed layers. Its indexing workflow and supported registry behavior can vary with the project version; consult its current setup and compatibility documentation before standardizing on it. A workload may use a mixture of indexed and non-indexed layers, so the presence of an index does not mean every byte in an image will be lazy-loaded.
Both approaches add runtime and artifact-management requirements. Teams should compare startup time and application latency using their own images and workload traces, not assume that one snapshotter or index type is fastest for every registry and filesystem access pattern.
Real-World Use Cases
Autoscaled application nodes can start containers sooner when images are large but processes initially read a small subset of files. This can help deployments and burst capacity become useful sooner, although readiness probes still have to wait for required application data and dependencies.
Short-lived CI workers and batch jobs may use only a small part of a broad build or analysis image before exiting. Avoiding unnecessary downloads can reduce wasted transfer on ephemeral hosts. If the job reads most of the image, a warm local cache or a smaller image may be a better solution.
Model-serving containers sometimes bundle large model files alongside a smaller service process. Lazy access can make the process available before all model data has arrived, but inference cannot use a weight shard until it is fetched. The result depends on which files the service opens during initialization and whether remote-read latency fits the serving target.
These are not substitutes for image hygiene. Removing unused packages, separating rarely used assets into distinct layers, and reusing a node-level cache can improve both ordinary pulls and lazy-loading behavior.
Getting Started with SOCI on Linux
The following example demonstrates the main SOCI workflow on a rootful Linux host running containerd and nerdctl. Install versions supported by the SOCI project and your distribution; the pinned release below is an example for Linux amd64. Use the matching release asset for a different architecture, and configure the snapshotter as a managed service for production rather than leaving it in a shell.
Download the SOCI CLI and plugin binaries from the project’s release assets:
VERSION=0.16.1
curl -fL "https://github.com/awslabs/soci-snapshotter/releases/download/v${VERSION}/soci-snapshotter-${VERSION}-linux-amd64.tar.gz" \
-o "soci-snapshotter-${VERSION}-linux-amd64.tar.gz"
sudo tar -C /usr/local/bin -xzf "soci-snapshotter-${VERSION}-linux-amd64.tar.gz" soci soci-snapshotter-grpc
sudo soci --help
Add the following to /etc/containerd/config.toml, merging it with the existing configuration rather than replacing the file. The proxy plugin entry registers SOCI’s gRPC socket; the exports supply the socket and snapshotter root used by current SOCI releases.
[proxy_plugins]
[proxy_plugins.soci]
type = "snapshot"
address = "/run/soci-snapshotter-grpc/soci-snapshotter-grpc.sock"
[proxy_plugins.soci.exports]
address = "/run/soci-snapshotter-grpc/soci-snapshotter-grpc.sock"
enable_remote_snapshot_annotations = "true"
root = "/var/lib/soci-snapshotter-grpc/"
Restart containerd, then start soci-snapshotter-grpc using a service unit configured for your host. For a foreground test, run the daemon in a second terminal and inspect the registered plugins:
sudo systemctl restart containerd
sudo ctr plugins ls | grep snapshotter
sudo nerdctl system info
To generate and publish an index, authenticate to a registry where you can push an image. Set REGISTRY to your own registry and repository; both the runtime and the SOCI CLI need working registry credentials. The following pattern uses the commands in the SOCI getting-started documentation:
SOURCE=docker.io/library/rabbitmq:latest
TARGET="${REGISTRY}/rabbitmq:latest"
sudo nerdctl pull "$SOURCE"
sudo soci convert "$SOURCE" "$TARGET"
sudo nerdctl push "$TARGET"
sudo soci index list
After the index is published and the SOCI daemon is running, test lazy pulling and start the image with the SOCI snapshotter:
sudo nerdctl pull --snapshotter soci "$TARGET"
sudo nerdctl run --snapshotter soci --net host --rm "$TARGET"
The run command starts the image’s default process; stop it when the test is complete. Confirm that indexed layer mounts appear with mount | grep fuse. If the pull downloads all image layers, check that the index was pushed, the registry is on SOCI’s compatibility list, and the node is selecting soci rather than its default snapshotter.
Common Misconceptions
Lazy loading does not shrink an image. It postpones some reads and may avoid fetching data that is never accessed. Image size, registry storage, and eventual transfer are not automatically reduced.
A container starting sooner does not guarantee faster application response. A process that immediately reads many uncached files can block on registry requests. Startup time, readiness, first-request latency, and steady-state performance should be measured separately.
OCI compatibility does not mean lazy loading works everywhere. An OCI image can run on many runtimes, but a host still needs a compatible snapshotter and any required image index or format. Authentication, filesystem access, cache behavior, and registry features also need verification.
Related Articles
- Container storage concepts and runtime filesystems
- Cloud-native application development
- Kubernetes cluster autoscaling and node provisioning
Changelog
- First publication.

