Prometheus Remote Write and Long-Term Metrics Storage

Updated on
10 min read

Prometheus remote write lets a monitoring server send the samples it scrapes to another system for centralized querying and longer retention. It is useful when a single Prometheus server’s local disk is not the right place to keep every metric, or when teams need data from multiple clusters in one view. This guide explains the protocol, its queue and failure limits, and how to connect a Prometheus server to a remote metrics backend without mistaking data shipping for a complete backup strategy.

What Is Prometheus Remote Write?

Prometheus is an open-source monitoring and alerting system built around labeled time series. A Prometheus server normally discovers targets, scrapes their metrics, stores samples in its local time-series database (TSDB), and evaluates queries and alerting rules. Remote write is an optional outbound path that forwards ingested samples to a compatible receiver over HTTP.

Each sample belongs to a series identified by its metric name and labels, and carries a timestamp and value. The receiver accepts those samples and stores them according to its own design. The Prometheus Remote Write specification defines the stable 1.0 wire protocol; requests use an HTTP POST with a Snappy-compressed Protocol Buffers payload. Prometheus also documents an experimental 2.0 message format with additional metadata and native-histogram support, but a sender and receiver must agree on a supported format. Check the remote-write configuration reference and the receiver’s compatibility requirements before changing the protocol message.

Remote write does not replace scraping. The Prometheus server still gathers metrics from targets and maintains local storage; it additionally ships samples to the configured destination. The receiver may be a self-managed or hosted service, but accepting remote-write requests alone does not define its retention, replication, query, or backup behavior.

The Problem Remote Write Solves

A single Prometheus server’s local TSDB is efficient for recent operational data, but it is tied to that server’s storage and failure domain. Local retention is bounded by disk capacity and configured retention limits. Keeping more history on the same node does not make the data replicated or protect it from a failed disk, host, or region. The Prometheus storage documentation describes these local-storage limits and the separate remote-storage interfaces.

Organizations may also run independent Prometheus servers for different clusters, teams, or regions. Querying each separately makes cross-system comparisons and long-range analysis harder. A remote-write destination can collect those streams in a central place, where the backend can provide its own retention and query layer.

The trade-off is another network and service dependency. A receiver can reject requests, become slow, or be unreachable. Samples waiting to be sent consume resources, and the sender’s WAL is a bounded recovery window rather than an archival queue. Remote write therefore extends a metrics architecture; it does not remove the need to design for delivery failures, capacity, and recovery.

How Prometheus Remote Write Works

The usual data path has four stages:

  1. Prometheus scrapes a target and appends the resulting samples to its local TSDB write-ahead log (WAL).
  2. A remote-write queue reads samples from the WAL and places them into in-memory queues, divided into shards for parallel sending.
  3. Each shard batches samples into protocol requests and posts them to the configured receiver.
  4. The receiver responds and stores accepted samples. Prometheus advances its sending work and continues reading new data.

The remote-write tuning guide describes this WAL-to-queue-to-shard flow. Sharding lets a destination send several requests concurrently, but more shards also mean more in-flight work and memory. Prometheus adjusts the shard count based on input and sending rates within the configured bounds. The queue is per destination, so adding another remote-write target adds another delivery path to monitor.

The receiver’s HTTP response is part of delivery behavior. A successful response acknowledges a request; an error can trigger retries or indicate a request that needs configuration changes. The Remote Write specification defines protocol-specific handling, while RFC 9110 defines general HTTP semantics and status codes. A 400 caused by an invalid request, for example, is not fixed by increasing the queue, while a receiver outage can create a growing backlog.

Concern Local Prometheus TSDB Remote-write destination
Data path Written by the scraping Prometheus server Sent over HTTP by its remote-write queue
Retention Controlled by the local server’s storage settings and disk Defined by the receiving system and its policies
Failure domain Local disk and Prometheus host Backend, network, credentials, and its own storage
Primary role Recent queries, rules, and operational troubleshooting Centralized ingestion, longer history, or multi-cluster views
Query behavior Queried directly by that Prometheus server Depends on the backend or a separately configured remote-read path
Recovery WAL replay and local TSDB recovery Receiver-specific buffering, replication, and backup

An important limit is the WAL’s lifetime. Prometheus documents that remote-write failures can be retried without data loss while the needed samples remain in the WAL, but if an endpoint stays down for more than about two hours, WAL compaction can remove unsent data. The exact failure mode depends on ingestion and recovery rates, but an indefinitely growing backlog is not guaranteed. Alert on pending samples and test an outage-and-recovery scenario rather than treating the sender as a durable message broker.

Components and Key Concepts

Remote-write endpoint: The receiver URL and path are determined by the backend. Prometheus can send to compatible systems, but the built-in Prometheus receiver is not automatically enabled; enabling it requires a command-line flag. A receiving endpoint should be exposed only where needed and protected with TLS and appropriate authentication.

Queue and shards: Each destination has a queue with settings such as minimum and maximum shards, queue capacity, and maximum samples per request. Increasing capacity can help absorb bursts, but it also increases memory use across shards. First observe the pending-sample backlog, retries, send latency, and receiver capacity. Tune only after identifying whether the sender, network, or backend is the bottleneck.

Labels, cardinality, and filtering: Labels identify series, so high-cardinality values multiply the volume a backend must ingest and index. Write relabeling can exclude samples from a particular destination, but it does not reduce what Prometheus already scraped or stored locally. Decide which data can be omitted before filtering: a sample excluded from remote write will not appear in that destination’s history. For related instrumentation controls, see OpenTelemetry metric cardinality limits.

Retention and high availability: The receiver determines how long samples remain queryable and whether it replicates them. With multiple Prometheus replicas, both may send overlapping series. Some backends deduplicate replicas when configured with appropriate cluster and replica labels; this behavior is backend-specific, not a guarantee of the remote-write protocol. Keep replica identity and deduplication rules consistent with the backend’s documented model.

Remote read is separate: Remote write sends samples outward. Remote read is a different integration for retrieving data, and it is not automatically enabled just because a server has a remote-write destination. Prometheus documents remote read as fetching selected samples for local PromQL evaluation, rather than distributing query execution to the storage backend.

Real-World Use Cases

  • Longer retention: Keep a shorter local window for fast incident response while a backend retains history for capacity planning or compliance requirements. Confirm that the backend’s retention and deletion policies meet the actual requirement.
  • Central views across clusters: Send data from multiple Prometheus servers to a shared backend so operators can compare services and regions. Preserve cluster labels so queries can distinguish sources.
  • Separating collection from analysis: A team can keep local scraping and alert evaluation near workloads while using a central system for long-range dashboards. If metrics pass through an OpenTelemetry Collector instead, its receiver, processor, and exporter pipeline has separate buffering and delivery semantics; see how OpenTelemetry Collector pipelines work.
  • Storage planning: A backend can provide downsampling or other storage tiers, but those are backend features. Estimate series growth and ingestion rate before setting retention or pricing expectations.

Getting Started with Remote Write

Add a remote_write entry to the Prometheus configuration. Replace the example URL with the exact endpoint from the chosen backend, and provision the token file through the host’s secret-management process rather than committing credentials:

global:
  scrape_interval: 15s

scrape_configs:
  - job_name: prometheus
    static_configs:
      - targets: ["localhost:9090"]

remote_write:
  - name: primary
    url: https://metrics.example.net/api/v1/write
    bearer_token_file: /etc/prometheus/secrets/remote-write-token

The URL path, supported authentication method, tenant headers, and protocol message depend on the receiving service. Use HTTPS for traffic that leaves a trusted local boundary, restrict the token to the required tenant or write scope, and mount the secret read-only where possible. Avoid copying queue-tuning values from another deployment; start with defaults and change them only after observing the sender and receiver under realistic load.

Check the configuration before reloading or restarting Prometheus:

promtool check config prometheus.yml

After the server starts with the configuration, query its own metrics endpoint or Prometheus UI for queue health:

prometheus_remote_storage_samples_pending
rate(prometheus_remote_storage_samples_retried_total[5m])
increase(prometheus_remote_storage_samples_failed_total[1h])
increase(prometheus_remote_storage_samples_dropped_total[1h])

A sustained rise in pending samples means the sender is falling behind; retries and failures help distinguish receiver or network problems from a healthy queue. Then query the remote backend for a known series, such as up for a configured target, and compare timestamps and labels with the source Prometheus. Check both ends: a successful configuration check proves syntax, not that the destination accepted and retained the samples.

For troubleshooting, inspect Prometheus logs for remote-write errors, verify DNS and TLS connectivity from the Prometheus host, confirm credentials and tenant routing, and check receiver limits and ingestion errors. If the backlog grows continuously, reduce avoidable volume only after evaluating the effect of write relabeling, then address the actual network, sender-capacity, or backend bottleneck. Test recovery from a planned short interruption before depending on remote storage.

Common Misconceptions

“Remote write is a backup.” It is a streaming integration, not a complete backup plan. It can lose samples during a sufficiently long outage, and it does not guarantee that the receiver is replicated or recoverable. Use backend-specific snapshots, replication, and restore tests for backup objectives.

“More local retention means more durability.” Longer retention preserves more history on the same storage system, but does not protect it from host or disk failure. Local retention and remote retention are independent policies with different costs and failure modes.

“The backend automatically deduplicates high-availability Prometheus replicas.” Deduplication depends on the backend and its configuration. Without a supported replica-label strategy, duplicate streams can consume extra storage or appear as separate series.

“Remote write moves Prometheus queries to the backend.” It ships samples; it does not by itself change where PromQL runs. A backend may offer its own query API, while Prometheus remote read is a separate path with different behavior and limits.

TBO Editorial

About the Author

TBO Editorial writes about the latest updates about products and services related to Technology, Business, Finance & Lifestyle. Do get in touch if you want to share any useful article with our community.