Skip to content
Muhammet Şafak
tr
Asked by: Umut Answered:

Should I instrument my services with RED metrics or USE metrics first?


Question

I'm adding Prometheus to a Laravel API and a Go worker fleet. I've read about the RED (Rate, Errors, Duration) and USE (Utilization, Saturation, Errors) models but can't decide which to apply where. Since the API is request/response, RED seems to fit there better, but the workers aren't request/response—they process jobs in the background. Is USE more appropriate for the workers? Or should I set up both, and which first?

Answer

Short answer: this isn’t an either/or choice.

Short answer

RED is for request-driven surfaces, USE is for resources. Both fleets ultimately need both; only the weighting and the starting order differ.

Why

  1. RED = Rate, Errors, Duration. A per-request/service view. It answers “are my users being served well?” That’s a perfect fit for the Laravel API’s HTTP endpoints; start there.

  2. USE = Utilization, Saturation, Errors. A per-resource view (CPU, memory, queue depth, connection pool). It answers “which resource is the bottleneck?” It feeds your capacity and saturation decisions.

  3. The worker fleet is where the two blend. A queue worker isn’t request/response, so pure RED sits awkwardly. The fix: model the job as a “request” — rate = jobs per second, errors = failed jobs, duration = processing time. Then RED applies to the worker too.

  4. USE, especially queue depth, is a leading indicator. On workers, queue backlog (saturation) tells you “you’re under-provisioned” before latency even rises. DB connection-pool saturation is the classic hidden killer; without USE you notice it too late.

  5. RED alerts on symptoms, USE on causes. This is where the distinction earns its keep: what actually affects users is RED, so build your paging alerts on RED (error rate, p99 latency). Use USE metrics not for paging but for diagnostic dashboards that answer “why” — they warn you when a resource nears saturation, but at night it should be the degraded RED that pages you, not USE.

    # RED (histogram) — both API and worker (count the job as a request)
    http_request_duration_seconds_bucket{route="/checkout", le="0.5"}
    job_processing_duration_seconds_bucket{queue="emails", le="2"}
    # USE (gauge) — resource saturation
    queue_depth{queue="emails"}
    db_pool_in_use_connections{pool="primary"}

What to do

  1. RED first on the API — because it maps directly to SLOs. Your user-facing SLOs are latency + error rate, which is RED directly. Instrument that first; it’s also the thing that pages you at night.

  2. Don’t blow up cardinality. A label per endpoint is fine, but an unbounded label like user_id will bloat Prometheus. Keep label sets bounded, or the metrics infrastructure itself becomes the problem.

  3. Collect duration as a histogram, not an average. The “D” in RED means p95/p99, not average latency; an average hides the worst experience in the tail. Use a Prometheus histogram and derive percentiles on dashboards with histogram_quantile.

Bottom line: personally I’d start with RED on the API (fastest value against SLOs), model jobs as requests to run RED on the workers too, then layer USE (especially queue depth and pool saturation) onto both. The order is RED → USE; but without both you only see half the picture. One note: Google’s “four golden signals” (latency, traffic, errors, saturation) is a nice bridge that unifies the two models — worth a look when you’re stuck deciding.

Comments

Sign in with your GitHub account to join the discussion. Comments are stored in GitHub Discussions.

More Questions

All questions

Search the site

Start typing to search posts, projects and pages.

Esc to close Powered by Pagefind