Back to production requirements

Cold cache during deployments

Status: Complete (via workaround)
Priority: Nice to have
Audience: SRE, search platform


Why it matters

Control-plane and node replacement events can clear caches. After deploy or scale events, the first wave of queries may see higher latency until caches warm again.

Current state

Accepted as complete via workaround for GA. Continue to follow platform offline / proactive warming improvements as they ship.

What to do

  1. Keep deploy / scale windows and expected latency behavior in the runbook.
  2. Follow Elastic guidance on offline / proactive warming as it becomes available.
  3. Prefer canary or staged traffic after control-plane changes when practical.
  4. Alert on latency regressions tied to project version / capacity events.

Acceptance criteria

  • Accepted via workaround for GA
  • Warming approach kept current as platform improvements ship
  • On-call knows expected behavior during Serverless upgrades

Related

  • Autoscaling, reindex, and ingest stability (Underway — pending Central US)
  • Snapshot restore → Point-in-time restore