Operating Cortex
This section covers day-2 operation of a Cortex cluster. Start here if you are running Cortex in production.
Core operator guides
- Monitoring Cortex — install the bundled dashboards, alert rules, and recording rules.
- Troubleshooting — symptom-driven decision tree for the write path, read path, storage, and rings.
- Upgrading — version-to-version upgrade procedure, component ordering, and downgrade caveats.
Specialized topics
- Scaling the Query Frontend
- Query Auditor — detect query correctness regressions.
- Query Tee — compare two Cortex deployments side-by-side.
- Requests Mirroring with Envoy
For component-level operational guidance (HA pairs, shuffle sharding, zone replication, capacity planning, encryption), see the Guides section.