---
title: "I've operated petabyte-scale ClickHouse clusters for 5 years"
subtitle: "What I learned operating ClickHouse at scale: the wins, the failures, and the lessons that only come from production experience."
slug: ive-operated-petabyte-scale-clickhouse-clusters-for-5-years
url: https://listedarticles.com/articles/ive-operated-petabyte-scale-clickhouse-clusters-for-5-years
canonical_url: https://www.tinybird.co/blog/what-i-learned-operating-clickhouse
content_type: blog_post
language: en
published_at: 2026-01-15T00:00:00.000Z
updated_at: 2026-09-16T16:11:04.341Z
author: "Javi Santana"
authored_by: agent
publisher: "Tinybird"
publisher_url: https://www.tinybird.co
topics: ["ClickHouse", "Databases", "Data Engineering", "Infrastructure", "Scalability"]
license: all-rights-reserved
word_count: 280
reading_minutes: 1
citation: "Javi Santana, Tinybird. \"I've operated petabyte-scale ClickHouse clusters for 5 years.\" 15 Jan 2026. https://www.tinybird.co/blog/what-i-learned-operating-clickhouse (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# I've operated petabyte-scale ClickHouse clusters for 5 years

*What I learned operating ClickHouse at scale: the wins, the failures, and the lessons that only come from production experience.*

> Tinybird co-founder Javi Santana shares five years of operational experience running petabyte-scale ClickHouse clusters, covering architecture decisions, the challenges of storage-compute separation, zero-copy replication trade-offs, and how the upgrade process evolved from a three-hour ordeal into a CI/CD-integrated routine.

> **Indexed summary.** This entry is an agent-written synopsis of an article first published at [tinybird.co](https://www.tinybird.co/blog/what-i-learned-operating-clickhouse). Read the original for the full text.

Santana frames the post around a central observation: setting up ClickHouse is easy; keeping it running at petabyte scale is genuinely hard. The essay moves through architecture, storage, and upgrade concerns in roughly chronological order, reflecting lessons accumulated since running ClickHouse version 18.4.

## Key points

- The standard shards-and-replicas architecture becomes expensive quickly: a 300 TB table replicated across 10 nodes for query capacity requires 3 PB of storage.
- Open-source ClickHouse lacks mature cloud storage support; Tinybird uses a modified zero-copy replication in a private fork, but notes that the upstream feature has been buggy and nearly removed by ClickHouse Inc.
- For latency-sensitive customers, a hot/cold architecture with local SSDs and S3 outperforms pure S3 storage; Tinybird also keeps a dedicated write-only replica to isolate ingestion from query traffic.
- Upgrades were originally three hours with two weeks of preparation; the team eventually built a backward-compatible rolling upgrade process integrated into CI/CD.
- Safe upgrade checklist: add a replica on the new version, monitor logs (many alarming messages are harmless), avoid DDL changes during the process, send read traffic first before shifting writes.
- Large part sizes improve read performance but make merges expensive; balancing merge aggressiveness against query performance is an ongoing operational concern.

## Why it matters

ClickHouse has become a popular foundation for real-time analytics products, but most public documentation covers setup rather than sustained operation. Santana's account of concrete failures, architectural reversals, and hard-won process improvements fills a gap for teams running ClickHouse beyond the initial deployment phase.

---

*Source: [I've operated petabyte-scale ClickHouse clusters for 5 years](https://www.tinybird.co/blog/what-i-learned-operating-clickhouse)*
