Operator Issues
CR changes not being applied
Symptom: Changes made to the ClickHouseCluster CR are not being reconciled or applied to the cluster for an extended period. Cause: One of three common reasons:- The operator itself is crashlooping.
- ClickHouse pods (keeper or server) are crashlooping, which blocks reconciliation.
- A
clickhouse.com/skip-reconcileannotation is present on the CR, which instructs the operator to skip reconciliation entirely.
- Check operator logs for crashes or errors:
- Check keeper and server pods for crashloops:
- Check for the skip-reconcile annotation on the CR:
If present and no longer needed, remove it:
How to drop a server replica
Symptom: You need to remove a specific replica from the cluster without scaling in (reducing the desired replica count). Cause: A replica is unhealthy, stuck, or otherwise needs to be manually removed. Resolution:-
Add the skip-reconcile annotation to prevent the operator from interfering:
Confirm the operator has noticed by checking logs for:
-
Remove the replica from the ReplicaStateMap. First, inspect the current map:
Then edit the status subresource to remove the target replica entry:
-
Delete the StatefulSet for the replica:
Wait for the pod to terminate after deleting the StatefulSet.
-
Remove the skip-reconcile annotation:
The operator will launch a new replica and clean up the removed replica from ClickHouse.
- Verify by logging in to the ClickHouse cluster and confirming the old replica has been removed. If any leases the replica holds have not expired, the operator will retry removal. Cleanup should complete within 5 minutes.
Server pod hanging on termination
Symptom: Server pods remain inTerminating status for an extended period:
- Check the PreStop hook log for details on what the hook is waiting for:
- Check server logs for any received signal log messages that indicate the shutdown process state.
Multiple ClickHouseClusters in the same namespace
Symptom: Unexpected reconciliation behavior, resources being modified or deleted unexpectedly, or the operator oscillating between cluster states. Cause: More than oneClickHouseCluster custom resource exists in the same namespace. The operator assumes a single ClickHouseCluster per namespace and cannot correctly reconcile when multiple instances share a namespace.
Resolution:
- Check how many
ClickHouseClusterresources exist in the namespace: - If more than one exists, move the additional clusters to their own namespaces. Each
ClickHouseClustermust be the only instance in its namespace. - The
onprem-clickhouse-clusterHelm chart creates aResourceQuotaby default that prevents this situation. Verify the quota is in place:If the quota is missing, ensureresourceQuota.enabledis set totrue(the default) in your Helm values.
CR not in healthy Running state
Symptom: The ClickHouseCluster CR shows a state other thanRunning, or pods are restarting.
Cause: Various issues can prevent a healthy state — crashlooping pods, resource constraints, misconfigurations, or scheduling failures.
Resolution:
- Check cluster status across all namespaces:
- Describe the problem pod for detailed status and events:
- Check logs from the previously terminated container:
- Check namespace events for scheduling or resource issues:
ClickHouse Server Issues
Crashlooping server pods
Symptom: ClickHouse server pods are in aCrashLoopBackOff state.
Cause: The ClickHouse server process is crashing on startup or shortly after. This could be due to configuration errors, corrupt data, memory pressure, or OOM kills by Kubernetes.
Resolution:
- Check the ClickHouse server pod logs for the crash reason:
- If the crash is due to memory pressure or Kubernetes-initiated termination, check Kubernetes events:
Data loss/corruption (ClickHouseBrokenPartDetectedOnSelect)
Symptom: TheClickHouseBrokenPartDetectedOnSelect alert fires, or SELECT queries fail with POTENTIALLY_BROKEN_DATA_PART errors.
Cause: An SMT data part read failed with a non-retriable error. The POTENTIALLY_BROKEN_DATA_PART exception is thrown when a data part is found to be broken during a SELECT operation.
Resolution:
-
Examine logs for the
POTENTIALLY_BROKEN_DATA_PARTexception. If not found in logs, also checksystem.errors. -
Understand what data parts are lost. Query
system.replicasto find affected tables: -
Find logs related to lost parts. Check
/var/log/clickhouse-server/on the pod (usezgrepfor archived logs), or querysystem.text_log:Add a predicate forevent_timerange to speed up the query. -
Understand the history of lost parts. Pick a lost part name and find all related logs:
Focus on log messages before
Part * is lost forever. Messages after that point are irrelevant (any “found” part is actually an empty replacement). -
Check for false positives:
- Check if the table has TTL and the lost part should have been dropped anyway.
- Check
system.query_logfor TRUNCATE or DROP PARTITION queries that should have dropped the lost parts.
-
Additional investigation:
- If the part was detached as broken, determine why it was broken.
- If you see
The specified key does not exist, search all logs with the blob name to find when and why it was removed. Also check log messages about zero-copy locks.
- Contact ClickHouse support with your findings.
Table replicas read-only (ClickHouseTableReplicasReadOnly)
Symptom: TheClickHouseTableReplicasReadOnly alert fires, or writes to tables fail because they are in read-only mode.
Cause: A table has been in read-only mode for at least one hour. This excludes tables in *_broken_replicated_tables and *_broken_tables databases. It could be caused by a DROP operation that went badly, or a keeper connectivity issue.
Resolution:
-
Check if there are still read-only tables in the cluster:
-
Investigate the cause. Having read-only tables typically indicates that
StorageSharedMergeTree::shutdownwas run but the storage object was kept alive. Search text logs using the table name as the logger name. -
Try restarting the replica for each affected table:
You can get the table names from the query in step 1. Sometimes the problem is trivial and a restart resolves it.
- If the issue persists, check keeper logs for potential connectivity or coordination issues.
Replica already exists (ClickHouseReplicaAlreadyExists)
Symptom: TheClickHouseReplicaAlreadyExists alert fires, or replica creation fails with a REPLICA_ALREADY_EXISTS error.
Cause: A replicated table (SMT or RMT) could not be created because an existing replica is already associated with the ZooKeeper path. This is unlikely to be caused by user error (explicit UUID reuse is now prohibited via database_replicated_allow_explicit_uuid). This is likely a bug in the Replicated database or Shared Catalog.
Resolution:
Contact ClickHouse support. This is likely a bug that requires investigation by the engineering team.
Cannot write to file descriptor (ClickHouseCannotWriteToFileDescriptor)
Symptom: TheClickHouseCannotWriteToFileDescriptor alert fires, or errors such as CANNOT_WRITE_TO_FILE_DESCRIPTOR or no space left on device appear in logs.
Cause: The cache disk is full. The exception is thrown when there is not enough space for a new cache entry or for external data processing (e.g., external aggregation, external joins). This may be due to a misconfiguration where the disk was created with less space than requested in the CR config.
Resolution:
-
Check for the known
partial_mergeissue first. There is a known bug in tracking cache disk usage when thejoin_algorithm = 'partial_merge'query setting is specified. Check if this setting is in use. -
Connect to the pod:
-
Check actual disk size:
-
Check required cache disk size in ClickHouse:
Note that multiple caches (e.g.,
s3diskWithCache,diskPlainRewritableForSystemTablesWithCache) may share the same path (/mnt/clickhouse-cache/sharedS3DiskCache). -
Compare actual vs. required:
- If the actual disk size is smaller than required, the issue is a misconfiguration. Contact ClickHouse support.
- If the disk size is sufficient, it is likely a bug in cache disk usage tracking. Investigate via
system.filesystem_cache.
Storage and Infrastructure Issues
Zone details could not be found for any PV
Symptom: The ClickHouse operator logs show the errorzone details couldn't be found for any PV, and reconciliation cannot complete.
Cause: The operator expects topology-aware node affinities to be automatically populated on PersistentVolumes using the topology.kubernetes.io/zone label. Normally this is handled by topology-aware volume provisioners (e.g., AWS EBS). However, for certain cloud providers (e.g., IBM) or on-premise environments, this label is not automatically populated or non-standard labels are used.
Resolution:
-
Add
allowedTopologiesto your StorageClass to ensure volumes are created with the correct node affinities: -
If non-standard labels are used for topology, configure the operator’s
additionalZoneLabelRegexesproperty. For example, when using the Helm chart, set theoperator.additionalZoneLabelRegexesHelm value to a regex matching your labels (e.g.,directpv.*zone).
NVMe disk questions
Symptom: Uncertainty about whether NVMe-attached instances are required, or questions about using alternative storage for the cache disk. Cause: The NVMe disk serves as a cache volume for data coming from S3 (or equivalent object storage). The operator mounts it in ClickHouse server pods and configures it as an S3 cache disk:- Ensure the disks are attached to the instances used for ClickHouse server Kubernetes pods.
- Alter the launch template to reflect changes in disk architecture. You can still create a RAID disk and format as ext4 or xfs, but use the appropriate tools (e.g.,
lsblk) to list devices. - If the disks are not mounted at
/nvme/disk, set thehostPathBaseDirectoryin the ClickHouseCluster Helm chart to the actual mount point.
Configuration Issues
loadBalancerType field behavior
Symptom: Setting theloadBalancerType field on the ClickHouseCluster CRD does not create load balancer annotations on the Service.
Cause: The loadBalancerType field does not add load balancer annotations to the Service of the ClickHouse cluster. The ClickHouse operator does not support creating a load balancer for the ClickHouse cluster. This field only serves a purpose in ClickHouse Cloud, where it is used to help create users for the Cloud SQL Console.
Resolution:
No action is needed unless you are running in ClickHouse Cloud. If you need a load balancer for your ClickHouse cluster, configure it separately through standard Kubernetes Service annotations and configurations outside of the ClickHouseCluster CRD.
Parent instance stuck in Degraded during Shared Catalog migration
Symptom: During a Shared Catalog migration, a parent instance stays Degraded and never reaches Running:
Shared engine, and DDL queries do not work on the parent or on any of its children.
Cause: The migration completes only once every replica of the parent and of every child is running Shared Catalog. A parent whose children have not been migrated stays Degraded, and replicated DDL processing stays disabled across the whole group.
Resolution:
- Check which instances in the group are still on the
Replicatedengine: - Migrate every child that has not been migrated, following Migrating instances with compute-compute separation. The parent moves to
Runningonce the last child completes. - Do not remove
featureFlags.migrateToSharedCatalogfrom the parent while any child is unmigrated. If it has already been removed, restore it totrueso the Operator can finish the migration.
DDL with ON CLUSTER fails on Shared Catalog
Symptom: A DDL statement fails on an instance running Shared Catalog when it carries an ON CLUSTER clause:
ON CLUSTER sends it to each replica through the distributed DDL path instead, and Shared Catalog rejects the copy that arrives there as a non initial query.
Resolution:
Remove the ON CLUSTER clause from the statement:
ON CLUSTER from DDL queries for a query that helps find the affected clients.
Diagnostic Commands Quick Reference
Cluster status
Operator logs
Pod events and details
Previous container logs (after a crash)
Replica state map
Replication queue size per table (SQL)
system.replicas: