TL;DR
Real-time systems differ dramatically in the cost and efficiency of making incoming data query-ready. In this series, we trace the path from data arrival to answers, showing how it shapes preparation cost, query cost, and query runtime.
Here, we compare ClickHouse Cloud and Databricks while streaming 113.2 billion quotes and running queries. ClickHouse Cloud delivered 752× better end-to-end performance per dollar than the tested Databricks stack, whose summaries could lag behind incoming data.
From fresh data to fast answers
Real-time analytics has to stay fast as fresh data arrives. CostBench tests this under continuous ingestion: queries run while the dataset grows. Efficient preparation lowers the cost of making that data query-ready and reduces the work left for queries, lowering their cost and runtime.
Columnar storage lets queries skip unused fields; ordering lets drill-downs skip unrelated quotes; current pre-aggregations let interactive aggregations combine prepared summaries. The diagram below shows how this fresh-data path shapes the scanning and aggregation left for queries as preparation keeps up or falls behind.
CostBench measures ① preparation cost, ② query cost, and ③ runtime, then combines them into one end-to-end performance-per-dollar score.
This article follows those three effects through the ClickHouse–Databricks comparison, from preparing incoming data to serving queries.
This comparison uses Databricks Serverless SQL for query serving. Lakehouse//RT, Databricks’ new real-time warehouse currently in Beta, is not included.
The workload and the result
ClickHouse Cloud delivered 752× better end-to-end performance per dollar than Databricks in this continuous-ingestion workload.
Both systems used their recommended low-latency ingestion paths and native preparation features: ClickHouse asynchronous inserts; Databricks Zerobus Ingest, liquid clustering, and incremental materialized views. CostBench streamed 113.2 billion stock-market quotes at a target of 1 million rows per second, using matching sorting or clustering keys and daily pre-aggregations by stock symbol. During ingestion, four interactive aggregate queries ran every ten minutes; two drill-down queries ran every hour. Queries covered the growing history through the raw tables or maintained summaries. Read-side compute was matched at 16 CPUs, with query-result caching disabled.
CostBench calculates its end-to-end performance-per-dollar score using the formula shown below: Score = (① preparation cost + ② normalized query cost) × ③ accumulated query runtime.
Preparation cost covers the complete fresh-data path. Normalized query cost prices recorded query runtimes at the applicable read-side compute rates. Accumulated query runtime is the sum of query durations. Lower scores are better.
HOW DID WE MEASURE ① ② ③? (click to expand)
SHARED HARNESS
Both systems used the same source dataset, schema, workload definitions, pacing model, query schedule, and result format. Destination-specific adapters handled delivery and timing. Part 1 explains the shared methodology; the Databricks benchmark contract records this provider’s implementation.
DATASET AND PREPARATION
The complete NBBO dataset contains 113,219,565,734 narrow, 12-column rows. Raw tables used (sym, t), stock symbol and event timestamp; daily summaries used (sym, day), stock symbol and UTC day. The summaries maintained the counts, sums, and price minima and maxima used by the aggregate queries.
CONTINUOUS WORKLOAD
Fresh rows were paced toward 1 million per second while aggregate and drill-down queries ran on their ten-minute and hourly schedules. This models data being generated continuously and prepared as it arrives, rather than measuring how quickly an existing dataset can be bulk-loaded or backfilled. The Databricks MV path retained its native snapshot semantics; matching ingestion progress does not make both systems’ summaries equally fresh.
RECOMMENDED STACK AND SIZING
ClickHouse Cloud used a 16-CPU, 64-GiB read node. Databricks used an X-Small Serverless SQL warehouse fixed at one cluster, approximately 16 worker vCPUs; this sizing does not imply identical hardware or memory. Liquid clustering, Predictive Optimization, and triggered incremental MV refresh were enabled. Lakehouse//RT was unavailable in the test workspace, so the accepted comparison uses the Serverless SQL baseline.
① PREPARATION COST
The full ingest includes ClickHouse’s complete two-node HA ingest service and Databricks’ Zerobus ingestion, asynchronous clustering, and MV refresh. Databricks ingestion and refresh use allocated DBU usage; clustering uses the accepted Predictive Optimization DBU allocation. All three appear in the preparation-cost model.
② QUERY COST AND ③ RUNTIME
Each accepted query’s recorded duration is priced at its read-side compute rate and added to accumulated runtime. Databricks uses Query History total duration, excluding result fetch; ClickHouse uses its recorded runner duration. These normalized costs represent accepted query work, not complete warehouse invoices. The query-comparison and pricing notes below detail the timing and boundaries.
TARGET VERSUS OBSERVED RATE
Databricks completed the durable-ingest window in about 31.5 hours, averaging approximately 1 million rows/s. ClickHouse completed its ingest in about 36.75 hours, averaging approximately 0.86 million rows/s. The target was shared; observed progress differed. Databricks’ final reconciliation confirmed all 113,219,565,734 rows with no duplicate surplus. Durability acknowledgment and table query visibility are separate milestones.
REPORTING WINDOW
Per-query latency charts stop at 100 billion rows. Preparation costs cover the full active ingest; accumulated query cost and runtime use the accepted active-ingestion observations, ending at each workload’s shared comparison horizon. Post-ingestion queries and maintenance are outside the headline score. Displayed totals are rounded.
ClickHouse Cloud led in all three measured components:
① Preparation cost: $28.69 versus $695.60.
② Normalized query cost: about $0.05 versus $2.65.
③ Accumulated query runtime: 56.39 seconds versus 29.10 minutes.
The charts below track those costs and runtimes as ingestion progresses.
The score breakdown below combines the two advantages: 24.3× lower combined preparation and normalized query cost, multiplied by 31× lower accumulated query runtime, gives ClickHouse Cloud its overall performance-per-dollar lead.
PRICING, FORMULAS, AND THE 752× SCORE (click to expand)
PRICING BASIS
ClickHouse Cloud uses $0.3903 per compute unit (CU) per hour. Databricks uses checked-in public list rates for AWS eu-west-1, Premium: $0.39 per DBU for allocated ingestion and maintenance, and $0.91 per DBU for Serverless SQL. The accepted summary preserves full precision; these are modeled list-price costs, not a historical invoice reconstruction.
① COMPLETE PREPARATION COST
This covers the active ingest of all 113,219,565,734 rows, including the required preparation components.
CLICKHOUSE CLOUD
2 CUs × 36.75 hours × $0.3903/CU-hour = $28.68705, including both HA ingest nodes.
DATABRICKS INGESTION
Allocated Zerobus usage, 1,098.9261167 DBUs × $0.39 = $428.58119. This accepted DBU allocation is used in the headline rather than converting the source’s byte volume into a pricing proxy.
DATABRICKS CLUSTERING
The Predictive Optimization allocation contains 151.8613011 DBUs × $0.39 = $59.22591. This component uses Predictive Optimization’s operation-level DBU allocation. It is a modeled maintenance charge rather than an invoiced amount.
DATABRICKS MV REFRESH
532.8145064 allocated DBUs × $0.39 = $207.79766. Together, ingestion, clustering, and refresh total $695.60475 in the accepted preparation model.
② NORMALIZED QUERY COST
Each accepted query’s recorded duration is multiplied by the per-second read-side compute rate. This models accepted work; it excludes idle capacity and minimum warehouse billing.
CLICKHOUSE CLOUD
8 CUs × $0.3903/CU-hour × (10.312s aggregate + 46.074s drill-down) ÷ 3,600 = $0.04890.
DATABRICKS SERVERLESS SQL
X-Small is priced at 6 DBUs/hour × $0.91/DBU = $5.46/hour. Aggregate work: 613.164s × $5.46/hour ÷ 3,600 = $0.92997.
DATABRICKS DRILL-DOWN WORK
1,133.002s × $5.46/hour ÷ 3,600 = $1.71839. Total normalized query cost: $2.64835.
ALLOCATION BOUNDARY
Databricks ingestion and maintenance are allocated to the producer-active window, 2026-09-18 17:04:53 UTC through 2026-09-20 00:33:42.627 UTC, with an exclusive end. Post-ingestion maintenance, validation/control SQL, database storage, producer infrastructure, network charges, and idle/minimum warehouse charges are excluded. Maintenance was observed after ingestion but that later work is not charged to this score.
③ ACCUMULATED QUERY RUNTIME
ClickHouse: 10.312s aggregate + 46.074s drill-down = 56.386s. Databricks: 613.164s aggregate + 1,133.002s drill-down = 1,746.166s.
COMPARISON BOUNDARY
① covers the complete active ingest. ② and ③ use 189 four-query aggregate batches and 32 two-query drill-down batches. Their row-count comparison horizons differ from the 100-billion-row latency plots. Durability alignment does not establish identical query-visible rows or MV freshness. This is the complete modeled preparation path plus normalized query work, not the full provider bill.
FINAL SCORE
ClickHouse: ($28.68705 + $0.04890) × 56.386 = 1,620.30528. Databricks: ($695.60475 + $2.64835) × 1,746.166 = 1,219,265.82646. The ratio is 752.491×, rounded to 752×, using the full-precision inputs.
752× better real-time performance per dollar
Interested in seeing how ClickHouse works on your data? Get started with ClickHouse Cloud in minutes and receive $300 in free credits.
Sign upAll benchmark code and results are available in the CostBench repository. The stock-quotes dataset requires a separate data license, so the data itself cannot be redistributed.
To explain that result, we follow the same path as the data: preparation first, then query serving.
How ClickHouse Cloud prepares incoming data
The benchmark used a dedicated ClickHouse Cloud ingest service with two 2-CPU nodes for high availability, the smallest HA configuration that sustained the target during calibration. The shared client sent rows through asynchronous inserts built into each node.
The diagram below shows those same nodes handling ① columnar storage, ② ordering raw data, and ③ ordered pre-aggregation, plus background merges, on four CPU cores in total.
① Store in columns and ② order data: When an ingest node flushes its asynchronous-insert buffer, it sorts raw quotes by the MergeTree table’s (sym, t) key, stock symbol and event timestamp, and writes an ordered data part.
③ Order and pre-aggregate data: The same node runs the incremental MV query over the incoming block in memory, computes aggregate states, and sorts the results by the AggregatingMergeTree table’s (sym, day) key before writing an ordered part containing the summaries.
Background merges consolidate parts in both tables.
The result: raw data and pre-aggregations advance together from the same flushed insert block, without a separate refresh cycle.
HOW MUCH WORK DID THE FOUR-CORE INGEST SERVICE HANDLE? (click to expand)
The ClickHouse writer-utilization results show ingestion continuing alongside background maintenance. CPU usage remained around 3.1 of the 4 available cores, tracked memory stayed within capacity, background merges processed several million rows per second, and the maximum active-part count per partition stayed around 100. These diagnostics describe the ClickHouse source run reused for this comparison.
Both nodes were retained for high availability, so the preparation cost includes the complete two-node deployment.
Now follow the same three preparation tasks through Databricks, where ingestion, clustering, and MV refresh use separate services.
How Databricks prepares incoming data
The benchmark used Zerobus Ingest, Databricks’ managed path for pushing rows directly into Delta tables, through Arrow Flight streams.
As shown below, Zerobus handles ① columnar storage; asynchronous serverless clustering handles ② ordering raw data; a separate serverless MV pipeline handles ③ pre-aggregation.
① Store in columns: Zerobus buffers incoming rows and publishes them into a managed Delta table backed by columnar Parquet files. A durability acknowledgment confirms that Zerobus has durably accepted the data; it does not guarantee immediate visibility in the table.
② Order data: The raw table used liquid clustering on (sym, t), matching ClickHouse’s sorting key. Predictive Optimization ran clustering asynchronously on serverless compute, so newly published files could be queried before that layout work finished.
③ Order and pre-aggregate data: A separate serverless pipeline incrementally refreshed daily summaries, clustered by (sym, day). The tested definition used REFRESH POLICY INCREMENTAL STRICT and TRIGGER ON UPDATE AT MOST EVERY INTERVAL 1 MINUTE. The trigger uses Databricks’ minimum supported interval of 1 minute; it limits refresh starts, not completion time.
WHY THESE INGESTION PATHS, AND HOW WERE THEY BUFFERED? (click to expand)
THE SAME INGESTION PROBLEM
Applications produce frequent writes that must become efficient storage writes. Both destinations used direct application-to-table ingestion with managed buffering. No Kafka broker, file landing stage, or Auto Loader job was added to the measured path.
DATABRICKS’ MANAGED PATH
The accepted runner used the Zerobus Arrow Flight DoPut interface, with 16 concurrent streams and 50,000-row client batches. Streams rotated every ten minutes; SDK recovery and bounded retries were enabled. Zerobus ingestion compute was allocated and priced separately from clustering, MV refresh, and SQL serving.
CLICKHOUSE’S NATIVE PATH
Ordinary INSERT requests with async_insert enabled were buffered on each receiving node. An adaptive flush timeout responded to incoming traffic. Ordering and pre-aggregation ran on the ingest nodes; the dedicated ingest service shared storage with the independently sized read service.
SHARED PACING
Both adapters read the same Parquet dataset and decoded the same schema under the same target-rate model. Parallelism and client batch sizes were adapted to each destination’s API; they were not forced to be identical.
CLICKHOUSE BATCHES
The client sent 3,000-row batches through asynchronous inserts. Default buffering flushed at the first of three thresholds: 100 MiB buffered, an adaptive timeout between 50 milliseconds and 1 second, or 450 queued insert queries.
DATABRICKS BATCHES
Each Arrow Flight stream sent 50,000-row batches with IPC compression disabled. The target rate was based on logical source rows, with acknowledged durable progress recorded separately from table visibility. Those batches were not staged as files for a later bulk load.
BATCH-SIZE BOUNDARY
These are client delivery batch sizes. Both systems buffer delivery into storage writes, so 3,000 and 50,000 rows do not specify final part or Parquet-file sizes.
RETRIES AND RECONCILIATION
Zerobus used durable stream offsets and SDK recovery. The settled validation confirmed exact row reconciliation and no duplicate surplus. ClickHouse supports insert deduplication for retries, including dependent materialized views. The final row check validates delivery, not equal summary freshness during ingestion.
INCREMENTAL MAINTENANCE
The raw Delta table enabled row tracking, change data feed, and deletion vectors. The MV passed incremental eligibility checks and used INCREMENTAL STRICT. The refresh evidence recorded 1,028 incremental group-aggregate refreshes and zero full refreshes across the collected run, including its post-ingestion observation period.
WHAT WAS TESTED
The serving layer was Serverless SQL X-Small. Lakehouse//RT was not available in the workspace and was not tested. This article describes the accepted configuration and its results, rather than projecting the performance of an unavailable serving product.
The result: Databricks can expose raw data before clustering finishes, while its pre-aggregations advance through a separate refresh cycle.
What this means for freshness and preparation cost
Pre-aggregation lag
ClickHouse updates summaries during inserts; Databricks refreshes them asynchronously. The chart below follows Databricks’ refresh cadence: observed completed refreshes were about 1.8 minutes apart on average, reaching about 2.0 minutes in the plotted trend.
These are intervals between observed refresh completions, rather than an exact measurement of how far every summary trails the raw table. Between refreshes, queries against the Databricks MV can return older summaries.
HOW WAS PRE-AGGREGATION FRESHNESS MEASURED? (click to expand)
DATABRICKS SERIES
The freshness history was polled once a minute. The plotted metric is the time between distinct observed completed refresh IDs: floor(current completed_at − previous observed completed_at), in seconds. The selected active-ingestion evidence contains 1,021 such intervals; compact polling can miss intermediate completions.
SMOOTHING AND READOUTS
A centered 61-observation rolling mean averages about 1.8 minutes across the interval observations. An additional 11-observation display mean and shape-preserving interpolation produce a plotted peak of about 2.0 minutes. Readouts follow that displayed curve. This is a refresh-interval proxy, not a directly polled raw-to-MV watermark lag or a per-query stale-row count.
CLICKHOUSE BASELINE
The zero line is a benchmark-semantic baseline, not a separately polled provider metric. The raw table and incremental MV consume the same flushed insert block, with no independent summary-refresh interval between them. It does not claim zero source-to-query ingestion delay.
ClickHouse kept summaries current during ingestion; Databricks’ observed refreshes were roughly 1.8 minutes apart, leaving queries able to read older summaries.
That freshness difference follows the data into queries. First, the next chart accounts for what each preparation path cost.
Cost of keeping incoming data query-ready
ClickHouse handles ① columnar storage, ② ordering raw data, and ③ ordered pre-aggregation within its engine. Databricks spreads that work across ingestion, serverless clustering, and MV refresh. The chart below compares their accumulated preparation costs.
ClickHouse’s complete ingest service cost $28.69. Databricks’ modeled preparation cost totaled $695.60: $428.58 for Zerobus ingestion (①), $59.23 for asynchronous clustering (②), and $207.80 for MV refresh (③). Ingestion and refresh use allocated DBU usage; the folded pricing note gives the boundaries.
ClickHouse handled columnar storage, ordering, and pre-aggregation at 24.2× lower preparation cost in this run.
Now follow the raw data and summaries into query serving.
How ClickHouse Cloud serves queries
Ordered raw data lets drill-downs skip unrelated rows; pre-aggregations let interactive aggregations combine prepared results instead of recalculating them from individual quotes. The diagram below follows both paths through ClickHouse Cloud’s read service: one node with 16 CPUs and 64 GiB of memory, matching Databricks’ X-Small worker CPU count.
Drill-downs: Queries prune ordered raw data in the event-level MergeTree table. Filtering by symbol, the leading column of its (sym, t) sorting key, lets the read service skip unrelated quotes.
Interactive aggregations: Queries read current pre-aggregations from the AggregatingMergeTree table, combining prepared counts, sums, minima, and maxima from a much smaller set of daily summary rows. There is no independent refresh to wait for.
The result: drill-downs prune ordered raw data; interactive aggregations read summaries kept current during ingestion.
Databricks uses the same two query paths, but asynchronous preparation changes what queries can read.
How Databricks serves queries
The diagram below shows an X-Small Serverless SQL warehouse serving both workloads, fixed at one cluster with 16 worker vCPUs.
WHY WASN’T LAKEHOUSE//RT INCLUDED? (click to expand)
Lakehouse//RT is currently in Beta. We have requested access but have not yet received it.
Once Lakehouse//RT is generally available and we have access, we plan to repeat the full CostBench workload with Lakehouse//RT serving the queries and publish the results.
Drill-downs: Queries read the Delta table directly and filter by symbol. Liquid clustering on (sym, t) helps skip unrelated files after clustering runs, but newly published files can remain unclustered until asynchronous maintenance catches up.
Interactive aggregations: Queries read daily summaries from the incremental MV. They see the last completed refresh snapshot; newer raw rows are not automatically combined with it at query time. The answer can therefore omit data that has already reached the raw table.
The distinction shown by the two paths: Drill-downs can scan data whose clustering is unfinished; interactive aggregations can read summaries whose refresh is unfinished. Both arise from preparation that advances separately from ingestion.
WHAT DO DATABRICKS’ REFRESH AND VISIBILITY GUARANTEES MEAN? (click to expand)
SNAPSHOT CONTRACT
A refresh updates the MV from its source table. The measured aggregate queries read that materialized result directly. They did not force a synchronous refresh before each query or merge a raw-table delta into the answer. The measured latency therefore includes Databricks’ native allowance for stale summaries.
TRIGGER CONTRACT
TRIGGER ON UPDATE schedules refreshes when upstream data changes. AT MOST EVERY INTERVAL 1 MINUTE imposes a minimum interval between triggers, not a one-minute maximum on data age or refresh completion.
INCREMENTAL CONTRACT
INCREMENTAL STRICT controls the maintenance method: eligible changes are processed incrementally rather than silently falling back to a full rebuild. It does not tie MV visibility to each ingested block or make the MV continuously current.
VISIBILITY CONTRACT
The runner’s raw_rows records acknowledged durable ingestion progress. Zerobus can durably accept rows before they are published into the Delta table. A query pair aligned by durable ingestion progress is therefore not proof that both queries saw exactly the same raw-row count or an equally fresh summary snapshot.
COMPARISON BOUNDARY
The headline preserves the tested MV-serving path, including its freshness behavior. A design that forced a refresh or queried and re-aggregated the raw table would have a different latency and cost profile; that work is not included in this result.
The result: Databricks drill-downs may scan newly arrived, unclustered data; interactive aggregations may read summaries from an earlier refresh.
What this means for query performance and cost
Drill-downs over ordered raw data
Both systems used (sym, t) keys, but ClickHouse ordered rows during inserts while Databricks clustered files asynchronously. D1 and D2 query one stock’s growing history directly: hourly price summaries and a risk-and-liquidity profile. The chart below follows that raw-data path, bypassing the materialized views.
ClickHouse Cloud is yellow; Databricks is red.
HOW WERE QUERY LATENCIES AND ACCUMULATED RESULTS COMPARED? (click to expand)
QUERY DEFINITIONS
Part 1 describes A1–A4 and D1–D2. Both systems used the same analytical workload and schedule, with SQL adapted to each engine. Counts, sums, and extrema were compared exactly; the approximate-percentile portion of D2 used the contract’s explicit tolerance. MV freshness was not assumed to be identical.
CACHE POLICY
Query-result caching was disabled. Databricks’ compact query evidence reported no result-cache hits for accepted queries. Underlying data caches were allowed to warm normally; the benchmark did not flush them between queries.
TIMING
Databricks latency uses Query History total_duration_ms ÷ 1,000, excluding result-fetch time. ClickHouse uses its recorded query duration. These timing conventions are retained in normalized query cost. Median, P99, and maximum statistics pool unsmoothed per-query observations through 100 billion rows; P99 uses linear interpolation.
LATENCY CHARTS
Each system is plotted at its own observed ingestion progress. Aggregate trends use a centered seven-observation rolling median; drill-downs use five observations, with shrinking edge windows. No outliers are removed. The curves are display smoothing; the statistics and accumulated totals use the original recorded durations.
INTERACTIVE READOUTS
Values below the plots follow the displayed trends at the selected row count. Logarithmic scales keep milliseconds and seconds readable together; Linear shows absolute differences. Each chart has its own playback and row-position controls.
ACCEPTED ACCUMULATED WORKLOAD
The accepted selection contains 189 four-query aggregate batches and 32 two-query drill-down batches per system: 756 + 64 = 820 executions. ClickHouse observations were aligned with Databricks durability progress, through 112.849 billion rows for aggregation and 111.649 billion for drill-downs. The first full-dataset observation and later post-ingestion samples are excluded.
ACCUMULATED CHARTS
Query durations are added at corresponding ingestion counts as step sums, without smoothing or interpolating future executions into earlier positions. Preparation-cost lines allocate the complete-ingest total proportionally by row progress; they are not metered cost-at-time traces. The score sums query durations, not elapsed ingestion time.
ClickHouse Cloud:
- 722.5 ms median - 22.1× faster than Databricks
- 1.15 s P99 - 42.7× faster
- 1.18 s maximum - 48.2× faster
Accumulated runtime was 46.07 seconds across the 32 two-query batches. The ordered raw-data path kept both drill-downs near or below one second as the table grew.
Databricks:
- 15.99 s median
- 49.11 s P99
- 56.90 s maximum
Accumulated runtime was 18.88 minutes across the same 32 batches. The curves show substantially more time spent reading the growing raw table; asynchronous clustering can leave new files less organized for symbol filters.
Across the drill-down workload, ClickHouse Cloud was 24.6× faster overall.
Interactive aggregations over ordered pre-aggregated data
Daily pre-aggregations let A1–A4 combine prepared counts, sums, and price minima and maxima. A1–A2 filter summaries by symbol; A3–A4 read across all symbols. ClickHouse reads summaries maintained during inserts; Databricks reads the last refreshed snapshot. The chart below compares latency as ingestion continues.
ClickHouse Cloud:
- 13 ms median - 60.2× faster than Databricks
- 42 ms P99 - 48.4× faster
- 111 ms maximum - 47.7× faster
Accumulated runtime was 10.31 seconds across the 189 four-query batches.
Databricks:
- 782 ms median
- 2.02 s P99
- 5.30 s maximum
Accumulated runtime was 10.22 minutes across the same 189 batches. These queries read the MV snapshot without waiting for its next refresh. ClickHouse was faster even while maintaining summaries with the raw-data insert path; Databricks’ measured speed did not come with the same freshness.
Across the interactive-aggregation workload, ClickHouse Cloud was 59.5× faster overall, with summaries maintained during inserts.
Accumulated query cost and runtime
The two query paths add up differently: drill-downs account for most of Databricks’ accumulated runtime. The charts below combine both workloads and price their recorded durations using the benchmark’s read-side cost model.
ClickHouse Cloud accumulated 56.39 seconds of query runtime and Databricks 29.10 minutes. Their normalized query-serving costs were about $0.05 and $2.65, respectively. These are costs for the query work, not complete warehouse invoices.
Across the query workload, ClickHouse Cloud was 31× faster overall, with 54.2× lower normalized query cost.
From the field: Gala
Gala’s migration from Databricks to ClickHouse Cloud shows the production benefits of faster analytics at lower cost: more data available for analysis and self-service analytics accessible to more teams. On AWS, the team added continuous ingestion from more sources and rolled out Metabase so business teams could explore data themselves. By the end of the migration, it had decommissioned Databricks.
With the new layout configured correctly from the start, queries on previously unoptimized tables went from minutes to under 1 second. The data available for analysis grew from 3 TB to 9 TB during the migration, while initial costs were 30% lower. “We just don’t think about our data infrastructure as much anymore,” says Mike Rexford, Gala’s lead data analyst. Read Gala’s migration story.
Databricks can’t match ClickHouse Cloud for real-time analytics
While fresh data kept arriving, ClickHouse Cloud led across the complete tested path:
① Preparation cost: 24.2× lower for ingestion, ordering, and pre-aggregation.
② Normalized query cost: 54.2× lower for the queries.
③ Accumulated query runtime: 31× lower across both query workloads.
The final chart brings these results together, plotting combined preparation and normalized query cost against accumulated query runtime as fresh data arrives.
Up is lower total modeled cost; right is lower accumulated query runtime.
ClickHouse Cloud delivered 752× better end-to-end performance per dollar while continuously ingesting data and serving queries, with summaries kept current during inserts.
The difference begins as data arrives. ClickHouse prepares ordered raw data and current summaries on the insert path. Databricks adds separate clustering and refresh services, and its queries can encounter unfinished layout work or older summaries. In this workload, those services cost more to run, and both query paths still took longer.
That is why Databricks can’t match ClickHouse Cloud for real-time analytics.



