CertSafari
    Databricks Certified Data Analyst Associate· Lessons

    Domain 5 · Lesson 22/39

    Liquid Clustering: When It Speeds Up Filtered Queries

    Apply Liquid Clustering to improve query speed when filtering large tables on specific columns.

    9 min read
    2.56% of exam
    5 sources
    Published 3 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Explain what liquid clustering is and which older layout techniques it replaces
    • Describe how a clustered layout lets filtered queries skip files
    • Pick out the table and query patterns that benefit most from liquid clustering
    • Choose between liquid clustering and partitioning, based on table size and how the table is accessed

    Key concept

    Liquid clustering — A Delta Lake table layout where Databricks groups data files by the clustering keys you choose. Queries that filter on those keys can then skip files that cannot match. Unlike partitioning, you can change the keys later without rewriting the existing data.

    1.What liquid clustering is

    Say a table holds billions of sales rows and most dashboards filter it by customer_id or order_date. If those rows are scattered across every data file, each query has to read almost the whole table. The fix is to change how the data is laid out on storage, not to rewrite the query.

    Liquid clustering is Databricks' layout optimization for this. You name one or more clustering keys, and Databricks organizes the table's data around them. Databricks says liquid clustering "simplifies table management and optimizes query performance by automatically organizing data based on clustering keys."

    It replaces two older techniques: partitioning a table into fixed directories, and running ZORDER to co-locate values. The big practical difference is flexibility. With partitioning, you are stuck with the layout you picked at creation. With liquid clustering, you can redefine the keys later without rewriting the data already in the table, so the layout can change as your queries change.

    Databricks now treats liquid clustering as the default choice, not a special-case tuning step. It recommends liquid clustering for all new tables, including streaming tables and materialized views. For Delta Lake tables, it is generally available in Databricks Runtime 15.4 LTS and above.

    Checkpoint 1 of 5· Check yourself

    Which two older layout techniques does liquid clustering replace?

    Sources1

    2.Why a clustered layout makes filters faster

    The speed-up depends on a mechanism that runs whether or not you cluster: data skipping. When data is written to a Delta Lake table, Databricks automatically collects statistics for each file: the minimum and maximum values, null counts and total records. At query time, Databricks reads those statistics to skip irrelevant files and speed up queries.

    Skipping only works well when a file's minimum and maximum for the filtered column are close together. If every file covers almost the full range of customer_id, no file can be ruled out. Clustering groups rows with similar key values into the same files, so a filter on a clustering key matches far fewer files. Databricks' BI data-preparation guidance describes liquid clustering as improving "query performance with file and data skipping." Its action item for analysts is to apply it to large tables with filter patterns.

    Checkpoint 2 of 5· Check yourself

    What does Databricks use at query time to decide which files a filtered query can ignore?

    Sources23

    3.Tables that benefit most

    Databricks recommends liquid clustering for every new table, but it names some scenarios that gain the most. For an analyst, the first one matters most: queries that filter on high-cardinality columns, meaning columns with many distinct values, such as IDs.

    Scenarios Databricks lists as particularly benefiting from liquid clustering
    ScenarioWhat it looks like in practice
    Queries that filter on high cardinality columnsDashboards filtering on IDs or other columns with many distinct values
    Tables with heavy data skewA few key values account for most of the rows
    Fast growing tables that require maintenance and tuning effortTables whose layout would otherwise need constant manual tuning
    Tables with concurrent write requirementsSeveral writers updating the same table
    Tables with varied or changing access patternsFilter columns that differ between teams or change over time
    Tables where a typical partition key might return results from too many or too few partitionsPartitioning would produce either a few huge partitions or many tiny ones

    There is one more case where liquid clustering is the only option. If you filter on a field inside a struct column, such as struct_col.field, you cannot partition by it. Liquid clustering accepts a struct field as a clustering key, which makes it the only way to data-skip on that field without first extracting it into a top-level column.

    Checkpoint 3 of 5· Check yourself

    An analyst's slowest dashboard query filters a multi-terabyte table on transaction_id, which has millions of distinct values. Which description best explains why this table is a strong candidate for liquid clustering?

    Checkpoint 4 of 5· Exam question

    A data analyst manages a 40 TB Delta table that is queried mainly through filters on `customer_id`, a column with millions of distinct values. Query Profile shows most runtime is spent scanning files that do not match the filter. Which change to the table's physical layout is most likely to speed up these queries?

    Sources14

    4.Liquid clustering versus partitioning

    Many people assume that a big table needs partitioning. Databricks' guidance points the other way. Partitioning has a hidden cost: an ineffective strategy can hurt query performance, and fixing it means a full rewrite of the data, which is expensive and slow on a large table. Liquid clustering works for both low- and high-cardinality columns and avoids the fixed partition boundaries and small-file problems of static partitioning.

    Databricks' layout guidance by table size
    Table sizeRecommendation
    Less than 1 TBDon't partition
    More than 1 TB to 100 TBUse liquid clustering instead of partitioning
    100 TB or morePartitioning might help, but use liquid clustering first and verify performance improvements

    You also cannot mix the approaches on one table. Clustering is not compatible with partitioning or ZORDER, and the ZORDER BY clause of OPTIMIZE cannot be used on a table with liquid clustering. Choosing liquid clustering means letting it handle the whole layout.

    Checkpoint 5 of 5· Check yourself

    A table already uses liquid clustering. A colleague suggests also running OPTIMIZE ... ZORDER BY (eventType) to speed up filters on a second column. What happens?

    Sources4

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A very large table should be partitioned on its main filter column to speed up queries.Why is that wrong?

      For tables between 1 TB and 100 TB, Databricks says to use liquid clustering instead of partitioning. Above 100 TB, it still recommends trying clustering first.

      Covered in Liquid clustering versus partitioning

    2. 2.You can combine liquid clustering with partitioning or ZORDER for extra speed.Why is that wrong?

      Liquid clustering cannot be combined with partitioning or ZORDER on the same table. Put the extra filter columns in the clustering keys instead.

      Covered in Liquid clustering versus partitioning

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Liquid clustering is a data layout optimization technique that replaces table partitioning and ZORDER.”
      ↩︎ What liquid clustering is
      “Databricks recommends liquid clustering for all new tables, including streaming tables and materialized views.”
      ↩︎ What liquid clustering is
      “Queries that filter on high cardinality columns.”
      ↩︎ Tables that benefit most
      “Liquid clustering is a data layout optimization technique that replaces table partitioning and ZORDER.”
      ↩︎ Key concept
      “Clustering is not compatible with partitioning or ZORDER.”
      ↩︎ Exam trap 2
      “you can redefine clustering keys without rewriting existing data”
      ↩︎ Prediction
    2. 2.
      “at query time to skip irrelevant files and speed up queries”
      ↩︎ Why a clustered layout makes filters faster
    3. 4.
      “Liquid clustering is the only way to data-skip on a struct field without first extracting it into a top-level column.”
      ↩︎ Tables that benefit most
      “An ineffective partitioning strategy might negatively affect query performance and require a full rewrite of data to fix.”
      ↩︎ Liquid clustering versus partitioning
      “avoids the fixed partition boundaries and small-file issues common with static partitioning”
      ↩︎ Liquid clustering versus partitioning
      “With more than 1 TB to 100 TB of data, use liquid clustering instead of partitioning.”
      ↩︎ Exam trap 1
      “With more than 1 TB to 100 TB of data, use liquid clustering instead of partitioning.”
      ↩︎ Prediction

    Also cited

    Continue to page 2 of 2

    Liquid Clustering in SQL: CLUSTER BY and OPTIMIZE

    Spotted a mistake, or was something unclear? Tell us.