CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 19/48

    Spark DataFrame summary statistics with .summary() and dbutils.data.summarize

    Compute summary statistics on a Spark DataFrame using .summary() or dbutils data summaries

    15 min read
    2.08% of exam
    5 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Compute default and chosen summary statistics on a Spark DataFrame with .summary()
    • Explain how .summary() differs from .describe()
    • Generate a data summary with dbutils.data.summarize() or the notebook Data Profile, and choose the precise setting
    • Weigh what each method costs and what it returns

    Key concept

    Summary statistics for exploratory data analysis — Summary statistics boil each column down to a handful of numbers: count, mean, spread, extremes and percentiles. You use them to understand a dataset before you clean or model it. In Databricks you get them as a DataFrame from .summary(), or as an interactive visual report from dbutils.data.summarize().

    1.DataFrame.summary(): the default report and choosing your own statistics

    Before you train anything on a new table, you need to know what is in it. How many values does each column hold? Where is the centre, and how wide is the spread? Are the minimum and maximum believable? The quickest way to answer these on a Spark DataFrame is the summary() method. It is a DataFrame method, so it runs on Spark and works on any Spark DataFrame, in a Databricks notebook or anywhere else. It returns a new DataFrame. The first column is called summary and names each statistic. Every other column is one of the input columns.

    Default summary(): eight statistics, including the quartilespython
    df.select("age", "weight", "height").summary().show()
    # +-------+----+------------------+-----------------+
    # |summary| age|            weight|           height|
    # +-------+----+------------------+-----------------+
    # |  count|   3|                 3|                3|
    # |   mean|12.0| 40.73333333333333|            145.0|
    # | stddev| 1.0|3.1722757341273704|4.763402145525822|
    # |    min|  11|              37.8|            142.2|
    # |    25%|  11|              37.8|            142.2|
    # |    50%|  12|              40.3|            142.3|
    # |    75%|  13|              44.1|            150.5|
    # |    max|  13|              44.1|            150.5|
    # +-------+----+------------------+-----------------+

    Notice the pattern. You choose the columns with select() first. The arguments you pass to summary() itself are names of statistics, not column names. Pass any subset of count, mean, stddev, min and max. You can also pass any percentile written as a percentage string, such as "25%", "90%" or "99%". The output rows come back in the order you asked for them. Asking only for the range and the quartiles is a cheap way to look for suspicious tails: a max far above the 75% value suggests the column is skewed or has extreme values.

    Passing statistic names restricts the output to just those rowspython
    df.select("age", "weight", "height").summary("count", "min", "25%", "75%", "max").show()
    # +-------+---+------+------+
    # |summary|age|weight|height|
    # +-------+---+------+------+
    # |  count|  3|     3|     3|
    # |    min| 11|  37.8| 142.2|
    # |    25%| 11|  37.8| 142.2|
    # |    75%| 13|  44.1| 150.5|
    # |    max| 13|  44.1| 150.5|
    # +-------+---+------+------+
    What the two summary() calls above return
    CallRows in the result
    summary()count, mean, stddev, min, 25%, 50%, 75%, max
    summary("count", "min", "25%", "75%", "max")count, min, 25%, 75%, max (only the requested statistics)

    Checkpoint 1 of 7· Check yourself

    You want only the minimum, the median and the 90th percentile of the price and carat columns. Which call does that?

    Checkpoint 2 of 7· Exam question

    A data scientist runs `df.describe()` on a Spark DataFrame with a `price` column and gets `count`, `mean`, `stddev`, `min`, and `max`, but also needs the 25th, 50th, and 75th percentiles before deciding whether to log-transform the column. Which approach adds the quartiles without writing a custom aggregation?

    Sources1

    2.describe() versus summary()

    Spark has an older sibling of this method, describe(). Exam questions like to put the two side by side. describe() takes column names as its arguments and always returns the same five statistics: count, mean, stddev, min and max. It gives you no percentiles, so it cannot show you a median or quartiles. summary() returns those same five plus the 25%, 50% and 75% percentiles by default. It also lets you pick exactly which statistics to compute, including any percentile you like. The Databricks reference points you from one to the other.

    describe() on several columns: five statistics, no percentilespython
    df.describe(['age', 'weight', 'height']).show()
    # +-------+----+------------------+-----------------+
    # |summary| age|            weight|           height|
    # +-------+----+------------------+-----------------+
    # |  count|   3|                 3|                3|
    # |   mean|12.0| 40.73333333333333|            145.0|
    # | stddev| 1.0|3.1722757341273704|4.763402145525822|
    # |    min|  11|              37.8|            142.2|
    # |    max|  13|              44.1|            150.5|
    # +-------+----+------------------+-----------------+

    Both methods handle numeric and string columns. Both return an ordinary Spark DataFrame, so you can show() it, display() it, or filter it like any other DataFrame. Both also carry the same warning: the layout of the result DataFrame is not a stable contract. That is fine for looking at data. Do not build a production pipeline that relies on the exact layout of the result.

    Checkpoint 3 of 7· Check yourself

    Which statement about describe() and summary() is correct?

    Sources2

    3.dbutils.data.summarize() and the notebook Data Profile

    summary() gives you numbers in a grid. Sometimes you want a visual overview instead: a histogram for each column, the most frequent values of each categorical column, and a count of missing values. Databricks provides this through the data module of Databricks Utilities. Its single command, summarize, is described as a way to summarize a Spark DataFrame and visualize the statistics. The data module is marked EXPERIMENTAL in the dbutils module list. The command is in Public Preview and needs Databricks Runtime 9.0 or above.

    Summarizing a Spark DataFrame read from a sample CSVpython
    df = spark.read.format('csv').load(
      '/databricks-datasets/Rdatasets/data-001/csv/ggplot2/diamonds.csv',
      header=True,
      inferSchema=True
    )
    dbutils.data.summarize(df)

    Two details settle the predict. First, summarize accepts either an Apache Spark DataFrame or a pandas DataFrame, and you can call it from Python, Scala or R. Second, its signature is summarize(df: Object, precise: boolean): void. It shows its results in the notebook and returns nothing. That is the practical difference from summary(). If you need the statistics as data you can filter, join or save, use summary(). If you want to look at distributions quickly, use summarize.

    You can get the same report without writing code. After display(df), click + > Data Profile next to the Table tab in the output. This runs a new command that profiles the DataFrame. The report covers numeric, string and date columns, with a histogram for each. The UI Data Profile is available in Databricks Runtime 9.1 LTS and above. The documentation describes dbutils.data.summarize as the way to generate the same profile programmatically.

    Checkpoint 4 of 7· Fill the gap

    Which command completes this sample so it shows an interactive statistical summary of the DataFrame?

    df = spark.read.format('csv').load(
      '/databricks-datasets/Rdatasets/data-001/csv/ggplot2/diamonds.csv',
      header=True,
      inferSchema=True
    )
    dbutils.data. ? (df)

    Checkpoint 5 of 7· Exam question

    Before choosing between mean and median imputation for a categorical feature with a very high number of distinct string values, an ML engineer needs the exact distinct-value count and the exact most-frequent values for that column, not an estimate that could shift between runs. Which call satisfies this requirement?

    Sources34

    4.Precision, cost and choosing the right tool

    A summary of a few thousand rows is instant. A summary of a billion-row table is not. dbutils.data.summarize analyzes the complete contents of the DataFrame, and the docs warn that this can be very expensive on very large DataFrames. To keep run time down, Databricks Runtime 10.4 LTS and above adds a precise parameter. By default (precise set to false), some statistics are approximations. Setting it to true makes most statistics exact, but not all of them.

    What the precise parameter changes in dbutils.data.summarize
    Statisticprecise = false (default)precise = true
    Distinct values for categorical columns~5% relative error for high-cardinality columnsExact
    Frequent value countsError up to 0.01% when distinct values exceed 10000Exact
    Histograms and percentile estimates (numeric)Error up to 0.01% of total rowsError up to 0.0001% of total rows (still approximate)

    A tooltip at the top of the output tells you which mode the run used, so you can tell whether a summary was approximate. Exact distinct counts and frequent-value counts matter most for high-cardinality categorical columns. Turning on precise costs extra compute, so use it only when an approximate count could mislead you. Note that summary() percentiles are also approximate. When you need control over the accuracy of a single quantile, approxQuantile takes an explicit relative error, and a relative error of zero gives exact quantiles at a high cost.

    Here is a simple way to choose. To see what a new dataset looks like, run display(df) and open the Data Profile, or call dbutils.data.summarize. When you need specific statistics as a DataFrame, for example quartiles to feed into a later cleaning step, use select(...).summary(...).

    Checkpoint 6 of 7· Check yourself

    You run dbutils.data.summarize(df, precise=True). Which results can still be approximate?

    Checkpoint 7 of 7· Exam question

    A notebook needs a compact statistics table for a `revenue` column showing only `count`, `min`, `50%`, and `max` — omitting `mean`, `stddev`, and the other default percentiles — to keep a printed report narrow. Which call produces exactly that output?

    Sources35

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.describe() gives you the median and quartiles, just like summary().Why is that wrong?

      describe() returns only count, mean, stddev, min and max. You need summary() for percentiles and for choosing which statistics to compute.

      Covered in describe() versus summary()

    2. 2.dbutils.data.summarize() returns a statistics DataFrame you can assign to a variable and save.Why is that wrong?

      summarize shows its results in the notebook and returns void. Use DataFrame.summary() when you need the statistics as a DataFrame.

      Covered in dbutils.data.summarize() and the notebook Data Profile

    3. 3.dbutils.data.summarize samples the data, so it is cheap on any size of table.Why is that wrong?

      It analyzes the whole DataFrame, which can be very expensive on large data. The precise=false default approximates some statistics to cut run time, but it still reads all of the data.

      Covered in Precision, cost and choosing the right tool

    4. 4.Setting precise=True makes every statistic in the data summary exact.Why is that wrong?

      Numeric histograms and percentiles stay approximate, though with a much smaller error.

      Covered in Precision, cost and choosing the right tool

    Practise it for real

    Compare summary(), describe() and dbutils.data.summarize() on the same sample dataset in a Databricks notebook.

    1. 1.Load diamonds.csv from /databricks-datasets/Rdatasets/data-001/csv/ggplot2/ with spark.read.format('csv').load(..., header=True, inferSchema=True).

      Why: All three tools need the same Spark DataFrame so you can compare their output.

      You should see: A Spark DataFrame with numeric columns such as carat and price and string columns such as cut.

    2. 2.Run df.select("carat", "price").describe().show(), then df.select("carat", "price").summary().show().

      Why: Shows the fixed five-statistic output of describe() next to the default eight statistics of summary().

      You should see: summary() has extra 25%, 50% and 75% rows. In price, the mean should sit above the 50% value, a sign of right skew.

    3. 3.Run df.select("price").summary("min", "50%", "90%", "max").show().

      Why: Practises passing statistic names, including an arbitrary percentile.

      You should see: Four rows, in the order you asked for them.

    4. 4.Run dbutils.data.summarize(df).

      Why: Produces the visual data summary with a histogram per column and top values for categorical columns.

      You should see: An interactive report in the cell output. The tooltip at the top shows the approximate mode.

    5. 5.Run dbutils.data.summarize(df, precise=True) (Databricks Runtime 10.4 LTS or above).

      Why: Shows the effect of precise on accuracy and run time.

      You should see: The tooltip reports the precise mode. Distinct counts are now exact, and the run may take longer.

    Stuck? Get a nudge

    If summarize isn't available, check that the compute runs a Databricks Runtime that supports the data utility (9.0+, and 10.4 LTS+ for precise).

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Available statistics are: count, mean, stddev, min, max, arbitrary approximate percentiles specified as a percentage (e.g., 75%).”
      ↩︎ DataFrame.summary(): the default report and choosing your own statistics
      “DataFrame: A new DataFrame that provides statistics for the given DataFrame.”
      ↩︎ DataFrame.summary(): the default report and choosing your own statistics
      “This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting DataFrame.”
      ↩︎ Key concept
    2. 2.
      “Computes basic statistics for numeric and string columns.”
      ↩︎ describe() versus summary()
      “Use summary for expanded statistics and control over which statistics to compute.”
      ↩︎ describe() versus summary()
      “Use summary for expanded statistics and control over which statistics to compute.”
      ↩︎ Exam trap 1
    3. 3.
      “Calculates and displays summary statistics of an Apache Spark DataFrame or pandas DataFrame.”
      ↩︎ dbutils.data.summarize() and the notebook Data Profile
      “Summarize a Spark DataFrame and visualize the statistics to get quick insights”
      ↩︎ dbutils.data.summarize() and the notebook Data Profile
      “In Databricks Runtime 10.4 LTS and above, you can use the additional precise parameter to adjust the precision of the computed statistics.”
      ↩︎ Precision, cost and choosing the right tool
      “When precise is set to false (the default), some returned statistics include approximations to reduce run time.”
      ↩︎ Precision, cost and choosing the right tool
      “The tooltip at the top of the data summary output indicates the mode of the current run.”
      ↩︎ Precision, cost and choosing the right tool
      “summarize(df: Object, precise: boolean): void”
      ↩︎ Exam trap 2
      “This command analyzes the complete contents of the DataFrame. Running this command for very large DataFrames can be very expensive.”
      ↩︎ Exam trap 3
      “All statistics except for the histograms and percentiles for numeric columns are now exact.”
      ↩︎ Exam trap 4
      “All statistics except for the histograms and percentiles for numeric columns are now exact.”
      ↩︎ Checkpoint
    4. 4.
      “The data profile includes summary statistics for numeric, string, and date columns as well as histograms of the value distributions for each column.”
      ↩︎ dbutils.data.summarize() and the notebook Data Profile
      “You can also generate data profiles programmatically; see summarize command (dbutils.data.summarize).”
      ↩︎ dbutils.data.summarize() and the notebook Data Profile
    5. 5.
      “If set to zero, the exact quantiles are computed, which could be very expensive.”
      ↩︎ Precision, cost and choosing the right tool

    Spotted a mistake, or was something unclear? Tell us.