What you will be able to do
- Compute default and chosen summary statistics on a Spark DataFrame with .summary()
- Explain how .summary() differs from .describe()
- Generate a data summary with dbutils.data.summarize() or the notebook Data Profile, and choose the precise setting
- Weigh what each method costs and what it returns
Key concept
Summary statistics for exploratory data analysis — Summary statistics boil each column down to a handful of numbers: count, mean, spread, extremes and percentiles. You use them to understand a dataset before you clean or model it. In Databricks you get them as a DataFrame from .summary(), or as an interactive visual report from dbutils.data.summarize().
1.DataFrame.summary(): the default report and choosing your own statistics
Before you train anything on a new table, you need to know what is in it. How many values does each column hold? Where is the centre, and how wide is the spread? Are the minimum and maximum believable? The quickest way to answer these on a Spark DataFrame is the summary() method. It is a DataFrame method, so it runs on Spark and works on any Spark DataFrame, in a Databricks notebook or anywhere else. It returns a new DataFrame. The first column is called summary and names each statistic. Every other column is one of the input columns.
df.select("age", "weight", "height").summary().show()
# +-------+----+------------------+-----------------+
# |summary| age| weight| height|
# +-------+----+------------------+-----------------+
# | count| 3| 3| 3|
# | mean|12.0| 40.73333333333333| 145.0|
# | stddev| 1.0|3.1722757341273704|4.763402145525822|
# | min| 11| 37.8| 142.2|
# | 25%| 11| 37.8| 142.2|
# | 50%| 12| 40.3| 142.3|
# | 75%| 13| 44.1| 150.5|
# | max| 13| 44.1| 150.5|
# +-------+----+------------------+-----------------+Notice the pattern. You choose the columns with select() first. The arguments you pass to summary() itself are names of statistics, not column names. Pass any subset of count, mean, stddev, min and max. You can also pass any percentile written as a percentage string, such as "25%", "90%" or "99%". The output rows come back in the order you asked for them. Asking only for the range and the quartiles is a cheap way to look for suspicious tails: a max far above the 75% value suggests the column is skewed or has extreme values.
df.select("age", "weight", "height").summary("count", "min", "25%", "75%", "max").show()
# +-------+---+------+------+
# |summary|age|weight|height|
# +-------+---+------+------+
# | count| 3| 3| 3|
# | min| 11| 37.8| 142.2|
# | 25%| 11| 37.8| 142.2|
# | 75%| 13| 44.1| 150.5|
# | max| 13| 44.1| 150.5|
# +-------+---+------+------+| Call | Rows in the result |
|---|---|
| summary() | count, mean, stddev, min, 25%, 50%, 75%, max |
| summary("count", "min", "25%", "75%", "max") | count, min, 25%, 75%, max (only the requested statistics) |
Checkpoint 1 of 7· Check yourself
You want only the minimum, the median and the 90th percentile of the price and carat columns. Which call does that?
Pick the columns with select(), then pass statistic names to summary(). Percentiles are written as percentage strings such as "90%".
“Available statistics are: count, mean, stddev, min, max, arbitrary approximate percentiles specified as a percentage (e.g., 75%).”Source: docs.databricks.com
Checkpoint 2 of 7· Exam question
A data scientist runs `df.describe()` on a Spark DataFrame with a `price` column and gets `count`, `mean`, `stddev`, `min`, and `max`, but also needs the 25th, 50th, and 75th percentiles before deciding whether to log-transform the column. Which approach adds the quartiles without writing a custom aggregation?
Correct answer: A — Call `df.summary()`, since its default statistic set adds the 25%, 50%, and 75% percentiles alongside the same count, mean, stddev, min, and max that `.describe()` already returns.
- A. `.summary()` is the built-in Spark method whose default statistic list already extends `.describe()` with the 25%, 50%, and 75% percentiles, so a single call returns everything needed with no custom aggregation.
- B. `describe()` only accepts column names as arguments, not a percentiles list; it does not expose the optional statistic-name interface that `summary()` provides, so this call would fail or be ignored.
- C. Pulling all rows to the driver with `.toPandas()` before computing quantiles works but is unnecessary and does not scale, and the claim that neither Spark method supports percentiles is false since `.summary()` already does.
- D. This combination of `approxQuantile` and `.agg()` would eventually produce the same numbers, but it needlessly splits one statistic call that `.summary()` already handles in a single default invocation.
Sources1
2.describe() versus summary()
Spark has an older sibling of this method, describe(). Exam questions like to put the two side by side. describe() takes column names as its arguments and always returns the same five statistics: count, mean, stddev, min and max. It gives you no percentiles, so it cannot show you a median or quartiles. summary() returns those same five plus the 25%, 50% and 75% percentiles by default. It also lets you pick exactly which statistics to compute, including any percentile you like. The Databricks reference points you from one to the other.
df.describe(['age', 'weight', 'height']).show()
# +-------+----+------------------+-----------------+
# |summary| age| weight| height|
# +-------+----+------------------+-----------------+
# | count| 3| 3| 3|
# | mean|12.0| 40.73333333333333| 145.0|
# | stddev| 1.0|3.1722757341273704|4.763402145525822|
# | min| 11| 37.8| 142.2|
# | max| 13| 44.1| 150.5|
# +-------+----+------------------+-----------------+Both methods handle numeric and string columns. Both return an ordinary Spark DataFrame, so you can show() it, display() it, or filter it like any other DataFrame. Both also carry the same warning: the layout of the result DataFrame is not a stable contract. That is fine for looking at data. Do not build a production pipeline that relies on the exact layout of the result.
They can't see the median or the quartiles. A mean pulled up by a long right tail looks fine on its own, but summary() would show the 50% value sitting well below the mean and a big gap between the 75% value and the max. Those are the usual signs of skew.
Checkpoint 3 of 7· Check yourself
Which statement about describe() and summary() is correct?
describe() computes a fixed set of basic statistics for the columns you name. The docs point you to summary() when you want more statistics and control over which ones are computed.
“Use summary for expanded statistics and control over which statistics to compute.”Source: docs.databricks.com
Sources2
3.dbutils.data.summarize() and the notebook Data Profile
summary() gives you numbers in a grid. Sometimes you want a visual overview instead: a histogram for each column, the most frequent values of each categorical column, and a count of missing values. Databricks provides this through the data module of Databricks Utilities. Its single command, summarize, is described as a way to summarize a Spark DataFrame and visualize the statistics. The data module is marked EXPERIMENTAL in the dbutils module list. The command is in Public Preview and needs Databricks Runtime 9.0 or above.
df = spark.read.format('csv').load(
'/databricks-datasets/Rdatasets/data-001/csv/ggplot2/diamonds.csv',
header=True,
inferSchema=True
)
dbutils.data.summarize(df)Two details settle the predict. First, summarize accepts either an Apache Spark DataFrame or a pandas DataFrame, and you can call it from Python, Scala or R. Second, its signature is summarize(df: Object, precise: boolean): void. It shows its results in the notebook and returns nothing. That is the practical difference from summary(). If you need the statistics as data you can filter, join or save, use summary(). If you want to look at distributions quickly, use summarize.
You can get the same report without writing code. After display(df), click + > Data Profile next to the Table tab in the output. This runs a new command that profiles the DataFrame. The report covers numeric, string and date columns, with a histogram for each. The UI Data Profile is available in Databricks Runtime 9.1 LTS and above. The documentation describes dbutils.data.summarize as the way to generate the same profile programmatically.
Checkpoint 4 of 7· Fill the gap
Which command completes this sample so it shows an interactive statistical summary of the DataFrame?
df = spark.read.format('csv').load(
'/databricks-datasets/Rdatasets/data-001/csv/ggplot2/diamonds.csv',
header=True,
inferSchema=True
)
dbutils.data. ? (df)The data utility module's only command is summarize. summary and describe are Spark DataFrame methods, not dbutils commands.
Source: docs.databricks.comCheckpoint 5 of 7· Exam question
Before choosing between mean and median imputation for a categorical feature with a very high number of distinct string values, an ML engineer needs the exact distinct-value count and the exact most-frequent values for that column, not an estimate that could shift between runs. Which call satisfies this requirement?
Correct answer: A — `dbutils.data.summarize(df, precise=True)`, because setting `precise` to `True` computes exact statistics for every column except histograms and percentiles, which still carry a tiny bounded error.
- A. Setting `precise=True` on `dbutils.data.summarize()` switches most statistics, including distinct-value counts and most-frequent values, from approximate to exact, leaving only histograms and percentiles with a very small residual error.
- B. The default `precise=False` mode is the approximate one; distinct-value counts for high-cardinality columns carry roughly 5% relative error by default, so this option describes the opposite of the actual default behavior.
- C. `summary()` never computes distinct-value counts at all regardless of which statistic names are passed to it, and its precision is unrelated to how many statistic strings are requested.
- D. Chaining `distinct().count()` does yield an exact distinct count, but `describe()` does not report most-frequent values for a column, so this pairing still leaves part of the requirement unmet.
4.Precision, cost and choosing the right tool
A summary of a few thousand rows is instant. A summary of a billion-row table is not. dbutils.data.summarize analyzes the complete contents of the DataFrame, and the docs warn that this can be very expensive on very large DataFrames. To keep run time down, Databricks Runtime 10.4 LTS and above adds a precise parameter. By default (precise set to false), some statistics are approximations. Setting it to true makes most statistics exact, but not all of them.
| Statistic | precise = false (default) | precise = true |
|---|---|---|
| Distinct values for categorical columns | ~5% relative error for high-cardinality columns | Exact |
| Frequent value counts | Error up to 0.01% when distinct values exceed 10000 | Exact |
| Histograms and percentile estimates (numeric) | Error up to 0.01% of total rows | Error up to 0.0001% of total rows (still approximate) |
A tooltip at the top of the output tells you which mode the run used, so you can tell whether a summary was approximate. Exact distinct counts and frequent-value counts matter most for high-cardinality categorical columns. Turning on precise costs extra compute, so use it only when an approximate count could mislead you. Note that summary() percentiles are also approximate. When you need control over the accuracy of a single quantile, approxQuantile takes an explicit relative error, and a relative error of zero gives exact quantiles at a high cost.
Here is a simple way to choose. To see what a new dataset looks like, run display(df) and open the Data Profile, or call dbutils.data.summarize. When you need specific statistics as a DataFrame, for example quartiles to feed into a later cleaning step, use select(...).summary(...).
Checkpoint 6 of 7· Check yourself
You run dbutils.data.summarize(df, precise=True). Which results can still be approximate?
With precise set to true, everything except the numeric histograms and percentiles is exact. Those two shrink to an error of at most 0.0001% of the total number of rows.
“All statistics except for the histograms and percentiles for numeric columns are now exact.”Source: docs.databricks.com
Checkpoint 7 of 7· Exam question
A notebook needs a compact statistics table for a `revenue` column showing only `count`, `min`, `50%`, and `max` — omitting `mean`, `stddev`, and the other default percentiles — to keep a printed report narrow. Which call produces exactly that output?
Correct answer: A — `df.summary("count", "min", "50%", "max")`, because passing statistic names as string arguments overrides the default list and returns only the requested rows in that order.
- A. `summary()` treats its string arguments as an explicit override of the default statistic list, so passing exactly these four names returns only those rows, in the order supplied.
- B. The statistic names in `summary()` output appear as row values in a `summary` column rather than as separate output columns, so `.select()` on those names would not select the intended rows and would error or return nothing useful.
- C. `describe()` only accepts column names as arguments and has no optional statistic-name parameter, so it cannot be narrowed to a custom subset of statistics the way `summary()` can.
- D. This combination of `agg()` expressions can produce the same four values, but it is not the only way to control the output rows since `summary()` already supports the same narrowing with far less code.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.describe() gives you the median and quartiles, just like summary().Why is that wrong?
describe() returns only count, mean, stddev, min and max. You need summary() for percentiles and for choosing which statistics to compute.
Covered in describe() versus summary()
2.dbutils.data.summarize() returns a statistics DataFrame you can assign to a variable and save.Why is that wrong?
summarize shows its results in the notebook and returns void. Use DataFrame.summary() when you need the statistics as a DataFrame.
Covered in dbutils.data.summarize() and the notebook Data Profile
3.dbutils.data.summarize samples the data, so it is cheap on any size of table.Why is that wrong?
It analyzes the whole DataFrame, which can be very expensive on large data. The precise=false default approximates some statistics to cut run time, but it still reads all of the data.
Covered in Precision, cost and choosing the right tool
4.Setting precise=True makes every statistic in the data summary exact.Why is that wrong?
Numeric histograms and percentiles stay approximate, though with a much smaller error.
Covered in Precision, cost and choosing the right tool
Practise it for real
Compare summary(), describe() and dbutils.data.summarize() on the same sample dataset in a Databricks notebook.
1.Load diamonds.csv from /databricks-datasets/Rdatasets/data-001/csv/ggplot2/ with spark.read.format('csv').load(..., header=True, inferSchema=True).
Why: All three tools need the same Spark DataFrame so you can compare their output.
You should see: A Spark DataFrame with numeric columns such as carat and price and string columns such as cut.
2.Run df.select("carat", "price").describe().show(), then df.select("carat", "price").summary().show().
Why: Shows the fixed five-statistic output of describe() next to the default eight statistics of summary().
You should see: summary() has extra 25%, 50% and 75% rows. In price, the mean should sit above the 50% value, a sign of right skew.
3.Run df.select("price").summary("min", "50%", "90%", "max").show().
Why: Practises passing statistic names, including an arbitrary percentile.
You should see: Four rows, in the order you asked for them.
4.Run dbutils.data.summarize(df).
Why: Produces the visual data summary with a histogram per column and top values for categorical columns.
You should see: An interactive report in the cell output. The tooltip at the top shows the approximate mode.
5.Run dbutils.data.summarize(df, precise=True) (Databricks Runtime 10.4 LTS or above).
Why: Shows the effect of precise on accuracy and run time.
You should see: The tooltip reports the precise mode. Distinct counts are now exact, and the run may take longer.
Stuck? Get a nudge
If summarize isn't available, check that the compute runs a Databricks Runtime that supports the data utility (9.0+, and 10.4 LTS+ for precise).
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Available statistics are: count, mean, stddev, min, max, arbitrary approximate percentiles specified as a percentage (e.g., 75%).”
↩︎ DataFrame.summary(): the default report and choosing your own statistics“DataFrame: A new DataFrame that provides statistics for the given DataFrame.”
↩︎ DataFrame.summary(): the default report and choosing your own statistics“This function is meant for exploratory data analysis, as we make no guarantee about the backward compatibility of the schema of the resulting DataFrame.”
↩︎ Key concept - 2.
“Computes basic statistics for numeric and string columns.”
↩︎ describe() versus summary()“Use summary for expanded statistics and control over which statistics to compute.”
↩︎ describe() versus summary()“Use summary for expanded statistics and control over which statistics to compute.”
↩︎ Exam trap 1 - 3.
“Calculates and displays summary statistics of an Apache Spark DataFrame or pandas DataFrame.”
↩︎ dbutils.data.summarize() and the notebook Data Profile“Summarize a Spark DataFrame and visualize the statistics to get quick insights”
↩︎ dbutils.data.summarize() and the notebook Data Profile“In Databricks Runtime 10.4 LTS and above, you can use the additional precise parameter to adjust the precision of the computed statistics.”
↩︎ Precision, cost and choosing the right tool“When precise is set to false (the default), some returned statistics include approximations to reduce run time.”
↩︎ Precision, cost and choosing the right tool“The tooltip at the top of the data summary output indicates the mode of the current run.”
↩︎ Precision, cost and choosing the right tool“summarize(df: Object, precise: boolean): void”
↩︎ Exam trap 2“This command analyzes the complete contents of the DataFrame. Running this command for very large DataFrames can be very expensive.”
↩︎ Exam trap 3“All statistics except for the histograms and percentiles for numeric columns are now exact.”
↩︎ Exam trap 4“All statistics except for the histograms and percentiles for numeric columns are now exact.”
↩︎ Checkpoint - 4.
“The data profile includes summary statistics for numeric, string, and date columns as well as histograms of the value distributions for each column.”
↩︎ dbutils.data.summarize() and the notebook Data Profile“You can also generate data profiles programmatically; see summarize command (dbutils.data.summarize).”
↩︎ dbutils.data.summarize() and the notebook Data Profile - 5.
“If set to zero, the exact quantiles are computed, which could be very expensive.”
↩︎ Precision, cost and choosing the right tool