What you will be able to do
- Use a box chart to compare the spread and skew of a continuous feature across categories
- Use a scatter or bubble chart to look at two continuous features together
- Create, aggregate and inspect visualizations from a notebook result, and know which chart types are cut off at 64,000 rows
1.Continuous by categorical: the box chart
Numeric features usually need to be compared across groups. A box chart summarises the distribution of a numeric column separately for each level of a categorical column. Each box is built from quartiles, so at a glance you can compare value ranges across categories and see where each group's values sit, how spread out they are, and how skewed. In the Databricks rendering, the darker line in each box marks the interquartile range.
The column roles are fixed. In a vertical box chart, the X column must be categorical and the Y column numeric. If you select Horizontal Chart, the roles swap. The documentation's example puts l_returnflag on X and l_extendedprice on Y, and also groups by l_shipmode.
| Setting | Vertical chart | Horizontal Chart selected |
|---|---|---|
| X column | A categorical column | A number column |
| Y columns | A number column | A categorical column |
| Default grouping | Grouped by the X axis | Grouped by the Y axis |
| Show all points | Show each X axis value separately instead of a box | Same option |
| Missing and NULL values | Hide them, or convert to 0 and show them | Same option |
This limit changes how you read a box chart on a large table: the quartiles describe at most the first 64,000 rows, not the whole dataset. Two other settings change what the box shows. Under Missing and NULL values, you can either hide missing values or convert them to 0 and plot them. Converting them adds a cluster of zeros that pulls the lower quartile down. The Y axis Scale includes Logarithmic, which helps when one category's values are far larger than another's.
Checkpoint 1 of 5· Check yourself
You want a vertical box chart showing how l_extendedprice varies by l_returnflag. How should you assign the columns?
In a vertical chart, the X column is categorical and the Y column is numeric. The roles swap only when you select Horizontal Chart.
“X column: X axis values. For a vertical chart, choose a categorical column. For a horizontal chart, choose a number column.”Source: docs.databricks.com
Checkpoint 2 of 5· Exam question
A machine learning engineer wants to understand the overall shape of the distribution of a continuous `transaction_amount` column — whether it is roughly symmetric, skewed, or bimodal — before choosing an imputation strategy. Which visualization is most appropriate?
Correct answer: A — Plot a histogram of `transaction_amount`, dividing the value range into bins and showing the count of rows in each bin.
- A. A histogram bins a continuous variable's range into intervals and shows the count in each bin, which is exactly what reveals symmetry, skew, or multiple peaks in `transaction_amount`.
- B. Treating every unique dollar amount as its own bar produces an unreadable chart with as many bars as distinct values, since it ignores the continuous nature of the variable instead of grouping nearby values into ranges.
- C. Connecting values in row order shows a sequence over an arbitrary index, not a distribution; it reveals nothing about how the values cluster or spread across their range.
- D. A pie chart divides a whole into proportional wedges and is meant for categorical composition, not for showing the shape of a continuous variable's distribution across its range.
2.Two continuous features: scatter and bubble charts
When both features are continuous, plot one against the other. A scatter plot places each row at its (x, y) position. This makes correlations visible and, as the EDA tutorial points out, can highlight outliers. The tutorial plots renewable energy share against greenhouse gas emissions for the top emitters. It colours the points by country, which adds a categorical feature as a third dimension.
# Plot renewable share vs. greenhouse gas emissions over time
fig = px.scatter(top_emitters_data, x='renewable_share', y='greenhouse_gas_emissions',
color='country', title="Impact of Renewable Energy on Emissions for Top Emitters")
fig.show()No. The tutorial explains it this way: total energy consumption rises faster than renewable consumption. A scatter plot shows that two features move together. Explaining why they do still takes knowledge of the subject.
In the built-in editor, the bubble chart takes the scatter idea further: the size of each marker shows another metric. The Databricks example plots l_quantity against l_extendedprice, groups by l_returnflag, and sizes the bubbles by l_tax. One chart shows three continuous features and one categorical feature.
Checkpoint 3 of 5· Check yourself
You need one built-in chart that puts two continuous features on the axes and uses a third continuous feature as marker size. Which type do you choose?
A bubble chart is a scatter chart where marker size shows another metric. A heatmap uses colour across category axes instead.
“Bubble charts are scatter charts where the size of each point marker reflects a relevant metric.”Source: docs.databricks.com
Sources3
3.Building and inspecting charts in the notebook
You can create every chart type above without writing plotting code. Run display() on a DataFrame, click + above the result and choose Visualization. Then pick a type and columns and click Save. The chart appears as a tab in the cell output. The same + menu also offers Data Profile. It produces summary statistics for numeric, string and date columns plus a histogram for every column, which is a quick first look at all feature distributions together.
from pyspark.sql.functions import hour, col
pickupzip = '10001' # Example value for pickupzip
df = spark.table("samples.nyctaxi.trips")
result_df = df.filter(col("pickup_zip") == pickupzip) \
.groupBy(hour(col("tpep_dropoff_datetime")).alias("dropoff_hour")) \
.count() \
.withColumnRenamed("count", "num")
display(result_df)For bar, line, area, pie and heatmap charts, set the aggregation in the editor instead of in your query. Numeric Y columns can use Sum (the default), Average, Count, Count Distinct, Max, Min or Median. String columns can only use Count or Count Distinct. Aggregation set in the editor covers the whole dataset, not only the first 64,000 rows shown in the table. Filters are shared: a filter on the chart also applies to the results table, and a filter on the table also applies to the chart. On crowded charts, click and drag to zoom in on individual points. This helps you look at details and crop out outliers.
Checkpoint 4 of 5· Put it in order
Put these steps for adding a visualization to a notebook result in order
- 1.Click Save to add the chart as a tab in the cell output
- 2.Click + above the result and select Visualization
- 3.Choose a type in the Visualization Type drop-down and select the data
- 4.Run display() on the DataFrame to produce a results table
The editor opens from the + menu on a result that already exists. You choose the type and columns, then save.
“To create a visualization, click + above a result and select Visualization to open the visualization editor.”Source: docs.databricks.com
Checkpoint 5 of 5· Exam question
An analyst suspects the continuous `delivery_time_hours` column has a handful of unusually large values that could be outliers, and wants a single plot that shows the median, the interquartile range, and any points that fall far outside that range. Which visualization satisfies this?
Correct answer: A — A box plot of `delivery_time_hours`, drawing the median line, the interquartile box, and points beyond the whiskers as outliers.
- A. A box plot is built specifically from the median and the interquartile range, and it plots any value beyond 1.5 times the IQR from the box as an individual outlier point, matching every requirement in the scenario.
- B. A histogram summarizes the count of values per bin, which shows overall shape but does not compute or draw a median line or explicitly flag individual outlier points.
- C. A bar chart of a continuous column like delivery time treats each raw value as a separate category and provides no summary statistic such as a median or interquartile range.
- D. A pie chart shows proportional shares of a total and has no mechanism for representing quartiles, a median, or outlier points along a numeric range.
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Every built-in Databricks chart aggregates the full dataset, so a box chart on a large table summarises all rows.Why is that wrong?
Box charts stop at 64,000 rows and cut off larger datasets. Histograms and bar charts support backend aggregation over the full data.
Covered in Continuous by categorical: the box chart
2.Aggregation chosen in the visualization editor only covers the rows shown in the results table.Why is that wrong?
For bar, line, area, pie and heatmap charts, aggregation set in the editor covers the whole dataset.
Practise it for real
Turn a notebook DataFrame into a bar chart, a histogram and a data profile with the built-in visualization editor
1.In a Python notebook cell, run the samples.nyctaxi.trips example that groups trips by dropoff_hour and ends with display(result_df).
Why: The visualization editor works on a result that has been displayed.
You should see: A results table with dropoff_hour and num columns.
2.Click + above the result and choose Visualization. Set the type to Bar, put dropoff_hour on X and num on Y, then click Save.
Why: Each dropoff_hour value works as a separate level, so one bar per hour compares trip counts across hours.
You should see: A new tab in the cell output with one bar per hour.
3.Add another visualization of type Histogram with num as the X Column, and try several Number of Bins values.
Why: A histogram shows how the count values are distributed, and the bin count changes how that shape looks.
You should see: A frequency chart whose shape changes as you change the bin count.
4.Click + and choose Data Profile.
Why: The profile gives summary statistics and a histogram for every column at once.
You should see: A profile tab with statistics and a distribution histogram for each column.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“you can quickly compare the value ranges across categories and visualize the locality, spread and skewness groups of the values through their quartiles.”
↩︎ Continuous by categorical: the box chart“Box charts only support aggregation for up to 64,000 rows. If a dataset is larger than 64,000 rows, data will be truncated.”
↩︎ Exam trap 1“Box charts only support aggregation for up to 64,000 rows. If a dataset is larger than 64,000 rows, data will be truncated.”
↩︎ Prediction“Bubble charts are scatter charts where the size of each point marker reflects a relevant metric.”
↩︎ Checkpoint - 2.
“Missing and NULL values: Whether to hide missing or NULL values or to convert them to 0 and show them in the visualization.”
↩︎ Continuous by categorical: the box chart“X column: X axis values. For a vertical chart, choose a categorical column. For a horizontal chart, choose a number column.”
↩︎ Checkpoint - 3.
“For instance, scatter plots can highlight outliers, while time series plots can reveal trends and seasonality.”
↩︎ Two continuous features: scatter and bubble charts“meaning that its total energy consumption rises faster than its renewable consumption.”
↩︎ Two continuous features: scatter and bubble charts“The data profile includes summary statistics for numeric, string, and date columns as well as histograms of the value distributions for each column.”
↩︎ Building and inspecting charts in the notebook - 4.https://docs.databricks.com/aws/en/visualizationsOfficial docs
“zooming in on individual data points can be helpful to investigate details and to crop outliers.”
↩︎ Building and inspecting charts in the notebook“The aggregation applies to the entire dataset, not just the first 64,000 rows displayed in a table.”
↩︎ Exam trap 2“To create a visualization, click + above a result and select Visualization to open the visualization editor.”
↩︎ Checkpoint