CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 21/48

    Box Charts and Scatter Plots for Feature Relationships

    Create visualizations for categorical or continuous features

    9 min read
    2.08% of exam
    4 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Use a box chart to compare the spread and skew of a continuous feature across categories
    • Use a scatter or bubble chart to look at two continuous features together
    • Create, aggregate and inspect visualizations from a notebook result, and know which chart types are cut off at 64,000 rows

    1.Continuous by categorical: the box chart

    Numeric features usually need to be compared across groups. A box chart summarises the distribution of a numeric column separately for each level of a categorical column. Each box is built from quartiles, so at a glance you can compare value ranges across categories and see where each group's values sit, how spread out they are, and how skewed. In the Databricks rendering, the darker line in each box marks the interquartile range.

    The column roles are fixed. In a vertical box chart, the X column must be categorical and the Y column numeric. If you select Horizontal Chart, the roles swap. The documentation's example puts l_returnflag on X and l_extendedprice on Y, and also groups by l_shipmode.

    Box chart general settings and what each one expects
    SettingVertical chartHorizontal Chart selected
    X columnA categorical columnA number column
    Y columnsA number columnA categorical column
    Default groupingGrouped by the X axisGrouped by the Y axis
    Show all pointsShow each X axis value separately instead of a boxSame option
    Missing and NULL valuesHide them, or convert to 0 and show themSame option

    This limit changes how you read a box chart on a large table: the quartiles describe at most the first 64,000 rows, not the whole dataset. Two other settings change what the box shows. Under Missing and NULL values, you can either hide missing values or convert them to 0 and plot them. Converting them adds a cluster of zeros that pulls the lower quartile down. The Y axis Scale includes Logarithmic, which helps when one category's values are far larger than another's.

    Checkpoint 1 of 5· Check yourself

    You want a vertical box chart showing how l_extendedprice varies by l_returnflag. How should you assign the columns?

    Checkpoint 2 of 5· Exam question

    A machine learning engineer wants to understand the overall shape of the distribution of a continuous `transaction_amount` column — whether it is roughly symmetric, skewed, or bimodal — before choosing an imputation strategy. Which visualization is most appropriate?

    Sources12

    2.Two continuous features: scatter and bubble charts

    When both features are continuous, plot one against the other. A scatter plot places each row at its (x, y) position. This makes correlations visible and, as the EDA tutorial points out, can highlight outliers. The tutorial plots renewable energy share against greenhouse gas emissions for the top emitters. It colours the points by country, which adds a categorical feature as a third dimension.

    Scatter plot of two continuous features, coloured by a categorical featurepython
    # Plot renewable share vs. greenhouse gas emissions over time
    fig = px.scatter(top_emitters_data, x='renewable_share', y='greenhouse_gas_emissions',
                    color='country', title="Impact of Renewable Energy on Emissions for Top Emitters")
    fig.show()

    In the built-in editor, the bubble chart takes the scatter idea further: the size of each marker shows another metric. The Databricks example plots l_quantity against l_extendedprice, groups by l_returnflag, and sizes the bubbles by l_tax. One chart shows three continuous features and one categorical feature.

    Checkpoint 3 of 5· Check yourself

    You need one built-in chart that puts two continuous features on the axes and uses a third continuous feature as marker size. Which type do you choose?

    Sources3

    3.Building and inspecting charts in the notebook

    You can create every chart type above without writing plotting code. Run display() on a DataFrame, click + above the result and choose Visualization. Then pick a type and columns and click Save. The chart appears as a tab in the cell output. The same + menu also offers Data Profile. It produces summary statistics for numeric, string and date columns plus a histogram for every column, which is a quick first look at all feature distributions together.

    Producing a result set to visualize in a Python notebook cellpython
    from pyspark.sql.functions import hour, col
    
    pickupzip = '10001'  # Example value for pickupzip
    df = spark.table("samples.nyctaxi.trips")
    result_df = df.filter(col("pickup_zip") == pickupzip) \
                  .groupBy(hour(col("tpep_dropoff_datetime")).alias("dropoff_hour")) \
                  .count() \
                  .withColumnRenamed("count", "num")
    display(result_df)

    For bar, line, area, pie and heatmap charts, set the aggregation in the editor instead of in your query. Numeric Y columns can use Sum (the default), Average, Count, Count Distinct, Max, Min or Median. String columns can only use Count or Count Distinct. Aggregation set in the editor covers the whole dataset, not only the first 64,000 rows shown in the table. Filters are shared: a filter on the chart also applies to the results table, and a filter on the table also applies to the chart. On crowded charts, click and drag to zoom in on individual points. This helps you look at details and crop out outliers.

    Checkpoint 4 of 5· Put it in order

    Put these steps for adding a visualization to a notebook result in order

    1. 1.Click Save to add the chart as a tab in the cell output
    2. 2.Click + above the result and select Visualization
    3. 3.Choose a type in the Visualization Type drop-down and select the data
    4. 4.Run display() on the DataFrame to produce a results table

    Checkpoint 5 of 5· Exam question

    An analyst suspects the continuous `delivery_time_hours` column has a handful of unusually large values that could be outliers, and wants a single plot that shows the median, the interquartile range, and any points that fall far outside that range. Which visualization satisfies this?

    Sources34

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.Every built-in Databricks chart aggregates the full dataset, so a box chart on a large table summarises all rows.Why is that wrong?

      Box charts stop at 64,000 rows and cut off larger datasets. Histograms and bar charts support backend aggregation over the full data.

      Covered in Continuous by categorical: the box chart

    2. 2.Aggregation chosen in the visualization editor only covers the rows shown in the results table.Why is that wrong?

      For bar, line, area, pie and heatmap charts, aggregation set in the editor covers the whole dataset.

      Covered in Building and inspecting charts in the notebook

    Practise it for real

    Turn a notebook DataFrame into a bar chart, a histogram and a data profile with the built-in visualization editor

    1. 1.In a Python notebook cell, run the samples.nyctaxi.trips example that groups trips by dropoff_hour and ends with display(result_df).

      Why: The visualization editor works on a result that has been displayed.

      You should see: A results table with dropoff_hour and num columns.

    2. 2.Click + above the result and choose Visualization. Set the type to Bar, put dropoff_hour on X and num on Y, then click Save.

      Why: Each dropoff_hour value works as a separate level, so one bar per hour compares trip counts across hours.

      You should see: A new tab in the cell output with one bar per hour.

    3. 3.Add another visualization of type Histogram with num as the X Column, and try several Number of Bins values.

      Why: A histogram shows how the count values are distributed, and the bin count changes how that shape looks.

      You should see: A frequency chart whose shape changes as you change the bin count.

    4. 4.Click + and choose Data Profile.

      Why: The profile gives summary statistics and a histogram for every column at once.

      You should see: A profile tab with statistics and a distribution histogram for each column.

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “you can quickly compare the value ranges across categories and visualize the locality, spread and skewness groups of the values through their quartiles.”
      ↩︎ Continuous by categorical: the box chart
      “Box charts only support aggregation for up to 64,000 rows. If a dataset is larger than 64,000 rows, data will be truncated.”
      ↩︎ Exam trap 1
      “Box charts only support aggregation for up to 64,000 rows. If a dataset is larger than 64,000 rows, data will be truncated.”
      ↩︎ Prediction
      “Bubble charts are scatter charts where the size of each point marker reflects a relevant metric.”
      ↩︎ Checkpoint
    2. 2.
      “Missing and NULL values: Whether to hide missing or NULL values or to convert them to 0 and show them in the visualization.”
      ↩︎ Continuous by categorical: the box chart
      “X column: X axis values. For a vertical chart, choose a categorical column. For a horizontal chart, choose a number column.”
      ↩︎ Checkpoint
    3. 3.
      “For instance, scatter plots can highlight outliers, while time series plots can reveal trends and seasonality.”
      ↩︎ Two continuous features: scatter and bubble charts
      “meaning that its total energy consumption rises faster than its renewable consumption.”
      ↩︎ Two continuous features: scatter and bubble charts
      “The data profile includes summary statistics for numeric, string, and date columns as well as histograms of the value distributions for each column.”
      ↩︎ Building and inspecting charts in the notebook
    4. 4.
      “zooming in on individual data points can be helpful to investigate details and to crop outliers.”
      ↩︎ Building and inspecting charts in the notebook
      “The aggregation applies to the entire dataset, not just the first 64,000 rows displayed in a table.”
      ↩︎ Exam trap 2
      “To create a visualization, click + above a result and select Visualization to open the visualization editor.”
      ↩︎ Checkpoint

    Ready to test yourself?

    Practise the 7 questions on this subdomain.

    Spotted a mistake, or was something unclear? Tell us.