CertSafari
    Databricks Certified Machine Learning Associate· Lessons

    Domain 2 · Lesson 21/48

    Histograms, Bar Charts and Heatmaps for Feature Types

    Create visualizations for categorical or continuous features

    9 min read
    2.08% of exam
    5 sources
    Published 2 Oct 2026
    Docs as of 30 Sep 2026

    What you will be able to do

    • Decide whether a feature is continuous or categorical and pick a chart that shows it honestly
    • Configure a Databricks histogram, including bin count and a logarithmic axis scale
    • Plot a categorical feature as a bar chart with Plotly and compare two categorical features with a heatmap or pivot table

    Key concept

    Match the chart to the feature type — To explore a continuous feature, look at the shape of its distribution with a histogram or box chart. To explore a categorical feature, compare an aggregated metric across its levels with a bar chart, pie chart or heatmap. So the first step in choosing a chart is to ask which kind of feature you have.

    1.Why visualize features before modelling

    Exploratory data analysis (EDA) is where you learn what a dataset looks like before you choose how to prepare it or which model to use. Databricks describes EDA as analyzing and visualizing data for four purposes: to find its main characteristics, spot patterns and trends, detect anomalies, and understand how variables relate to each other. The last purpose matters most for an ML associate, because the docs note that what you learn here can affect which algorithms you train.

    Summary numbers alone hide a lot. Two features can have the same mean but completely different shapes, and an average can hide an outlier. Charts reveal patterns that numbers alone can miss. A Databricks notebook gives you two ways to make them. The first is the built-in visualization editor, which you can open on any display() result. The second is a Python library such as Plotly or Matplotlib. The EDA tutorial uses both.

    Checkpoint 1 of 5· Check yourself

    According to the Databricks EDA guidance, which result of EDA directly affects model training?

    Sources12

    2.Continuous features: the histogram

    A continuous feature, such as a price, a duration or an energy consumption figure, can take any value in a range, so counting each separate value tells you little. A histogram instead groups values into ranges and plots how often each range occurs. Databricks says a histogram shows whether values are clustered in a few ranges or spread out. It is drawn as a bar chart where you set the number of bars, called bins.

    The bin count is the main setting to adjust. With too few bins, a distribution with two peaks flattens into one block. With too many, it turns into noise. The documentation's example plots o_totalprice from samples.tpch.orders with 20 bins.

    The X axis also has a Scale option. Some features have values that span several orders of magnitude, so most of the data gets squashed into the first bar. Setting the scale to Logarithmic spreads those values out so you can see the shape. Histograms also support backend aggregation, so the chart covers more than the first 64K rows instead of a cut-off sample.

    Histogram settings in the notebook and SQL editor visualization editor
    TabSettingWhat it controls
    General (required)X ColumnThe results column from the dataset to display
    General (required)Number of BinsNumber of bins in which to display the data
    X axisScaleAutomatic, Datetime, Linear, Logarithmic, or Categorical
    Y axisStart Value / End ValueShow only values higher or lower than a given value
    Data labelsNumber values formatFormat used for labels of numeric values

    Checkpoint 2 of 5· Check yourself

    A feature's values range from 1 to 1,000,000, and most rows are below 100. Which histogram setting lets you see the distribution's shape across the whole range?

    Checkpoint 3 of 5· Exam question

    A data scientist loads a table into a pandas DataFrame in a Databricks notebook and wants to see how many rows fall into each of the five values in the `region` column. Which visualization best shows this?

    Sources3

    3.Categorical features: bar and pie charts

    A categorical feature, such as region, order priority or return flag, has a fixed set of levels. The useful question is not how the values are distributed, but how some metric compares across the levels. A bar chart answers that question. It draws one bar per category, and each bar's height is an aggregate such as a sum or a count. Databricks describes bar charts as showing change over time or proportionality, like a pie chart.

    The EDA tutorial does this with Plotly. It filters the data to six regions, totals the emissions for each region with groupby(...).sum(), and passes the resulting Series to px.bar.

    Totalling a continuous metric for each category, then plotting the categories as a bar chartpython
    # Calculate total emissions for each region
    regional_emissions = df[df['country'].isin(regions)].groupby('country')['greenhouse_gas_emissions'].sum()
    
    # Plot the comparison
    fig = px.bar(regional_emissions, title="Greenhouse Gas Emissions by Region")
    fig.show()

    The tutorial reads a ranking from this chart: Asia is highest, and Oceania, South America and Africa are lowest. A pie chart shows the same kind of proportions between categories, but the docs add a limit: pie charts are not meant for time series data. If you build either chart in the Databricks editor, string columns can only be aggregated with Count or Count Distinct. That is exactly what you need when each bar should show how often its category occurs.

    Checkpoint 4 of 5· Fill the gap

    Which Plotly Express function completes this sample so that it compares total emissions across regions?

    # Plot the comparison
    fig = px. ? (regional_emissions, title="Greenhouse Gas Emissions by Region")
    fig.show()

    Sources34

    4.Two categorical features at once: heatmaps and pivot tables

    You often want to see how two categorical features interact, for example whether order status depends on order priority. A heatmap puts one category on each axis and colours each cell by a number. In the Databricks example, o_orderpriority is the X column, o_orderstatus is the Y column, and the colour shows the average of o_totalprice. Warmer colours usually mark the highest values and cooler colours the lowest.

    A pivot table visualization shows the same cross-tabulation as numbers. Its rows come from one column (l_returnflag) and its columns from another (l_shipmode). Each cell holds a summed value and can be coloured by that value. It works like a SQL PIVOT or GROUP BY.

    Checkpoint 5 of 5· Match them up

    Match each chart to the question about features that it answers best

    Tap a term, then the definition that fits it.

    Sources3

    Exam traps

    Each one states something that sounds right. Open it to see what is actually true.

    1. 1.A pie chart is a fine way to show how category shares change from year to year.Why is that wrong?

      Pie charts show proportions between metrics at a single point. Databricks states they are not meant for time series data, so use a bar or line chart instead.

      Covered in Categorical features: bar and pie charts

    2. 2.A histogram is just a bar chart of a categorical column.Why is that wrong?

      A histogram plots how often values occur in a column, usually a continuous one, by grouping them into bins whose number you set. A bar chart compares an aggregate across categories.

      Covered in Continuous features: the histogram

    Sources

    Every claim above is drawn from one of these pages, quoted as it was written on the date shown.

    1. 1.
      “Use libraries such as Plotly or Matplotlib for common visualization techniques including scatter plots, bar charts, line graphs, and histograms.”
      ↩︎ Why visualize features before modelling
      “These visual tools allow data scientists to identify anomalies, understand data distributions, and observe correlations between variables.”
      ↩︎ Key concept
    2. 2.
      “EDA can also influence which algorithms you choose to apply for training ML models.”
      ↩︎ Why visualize features before modelling
    3. 3.
      “whether a dataset has values that are clustered around a small number of ranges or are more spread out”
      ↩︎ Continuous features: the histogram
      “Histogram charts support backend aggregations, providing support for queries returning more than 64K rows of data without truncation of the result set.”
      ↩︎ Continuous features: the histogram
      “Bar charts represent the change in metrics over time or to show proportionality, similar to a pie chart.”
      ↩︎ Categorical features: bar and pie charts
      “A pivot table visualization aggregates records from a query result into a new tabular display.”
      ↩︎ Two categorical features at once: heatmaps and pivot tables
      “Pie charts show proportionality between metrics. They are not meant for conveying time series data.”
      ↩︎ Exam trap 1
      “A histogram is displayed as a bar chart in which you control the number of distinct bars (also called bins).”
      ↩︎ Exam trap 2
      “Heatmap charts blend features of bar charts, stacking, and bubble charts allowing you to visualize numerical data using colors.”
      ↩︎ Checkpoint
    4. 4.
      “Or from the following for string types: Count Count Distinct”
      ↩︎ Categorical features: bar and pie charts

    Also cited

    Continue to page 2 of 2

    Box Charts and Scatter Plots for Feature Relationships

    Spotted a mistake, or was something unclear? Tell us.