What you will be able to do
- Decide whether a feature is continuous or categorical and pick a chart that shows it honestly
- Configure a Databricks histogram, including bin count and a logarithmic axis scale
- Plot a categorical feature as a bar chart with Plotly and compare two categorical features with a heatmap or pivot table
Key concept
Match the chart to the feature type — To explore a continuous feature, look at the shape of its distribution with a histogram or box chart. To explore a categorical feature, compare an aggregated metric across its levels with a bar chart, pie chart or heatmap. So the first step in choosing a chart is to ask which kind of feature you have.
1.Why visualize features before modelling
Exploratory data analysis (EDA) is where you learn what a dataset looks like before you choose how to prepare it or which model to use. Databricks describes EDA as analyzing and visualizing data for four purposes: to find its main characteristics, spot patterns and trends, detect anomalies, and understand how variables relate to each other. The last purpose matters most for an ML associate, because the docs note that what you learn here can affect which algorithms you train.
Summary numbers alone hide a lot. Two features can have the same mean but completely different shapes, and an average can hide an outlier. Charts reveal patterns that numbers alone can miss. A Databricks notebook gives you two ways to make them. The first is the built-in visualization editor, which you can open on any display() result. The second is a Python library such as Plotly or Matplotlib. The EDA tutorial uses both.
Checkpoint 1 of 5· Check yourself
According to the Databricks EDA guidance, which result of EDA directly affects model training?
EDA guides how you prepare data and can shape which algorithm you pick. By itself it does not encode, validate or clean data.
“EDA can also influence which algorithms you choose to apply for training ML models.”Source: docs.databricks.com
2.Continuous features: the histogram
A continuous feature, such as a price, a duration or an energy consumption figure, can take any value in a range, so counting each separate value tells you little. A histogram instead groups values into ranges and plots how often each range occurs. Databricks says a histogram shows whether values are clustered in a few ranges or spread out. It is drawn as a bar chart where you set the number of bars, called bins.
The bin count is the main setting to adjust. With too few bins, a distribution with two peaks flattens into one block. With too many, it turns into noise. The documentation's example plots o_totalprice from samples.tpch.orders with 20 bins.
The X axis also has a Scale option. Some features have values that span several orders of magnitude, so most of the data gets squashed into the first bar. Setting the scale to Logarithmic spreads those values out so you can see the shape. Histograms also support backend aggregation, so the chart covers more than the first 64K rows instead of a cut-off sample.
| Tab | Setting | What it controls |
|---|---|---|
| General (required) | X Column | The results column from the dataset to display |
| General (required) | Number of Bins | Number of bins in which to display the data |
| X axis | Scale | Automatic, Datetime, Linear, Logarithmic, or Categorical |
| Y axis | Start Value / End Value | Show only values higher or lower than a given value |
| Data labels | Number values format | Format used for labels of numeric values |
Checkpoint 2 of 5· Check yourself
A feature's values range from 1 to 1,000,000, and most rows are below 100. Which histogram setting lets you see the distribution's shape across the whole range?
The histogram's X axis has a Logarithmic scale, which spreads out values that span several orders of magnitude. Histograms have no Y aggregation setting, and two bins would hide the shape.
“Scale: Select Automatic, Datetime, Linear, Logarithmic, or Categorical.”Source: docs.databricks.com
Checkpoint 3 of 5· Exam question
A data scientist loads a table into a pandas DataFrame in a Databricks notebook and wants to see how many rows fall into each of the five values in the `region` column. Which visualization best shows this?
Correct answer: A — Create a bar chart with one bar per `region` value, so each bar's height shows the count of rows in that category.
- A. A bar chart with one bar per category is the standard way to display counts of a categorical feature like `region`, since each bar's height directly encodes that category's frequency.
- B. Histograms bin continuous numeric values into equal-width intervals; `region` is a categorical column with discrete labels, not a continuous quantity, so binning it numerically does not produce a meaningful chart.
- C. Plotting a category against row index does not summarize frequency at all — it scatters points by the arbitrary order rows appear in, which tells the viewer nothing about how many rows belong to each region.
- D. Box plots summarize the spread and outliers of a continuous numeric variable using quartiles; `region` has no numeric quartiles to compute, so a box plot cannot be built directly from a categorical label like this.
Sources3
3.Categorical features: bar and pie charts
A categorical feature, such as region, order priority or return flag, has a fixed set of levels. The useful question is not how the values are distributed, but how some metric compares across the levels. A bar chart answers that question. It draws one bar per category, and each bar's height is an aggregate such as a sum or a count. Databricks describes bar charts as showing change over time or proportionality, like a pie chart.
The EDA tutorial does this with Plotly. It filters the data to six regions, totals the emissions for each region with groupby(...).sum(), and passes the resulting Series to px.bar.
# Calculate total emissions for each region
regional_emissions = df[df['country'].isin(regions)].groupby('country')['greenhouse_gas_emissions'].sum()
# Plot the comparison
fig = px.bar(regional_emissions, title="Greenhouse Gas Emissions by Region")
fig.show()The tutorial reads a ranking from this chart: Asia is highest, and Oceania, South America and Africa are lowest. A pie chart shows the same kind of proportions between categories, but the docs add a limit: pie charts are not meant for time series data. If you build either chart in the Databricks editor, string columns can only be aggregated with Count or Count Distinct. That is exactly what you need when each bar should show how often its category occurs.
Checkpoint 4 of 5· Fill the gap
Which Plotly Express function completes this sample so that it compares total emissions across regions?
# Plot the comparison
fig = px. ? (regional_emissions, title="Greenhouse Gas Emissions by Region")
fig.show()regional_emissions holds one total per region, so a bar chart compares them. A histogram would sort the six totals into bins instead of labelling each region.
4.Two categorical features at once: heatmaps and pivot tables
You often want to see how two categorical features interact, for example whether order status depends on order priority. A heatmap puts one category on each axis and colours each cell by a number. In the Databricks example, o_orderpriority is the X column, o_orderstatus is the Y column, and the colour shows the average of o_totalprice. Warmer colours usually mark the highest values and cooler colours the lowest.
A pivot table visualization shows the same cross-tabulation as numbers. Its rows come from one column (l_returnflag) and its columns from another (l_shipmode). Each cell holds a summed value and can be coloured by that value. It works like a SQL PIVOT or GROUP BY.
Checkpoint 5 of 5· Match them up
Match each chart to the question about features that it answers best
Tap a term, then the definition that fits it.
A histogram shows how often values occur. Bar and pie charts compare categories. A heatmap shows a number as colour across two category axes.
“Heatmap charts blend features of bar charts, stacking, and bubble charts allowing you to visualize numerical data using colors.”Source: docs.databricks.com
Sources3
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A pie chart is a fine way to show how category shares change from year to year.Why is that wrong?
Pie charts show proportions between metrics at a single point. Databricks states they are not meant for time series data, so use a bar or line chart instead.
Covered in Categorical features: bar and pie charts
2.A histogram is just a bar chart of a categorical column.Why is that wrong?
A histogram plots how often values occur in a column, usually a continuous one, by grouping them into bins whose number you set. A bar chart compares an aggregate across categories.
Covered in Continuous features: the histogram
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Use libraries such as Plotly or Matplotlib for common visualization techniques including scatter plots, bar charts, line graphs, and histograms.”
↩︎ Why visualize features before modelling“These visual tools allow data scientists to identify anomalies, understand data distributions, and observe correlations between variables.”
↩︎ Key concept - 2.
“EDA can also influence which algorithms you choose to apply for training ML models.”
↩︎ Why visualize features before modelling - 3.
“whether a dataset has values that are clustered around a small number of ranges or are more spread out”
↩︎ Continuous features: the histogram“Histogram charts support backend aggregations, providing support for queries returning more than 64K rows of data without truncation of the result set.”
↩︎ Continuous features: the histogram“Bar charts represent the change in metrics over time or to show proportionality, similar to a pie chart.”
↩︎ Categorical features: bar and pie charts“A pivot table visualization aggregates records from a query result into a new tabular display.”
↩︎ Two categorical features at once: heatmaps and pivot tables“Pie charts show proportionality between metrics. They are not meant for conveying time series data.”
↩︎ Exam trap 1“A histogram is displayed as a bar chart in which you control the number of distinct bars (also called bins).”
↩︎ Exam trap 2“Heatmap charts blend features of bar charts, stacking, and bubble charts allowing you to visualize numerical data using colors.”
↩︎ Checkpoint - 4.https://docs.databricks.com/aws/en/visualizationsOfficial docs
“Or from the following for string types: Count Count Distinct”
↩︎ Categorical features: bar and pie charts
Also cited
“Number of Bins: Number of bins in which to display the data.”
↩︎ Prediction“Scale: Select Automatic, Datetime, Linear, Logarithmic, or Categorical.”
↩︎ Checkpoint