What you will be able to do
- Choose a histogram, box chart or heatmap for distribution questions
- Use pivot, cohort and sankey visualizations for cross-tabulated, retention and flow insights
- Pick a choropleth, point map or path map based on the geographic data the query returns
1.How values are spread: histogram, box and heatmap
Some questions are not about totals or trends. They are about shape: are order values bunched together or spread out, and does one category vary more than another? The dashboard concepts page groups three types for this: box plots, histograms and heatmaps, listed together "for distribution analysis."
A histogram takes one numeric column and shows whether its values cluster in a few ranges or spread widely. It is drawn as bars, but you choose the number of bins. The documentation's example uses o_totalprice with 20 bins. Bin count matters because it changes how coarse or fine the distribution appears.
A box chart summarizes a distribution with quartiles, optionally grouped by category. That lets you "quickly compare the value ranges across categories" and see locality, spread and skew. The example places return flag on the x-axis and extended price on the y-axis, so each return flag gets its own box.
A heatmap uses color to show numbers in a grid of two categorical axes. It "blend[s] features of bar charts, stacked charts, and bubble charts." The example puts order priority on the x-axis and ship mode on the y-axis, and color intensity shows the summed order count. Heatmaps can display up to 64K rows or 10MB. The dataset behind that example is pre-aggregated in SQL:
SELECT
o.o_orderpriority AS priority,
l.l_shipmode AS ship_mode,
COUNT(*) AS order_count,
o.o_orderdate
FROM
samples.tpch.orders AS o
JOIN
samples.tpch.lineitem AS l
ON
o.o_orderkey = l.l_orderkey
GROUP BY
o.o_orderpriority,
l.l_shipmode,
o.o_orderdate
ORDER BY
priority,
ship_mode;Checkpoint 1 of 5· Check yourself
A quality team wants to compare the spread and skew of extended price across return flags in one view. Which type is designed for this?
Box charts summarize a numeric distribution with quartiles, grouped by category, which makes ranges, spread and skew easy to compare.
“The box chart visualization shows the distribution summary of numerical data, optionally grouped by category.”Source: docs.databricks.com
2.Cross-tabs, retention and flows: pivot, cohort and sankey
When readers need exact numbers broken down by two dimensions, a chart can be harder to read than a grid. A pivot visualization "aggregates records from a query result into a tabular display," much like PIVOT or GROUP BY in SQL. You place fields in Rows, Columns and Values, and you can turn on totals for each dimension. The example sums l_quantity, with return flag as rows and ship mode as columns. Pivot tables support rendering up to 1000 columns by 1000 rows.
Retention is a special case of a cross-tab. A cohort chart groups users by a shared starting point, such as sign-up date, and tracks how many remain active in later periods. In AI/BI dashboards there is no separate cohort type: you use a pivot. Cohort goes in Rows, active period goes in Columns, and the retention cell uses a Color Scale style, so darker cells show higher retention. The notebook and SQL editor visualizations do have their own cohort type, which only aggregates over dates. Any other aggregation must happen in the query.
For movement between states, such as where taxi trips start, where they end and which fare band they fall in, a sankey diagram "visualizes the flow from one set of values to another." You configure it with ordered stages plus a value, such as SUM(value).
A pivot visualization fed with retention data. Rows hold the cohort (yearly), Columns hold the active period, and the Retention cell uses a Color Scale, so darker colors mean higher retention.
Checkpoint 2 of 5· Check yourself
In an AI/BI dashboard, which visualization do you configure to build a cohort retention chart?
The AI/BI documentation builds cohort charts as a pivot visualization over retention data, with a Color Scale style on the cell.
“To create a cohort chart, use a pivot visualization with retention data.”Source: docs.databricks.com
3.Geography: choropleth, point map or path map
Map choice depends on the geographic data the query returns. A choropleth colors regions, such as countries, states or counties, by an aggregated value. That needs locations by name, or a GEOMETRY or GEOGRAPHY column. The documentation's example sums total_sales by U.S. state name. A point map places a marker at each coordinate, so the result must include latitude and longitude pairs (or a geometry column). A path map draws lines, either from a line geometry or by connecting points in order. It fits routes, transit lines, rivers and GPS tracks.
In short, use a choropleth for totals by region, a point map for individual locations, and a path map for routes and trajectories.
Checkpoint 3 of 5· Exam question
An operations dashboard needs a widget that prominently displays the current count of active daily users as a single large number, with a small comparison against the prior day's count. Which widget type should the analyst use?
Correct answer: A — Configure a counter visualization, since it displays a single value prominently and supports comparing that value against an offset such as the prior period.
- A. A counter visualization is purpose-built to display one value prominently and includes an offset comparison option, which is exactly the current-count-versus-prior-day requirement described. It avoids the clutter of a full chart when only one headline number matters.
- B. A pivot table aggregates results into a multi-dimensional tabular display meant for scanning many rows and columns, which is more detail than a single prominent headline figure needs. It would bury the one number the widget is supposed to highlight.
- C. A bar chart represents change across categories or time and would require plotting a full history of prior days as separate bars, rather than surfacing one large current value with a simple offset comparison.
- D. A heatmap encodes numerical values with color intensity across a grid, which suits comparing many combinations at once, not prominently surfacing a single current metric with one comparison point.
Checkpoint 4 of 5· Check yourself
A dataset contains one row per delivery with latitude and longitude, and no region names. The team wants to see where deliveries happen. Which map fits?
A point map plots individual locations from coordinates. A choropleth needs region names or geometry to color administrative areas.
“Markers are positioned using latitude and longitude coordinates, which must be included as part of the result set for this chart type.”Source: docs.databricks.com
4.Putting it together: insight to visualization
Each type covered here answers a question that trend, proportion and KPI charts cannot. The table maps each kind of question to the type built for it.
| Question the viewer asks | Visualization | What it needs |
|---|---|---|
| How are values of one measure spread out? | Histogram | One numeric column and a number of bins |
| How do value ranges compare across categories? | Box | Category on X, numeric measure on Y |
| Where do two categories combine to give high values? | Heatmap | Two categorical axes and a color column |
| What are the exact totals by two dimensions? | Pivot | Rows, Columns and Values fields |
| How does volume flow between stages? | Sankey | Stages and a value |
| How do totals vary by state or country? | Choropleth | Locations by name or a GEOMETRY/GEOGRAPHY column |
| Where are individual sites located? | Point map | Latitude and longitude pairs |
Checkpoint 5 of 5· Match them up
Match each visualization to the insight it communicates
Tap a term, then the definition that fits it.
Each pairing follows the purpose statement for that type on the AI/BI visualization types page.
“A sankey diagram visualizes the flow from one set of values to another.”Source: docs.databricks.com
Sources2
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.A histogram is just a bar chart of a measure by category.Why is that wrong?
A histogram plots how often values occur, grouped into bins you control. It shows a distribution, not a total per category.
Covered in How values are spread: histogram, box and heatmap
2.A choropleth can be drawn from raw latitude/longitude points.Why is that wrong?
A choropleth colors regions and needs locations by name or a GEOMETRY/GEOGRAPHY column. Latitude and longitude pairs go to a point map.
Covered in Geography: choropleth, point map or path map
3.AI/BI dashboards have a dedicated cohort visualization type, as in notebooks.Why is that wrong?
In AI/BI dashboards, a cohort chart is built as a pivot visualization over retention data, styled with a color scale.
Covered in Cross-tabs, retention and flows: pivot, cohort and sankey
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Box plots, histograms, heatmaps for distribution analysis”
↩︎ How values are spread: histogram, box and heatmap - 2.
“Using a box chart visualization, you can quickly compare the value ranges across categories”
↩︎ How values are spread: histogram, box and heatmap“Heatmap charts blend features of bar charts, stacked charts, and bubble charts, allowing you to visualize numerical data using colors.”
↩︎ How values are spread: histogram, box and heatmap“A pivot visualization aggregates records from a query result into a tabular display.”
↩︎ Cross-tabs, retention and flows: pivot, cohort and sankey“A sankey diagram visualizes the flow from one set of values to another.”
↩︎ Cross-tabs, retention and flows: pivot, cohort and sankey“Use them to visualize routes, transit lines, rivers, or point-by-point trajectories such as GPS tracks.”
↩︎ Geography: choropleth, point map or path map“A histogram plots the frequency that a given value occurs in a dataset.”
↩︎ Putting it together: insight to visualization“A histogram is displayed as a bar chart in which you control the number of distinct bars (also called bins).”
↩︎ Exam trap 1“To create a cohort chart, use a pivot visualization with retention data.”
↩︎ Exam trap 3“A histogram plots the frequency that a given value occurs in a dataset.”
↩︎ Prediction“The box chart visualization shows the distribution summary of numerical data, optionally grouped by category.”
↩︎ Checkpoint“To create a cohort chart, use a pivot visualization with retention data.”
↩︎ Checkpoint“Markers are positioned using latitude and longitude coordinates, which must be included as part of the result set for this chart type.”
↩︎ Checkpoint“A sankey diagram visualizes the flow from one set of values to another.”
↩︎ Checkpoint - 3.
“Pivot tables support rendering up to 1000 columns x 1000 rows.”
↩︎ Cross-tabs, retention and flows: pivot, cohort and sankey - 4.
“The cohort visualization only aggregates over dates (it allows for monthly aggregations).”
↩︎ Cross-tabs, retention and flows: pivot, cohort and sankey - 5.
“The query must return geographic locations by name (see Region lookup tables for supported names) or as a GEOMETRY or GEOGRAPHY column.”
↩︎ Geography: choropleth, point map or path map“The query result must return latitude and longitude pairs or a GEOMETRY or GEOGRAPHY column.”
↩︎ Exam trap 2