What you will be able to do
- Configure general, axis, series, and data-label options shared by the common chart types
- Know which chart types support backend aggregation and which truncate at 64,000 rows
- Format, pin, and conditionally color columns in a table visualization, and write numeric format strings
- Set up choropleth and marker map visualizations with the geographic data each one needs
1.Options shared by the common chart types
Bar, Line, Area, Pie, Scatter, Bubble, and Combo charts all use much the same configuration panel in the notebook and SQL editor visualization editor. General holds the required mapping: an X column, one or more Y columns (with an optional aggregation), and an optional Group by. Group by is the dimension that splits the data into series. Each group gets its own legend entry and its own color. The other panels change how the chart looks.
| Panel | Settings | What they do |
|---|---|---|
| General | Y column aggregation | SUM, COUNT, COUNT DISTINCT, AVERAGE, MEDIAN, MIN, MAX, STANDARD DEVIATION, VARIANCE |
| General | Disable aggregations | Applies no aggregation and keeps your query's sort order |
| General | Stacking / Normalize values to percentage | Stacked or grouped series; each series shown as a percentage of the total |
| General | Horizontal chart | Flips the X and Y axis |
| X axis / Y axis | Scale Type | Categorical, linear, or logarithmic |
| Options | Y axis assignment / Series type | Left or right axis; bar or line per series |
| Data labels | Number / Percent / Date values format | Formats data labels and tooltips |
For Scale Type, choose categorical when each value is a discrete category (a region, for example). Choose linear or logarithmic for continuous values such as temperatures. If you don't choose, Databricks picks a scale from the field's data type. Disable aggregations matters when your SQL already does the work: it stops the chart from aggregating again, and the chart keeps the sort order from your query.
Checkpoint 1 of 5· Check yourself
Your query returns monthly revenue already sorted by a custom business order, but the chart re-orders and re-sums the values. Which setting fixes both problems?
Disable aggregations stops the chart from aggregating and keeps the query's sort order in the visualization.
“This will ensure no aggregation is applied, and ensure any sort order in your query is maintained in the visualization.”Source: docs.databricks.com
Sources1
2.Chart types and the 64K-row boundary
The visualization-types page describes each chart's purpose and whether it can handle large results. Area, bar, bubble, and combo charts support backend aggregation, so a query that returns more than 64K rows isn't truncated. The box chart, which shows distributions through quartiles, has no such support and truncates past 64,000 rows. The cohort visualization works differently again: it aggregates only over dates (monthly), so every other aggregation has to be done in the query.
| Type | Used for | More than 64K rows? |
|---|---|---|
| Area | How groups' values change over time, e.g. sales funnel changes | Yes, backend aggregation |
| Bar | Change in metrics over time, or proportionality | Yes, backend aggregation |
| Bubble | Scatter where marker size reflects a metric | Yes, backend aggregation |
| Combo | Line and bar together: change over time with proportionality | Yes, backend aggregation |
| Box | Distribution summary through quartiles, optionally by category | No, truncated beyond 64,000 |
| Cohort | Outcomes of cohorts across stages | Aggregates only over dates; other aggregation in the query |
select * from samples.tpch.lineitem where l_quantity < 45Checkpoint 2 of 5· Exam question
An analyst has a query result showing revenue share across five product categories for the most recent quarter only, with no time trend involved. Which visualization choice correctly matches the documented guidance for this case?
Correct answer: A — A pie visualization, since it shows proportionality between metrics and is not recommended for time series data.
- A. A pie chart shows proportionality between metrics and is explicitly documented as unsuited to time series, which fits a single-quarter proportional breakdown with no trend dimension.
- B. A line visualization exists to plot a metric's change over a continuous progression such as time, which does not apply here since the scenario has no trend to plot.
- C. An area visualization tracks how grouped values change over a second variable, typically time, which again does not fit a single-quarter snapshot with no progression.
- D. A waterfall chart shows the cumulative effect of sequential positive and negative changes across stages, which suits contribution analysis rather than a static proportional breakdown.
Sources2
3.Table visualizations and number formats
A table visualization can be changed without affecting the cell's original results table. You can drag columns to reorder them and toggle columns to hide them. Each column's kebab menu has Copy column name, Filter, Format, and Pin column (which keeps the column visible while you scroll right). Under Format, the Display as field also covers three special types. Image renders image links inline. JSON shows collapsible elements. Link builds clickable links from URL, text, and title templates that take mustache-style {{column}} parameters.
Font Conditions color a column's text when its value passes a threshold. The threshold must have the same data type as the column, so a numeric threshold is written without thousands separators.
Checkpoint 3 of 5· Check yourself
You want amounts above five hundred thousand shown in red. Which font condition works?
The threshold must match the column's numeric type. The comma in 500,000 means it is not a numeric value.
“to colorize results whose values exceed the numeric value 500000, create the threshold > 500000, rather than > 500,000.”Source: docs.databricks.com
Numbers are formatted with format strings. A format applies to numbers in a table visualization and to the values shown when you hover over chart data points. It does not apply to axis values.
| Number | Format | Output |
|---|---|---|
| 10000.23 | '0,0' | 10,000 |
| 1230974 | '0.0a' | 1.2m |
| 1000.234 | '$0,000.00' | $1,000.23 |
| -1000.234 | '($0,0)' | ($1,000) |
| 97.4878234 | '0.000%' | 97.488% |
| 2048 | '0 ib' | 2 KiB |
| 100 | '0o' | 100th |
4.Map visualizations: choropleth and marker
There are two map visualizations, and each needs a different shape of data. A choropleth colors whole regions, such as countries or states, by an aggregated value. The query must return the locations by name. You pick a Map (Countries, USA, or Japan/Prefectures), a Geographic column, a Geographic type that matches how the values are written (for example full name, or 2- or 3-letter ISO code), and a Value column. If a value doesn't match the chosen format, that region shows no data. A marker map places points at coordinates, so the query must return latitude and longitude pairs. Its options include Cluster markers, which merges nearby markers into one marker showing a count.
Choropleth colors come from Steps spread between Min Color and Max Color. The Clustering mode decides how values are split into those steps.
Checkpoint 4 of 5· Match them up
Match each choropleth clustering mode to how it segments the results.
Tap a term, then the definition that fits it.
Equidistant divides by value range, Quantile by result count (and can end up with fewer steps), and K-mean by closeness to a segment mean.
“Quantile: Results are segmented into a number of segments less than or equal to the value you specify in the Steps field”Source: docs.databricks.com
Checkpoint 5 of 5· Check yourself
Your query returns store_id, latitude, and longitude for each store. Which map visualization fits this data?
A marker map needs latitude and longitude pairs. A choropleth needs geographic locations by name.
“The query result must return latitude and longitude pairs.”Source: docs.databricks.com
Sources5
Exam traps
Each one states something that sounds right. Open it to see what is actually true.
1.Every chart type in the notebook and SQL editor handles results of any size without truncation.Why is that wrong?
Bar, area, bubble, and combo charts support backend aggregation, but box charts truncate data beyond 64,000 rows.
Covered in Chart types and the 64K-row boundary
2.A numeric format string also reformats the chart's axis labels.Why is that wrong?
Format strings apply to table values and to values in hover tooltips, not to axis values.
Covered in Table visualizations and number formats
3.A choropleth will plot any region column, whatever format its values are in.Why is that wrong?
The values must match the selected Geographic type. Regions whose values don't match show no data.
Covered in Map visualizations: choropleth and marker
Practise it for real
Build the documentation's stacked bar chart of total order price by month and priority in the SQL editor, then change its aggregation without editing the query.
1.In the SQL editor, run: select * from samples.tpch.orders
Why: This raw result is the dataset the documentation uses for its bar chart example; the chart does the aggregating.
You should see: A results table of orders appears in the output pane.
2.Click + above the result, select Visualization, and choose Bar in the Visualization Type drop-down.
Why: The visualization editor is opened from the result set.
You should see: The visualization editor opens with bar chart options.
3.Set X column to o_orderdate with Date level Months, Y column to o_totalprice with aggregation Sum, and Group by to o_orderpriority; set Stacking to Stack.
Why: These are the documented configuration values for the example bar chart.
You should see: Monthly bars, stacked by order priority, each in its own color.
4.Rename the X axis to Order month and the Y axis to Total price, then click Save.
Why: Axis rename overrides the default axis names.
You should see: A new visualization tab appears next to the results table.
5.Edit the visualization and change the Y aggregation from Sum to Average, then save.
Why: Changing the aggregation in the chart lets you try different scenarios without changing the query.
You should see: The bars now show average order price per month and priority; the query is unchanged.
Stuck? Get a nudge
If the bars look too few, check whether a filter on the results table is also narrowing the chart.
Sources
Every claim above is drawn from one of these pages, quoted as it was written on the date shown.
- 1.
“Fields can optionally be aggregated by SUM, COUNT, COUNT DISTINCT, AVERAGE, MEDIAN, MIN, MAX, STANDARD DEVIATION, VARIANCE.”
↩︎ Options shared by the common chart types“Group by: Also known as color in some tools, select a dimension by which to group all other values.”
↩︎ Options shared by the common chart types“Scale Type: Options are categorical, linear, or logarithmic.”
↩︎ Options shared by the common chart types“This will ensure no aggregation is applied, and ensure any sort order in your query is maintained in the visualization.”
↩︎ Checkpoint - 2.
“Bar charts support backend aggregations, providing support for queries returning more than 64K rows of data without truncation of the result set.”
↩︎ Chart types and the 64K-row boundary“The cohort visualization only aggregates over dates (it allows for monthly aggregations).”
↩︎ Chart types and the 64K-row boundary“Bubble charts are scatter charts where the size of each point marker reflects a relevant metric.”
↩︎ Chart types and the 64K-row boundary“Box charts only support aggregation for up to 64,000 rows. If a dataset is larger than 64,000 rows, data will be truncated.”
↩︎ Exam trap 1“Box charts only support aggregation for up to 64,000 rows. If a dataset is larger than 64,000 rows, data will be truncated.”
↩︎ Prediction - 3.
“A table visualization can be manipulated independently of the original cell results table.”
↩︎ Table visualizations and number formats“Databricks supports the following special data types: image, JSON, and link.”
↩︎ Table visualizations and number formats“to colorize results whose values exceed the numeric value 500000, create the threshold > 500000, rather than > 500,000.”
↩︎ Checkpoint - 4.
“You control the format by supplying a format string.”
↩︎ Table visualizations and number formats“but not when formatting axis values”
↩︎ Exam trap 2 - 5.
“The query must return geographic locations by name.”
↩︎ Map visualizations: choropleth and marker“markers that are close to each other are clustered into a single marker”
↩︎ Map visualizations: choropleth and marker“If a geographic column value doesn't match one of the target field formats, no data is shown for that locality.”
↩︎ Exam trap 3“Quantile: Results are segmented into a number of segments less than or equal to the value you specify in the Steps field”
↩︎ Checkpoint“The query result must return latitude and longitude pairs.”
↩︎ Checkpoint