Skip to content
,
Geospatial for beginners · Part 1

When geospatial analysis helps

9 min read
Editorial featured image for When geospatial analysis helps. Title text reads When geospatial analysis helps.

Geospatial analysis, meaning analysis that uses location, is worth doing only when place changes the decision. If a bar chart by region name would lead to the same action, you do not need a map. This post is a filter you can use before you open any mapping tool: when location helps, when it wastes time, and what to write down first.

Say someone asks for “a map of our customers.” You could spend a week converting street addresses into coordinates (that step is called geocoding), picking a basemap, and shipping a glowing dot-density map that looks like a keynote slide. Then the real decision meeting turns out to be about whether to open a second warehouse in a metro area you already serve, and the map never answered travel time, capacity, or overlap. Pretty is not the same as useful.

This post opens Geospatial for beginners, a short series for analysts who can join tables and build dashboards but have not yet decided when latitude and longitude are worth the mess. The next post covers map design honesty. This one is the filter.

The decision test for geo

Ask this before you geocode anything: If two rows had the same attributes but different locations, would the recommended action change? If yes, space is part of the model. If no, a bar chart by region name might be enough, and a full spatial stack is optional flair.

The answer is usually yes for questions like these.

  • Which store should fulfill this order, given drive time and inventory?
  • Where do service calls cluster relative to technician home bases?
  • Does our retail coverage leave a dense residential pocket unserved?
  • Is the pilot marketing geo fence, a virtual boundary around a place, actually covering the campus we care about?

The answer is often no for these.

  • We want the dashboard to look modern for a board offsite.
  • Leadership likes maps more than tables, but taste is not a spatial join.
  • Region is already a clean category field, and the decision is budget by region name.
  • You only have country-level data, but someone asked for neighborhood heat.

Rule of thumb: If the insight survives when you shuffle points randomly inside the same city, you did not need geometry. You needed a group-by.

Good fit versus weak fit

Use this side-by-side as a gut check when a stakeholder says “can we map it?”

Filled geo fit: catchment and coverage versus tiny-n hexes
Filled geo fit: catchment and coverage versus tiny-n hexes

Good fit: catchment

Catchment analysis asks who or what falls inside a reachable area around a site, such as a drive-time polygon, a bike radius, or a sales territory. Retail, clinics, field service, and schools use catchments constantly. The spatial question is real here, because distance and access change demand and cost.

Good fit: delivery radius

Delivery and service promises are spatial contracts. “We deliver within 5 miles” is not the same as “within 20 minutes at rush hour.” Distance in miles can be a start, but time-based reach areas (called isochrones) are often closer to what the customer experiences. Either way, location is not optional.

Good fit: store or asset coverage

Coverage asks whether your assets sit where demand is. Assets can be stores, lockers, cell sites, electric-vehicle chargers, or warehouses. Overlaps and gaps are geometric facts, so a ranked list of cities will not show a hole between two stores on the same highway corridor the way a map will.

Weak fit: tiny counts by hex

Hex bins and grid cells are popular and often good. They become theater when each cell has two events and the color scale screams “hotspot.” Grouping by area does not create sample size. If the counts are tiny, widen the bin, lengthen the time window, or switch to a chart that is not a map and use careful language about uncertainty.

Weak fit: privacy risk

Precise home addresses, patient locations, or employee residences can identify real people even after you “remove names.” Rolling up to larger regions, hiding cells with very few people, and limiting who sees point maps are part of the analysis design, and they are not a polish step. If you cannot meet a privacy bar, geo may be the wrong way to present the work, even when space matters behind the scenes.

Weak fit: just pretty maps

Fancy basemaps, animated flights of dots, and 3D extrusions can hide a missing question. If the only action is “wow,” you built marketing material. That can be intentional, but you should not file it under decision analytics.

Beginner concepts that prevent rework

Unit of analysis

Are you studying stores, deliveries, customers, census tracts, or hex cells? Mixing units is how you double-count. A customer can place many orders, an order has one ship-to point, and a store has one location and many customers. Pick the level of detail the way you would in any warehouse model. The metrics series habit of naming what one row means applies here, with coordinates attached.

Join keys

Spatial analysis still needs keys to connect tables. Sometimes the key is a region id, such as a ZIP code, a county code from the Federal Information Processing Standards (FIPS), or an H3 index (a code for one hexagon on a world grid). Sometimes it is a point-in-polygon join, which places a latitude and longitude inside a boundary. Sometimes it is a nearest-neighbor join to the closest store. Write the join type on the project card, because “we mapped it” is not a join type.

Coordinate sanity

Some errors are classic. Latitude and longitude get swapped, minus signs go missing for the western hemisphere, geocodes land in the ocean, and default zeros show up at (0,0) in the Gulf of Guinea. Always plot a quick scatter or a map sample before you trust any totals. The empty spot at (0,0) even has a nickname, Null Island, and it is funny exactly once.

Projection (light version)

The earth is not flat. For city-scale work, many tools handle distance well enough if you use their geography types correctly. For country-scale comparisons of area, a bad projection distorts size. Beginners can go far by using library defaults designed for distance on a globe, and by not computing distances in degrees as if they were meters. When area rankings matter a lot, partner with someone who knows coordinate reference system (CRS) codes, and do not invent a custom formula late at night.

The geo project card

Before you touch any tool, fill in a card. The next post cares about colors, while this one cares about whether you should start at all.

Geo project card with unit example store, join key lat/long or region id, and privacy set to aggregate only
Geo project card with unit example store, join key lat/long or region id, and privacy set to aggregate only

Expand that card with a few more lines in your team wiki.

FieldPromptExample
DecisionWhat choice will this change?Where to place a locker
UnitWhat is one row in the fact?Store or order ship-to
Join keyHow does place attach?lat/long to H3 or ZIP
PrivacyPoint, aggregate, or suppressed?Aggregate only
SuccessHow will we know the map helped?Gap list with demand score
Non-goalWhat are we not optimizing?Pretty basemap branding

Worked example: store coverage sketch

Picture a fictional city coffee chain with five stores. Leadership asks whether the north side is “covered.” You refuse a vibes map and define coverage instead: the share of last-90-day orders whose ship-to address (or a home proxy) falls within a 1.5 km radius of at least one store. You also set a privacy rule. There is no public point map of homes, and you report counts by 1 km grid cell and hide any cell with fewer than 5 orders.

Here is the toy store list.

StoreLatLon
S140.74-73.99
S240.73-74.00
S340.76-73.98
S440.71-74.01
S540.75-73.97

Below is pseudo-logic for “is this order covered.” It is not production geospatial SQL, and it only shows the spirit of the approach.

-- Conceptual: flag orders within 1.5 km of any store
SELECT
  o.order_id,
  o.grid_cell_id,
  MIN(distance_km(o.lat, o.lon, s.lat, s.lon)) AS km_to_nearest_store,
  MIN(distance_km(o.lat, o.lon, s.lat, s.lon)) <= 1.5 AS is_covered
FROM orders o
CROSS JOIN stores s
GROUP BY o.order_id, o.grid_cell_id;

Then you roll the results up by grid cell.

SELECT
  grid_cell_id,
  COUNT(*) AS orders,
  AVG(CASE WHEN is_covered THEN 1.0 ELSE 0.0 END) AS coverage_rate
FROM order_coverage
GROUP BY grid_cell_id
HAVING COUNT(*) >= 5;

The output for leadership is not a constellation of home dots. It is a table of under-covered cells with order volume, plus an internal map of cells and not of people. Your decision language becomes something like “Cells A12 and B7 have high demand and coverage under 40%, so walk those blocks before approving store six.” That is geo helping.

Data you typically need

  • Points, such as stores, lockers, towers, and delivery drop-offs.
  • Addresses or coordinates, geocoded with a quality flag.
  • Boundaries, such as ZIP codes, cities, districts, and custom territories.
  • Optionally, network context like drive time, when road distance matters.
  • Attribute facts, such as revenue, orders, incidents, and population proxies.

Geocoding quality deserves its own column, with values like match score, partial match, and rooftop versus centroid. The centroid of a huge ZIP code is not a home. If your join key is weak, your map will be confidently wrong, so treat that as a data quality issue with the same seriousness you would bring to the data quality series.

Common mistakes

  • Starting in the map screen before writing down the decision.
  • Publishing point maps of sensitive data about people.
  • Coloring regions by raw counts without a population or opportunity denominator, which the next post covers in depth.
  • Using country-level data to answer street-level questions.
  • Ignoring geocode failures that silently drop rural or apartment addresses.
  • Treating region names as geometry when spellings differ across systems.
  • Building a hotspot story on a week of data with tiny counts.

How to practice

  1. Pick one real decision at work that mentions place, and write your answer to the decision test in two sentences.
  2. Fill in a geo project card covering the unit, join key, privacy, success, and non-goal.
  3. Plot 50 sample points from your data, or open sample city data, and fix any coordinate problems before you total anything.
  4. Compute a simple coverage or distance metric in a table first, and only then draw a map.
  5. List three stakeholder requests that should be turned down as pretty-only, and practice the polite no.

The next post in the series covers maps that inform without misleading, including rate maps, binning, and missing regions. For broader analytics learning paths, use the Learn hub. Stewardship and access control for location data sit next to general privacy habits in the data stewardship series.

Quick recap

  • Geo helps when location changes the decision, and it does not help when maps merely impress.
  • Good fits include catchment, delivery radius, and coverage problems.
  • Weak fits include heat maps built on tiny counts, points that are unsafe for privacy, and decoration.
  • Name the unit, the join key, and the privacy rule before you pick tools.
  • Tables and project cards beat sightseeing on a basemap.
  • Coordinate and geocode quality are data quality problems in a longitude and latitude costume.

Series notes

This is Part 1 of the Geospatial for beginners series. Related: metrics, data quality, and data stewardship.

Sources

Written by

Jose S

Founder & Lead Analyst · Analytics Made Simple

Hands-on data strategist, analytics engineering lead, and educator. Writing practical, no-fluff guides to help everyday teams, analysts, and engineers master SQL, AI systems, and modern data architectures.

Keep going

Same lessons in your feed

Short diagrams, hooks, and weekly tutorials on Substack, Instagram, X, and Facebook.

Google Search Prefer our practical guides in Google Search & Top Stories: