library(sf)
library(tidyverse)
library(ggplot2)
library(here)7 Community Data Discovery
7.1 Overview
Finding data about community context, the social determinants of health, & structural drivers of health outcomes is essential for research in opioid environments, but can be difficult. Data are often distributed across federal agencies, universities, local governments, research repositories, and other organizations. Even after identifying a useful dataset, additional work is usually required to understand the geographic scale, time period, file structure, identifiers, and documentation before the data can be integrated into a spatial analysis.
There are multiple options for discovering new datasets needed for your search goals. In this tutorial, we’ll share one resource that supports researchers in the discovery process, and work through the stages of inspection, download, and integration into a reproducible coding environment.
Specifically, we’ll use resources from the SDOH & Place Project, a free site that brings together spatial data science, health equity research, and human-centered design to help researchers, advocates, students, planners, and community organizations find and use place-based SDOH data. The Data Discovery Search Tool includes pre-screened, national-level data to help users identify relevant datasets and directs them to the organization or repository where the underlying data can be downloaded.
In this tutorial, our objectives will be to:
- Demonstrate the SDOH & Place Data Discovery Search tool
- Download a relevant socioeconomic dataset from its source repository
- Inspect, clean, and merge the downloaded files in R with spatial boundaries
- Map neighborhood disadvantage at local and national scales
The worked example uses the National Neighborhood Data Archive (NaNDA) socioeconomic and demographic dataset for 2018–2022 using 2022 Census tract geography.
7.2 Environment Setup
To replicate the code and functions illustrated in this tutorial, you will need to have R and RStudio installed on your system. This tutorial assumes some basic familiarity with the R programming language, such as creating objects and running code from an R script.
7.2.1 Input/Output
Place your input data in yout desired folder location and neame them so it is easier to find.The main external inputs for this tutorial are:
- the downloaded NaNDA socioeconomic and demographic dataset from ICPSR, as discovered at the SDOH & Place Search tool;
- the NaNDA codebook and documentation included with the download; and
- a 2020 Census tract boundary shapefile or GeoJSON, downloaded via Data.gov.
In this example, the downloaded NaNDA folder is stored as:
~/GCCP-Data/NaNDA-SocEco1990-2022/ICPSR_38528
The census tract boundary is stored as:
~/GCCP-Data/Boundaries/tract-2022-500k-shp
The main outputs will be:
- a spatial census tract dataset containing the NaNDA neighborhood disadvantage index;
- a neighborhood disadvantage map for Cook County, Illinois;
- a neighborhood disadvantage map for the contiguous United States; and
- optionally, a GeoPackage or GeoJSON file for use in other GIS software.
7.2.2 Load Libraries
We will use the following packages in this tutorial:
sf: to read, manipulate, transform, and write spatial objectstidyverse: tidy data wrangling toolsggplot2: to create maps (as a break from tmap)here: to help construct reproducible project file paths
We’ll be leveraging ‘dyplyr’ and ‘stringr’ within the tidyverse package, specifically.
Load the required libraries.
7.3 Discover Community Data
We can begin by using the SDOH & Place Data Discovery platform to identify a dataset that measures the community characteristic of interest.
The Data Discovery platform helps answer several questions before we begin coding:
- What dataset measures the concept we are interested in?
- At what geographic scale is the dataset available?
- What years or time periods are available?
- Does the dataset cover the community or region of interest?
- Where can the underlying data be downloaded?
7.3.1 Search by Topic
For this example, we are interested in neighborhood socioeconomic conditions that may influence community epxeriences that impact opioid risk environments.
Open the Data Discovery platform and search for:
economic
We’ll use the default search option, keyword search, in the Discovery Tool. For more complex tasks, you can expand your search with the AI option, and/or use the map to zoom into an area of interest and filter your data options accordingly.
Related searches might include:
socioeconomic
income
poverty
unemployment
neighborhood disadvantage
The search results contain datasets related to economic stability and neighborhood socioeconomic conditions.
7.3.2 Refine by Geography and Time
The Data Discovery platform can be used to refine datasets by geographic unit and time period.
For a neighborhood-level analysis, we are interested in census tract data rather than state-level or county-level data.
The selected dataset should therefore meet the following criteria:
Topic: Socioeconomic conditions
Geography: Census tract
Coverage: United States
Time period: Recent multi-year estimates
7.3.3 Follow the Dataset Source
One relevant dataset is:
National Neighborhood Data Archive (NaNDA): Socioeconomic Status and Demographic Characteristics of Census Tracts and ZIP Code Tabulation Areas, United States, 1990–2022
The underlying data are distributed through the Inter-university Consortium for Political and Social Research (ICPSR).
Dataset page:
https://www.icpsr.umich.edu/web/ICPSR/studies/38528
The Data Discovery platform helps us identify the resource. The actual data are then downloaded from the original repository.
7.4 Select the Appropriate NaNDA Dataset
The downloaded ICPSR package contains several dataset folders, identified as DS0001 through DS0008.
The datasets differ by:
- geographic unit;
- time period; and
- Census geography vintage.
The available folders include:
| Folder | Geography and Period |
|---|---|
DS0001 |
2010 Census tracts, 1990–2010 |
DS0002 |
2010 Census tracts, 2008–2017 |
DS0003 |
2010 ZCTAs, 2008–2017 |
DS0004 |
2020 Census tracts, 2016–2020 |
DS0005 |
2020 ZCTAs, 2016–2020 |
DS0006 |
2010 Census tracts, 2018–2022 |
DS0007 |
2020 Census tracts, 2018–2022 |
DS0008 |
2020 ZCTAs, 2018–2022 |
For this tutorial, we want recent socioeconomic data at the census tract level using 2020 Census geography.
We therefore use:
DS0007
This corresponds to:
2018–2022 socioeconomic and demographic data
2022 Census tract geography
The geographic vintage is important because the tabular dataset and boundary file should represent the same tract geography.
7.6 Explore Variables
Before selecting a variable for analysis, we should inspect the available fields.
names(nanda)Some of the available variables include:
TRACT_FIPS20
TOTPOP
POPDEN
PUNEMP
PPOV
PPUBAS
AFFLUENCE
DISADVANTAGE
MEDFAMINC
Rather than manually searching all variable names, we can use str_detect() from the stringr package.
7.6.1 Search for Socioeconomic Variables
names(nanda)[
str_detect(
names(nanda),
regex(
"disadv|afflu|poverty|income|assist",
ignore_case = TRUE
)
)
]Output:
[1] "AFFLUENCE" "DISADVANTAGE"
This search finds variables whose names contain those exact text patterns.
The codebook should always be consulted because several socioeconomic variables use abbreviations, such as:
PPOV
PPUBAS
MEDFAMINC
7.6.2 Find the Geographic Identifier
Search for possible tract or geographic identifiers.
names(nanda)[
str_detect(
names(nanda),
regex(
"geoid|tract|fips",
ignore_case = TRUE
)
)
]Output:
[1] "TRACT_FIPS20"
TRACT_FIPS20 is the geographic identifier used to connect the NaNDA data to 2020 Census tract boundaries.
7.7 Prepare the Neighborhood Disadvantage Data
For this tutorial, we only need:
- the census tract identifier; and
- the neighborhood disadvantage measure.
Before creating the smaller dataset, it is useful to understand the tract identifier.
7.7.1 Census Tract GEOID
A census tract GEOID contains 11 characters.
The general structure is:
SSCCCTTTTTT
where:
SS = state FIPS
CCC = county FIPS
TTTTTT = census tract code
For example:
17031839100
contains:
17 Illinois
031 Cook County
839100 Census tract code
Geographic identifiers should generally be stored as character variables rather than numbers because some identifiers begin with zero.
Create a simplified NaNDA dataset.
nanda_disadvantage <- nanda %>%
transmute(
GEOID = as.character(TRACT_FIPS20),
DISADVANTAGE = as.numeric(DISADVANTAGE)
)Inspect the result.
head(nanda_disadvantage)Output:
GEOID DISADVANTAGE
1 01001020100 0.17298892
2 01001020200 0.14308497
3 01001020300 0.10727025
4 01001020400 0.09424171
5 01001020501 0.08787425
8 01001020502 0.04378821
At this stage, the dataset is still a regular table. It does not yet contain spatial geometry.
7.8 Get Geometry
Geometry or geographic boundaries allow us to map and spatially analyze the NaNDA socioeconomic data.
The downloaded NaNDA table contains a tract identifier, but it does not contain the polygon geometry needed to draw each tract.
We therefore need to connect it to a census tract boundary file.
In this example, we use a generalized 2022 Census tract shapefile based on 2020 Census tract geography.
7.8.1 Read the Census Tract Boundary
Read the tract boundary using st_read() from the sf package.
tracts <- st_read(
"~/GCCP-Data/Boundaries/tract-2022-500k-shp"
)The output should resemble:
Reading layer `tract-2022-500k' ...
Simple feature collection with 85185 features and 20 fields
Geometry type: MULTIPOLYGON
Dimension: XY
Geodetic CRS: NAD83
Inspect the spatial object.
tractsImportant fields include:
STATEFP
COUNTYFP
TRACTCE
GEOID
STUSPS
STATE_NAME
geometry
Inspect the column names.
names(tracts)7.8.2 Prepare the Boundary GEOID
Make sure the boundary GEOID is stored as character text.
tracts <- tracts %>%
mutate(
GEOID = as.character(GEOID)
)Check that both join fields are character variables.
class(tracts$GEOID)
class(nanda_disadvantage$GEOID)Expected output:
[1] "character"
[1] "character"
7.9 Merge Community Data with Geometry
We can now join the NaNDA neighborhood disadvantage table to the census tract boundary using GEOID.
tracts_disadvantage <- tracts %>%
left_join(
nanda_disadvantage,
by = "GEOID"
)Because tracts is the object on the left side of left_join(), all tract polygons are retained and NaNDA attributes are added when the GEOIDs match.
Inspect the joined data without printing the large geometry field.
tracts_disadvantage %>%
st_drop_geometry() %>%
select(
GEOID,
DISADVANTAGE
) %>%
head()The output will contain one tract GEOID and its corresponding disadvantage value.
7.9.1 Check the Join
Before mapping, check how many tract polygons received a neighborhood disadvantage value.
sum(
!is.na(
tracts_disadvantage$DISADVANTAGE
)
)Check the total number of polygons.
nrow(tracts_disadvantage)Calculate the percent matched.
mean(
!is.na(
tracts_disadvantage$DISADVANTAGE
)
) * 100This provides a simple quality-control check before moving to visualization.
A small number of unmatched polygons does not necessarily indicate an error. Differences may occur because of data availability, special-use tracts, territorial coverage, or differences between the source datasets.
7.10 Select a Specific County
Community analyses often focus on a smaller geographic area rather than the entire United States.
For this example, we will select Cook County, Illinois.
Cook County is identified by:
State FIPS: 17
County FIPS: 031
Filter the national spatial dataset.
cook_disadvantage <- tracts_disadvantage %>%
filter(
STATEFP == "17",
COUNTYFP == "031"
)Inspect the result.
cook_disadvantage %>%
st_drop_geometry() %>%
select(
GEOID,
STATEFP,
COUNTYFP,
DISADVANTAGE
) %>%
head()To use another county, replace the state and county FIPS values.
For example:
state_fips <- "17"
county_fips <- "031"
county_disadvantage <- tracts_disadvantage %>%
filter(
STATEFP == state_fips,
COUNTYFP == county_fips
)7.11 Map Neighborhood Disadvantage
Once the neighborhood disadvantage values have been attached to the tract polygons, we can create a choropleth map using ggplot2.
7.11.1 County Level
Map Cook County.
ggplot(cook_disadvantage) +
geom_sf(
aes(fill = DISADVANTAGE),
color = NA
) +
scale_fill_viridis_c(
name = "Neighborhood\nDisadvantage",
na.value = "grey90"
) +
labs(
title = "Neighborhood Disadvantage by Census Tract",
subtitle = "Cook County, Illinois — NaNDA 2018–2022",
caption = "Source: National Neighborhood Data Archive"
) +
theme_void()The resulting map displays one polygon for each census tract.
Higher values of DISADVANTAGE represent greater neighborhood disadvantage according to the NaNDA index.
Missing values are shown separately in light gray.
7.11.2 Contiguous United States
If we map the complete national tract file directly, the plotting extent may include Alaska, Hawaii, Puerto Rico, and other U.S. territories. This can make the contiguous United States appear small.
We therefore first remove non-contiguous states and territories.
conus_disadvantage <- tracts_disadvantage %>%
filter(
!STUSPS %in% c(
"AK",
"HI",
"PR",
"GU",
"MP",
"AS",
"VI"
)
)Transform the data to a projection designed for the contiguous United States.
conus_disadvantage <- conus_disadvantage %>%
st_transform(5070)Check the coordinate reference system.
st_crs(conus_disadvantage)The output should identify:
EPSG: 5070
NAD83 / Conus Albers
Now create the national map.
ggplot(conus_disadvantage) +
geom_sf(
aes(fill = DISADVANTAGE),
color = NA
) +
scale_fill_viridis_c(
name = "Neighborhood\nDisadvantage",
na.value = "grey90"
) +
labs(
title = "Neighborhood Disadvantage Across the Contiguous United States",
subtitle = "NaNDA 2018–2022, 2020 Census Tracts",
caption = "Source: National Neighborhood Data Archive"
) +
coord_sf(
datum = NA
) +
theme_void()This produces a tract-level map focused on the 48 contiguous states and the District of Columbia.
7.12 Save Data
After completing the join, the resulting spatial object can be saved for use in another R project, ArcGIS Pro, QGIS, or a web mapping application.
7.12.1 Save as GeoPackage
GeoPackage is useful because it stores the complete spatial dataset in a single file.
st_write(
tracts_disadvantage,
"data/nanda_neighborhood_disadvantage_2020_tracts.gpkg",
delete_dsn = TRUE
)Save only Cook County.
st_write(
cook_disadvantage,
"data/cook_county_neighborhood_disadvantage.gpkg",
delete_dsn = TRUE
)7.12.2 Save as GeoJSON
GeoJSON is also convenient for web mapping and data sharing.
st_write(
cook_disadvantage,
"data/cook_county_neighborhood_disadvantage.geojson",
driver = "GeoJSON",
delete_dsn = TRUE
)7.12.3 Save Maps
Create a map object first.
cook_map <- ggplot(cook_disadvantage) +
geom_sf(
aes(fill = DISADVANTAGE),
color = NA
) +
scale_fill_viridis_c(
name = "Neighborhood\nDisadvantage",
na.value = "grey90"
) +
labs(
title = "Neighborhood Disadvantage by Census Tract",
subtitle = "Cook County, Illinois — NaNDA 2018–2022",
caption = "Source: National Neighborhood Data Archive"
) +
theme_void()Save it as a PNG.
ggsave(
"figures/cook_county_neighborhood_disadvantage.png",
plot = cook_map,
width = 8,
height = 8,
dpi = 300
)7.13 Using Community Data for Health Equity Research
The workflow illustrated in this tutorial can be adapted to many SDOH datasets identified through the Data Discovery platform.
Neighborhood disadvantage could be combined with:
- air pollution
- heat exposure
- traffic exposure
- housing conditions
- industrial land use
- food access
- tree canopy
- healthcare accessibility
- hospitalization
- chronic disease outcomes
For example, a researcher might ask:
Are census tracts with greater neighborhood disadvantage
also experiencing greater environmental burdens?
A community organization might use the resulting maps to identify areas for additional outreach, compare neighborhood conditions with services or infrastructure, support grant applications, or communicate local priorities.
The purpose of spatial data is not simply to produce a map. Spatial data allow researchers and communities to connect what is happening with where it is happening.
7.14 Considerations When Mapping Community Data
Community data should be interpreted carefully.
7.14.1 Area-Level Measures
Neighborhood disadvantage is an area-level measure.
A high tract-level disadvantage score does not mean that every person living within that tract has the same socioeconomic circumstances.
7.14.2 Geographic Units
Census tracts are statistical units and may not correspond to how residents define their neighborhoods.
The geographic scale selected for an analysis should therefore be connected to the research or community question.
7.14.3 Missing Data
Missing values should not automatically be replaced with zero.
A zero represents a measured value. A missing value means that the information was not available or did not match during the join.
7.14.4 Community Context
Measures such as neighborhood disadvantage describe structural and socioeconomic conditions. They should not be used to characterize communities as inherently deficient.
Where possible, pair measures of disadvantage with information about community assets, institutions, services, historical context, and resident priorities.
7.15 Appendix
7.15.1 Common R Errors
7.15.1.1 could not find function "mutate"
mutate() is part of dplyr.
library(dplyr)7.15.1.2 could not find function "str_detect"
str_detect() is part of stringr.
library(stringr)7.15.1.3 Join column GEOID is missing
Check the fields in both objects.
names(tracts)
names(nanda_disadvantage)Both objects should contain a field named:
GEOID
The cleaned NaNDA table used in this tutorial creates the field directly.
nanda_disadvantage <- nanda %>%
transmute(
GEOID = as.character(TRACT_FIPS20),
DISADVANTAGE = as.numeric(DISADVANTAGE)
)7.15.1.4 Very few records match
Check the identifier lengths.
table(
nchar(
nanda_disadvantage$GEOID
)
)
table(
nchar(
tracts$GEOID
)
)For census tracts, the GEOID should contain 11 characters.
Also confirm that the NaNDA file and boundary file use compatible Census geography.
7.15.2 Condensed Workflow
Once the steps above are familiar, the essential workflow can be shortened to:
library(sf)
library(dplyr)
library(ggplot2)
# Locate DS0007
nanda_root <- "~/GCCP-Data/NaNDA-SocEco1990-2022/ICPSR_38528"
ds7_folder <- file.path(
nanda_root,
"DS0007"
)
# Find and load the R data file
r_file <- list.files(
ds7_folder,
pattern = "\\.(rda|RData)$",
full.names = TRUE,
ignore.case = TRUE
)[1]
loaded_object <- load(r_file)
nanda <- get(loaded_object[1])
# Prepare neighborhood disadvantage
nanda_disadvantage <- nanda %>%
transmute(
GEOID = as.character(TRACT_FIPS20),
DISADVANTAGE = as.numeric(DISADVANTAGE)
)
# Read tract geometry
tracts <- st_read(
"~/GCCP-Data/Boundaries/tract-2022-500k-shp"
) %>%
mutate(
GEOID = as.character(GEOID)
)
# Join data to geometry
tracts_disadvantage <- tracts %>%
left_join(
nanda_disadvantage,
by = "GEOID"
)
# Select Cook County
cook_disadvantage <- tracts_disadvantage %>%
filter(
STATEFP == "17",
COUNTYFP == "031"
)
# Map Cook County
ggplot(cook_disadvantage) +
geom_sf(
aes(fill = DISADVANTAGE),
color = NA
) +
scale_fill_viridis_c(
name = "Neighborhood\nDisadvantage",
na.value = "grey90"
) +
theme_void()
# Select the contiguous United States
conus_disadvantage <- tracts_disadvantage %>%
filter(
!STUSPS %in% c(
"AK", "HI", "PR", "GU", "MP", "AS", "VI"
)
) %>%
st_transform(5070)
# Map contiguous United States
ggplot(conus_disadvantage) +
geom_sf(
aes(fill = DISADVANTAGE),
color = NA
) +
scale_fill_viridis_c(
name = "Neighborhood\nDisadvantage",
na.value = "grey90"
) +
coord_sf(datum = NA) +
theme_void()7.15.3 Resources
- SDOH & Place Project
- SDOH & Place Data Discovery
- National Neighborhood Data Archive Socioeconomic Dataset
When using published data in a research product, always be sure consult and cite the codebook, documentation, dataset version, and original data citation associated with the specific dataset used.