7  Community Data Discovery

7.1 Overview

Finding data about community context, the social determinants of health, & structural drivers of health outcomes is essential for research in opioid environments, but can be difficult. Data are often distributed across federal agencies, universities, local governments, research repositories, and other organizations. Even after identifying a useful dataset, additional work is usually required to understand the geographic scale, time period, file structure, identifiers, and documentation before the data can be integrated into a spatial analysis.

There are multiple options for discovering new datasets needed for your search goals. In this tutorial, we’ll share one resource that supports researchers in the discovery process, and work through the stages of inspection, download, and integration into a reproducible coding environment.

Specifically, we’ll use resources from the SDOH & Place Project, a free site that brings together spatial data science, health equity research, and human-centered design to help researchers, advocates, students, planners, and community organizations find and use place-based SDOH data. The Data Discovery Search Tool includes pre-screened, national-level data to help users identify relevant datasets and directs them to the organization or repository where the underlying data can be downloaded.

In this tutorial, our objectives will be to:

  • Demonstrate the SDOH & Place Data Discovery Search tool
  • Download a relevant socioeconomic dataset from its source repository
  • Inspect, clean, and merge the downloaded files in R with spatial boundaries
  • Map neighborhood disadvantage at local and national scales

The worked example uses the National Neighborhood Data Archive (NaNDA) socioeconomic and demographic dataset for 2018–2022 using 2022 Census tract geography.

7.2 Environment Setup

To replicate the code and functions illustrated in this tutorial, you will need to have R and RStudio installed on your system. This tutorial assumes some basic familiarity with the R programming language, such as creating objects and running code from an R script.

7.2.1 Input/Output

Place your input data in yout desired folder location and neame them so it is easier to find.The main external inputs for this tutorial are:

  • the downloaded NaNDA socioeconomic and demographic dataset from ICPSR, as discovered at the SDOH & Place Search tool;
  • the NaNDA codebook and documentation included with the download; and
  • a 2020 Census tract boundary shapefile or GeoJSON, downloaded via Data.gov.

In this example, the downloaded NaNDA folder is stored as:

~/GCCP-Data/NaNDA-SocEco1990-2022/ICPSR_38528

The census tract boundary is stored as:

~/GCCP-Data/Boundaries/tract-2022-500k-shp

The main outputs will be:

  • a spatial census tract dataset containing the NaNDA neighborhood disadvantage index;
  • a neighborhood disadvantage map for Cook County, Illinois;
  • a neighborhood disadvantage map for the contiguous United States; and
  • optionally, a GeoPackage or GeoJSON file for use in other GIS software.

7.2.2 Load Libraries

We will use the following packages in this tutorial:

  • sf: to read, manipulate, transform, and write spatial objects
  • tidyverse: tidy data wrangling tools
  • ggplot2: to create maps (as a break from tmap)
  • here: to help construct reproducible project file paths

We’ll be leveraging ‘dyplyr’ and ‘stringr’ within the tidyverse package, specifically.

Load the required libraries.

library(sf)
library(tidyverse)
library(ggplot2)
library(here)

7.3 Discover Community Data

We can begin by using the SDOH & Place Data Discovery platform to identify a dataset that measures the community characteristic of interest.

The Data Discovery platform helps answer several questions before we begin coding:

  1. What dataset measures the concept we are interested in?
  2. At what geographic scale is the dataset available?
  3. What years or time periods are available?
  4. Does the dataset cover the community or region of interest?
  5. Where can the underlying data be downloaded?

7.3.1 Search by Topic

For this example, we are interested in neighborhood socioeconomic conditions that may influence community epxeriences that impact opioid risk environments.

Open the Data Discovery platform and search for:

economic

We’ll use the default search option, keyword search, in the Discovery Tool. For more complex tasks, you can expand your search with the AI option, and/or use the map to zoom into an area of interest and filter your data options accordingly.

Related searches might include:

socioeconomic
income
poverty
unemployment
neighborhood disadvantage

The search results contain datasets related to economic stability and neighborhood socioeconomic conditions.

7.3.2 Refine by Geography and Time

The Data Discovery platform can be used to refine datasets by geographic unit and time period.

For a neighborhood-level analysis, we are interested in census tract data rather than state-level or county-level data.

The selected dataset should therefore meet the following criteria:

Topic:       Socioeconomic conditions
Geography:   Census tract
Coverage:    United States
Time period: Recent multi-year estimates

7.3.3 Follow the Dataset Source

One relevant dataset is:

National Neighborhood Data Archive (NaNDA): Socioeconomic Status and Demographic Characteristics of Census Tracts and ZIP Code Tabulation Areas, United States, 1990–2022

The underlying data are distributed through the Inter-university Consortium for Political and Social Research (ICPSR).

Dataset page:

https://www.icpsr.umich.edu/web/ICPSR/studies/38528

The Data Discovery platform helps us identify the resource. The actual data are then downloaded from the original repository.

7.4 Select the Appropriate NaNDA Dataset

The downloaded ICPSR package contains several dataset folders, identified as DS0001 through DS0008.

The datasets differ by:

  • geographic unit;
  • time period; and
  • Census geography vintage.

The available folders include:

Folder Geography and Period
DS0001 2010 Census tracts, 1990–2010
DS0002 2010 Census tracts, 2008–2017
DS0003 2010 ZCTAs, 2008–2017
DS0004 2020 Census tracts, 2016–2020
DS0005 2020 ZCTAs, 2016–2020
DS0006 2010 Census tracts, 2018–2022
DS0007 2020 Census tracts, 2018–2022
DS0008 2020 ZCTAs, 2018–2022

For this tutorial, we want recent socioeconomic data at the census tract level using 2020 Census geography.

We therefore use:

DS0007

This corresponds to:

2018–2022 socioeconomic and demographic data
2022 Census tract geography

The geographic vintage is important because the tabular dataset and boundary file should represent the same tract geography.


7.6 Explore Variables

Before selecting a variable for analysis, we should inspect the available fields.

names(nanda)

Some of the available variables include:

TRACT_FIPS20
TOTPOP
POPDEN
PUNEMP
PPOV
PPUBAS
AFFLUENCE
DISADVANTAGE
MEDFAMINC

Rather than manually searching all variable names, we can use str_detect() from the stringr package.

7.6.1 Search for Socioeconomic Variables

names(nanda)[
  str_detect(
    names(nanda),
    regex(
      "disadv|afflu|poverty|income|assist",
      ignore_case = TRUE
    )
  )
]

Output:

[1] "AFFLUENCE"    "DISADVANTAGE"

This search finds variables whose names contain those exact text patterns.

The codebook should always be consulted because several socioeconomic variables use abbreviations, such as:

PPOV
PPUBAS
MEDFAMINC

7.6.2 Find the Geographic Identifier

Search for possible tract or geographic identifiers.

names(nanda)[
  str_detect(
    names(nanda),
    regex(
      "geoid|tract|fips",
      ignore_case = TRUE
    )
  )
]

Output:

[1] "TRACT_FIPS20"

TRACT_FIPS20 is the geographic identifier used to connect the NaNDA data to 2020 Census tract boundaries.

7.7 Prepare the Neighborhood Disadvantage Data

For this tutorial, we only need:

  • the census tract identifier; and
  • the neighborhood disadvantage measure.

Before creating the smaller dataset, it is useful to understand the tract identifier.

7.7.1 Census Tract GEOID

A census tract GEOID contains 11 characters.

The general structure is:

SSCCCTTTTTT

where:

SS      = state FIPS
CCC     = county FIPS
TTTTTT  = census tract code

For example:

17031839100

contains:

17     Illinois
031    Cook County
839100 Census tract code

Geographic identifiers should generally be stored as character variables rather than numbers because some identifiers begin with zero.

Create a simplified NaNDA dataset.

nanda_disadvantage <- nanda %>%
  transmute(
    GEOID = as.character(TRACT_FIPS20),
    DISADVANTAGE = as.numeric(DISADVANTAGE)
  )

Inspect the result.

head(nanda_disadvantage)

Output:

        GEOID DISADVANTAGE
1 01001020100   0.17298892
2 01001020200   0.14308497
3 01001020300   0.10727025
4 01001020400   0.09424171
5 01001020501   0.08787425
8 01001020502   0.04378821

At this stage, the dataset is still a regular table. It does not yet contain spatial geometry.

7.8 Get Geometry

Geometry or geographic boundaries allow us to map and spatially analyze the NaNDA socioeconomic data.

The downloaded NaNDA table contains a tract identifier, but it does not contain the polygon geometry needed to draw each tract.

We therefore need to connect it to a census tract boundary file.

In this example, we use a generalized 2022 Census tract shapefile based on 2020 Census tract geography.

7.8.1 Read the Census Tract Boundary

Read the tract boundary using st_read() from the sf package.

tracts <- st_read(
  "~/GCCP-Data/Boundaries/tract-2022-500k-shp"
)

The output should resemble:

Reading layer `tract-2022-500k' ...
Simple feature collection with 85185 features and 20 fields
Geometry type: MULTIPOLYGON
Dimension: XY
Geodetic CRS: NAD83

Inspect the spatial object.

tracts

Important fields include:

STATEFP
COUNTYFP
TRACTCE
GEOID
STUSPS
STATE_NAME
geometry

Inspect the column names.

names(tracts)

7.8.2 Prepare the Boundary GEOID

Make sure the boundary GEOID is stored as character text.

tracts <- tracts %>%
  mutate(
    GEOID = as.character(GEOID)
  )

Check that both join fields are character variables.

class(tracts$GEOID)

class(nanda_disadvantage$GEOID)

Expected output:

[1] "character"
[1] "character"

7.9 Merge Community Data with Geometry

We can now join the NaNDA neighborhood disadvantage table to the census tract boundary using GEOID.

tracts_disadvantage <- tracts %>%
  left_join(
    nanda_disadvantage,
    by = "GEOID"
  )

Because tracts is the object on the left side of left_join(), all tract polygons are retained and NaNDA attributes are added when the GEOIDs match.

Inspect the joined data without printing the large geometry field.

tracts_disadvantage %>%
  st_drop_geometry() %>%
  select(
    GEOID,
    DISADVANTAGE
  ) %>%
  head()

The output will contain one tract GEOID and its corresponding disadvantage value.

7.9.1 Check the Join

Before mapping, check how many tract polygons received a neighborhood disadvantage value.

sum(
  !is.na(
    tracts_disadvantage$DISADVANTAGE
  )
)

Check the total number of polygons.

nrow(tracts_disadvantage)

Calculate the percent matched.

mean(
  !is.na(
    tracts_disadvantage$DISADVANTAGE
  )
) * 100

This provides a simple quality-control check before moving to visualization.

A small number of unmatched polygons does not necessarily indicate an error. Differences may occur because of data availability, special-use tracts, territorial coverage, or differences between the source datasets.

7.10 Select a Specific County

Community analyses often focus on a smaller geographic area rather than the entire United States.

For this example, we will select Cook County, Illinois.

Cook County is identified by:

State FIPS:  17
County FIPS: 031

Filter the national spatial dataset.

cook_disadvantage <- tracts_disadvantage %>%
  filter(
    STATEFP == "17",
    COUNTYFP == "031"
  )

Inspect the result.

cook_disadvantage %>%
  st_drop_geometry() %>%
  select(
    GEOID,
    STATEFP,
    COUNTYFP,
    DISADVANTAGE
  ) %>%
  head()

To use another county, replace the state and county FIPS values.

For example:

state_fips <- "17"
county_fips <- "031"

county_disadvantage <- tracts_disadvantage %>%
  filter(
    STATEFP == state_fips,
    COUNTYFP == county_fips
  )

7.11 Map Neighborhood Disadvantage

Once the neighborhood disadvantage values have been attached to the tract polygons, we can create a choropleth map using ggplot2.

7.11.1 County Level

Map Cook County.

ggplot(cook_disadvantage) +
  geom_sf(
    aes(fill = DISADVANTAGE),
    color = NA
  ) +
  scale_fill_viridis_c(
    name = "Neighborhood\nDisadvantage",
    na.value = "grey90"
  ) +
  labs(
    title = "Neighborhood Disadvantage by Census Tract",
    subtitle = "Cook County, Illinois — NaNDA 2018–2022",
    caption = "Source: National Neighborhood Data Archive"
  ) +
  theme_void()

The resulting map displays one polygon for each census tract.

Higher values of DISADVANTAGE represent greater neighborhood disadvantage according to the NaNDA index.

Missing values are shown separately in light gray.

7.11.2 Contiguous United States

If we map the complete national tract file directly, the plotting extent may include Alaska, Hawaii, Puerto Rico, and other U.S. territories. This can make the contiguous United States appear small.

We therefore first remove non-contiguous states and territories.

conus_disadvantage <- tracts_disadvantage %>%
  filter(
    !STUSPS %in% c(
      "AK",
      "HI",
      "PR",
      "GU",
      "MP",
      "AS",
      "VI"
    )
  )

Transform the data to a projection designed for the contiguous United States.

conus_disadvantage <- conus_disadvantage %>%
  st_transform(5070)

Check the coordinate reference system.

st_crs(conus_disadvantage)

The output should identify:

EPSG: 5070
NAD83 / Conus Albers

Now create the national map.

ggplot(conus_disadvantage) +
  geom_sf(
    aes(fill = DISADVANTAGE),
    color = NA
  ) +
  scale_fill_viridis_c(
    name = "Neighborhood\nDisadvantage",
    na.value = "grey90"
  ) +
  labs(
    title = "Neighborhood Disadvantage Across the Contiguous United States",
    subtitle = "NaNDA 2018–2022, 2020 Census Tracts",
    caption = "Source: National Neighborhood Data Archive"
  ) +
  coord_sf(
    datum = NA
  ) +
  theme_void()

This produces a tract-level map focused on the 48 contiguous states and the District of Columbia.

7.12 Save Data

After completing the join, the resulting spatial object can be saved for use in another R project, ArcGIS Pro, QGIS, or a web mapping application.

7.12.1 Save as GeoPackage

GeoPackage is useful because it stores the complete spatial dataset in a single file.

st_write(
  tracts_disadvantage,
  "data/nanda_neighborhood_disadvantage_2020_tracts.gpkg",
  delete_dsn = TRUE
)

Save only Cook County.

st_write(
  cook_disadvantage,
  "data/cook_county_neighborhood_disadvantage.gpkg",
  delete_dsn = TRUE
)

7.12.2 Save as GeoJSON

GeoJSON is also convenient for web mapping and data sharing.

st_write(
  cook_disadvantage,
  "data/cook_county_neighborhood_disadvantage.geojson",
  driver = "GeoJSON",
  delete_dsn = TRUE
)

7.12.3 Save Maps

Create a map object first.

cook_map <- ggplot(cook_disadvantage) +
  geom_sf(
    aes(fill = DISADVANTAGE),
    color = NA
  ) +
  scale_fill_viridis_c(
    name = "Neighborhood\nDisadvantage",
    na.value = "grey90"
  ) +
  labs(
    title = "Neighborhood Disadvantage by Census Tract",
    subtitle = "Cook County, Illinois — NaNDA 2018–2022",
    caption = "Source: National Neighborhood Data Archive"
  ) +
  theme_void()

Save it as a PNG.

ggsave(
  "figures/cook_county_neighborhood_disadvantage.png",
  plot = cook_map,
  width = 8,
  height = 8,
  dpi = 300
)

7.13 Using Community Data for Health Equity Research

The workflow illustrated in this tutorial can be adapted to many SDOH datasets identified through the Data Discovery platform.

Neighborhood disadvantage could be combined with:

  • air pollution
  • heat exposure
  • traffic exposure
  • housing conditions
  • industrial land use
  • food access
  • tree canopy
  • healthcare accessibility
  • hospitalization
  • chronic disease outcomes

For example, a researcher might ask:

Are census tracts with greater neighborhood disadvantage
also experiencing greater environmental burdens?

A community organization might use the resulting maps to identify areas for additional outreach, compare neighborhood conditions with services or infrastructure, support grant applications, or communicate local priorities.

The purpose of spatial data is not simply to produce a map. Spatial data allow researchers and communities to connect what is happening with where it is happening.

7.14 Considerations When Mapping Community Data

Community data should be interpreted carefully.

7.14.1 Area-Level Measures

Neighborhood disadvantage is an area-level measure.

A high tract-level disadvantage score does not mean that every person living within that tract has the same socioeconomic circumstances.

7.14.2 Geographic Units

Census tracts are statistical units and may not correspond to how residents define their neighborhoods.

The geographic scale selected for an analysis should therefore be connected to the research or community question.

7.14.3 Missing Data

Missing values should not automatically be replaced with zero.

A zero represents a measured value. A missing value means that the information was not available or did not match during the join.

7.14.4 Community Context

Measures such as neighborhood disadvantage describe structural and socioeconomic conditions. They should not be used to characterize communities as inherently deficient.

Where possible, pair measures of disadvantage with information about community assets, institutions, services, historical context, and resident priorities.

7.15 Appendix

7.15.1 Common R Errors

7.15.1.1 could not find function "mutate"

mutate() is part of dplyr.

library(dplyr)

7.15.1.2 could not find function "str_detect"

str_detect() is part of stringr.

library(stringr)

7.15.1.3 Join column GEOID is missing

Check the fields in both objects.

names(tracts)

names(nanda_disadvantage)

Both objects should contain a field named:

GEOID

The cleaned NaNDA table used in this tutorial creates the field directly.

nanda_disadvantage <- nanda %>%
  transmute(
    GEOID = as.character(TRACT_FIPS20),
    DISADVANTAGE = as.numeric(DISADVANTAGE)
  )

7.15.1.4 Very few records match

Check the identifier lengths.

table(
  nchar(
    nanda_disadvantage$GEOID
  )
)

table(
  nchar(
    tracts$GEOID
  )
)

For census tracts, the GEOID should contain 11 characters.

Also confirm that the NaNDA file and boundary file use compatible Census geography.

7.15.2 Condensed Workflow

Once the steps above are familiar, the essential workflow can be shortened to:

library(sf)
library(dplyr)
library(ggplot2)

# Locate DS0007
nanda_root <- "~/GCCP-Data/NaNDA-SocEco1990-2022/ICPSR_38528"

ds7_folder <- file.path(
  nanda_root,
  "DS0007"
)

# Find and load the R data file
r_file <- list.files(
  ds7_folder,
  pattern = "\\.(rda|RData)$",
  full.names = TRUE,
  ignore.case = TRUE
)[1]

loaded_object <- load(r_file)

nanda <- get(loaded_object[1])

# Prepare neighborhood disadvantage
nanda_disadvantage <- nanda %>%
  transmute(
    GEOID = as.character(TRACT_FIPS20),
    DISADVANTAGE = as.numeric(DISADVANTAGE)
  )

# Read tract geometry
tracts <- st_read(
  "~/GCCP-Data/Boundaries/tract-2022-500k-shp"
) %>%
  mutate(
    GEOID = as.character(GEOID)
  )

# Join data to geometry
tracts_disadvantage <- tracts %>%
  left_join(
    nanda_disadvantage,
    by = "GEOID"
  )

# Select Cook County
cook_disadvantage <- tracts_disadvantage %>%
  filter(
    STATEFP == "17",
    COUNTYFP == "031"
  )

# Map Cook County
ggplot(cook_disadvantage) +
  geom_sf(
    aes(fill = DISADVANTAGE),
    color = NA
  ) +
  scale_fill_viridis_c(
    name = "Neighborhood\nDisadvantage",
    na.value = "grey90"
  ) +
  theme_void()

# Select the contiguous United States
conus_disadvantage <- tracts_disadvantage %>%
  filter(
    !STUSPS %in% c(
      "AK", "HI", "PR", "GU", "MP", "AS", "VI"
    )
  ) %>%
  st_transform(5070)

# Map contiguous United States
ggplot(conus_disadvantage) +
  geom_sf(
    aes(fill = DISADVANTAGE),
    color = NA
  ) +
  scale_fill_viridis_c(
    name = "Neighborhood\nDisadvantage",
    na.value = "grey90"
  ) +
  coord_sf(datum = NA) +
  theme_void()

7.15.3 Resources

When using published data in a research product, always be sure consult and cite the codebook, documentation, dataset version, and original data citation associated with the specific dataset used.