6  Opioid Environment Data

6.1 Overview

The Opioid Environment Policy Scan Data Ecosystem, otherwise known as OEPS, is a data resource and warehouse with nationwide health, policy, and community data related to the opioid epidemic, available to download, analyze, and visualize at multiple spatial scales. It integrates social and structural determinants of health, which are grouped thematically. It currently includes hundreds variable constructs with the data available for download as CSV files which can be easily merged with geographic shapefiles (that are also provided).

This tutorial will show you how to access data from OEPS and directly integrate it within your own R environment. Objectives include:

  • Navigating the most recent OEPS data suite
  • Cleaning, merging, and mapping OEPS data in R
  • Integrating additional data to OEPS measures

6.2 Environment Setup

To replicate the code and functions illustrated in this tutorial, you will need to have R and RStudio installed on your system. This tutorial assumes some basic familiarity with the R programming language, such as creating objects and running code from an R script.

6.2.1 Input/Output

We will download a data package from OEPS directly, and extract data from the state.csv file for our analyses, and later merge with geospatial data from the same package, state-2020-500k-shp, for mapping and exploration.

We’ll also work with the data at the county level from the same data package, county.csv and county-2020-500k-shp, and bring in some measures from an external dataset, NC_cty_foodstamps_2025.csv.

Our output will be a cleaned county-level CSV file, final-county-data.csv, to be used in further research & analyses, as well as several images of maps from our session.

6.2.2 Install Libraries

We will use the following packages in this tutorial:

  • sf: to read, manipulate, transform, and write spatial objects
  • tidyverse: tidy data wrangling tools
  • tmap: to create maps

If you have not already installed tidyverse, sf, and tmap, install them first using install.packages("PackageName"). Then load them with the library() function.

6.3 Download OEPS Data

First, go to the OEPS website at https://oeps.healthyregions.org/.

If this is your first time visiting the site, take a moment to explore the About OEPS, Methodology, Data Inventory and Data Standards pages to get familiar with site.

Next, under the Data Access tab, click Download. On this page, Click Aggregated Data Packages.

This tutorial uses data from the 2023 ACS release, which is a 5-year pooled estimate covering 2019-2023.

Under the DSuite2023 heading, download both of the following files:

  • DSuite2023 data dictionary
  • DSuite2023 data package (100mb+).

See screenshot below.


After downloading the files, create a new folder on your computer called OEPS_Tutorial and save the files there. If the data package downloads as a .zip file, open your Downloads folder, and extract the contents into your OEPS_Tutorial folder before continuing.

We also encourage you to pin this folder to Quick Access.

See the screenshot below.


6.5 Set Up Your R Project

Next, open RStudio and create a new R script called OEPS_Tutorial_RScript.R. Save it in your OEPS_Tutorial folder.

In the Files pane (usually on the bottom right-hand side), click the three dots on the right side, navigate to your OEPS_Tutorial folder, and click Open. Then click the gear icon and select Set As Working Directory.

If the Files pane is not open, click View in the main toolbar and then click Show Files.

At this stage, load the libraries referenced at the start of the tutorial.

## Load required packages
library(tidyverse)
library(sf)
library(tmap)

6.6 Prepare the State-Level Data for Analysis and Visualization

6.6.1 Explore the State-Level Dataset

Now, we will read in the state-level CSV file and inspect its contents.

## Read state.csv
state_data <- read.csv("data/state.csv")

## Glimpse state_data
glimpse(state_data)
Rows: 56
Columns: 83
$ HEROP_ID     <chr> "040US35", "040US46", "040US06", "040US21", "040US01", "0…
$ FIPS         <int> 35, 46, 6, 21, 1, 13, 5, 42, 29, 8, 49, 40, 47, 56, 36, 1…
$ FqhcAvTmDr   <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, N…
$ FqhcCtTmDr   <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, N…
$ FqhcTmDrP    <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, N…
$ TotTracts    <int> 612, 242, NA, 1376, NA, 2791, NA, 3445, 1654, NA, 716, 29…
$ A50_74HcvD   <int> 92, 10, 1407, 218, 122, 205, 93, 328, 110, 289, 12, 457, …
$ AmInHcvD     <int> NA, 15, 22, NA, NA, NA, NA, NA, NA, 11, NA, 57, NA, NA, N…
$ AsHcvD       <dbl> NA, 0, 87, 0, 0, NA, NA, NA, NA, NA, 0, NA, NA, 0, 17, NA…
$ AvA50_74HcvD <dbl> 131, 6, 1703, 219, 132, 251, 115, 357, 163, 331, 43, 453,…
$ AvAmInHcvD   <dbl> 7, 2, 32, NA, NA, 0, NA, 0, 0, 4, 0, 58, NA, NA, NA, 0, N…
$ AvBlkHcvD    <dbl> NA, NA, 246, 34, 48, 96, 21, 97, 40, 36, NA, 44, 82, 0, 1…
$ AvFlHcvD     <dbl> 40, 0, 605, 88, 46, 84, 42, 114, 60, 112, 20, 163, 139, N…
$ AvHcvD       <dbl> 159, 33, 2085, 302, 161, 292, 135, 432, 197, 388, 64, 536…
$ AvHspHcvD    <dbl> 88, 0, 645, NA, NA, NA, NA, 36, NA, 101, 2, 21, NA, NA, 1…
$ AvMlHcvD     <dbl> 119, 25, 1479, 213, 115, 208, 94, 318, 137, 276, 43, 373,…
$ AvO75HcvD    <dbl> NA, NA, 263, NA, NA, 10, NA, 22, 1, 20, NA, 22, 17, NA, 7…
$ AvU50HcvD    <dbl> 1, NA, 108, 62, NA, 1, NA, 25, 4, 7, NA, 34, 66, NA, 14, …
$ MdMarijLaw   <chr> "True", "True", "True", "False", "True", "False", "True",…
$ AnyPdmpDt    <chr> "7/15/2004", "3/29/2010", "1/1/1939", "7/15/1998", "8/1/2…
$ AnyPdmpFr    <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ AnyPdmphDt   <chr> "7/1/2004", "3/1/2010", "1/1/1990", "7/1/1998", "11/1/200…
$ AnyPdmphFr   <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ ElcPdmpDt    <chr> "7/15/2004", "7/1/2010", "1/1/1998", "7/15/1998", "8/1/20…
$ ElcPdmpFr    <dbl> 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 0.05, 1.0…
$ MsAcPdmpDt   <chr> "9/28/2012", "", "10/2/2018", "7/20/2012", "3/9/2017", "1…
$ MsAcPdmpFr   <dbl> 1, 0, 1, 1, 1, 1, 1, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 0, 1, …
$ OpPdmpDt     <chr> "8/1/2005", "3/1/2012", "9/1/2009", "7/1/1999", "4/1/2006…
$ OpPdmpFr     <dbl> 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 0.05, 1.0…
$ CrrctExp     <int> 6359956, 783784, 112801895, 6670160, 12533999, 16334064, …
$ ExpnFedExp   <dbl> 2026725000, 11769600, 28775203200, 5189196600, NA, NA, 23…
$ ExpnSttExp   <dbl> 211870300, 960000, 3268018800, 578826600, NA, NA, 2771269…
$ HlthExp      <int> 4692974, 316744, 76926071, 4620478, 10626825, 11019466, 2…
$ MedcdExp     <dbl> 8173504300, 1167718100, 122791744500, 16172800500, 787630…
$ PlcFyrExp    <int> 1214640, 369042, 35511690, 1636492, 2145358, 4925933, 116…
$ TradFedExp   <dbl> 4739079800, 786090500, 50713624800, 8031034100, 611585800…
$ TradSttExp   <dbl> 1195829200, 368898000, 40034897800, 2373743200, 176044310…
$ WlfrExp      <int> 10049039, 1731185, 167181476, 16877514, 9596296, 17009307…
$ BachelorsP   <dbl> 16.59, 21.52, 22.40, 15.91, 16.96, 20.73, 15.88, 20.57, 1…
$ EduHsP       <dbl> 25.73, 29.48, 20.40, 32.71, 30.34, 26.89, 34.39, 33.21, 3…
$ EduNoHsP     <dbl> 12.28, 6.99, 15.40, 11.47, 11.87, 11.05, 11.44, 8.09, 8.4…
$ EngProf      <dbl> 91.34, 97.90, 82.72, 97.31, 97.64, 94.28, 96.75, 95.28, 9…
$ GradSclP     <dbl> 13.61, 9.58, 14.10, 11.08, 10.79, 13.47, 9.23, 13.92, 12.…
$ SomeCollegeP <dbl> 19.47, 17.07, 18.18, 16.81, 17.43, 16.33, 17.82, 12.87, 1…
$ FemP         <dbl> 50.33, 49.33, 50.04, 50.48, 51.46, 51.20, 50.67, 50.71, 5…
$ MaleP        <dbl> 49.67, 50.67, 49.96, 49.52, 48.54, 48.80, 49.33, 49.29, 4…
$ Ovr16        <dbl> 1705686, 704344, 31545603, 3605426, 4056609, 8582918, 240…
$ Ovr16P       <dbl> 80.66, 78.33, 80.39, 79.93, 80.26, 79.31, 79.46, 81.84, 8…
$ Ovr18        <dbl> 1648151, 679374, 30513773, 3487979, 3926252, 8280518, 232…
$ Ovr18P       <dbl> 77.94, 75.55, 77.76, 77.33, 77.68, 76.51, 76.72, 79.41, 7…
$ Ovr21        <dbl> 1443310, 595134, 26488214, 3028981, 3389617, 7120517, 202…
$ Ovr21P       <dbl> 68.25, 66.19, 67.50, 67.15, 67.06, 65.79, 66.69, 68.94, 6…
$ Ovr65        <dbl> 397156, 158103, 5994486, 767995, 884210, 1581263, 525644,…
$ Ovr65P       <dbl> 18.78, 17.58, 15.28, 17.03, 17.49, 14.61, 17.33, 19.07, 1…
$ SRatio       <dbl> 98.68, 102.71, 99.84, 98.11, 94.33, 95.32, 97.35, 97.20, …
$ SRatio18     <dbl> 97.27, 102.18, 98.16, 95.74, 91.27, 92.35, 94.85, 94.85, …
$ SRatio65     <dbl> 85.51, 89.35, 82.11, 81.51, 78.70, 78.77, 81.21, 80.64, 8…
$ TotPop       <int> 2114768, 899194, 39242785, 4510725, 5054253, 10822590, 30…
$ AmIndE       <dbl> 201346, 69514, 445219, 7610, 22491, 42058, 16893, 23921, …
$ AmIndP       <dbl> 9.52, 7.73, 1.13, 0.17, 0.44, 0.39, 0.56, 0.18, 0.27, 1.0…
$ AsianE       <dbl> 36065, 12554, 5997069, 68482, 71969, 472915, 47309, 48043…
$ AsianP       <dbl> 1.71, 1.40, 15.28, 1.52, 1.42, 4.37, 1.56, 3.70, 2.07, 3.…
$ BlackE       <dbl> 44709, 20149, 2173343, 355237, 1318507, 3391689, 452127, …
$ BlackP       <dbl> 2.11, 2.24, 5.54, 7.88, 26.09, 31.34, 14.91, 10.73, 11.13…
$ HisE         <dbl> 1018321, 41281, 15630830, 212163, 271640, 1158299, 265833…
$ HisP         <dbl> 48.15, 4.59, 39.83, 4.70, 5.37, 10.70, 8.77, 8.38, 5.06, …
$ OtherE       <dbl> 253794, 12547, 6820303, 67198, 107302, 448586, 91871, 445…
$ OtherP       <dbl> 12.00, 1.40, 17.38, 1.49, 2.12, 4.14, 3.03, 3.43, 1.73, 5…
$ PacIsE       <dbl> 2049, 606, 147827, 3737, 2564, 7338, 12040, 4614, 9681, 8…
$ PacIsP       <dbl> 0.10, 0.07, 0.38, 0.08, 0.05, 0.07, 0.40, 0.04, 0.16, 0.1…
$ TwoRaceE     <dbl> 442934, 50789, 6410245, 233880, 228050, 782473, 263525, 7…
$ TwoRaceP     <dbl> 20.94, 5.65, 16.33, 5.18, 4.51, 7.23, 8.69, 6.12, 6.30, 1…
$ WhiteE       <dbl> 1133871, 733035, 17248779, 3774581, 3303370, 5677531, 214…
$ WhiteP       <dbl> 53.62, 81.52, 43.95, 83.68, 65.36, 52.46, 70.86, 75.80, 7…
$ DsmAs        <dbl> 0.42, 0.24, 0.41, 0.52, 0.57, 0.47, 0.56, 0.47, 0.48, 0.3…
$ DsmBlk       <dbl> 0.39, 0.28, 0.52, 0.45, 0.40, 0.34, 0.42, 0.55, 0.44, 0.4…
$ DsmHsp       <dbl> 0.24, 0.19, 0.33, 0.40, 0.44, 0.34, 0.38, 0.40, 0.33, 0.2…
$ IntrAsWht    <dbl> 0.36, 0.53, 0.47, 0.76, 0.62, 0.54, 0.68, 0.82, 0.83, 0.6…
$ IntrBlkWht   <dbl> 0.38, 0.61, 0.45, 0.84, 0.54, 0.54, 0.66, 0.74, 0.84, 0.6…
$ IntrHspWht   <dbl> 0.37, 0.74, 0.45, 0.86, 0.60, 0.59, 0.68, 0.78, 0.86, 0.6…
$ IsoAs        <dbl> 0.03, 0.01, 0.13, 0.02, 0.02, 0.04, 0.02, 0.04, 0.02, 0.0…
$ IsoBlk       <dbl> 0.03, 0.01, 0.07, 0.07, 0.37, 0.35, 0.21, 0.13, 0.07, 0.0…
$ IsoHsp       <dbl> 0.53, 0.04, 0.40, 0.06, 0.08, 0.11, 0.10, 0.10, 0.06, 0.2…

The state-level dataset contains 56 rows, which represent the 50 states plus the District of Columbia and U.S. territories. It also contains 83 columns, each representing a different variable.

The glimpse() function shows the structure of the dataset, including each variable’s type and a preview of its values.

You can also view the loaded dataset by clicking on the name in the Environment tab.

6.6.2 Choose a Variable to Explore

For this tutorial, we will explore Hepatitis C (HCV) mortality as an example health outcome variable. The same workflow can be used to explore other health outcomes.

Open the HepatitisC_Rates.md in the metadata folder and read through the contents. The OEPS Hepatitis C (HCV) data was orginially sourced from HepVu. HepVu includes single-year data from 2018-2022. OEPS researchers cleaned and prepared this data for analysis by aggregating the multiple single year datasets into a single multi-year state-level dataset for 2018-2022.

Hepatitis C, also referred to as HCV, is relevant to opioid-environment research because injection drug use is one of the most common risk factors for HCV infection. People who use drugs are therefore at elevated risk for HCV. HCV can lead to chronic infection, serious liver disease, and/or death if left untreated. For these reasons, HCV is often used as a related health outcome in opioid research.

In the DSuite2023-data-dictionary.xlsx, in the Construct column, scroll to Hepatitis C Rates. You will see several variables related to HCV deaths by age, race, and ethnicity.

For simplicity, we will use AvHcvD, which is the average annual number of HCV deaths from 2018-2022. This variable is only available at the state-level in the OEPS database. While HepVu does include county-level HCV mortality, the data is a calculated estimation. For a complete description of the data, see HepVu’s Data Methods.

Return to RStudio and locate the AvHcvD variable in state_data. You will see that the row with FIPS code 35, which corresponds to New Mexico, has an average annual HCV death count of 159 people.

To check which FIPS code corresponds to which state or territory, visit this page from the U.S. Census Bureau.

6.6.3 Create a Population-Adjusted Measure

Next, we will prepare the data and generate descriptive statistics for the health outcome variable.

The AvHcvD variable is a raw count. Because states have different population sizes, we will standardize this measure by dividing AvHcvD by TotPop, the total population.

We will call this new variable PropHcvD. It represents the average annual proportion of the total state population that died from HCV.

We use total population as the denominator, rather than total HCV cases because of data availability constraints. Data for average annual HCV cases is only available as a 5-year pooled estimate covering 2013-2016. Whereas, this tutorial is using pooled mortality data from 2018-2022.

## Create a population-adjusted variable
state_data$PropHcvD <- (state_data$AvHcvD / state_data$TotPop)

## View the first few values
head(state_data$PropHcvD)
[1] 7.518555e-05 3.669953e-05 5.313079e-05 6.695154e-05 3.185436e-05
[6] 2.698060e-05

6.6.4 Summarize the Variable

Next, use the summary() function to examine the distribution of PropHcvD. This output includes the mean, median, minimum, maximum, first quartile, third quartile, and the count of missing values.

## Summarize new variable
summary(state_data$PropHcvD)
    Min.  1st Qu.   Median     Mean  3rd Qu.     Max.     NA's 
0.000019 0.000032 0.000039 0.000046 0.000054 0.000134        5 

Because these values are very small, they can be difficult to interpret directly. Additionally, the output shows five missing values, which we will address in a later section.

6.6.5 Re-express the Variable Per 100,000 People

To make the variable easier to interpret, multiply PropHcvD by 100,000. This gives the average annual HCV deaths per 100,000 people. We will call this new variable AvHcVD100K.

## Re-express the variable per 100,000 people
state_data$AvHcvD100K <- state_data$PropHcvD * 100000

6.6.6 Identify Missing Values

Now we will identify the rows with missing values for AvHcvD100K.

## View the rows with missing AvHcvD100K data
state_data[is.na(state_data$AvHcvD100K),]
   HEROP_ID FIPS FqhcAvTmDr FqhcCtTmDr FqhcTmDrP TotTracts A50_74HcvD AmInHcvD
30  040US60   60         NA         NA        NA      8706         NA       NA
35  040US72   72         NA         NA        NA       939         NA       NA
48  040US78   78         NA         NA        NA        29         NA       NA
54  040US66   66         NA         NA        NA        56         NA       NA
55  040US69   69         NA         NA        NA        23         NA       NA
   AsHcvD AvA50_74HcvD AvAmInHcvD AvBlkHcvD AvFlHcvD AvHcvD AvHspHcvD AvMlHcvD
30     NA           NA         NA        NA       NA     NA        NA       NA
35     NA           NA         NA        NA       NA     NA        NA       NA
48     NA           NA         NA        NA       NA     NA        NA       NA
54     NA           NA         NA        NA       NA     NA        NA       NA
55     NA           NA         NA        NA       NA     NA        NA       NA
   AvO75HcvD AvU50HcvD MdMarijLaw AnyPdmpDt AnyPdmpFr AnyPdmphDt AnyPdmphFr
30        NA        NA      False                  NA                    NA
35        NA        NA      False                  NA                    NA
48        NA        NA      False                  NA                    NA
54        NA        NA      False                  NA                    NA
55        NA        NA      False                  NA                    NA
   ElcPdmpDt ElcPdmpFr MsAcPdmpDt MsAcPdmpFr OpPdmpDt OpPdmpFr CrrctExp
30                  NA                    NA                NA       NA
35                  NA                    NA                NA       NA
48                  NA                    NA                NA       NA
54                  NA                    NA                NA       NA
55                  NA                    NA                NA       NA
   ExpnFedExp ExpnSttExp HlthExp MedcdExp PlcFyrExp TradFedExp TradSttExp
30         NA         NA      NA       NA        NA         NA         NA
35         NA         NA      NA       NA        NA         NA         NA
48         NA         NA      NA       NA        NA         NA         NA
54         NA         NA      NA       NA        NA         NA         NA
55         NA         NA      NA       NA        NA         NA         NA
   WlfrExp BachelorsP EduHsP EduNoHsP EngProf GradSclP SomeCollegeP FemP MaleP
30      NA         NA     NA       NA      NA       NA           NA   NA    NA
35      NA         NA     NA       NA      NA       NA           NA   NA    NA
48      NA         NA     NA       NA      NA       NA           NA   NA    NA
54      NA         NA     NA       NA      NA       NA           NA   NA    NA
55      NA         NA     NA       NA      NA       NA           NA   NA    NA
   Ovr16 Ovr16P Ovr18 Ovr18P Ovr21 Ovr21P Ovr65 Ovr65P SRatio SRatio18 SRatio65
30    NA     NA    NA     NA    NA     NA    NA     NA     NA       NA       NA
35    NA     NA    NA     NA    NA     NA    NA     NA     NA       NA       NA
48    NA     NA    NA     NA    NA     NA    NA     NA     NA       NA       NA
54    NA     NA    NA     NA    NA     NA    NA     NA     NA       NA       NA
55    NA     NA    NA     NA    NA     NA    NA     NA     NA       NA       NA
   TotPop AmIndE AmIndP AsianE AsianP BlackE BlackP HisE HisP OtherE OtherP
30     NA     NA     NA     NA     NA     NA     NA   NA   NA     NA     NA
35     NA     NA     NA     NA     NA     NA     NA   NA   NA     NA     NA
48     NA     NA     NA     NA     NA     NA     NA   NA   NA     NA     NA
54     NA     NA     NA     NA     NA     NA     NA   NA   NA     NA     NA
55     NA     NA     NA     NA     NA     NA     NA   NA   NA     NA     NA
   PacIsE PacIsP TwoRaceE TwoRaceP WhiteE WhiteP DsmAs DsmBlk DsmHsp IntrAsWht
30     NA     NA       NA       NA     NA     NA    NA     NA     NA        NA
35     NA     NA       NA       NA     NA     NA    NA     NA     NA        NA
48     NA     NA       NA       NA     NA     NA    NA     NA     NA        NA
54     NA     NA       NA       NA     NA     NA    NA     NA     NA        NA
55     NA     NA       NA       NA     NA     NA    NA     NA     NA        NA
   IntrBlkWht IntrHspWht IsoAs IsoBlk IsoHsp PropHcvD AvHcvD100K
30         NA         NA    NA     NA     NA       NA         NA
35         NA         NA    NA     NA     NA       NA         NA
48         NA         NA    NA     NA     NA       NA         NA
54         NA         NA    NA     NA     NA       NA         NA
55         NA         NA    NA     NA     NA       NA         NA

These missing values correspond to American Samoa (60), Guam (66), Northern Mariana Islands (69), Puerto Rico (72), and the U.S. Virgin Islands (78). For these locations, the dataset does not include HCV deaths or total population values.

6.6.7 Filter the Dataset

For the remainder of this tutorial, we will limit the analysis to the contiguous U.S., which will also make the mapping step easier. This excludes the territories listed above, as well as Alaska (02), and Hawaii (15).

## Filter to the contiguous U.S.
state_data_filtered <- state_data %>% filter(!FIPS %in% c(02, 15, 60, 66, 69, 72, 78))

6.6.8 Summary Statistics

Now that AvHcvD100K has been calculated and processed, use the summary() function again to examine the distribution.

## Summarize new variable
summary(state_data_filtered$AvHcvD100K)
   Min. 1st Qu.  Median    Mean 3rd Qu.    Max. 
  1.921   3.194   3.949   4.642   5.313  13.416 

From this output, we can see the typical range of values across states and identify whether any states have unusually high or low rates. The summary also helps us confirm that the missing values have been removed from the filtered dataset.

The average annual HCV death rate ranges from 1.921 to 13.416 deaths per 100,000 people across states in the contiguous U.S. The median rate is 3.949 and the mean rate is 4.642, which is higher than the median, suggesting that the distribution is slightly right-skewed, with a few states having relatively high rates.

The distribution suggests that HCV mortality is not uniform across states. But it is important to note that summary statistics alone cannot tell us why rates differ.

6.6.9 Load the State Shapefile

Now that the data has been prepared, we can create a map of the outcome variable.

Before we can map the data, we need to join the state-level dataset to a state shapefile. The shapefile is included in the OEPS data suite.

In your OEPS_Tutorial folder, open the data folder and locate state-2020-500k-shp.zip. If the file is still compressed, unzip it first. For best organization, create a folder named state-2020-500k-shp inside the data folder and extract the shapefile contents there.

See the screenshot below.

Now load the shapefile into R and display the geometry column to confirm that it imported correctly.

## Load in the state shapefile
state_geom <- st_read("data/state-2020-500k-shp/state-2020-500k.shp")
Reading layer `state-2020-500k' from data source 
  `/Users/maryniakolak/Code/opioid-environment-toolkit/data/state-2020-500k-shp/state-2020-500k.shp' 
  using driver `ESRI Shapefile'
Simple feature collection with 56 features and 16 fields
Geometry type: MULTIPOLYGON
Dimension:     XY
Bounding box:  xmin: -179.1467 ymin: -14.5487 xmax: 179.7785 ymax: 71.38782
Geodetic CRS:  NAD83
## Create a quick plot
plot(state_geom$geometry)

6.6.10 Filter to the Contiguous U.S.

Next, filter the shapefile to include only the contiguous U.S. This excludes Alaska, Hawaii, and the U.S. territories.

## Filter to the contiguous U.S.
state_geom_filtered <- state_geom %>% filter(!STATEFP %in% c("02", "15", "60", "66", "69", "72", "78"))

## Check plot
plot(state_geom_filtered$geometry)

6.6.11 Join the Data to the Shapefile

Now join the state_data_filtered dataframe to the filtered shapefile. We join using HEROP_ID which serves as a unified geographic identifier for joins. To read more about this variable click here.

## Join state_data_filtered to state_geom_filtered
data_to_map <- state_geom_filtered %>%
  left_join(state_data_filtered, by = "HEROP_ID")

6.6.12 Create a Map

Finally, map AvHcvD100K, which represents the average annual HCV death rate per 100,000 people.

# Map
tm_shape(data_to_map) +
  tm_polygons(
    fill = "AvHcvD100K",
    fill.scale = tm_scale_intervals(
      values = c("white", "#FFF7BC", "#FEC44F", "#FE9929", "#D95F0E", "#993404"),
      breaks = c(0, 0.0000001, 3.99999, 6.99999, 9.99999, 11.99999, 14.99999),
      labels = c("0", "1–3", "4–6", "7–9", "10–12", "12-14"),
      label.na = ""
    ),
    fill.legend = tm_legend(title = "")) +
  tm_title("Average Annual Hepatitis C (HCV) Deaths per 100,000 People")
[plot mode] fit legend/component: Some legend items or map compoments do not
fit well, and are therefore rescaled.
ℹ Set the tmap option `component.autoscale = FALSE` to disable rescaling.

The map suggests that HCV death rates are not evenly distributed across the U.S. Instead, several Southern and Western states appear to have higher rates compared to Northeastern states. Oklahoma appears to have the highest rate of annual HCV deaths.

6.7 Prepare the County-Level Data for Analysis and Visualization

OEPS includes data at multiple geographic levels. In this section, we will demonstrate how the state-level workflow can be applied to county-level data. Additionally, we will show how to subset the data to a single state for closer examination.

6.7.1 Explore the County-Level Dataset

Now, we will read in the county-level CSV file and inspect its contents.

## Read county.csv
county_data <- read.csv("data/county.csv")

## Glimpse state_data
glimpse(county_data)
Rows: 3,234
Columns: 100
$ HEROP_ID     <chr> "050US13031", "050US13121", "050US13179", "050US13189", "…
$ FIPS         <int> 13031, 13121, 13179, 13189, 13213, 13209, 13139, 13245, 1…
$ SviSmryRnk   <dbl> 0.8374, 0.6599, 0.9007, 0.6567, 0.6847, 0.7547, 0.8136, 0…
$ SviTh1       <dbl> 0.9297, 0.5571, 0.8167, 0.7353, 0.7852, 0.8858, 0.7897, 0…
$ SviTh2       <dbl> 0.2326, 0.2278, 0.8431, 0.5002, 0.6061, 0.5756, 0.6640, 0…
$ SviTh3       <dbl> 0.7521, 0.9259, 0.9322, 0.8540, 0.5250, 0.7191, 0.7941, 0…
$ SviTh4       <dbl> 0.8781, 0.8501, 0.8600, 0.4076, 0.4909, 0.4518, 0.7416, 0…
$ EssnWrkP     <dbl> 49.9, 27.6, 50.8, 54.9, 54.6, 57.4, 48.3, 53.4, 70.5, 61.…
$ HghRskP      <dbl> 18.0, 10.5, 14.8, 25.2, 41.6, 30.2, 28.9, 18.0, 28.0, 27.…
$ HltCrP       <dbl> 10.1, 10.2, 12.5, 15.5, 8.3, 14.6, 10.0, 16.3, 8.1, 14.0,…
$ RetailP      <dbl> 12.3, 9.0, 14.0, 11.9, 12.0, 9.4, 11.1, 12.3, 12.8, 18.7,…
$ TotWrkE      <int> 36964, 569609, 23579, 9334, 18970, 3565, 99845, 86611, 48…
$ GiniCoeff    <dbl> 0.4849, 0.5296, 0.4089, 0.4515, 0.4337, 0.4652, 0.4493, 0…
$ MedInc       <int> 35840, 61045, 36864, 36134, 41999, 41019, 41391, 35067, 3…
$ PciE         <int> 30450, 61438, 27648, 28071, 30205, 27633, 37271, 30209, 2…
$ PovP         <dbl> 22.5, 12.9, 14.5, 19.0, 13.0, 16.4, 12.7, 21.1, 22.5, 31.…
$ UnempP       <dbl> 7.9, 5.4, 8.0, 3.7, 5.0, 6.1, 3.9, 7.9, 10.1, 5.9, 3.4, 6…
$ LngTermP     <dbl> 14.2, 11.5, 12.3, 28.8, 19.1, 23.7, 13.6, 18.6, 22.4, 36.…
$ MobileP      <dbl> 15.4, 0.5, 16.7, 27.8, 28.9, 37.9, 8.4, 6.5, 37.1, 44.1, …
$ RentalP      <dbl> 111.0, 93.5, 139.0, 94.4, 68.5, 55.7, 89.4, 122.3, 50.0, …
$ TotUnits     <int> 33607, 500404, 27228, 9466, 16163, 3776, 79233, 92652, 47…
$ UnitDens     <dbl> 49.7, 950.1, 52.7, 36.8, 46.9, 15.7, 201.6, 285.7, 6.0, 1…
$ VacantP      <dbl> 9.9, 8.5, 14.0, 12.3, 5.6, 20.0, 10.1, 19.0, 13.2, 42.3, …
$ FqhcAvTmDr   <dbl> 10.51, 7.07, 11.77, 19.81, 12.37, 8.41, 10.97, 6.99, 18.7…
$ FqhcCtTmDr   <int> 20, 326, 16, 6, 9, 3, 50, 55, 2, 2, 17, 4, 8, 4, 8, 9, 2,…
$ FqhcTmDrP    <dbl> 100.00, 99.69, 94.12, 100.00, 90.00, 100.00, 100.00, 98.2…
$ TotTracts    <int> 20, 327, 17, 6, 10, 3, 50, 56, 3, 2, 17, 4, 8, 4, 8, 9, 2…
$ OpRxRt       <dbl> 72.4, 61.1, 27.5, 52.2, 28.3, 0.2, 79.0, 89.9, 58.6, 9.9,…
$ BachelorsP   <dbl> 18.6, 33.6, 13.5, 11.2, 8.8, 10.8, 16.7, 14.7, 8.4, 5.7, …
$ EduHsP       <dbl> 28.6, 15.4, 31.7, 40.4, 35.1, 37.3, 27.7, 31.8, 43.7, 39.…
$ EduNoHsP     <dbl> 10.4, 6.4, 8.0, 14.5, 21.3, 14.7, 18.6, 12.1, 14.8, 26.1,…
$ EngProf      <dbl> 0.007, 0.019, 0.011, 0.001, 0.026, 0.021, 0.080, 0.005, 0…
$ GradSclP     <dbl> 12.2, 24.4, 6.4, 6.6, 4.9, 9.5, 10.0, 9.4, 3.0, 2.7, 8.7,…
$ SomeCollegeP <dbl> 21.5, 14.4, 28.6, 18.6, 21.9, 16.9, 19.3, 22.3, 24.9, 18.…
$ CrowdHsng    <dbl> 0.01, 0.02, 0.02, 0.01, 0.03, 0.01, 0.05, 0.01, 0.03, 0.0…
$ DivrcdP      <dbl> 9.0, 8.9, 9.3, 10.8, 10.5, 10.0, 8.6, 12.5, 13.4, 6.3, 8.…
$ FamSize      <dbl> 3.05, 3.12, 3.39, 3.12, 3.00, 3.18, 3.37, 3.43, 3.71, 3.1…
$ HHSize       <dbl> 2.48, 2.26, 2.75, 2.58, 2.62, 2.59, 2.89, 2.62, 2.94, 2.4…
$ HhldFA       <dbl> 15.3, 21.0, 12.2, 16.0, 12.0, 14.7, 12.8, 19.5, 13.4, 20.…
$ HhldFC       <dbl> 6.6, 6.1, 8.6, 12.8, 4.3, 7.8, 4.5, 12.0, 4.1, 6.0, 3.7, …
$ HhldFS       <dbl> 6.1, 6.5, 4.4, 5.8, 6.6, 9.2, 6.7, 8.2, 6.1, 11.0, 5.5, 1…
$ HhldMA       <dbl> 14.0, 17.4, 15.1, 10.5, 8.0, 14.2, 8.9, 15.5, 14.8, 13.0,…
$ HhldMC       <dbl> 0.9, 1.1, 1.7, 1.3, 2.2, 1.5, 1.3, 1.3, 0.0, 0.0, 1.3, 2.…
$ HhldMS       <dbl> 3.9, 3.0, 2.6, 3.3, 2.9, 5.8, 3.0, 4.2, 3.7, 6.1, 2.3, 6.…
$ HsdTot       <dbl> 30276, 457791, 23406, 8298, 15258, 3021, 71263, 75023, 40…
$ HsdTypCo     <dbl> 7.4, 6.6, 6.8, 4.3, 5.4, 4.7, 5.8, 5.4, 8.5, 3.0, 5.4, 3.…
$ HsdTypM      <dbl> 37.0, 34.9, 45.8, 41.5, 55.8, 49.1, 55.6, 31.1, 41.6, 43.…
$ HsdTypMC     <dbl> 14.9, 14.3, 18.9, 10.5, 21.2, 14.0, 21.7, 9.1, 14.8, 6.2,…
$ MrrdP        <dbl> 37.2, 41.4, 47.3, 46.2, 56.7, 48.1, 55.3, 33.6, 34.5, 31.…
$ NonRelFhhP   <dbl> 7.0, 6.8, 7.4, 7.6, 8.2, 8.2, 9.7, 9.3, 9.5, 12.7, 7.6, 9…
$ NonRelNfhhP  <dbl> 10.3, 3.9, 2.0, 2.8, 1.8, 1.5, 3.1, 5.4, 3.4, 4.0, 2.3, 5…
$ NvMrrdP      <dbl> 49.5, 46.4, 38.0, 37.8, 28.1, 37.7, 32.3, 49.2, 48.2, 54.…
$ OccupantP    <dbl> 90.1, 91.5, 86.0, 87.7, 94.4, 80.0, 89.9, 81.0, 86.8, 57.…
$ SepartedP    <dbl> 1.6, 1.3, 3.5, 2.2, 1.6, 0.9, 1.5, 2.3, 1.9, 2.1, 1.1, 3.…
$ TotPopHh     <int> 75146, 1036112, 64336, 21421, 39993, 7827, 206163, 196186…
$ WidwdP       <dbl> 2.7, 2.0, 1.9, 3.0, 3.1, 3.3, 2.2, 2.5, 2.0, 6.1, 2.3, 2.…
$ DisbP        <dbl> 14.8, 10.2, 15.9, 12.4, 13.3, 14.0, 11.7, 19.1, 19.8, 23.…
$ FemP         <dbl> 51.5, 51.4, 48.4, 50.7, 49.9, 48.4, 50.1, 51.6, 42.0, 47.…
$ MaleP        <dbl> 48.5, 48.6, 51.6, 49.3, 50.1, 51.6, 49.9, 48.4, 58.0, 52.…
$ MedAge       <dbl> 30.2, 36.2, 28.7, 39.1, 39.3, 38.2, 38.2, 35.0, 38.0, 47.…
$ Ovr16        <dbl> 66867, 869503, 49840, 16947, 31927, 7071, 163982, 163638,…
$ Ovr16P       <dbl> 82.2, 81.4, 74.6, 78.1, 79.3, 81.5, 78.7, 79.4, 84.5, 86.…
$ Ovr18        <dbl> 64967, 842905, 48332, 16263, 30663, 6898, 158032, 158585,…
$ Ovr18P       <dbl> 79.8, 78.9, 72.3, 75.0, 76.1, 79.5, 75.8, 77.0, 80.3, 84.…
$ Ovr21        <dbl> 50293, 726211, 41973, 14025, 26256, 5695, 135491, 137391,…
$ Ovr21P       <dbl> 66.8, 74.7, 66.9, 70.8, 72.2, 71.6, 71.7, 71.9, 72.9, 81.…
$ Ovr62P       <dbl> 14.8, 15.7, 12.4, 22.6, 19.1, 21.8, 19.7, 18.6, 18.8, 27.…
$ Ovr65        <dbl> 9836, 132189, 6558, 4004, 5997, 1512, 33285, 30664, 1872,…
$ Ovr65P       <dbl> 12.1, 12.4, 9.8, 18.5, 14.9, 17.4, 16.0, 14.9, 14.7, 24.6…
$ SRatio       <dbl> 94.3, 94.7, 106.5, 97.1, 100.5, 106.4, 99.6, 93.8, 138.0,…
$ SRatio18     <dbl> 94.1, 92.6, 107.3, 85.3, 98.4, 104.6, 97.8, 89.8, 147.8, …
$ SRatio65     <dbl> 83.5, 74.0, 81.5, 72.1, 82.1, 81.7, 84.0, 73.8, 104.6, 87…
$ TotPop       <int> 81372, 1068507, 66826, 21687, 40282, 8675, 208395, 206040…
$ Und18P       <dbl> 20.2, 21.1, 27.7, 25.0, 23.9, 20.5, 24.2, 23.0, 19.7, 15.…
$ AmIndE       <dbl> 323, 2680, 310, 0, 337, 14, 896, 251, 75, 19, 151, 4, 10,…
$ AmIndP       <dbl> 0.4, 0.3, 0.5, 0.0, 0.8, 0.2, 0.4, 0.1, 0.6, 0.2, 0.2, 0.…
$ AsianE       <dbl> 1180, 81115, 1191, 96, 151, 53, 4351, 3687, 108, 47, 1679…
$ BlackE       <dbl> 23003, 460766, 28654, 8911, 246, 1991, 14159, 115506, 318…
$ BlackP       <dbl> 28.3, 43.1, 42.9, 41.1, 0.6, 23.0, 6.8, 56.1, 25.0, 71.4,…
$ HisE         <dbl> 4105, 86110, 8298, 777, 6287, 618, 59586, 11627, 1967, 55…
$ HisP         <dbl> 5.0, 8.1, 12.4, 3.6, 15.6, 7.1, 28.6, 5.6, 15.5, 0.6, 9.9…
$ OtherE       <dbl> 1128, 33908, 2140, 253, 1556, 349, 17409, 5171, 199, 36, …
$ OtherP       <dbl> 1.4, 3.2, 3.2, 1.2, 3.9, 4.0, 8.4, 2.5, 1.6, 0.4, 3.4, 1.…
$ PacIsE       <dbl> 30, 299, 295, 0, 136, 0, 15, 372, 0, 0, 34, 1, 0, 0, 37, …
$ PacIsP       <dbl> 0.0, 0.0, 0.4, 0.0, 0.3, 0.0, 0.0, 0.2, 0.0, 0.0, 0.0, 0.…
$ TwoRaceE     <dbl> 4625, 69252, 7686, 495, 2700, 401, 31113, 12304, 1918, 87…
$ TwoRaceP     <dbl> 5.7, 6.5, 11.5, 2.3, 6.7, 4.6, 14.9, 6.0, 15.1, 1.0, 6.8,…
$ WhiteE       <dbl> 51083, 420487, 26550, 11932, 35156, 5867, 140452, 68749, …
$ WhiteP       <dbl> 62.8, 39.4, 39.7, 55.0, 87.3, 67.6, 67.4, 33.4, 56.9, 26.…
$ DsmAs        <dbl> 0.42, 0.47, 0.38, 0.75, 0.55, 0.28, 0.47, 0.45, 0.66, 0.4…
$ DsmBlk       <dbl> 0.34, 0.71, 0.27, 0.34, 0.71, 0.45, 0.45, 0.45, 0.30, 0.2…
$ DsmHsp       <dbl> 0.30, 0.44, 0.28, 0.37, 0.35, 0.13, 0.48, 0.44, 0.46, 0.4…
$ IntrAsWht    <dbl> 0.65, 0.46, 0.34, 0.80, 0.80, 0.61, 0.57, 0.43, 0.50, 0.2…
$ IntrBlkWht   <dbl> 0.53, 0.15, 0.34, 0.46, 0.77, 0.54, 0.50, 0.24, 0.53, 0.2…
$ IntrHspWht   <dbl> 0.61, 0.38, 0.36, 0.56, 0.73, 0.66, 0.42, 0.32, 0.55, 0.2…
$ IsoAs        <dbl> 0.03, 0.23, 0.03, 0.03, 0.01, 0.01, 0.05, 0.05, 0.02, 0.0…
$ IsoBlk       <dbl> 0.37, 0.72, 0.46, 0.50, 0.03, 0.32, 0.11, 0.65, 0.35, 0.7…
$ IsoHsp       <dbl> 0.07, 0.18, 0.16, 0.07, 0.23, 0.07, 0.45, 0.11, 0.32, 0.0…
$ TotVetPop    <int> 64920, 841821, 40874, 16263, 30628, 6895, 157766, 151216,…
$ VetP         <dbl> 5.8, 4.7, 19.7, 8.1, 5.0, 5.6, 6.2, 11.5, 8.4, 5.8, 6.7, …

The county-level dataset contains 3,234 rows representing individual counties and 100 columns representing different variables.

The glimpse() function shows the structure of the dataset, including each variable’s type and a preview of its values.

6.7.2 Choose Variables to Explore

Next, we will identify a few variables to explore.

Open DSuite2023-data-dictionary.xlsx. Explore the variables that have a year listed under the county column. These are variables where data exists at the county-level.

For this tutorial, let’s focus on PovP, the number of individuals earning below the poverty income threshold as a percentage of the total population. Open Economic_Characteristics.md in the metadata folder to read about this variable.

Additionally, let’s look at FqhcAvTmDr, the average driving time (minutes) across tracts in state to nearest Federally Qualified Health Center (FQHC). Open Access_FQHCs.md in the metadata folder to read about this variable.

## Summary statistics for PovP
summary(county_data$PovP)
   Min. 1st Qu.  Median    Mean 3rd Qu.    Max.    NA's 
   1.70   10.00   13.30   14.26   17.50   52.80      99 

The county-level poverty percentage ranges from 1.7 to 52.8%. The median is 13.3% and the mean is 14.26%, suggesting that the distribution is slightly right-skewed, with few counties having high percentages. There are 99 counties with no percent poverty data.

## Summary statistics for FqhcAvTmDr
summary(county_data$FqhcAvTmDr)
   Min. 1st Qu.  Median    Mean 3rd Qu.    Max.    NA's 
   0.00    8.52   14.13   20.85   27.55   89.82     168 

The county-level average driving time to nearest FQHC ranges from 0 to 89.82 minutes. The median is 14.13 minutes and the mean is 20.85 minutes, which suggests that most counties are relatively close to an FQHC, but some are much further away. There are 168 counties with missing data, which corresponds to the worst access, where travel takes over 90 minutes in optimal conditions, or 180 minutes in normal conditions.

6.7.3 Filter the Dataset to a Single State

Next, we filter the county dataset to just counties in a single state using the HEROP_ID variable. In this tutorial, we will zoom in on North Carolina.

The HEROP_ID consists of four parts:

  • The 3-digit summary level code for counties: 050
  • The string: US
  • The 2-digit state FIPS code
  • The 3-digit county FIPS code

In the following code, we filter by the 6th and 7th positions of the HEROP_ID string which correspond to the 2-digit state FIPS code. North Carolina’s FIPS code is 37.

To check which FIPS code corresponds to which state or territory, visit this page from the U.S. Census Bureau.

## Filter to a single state
county_data_filtered <- county_data %>%
  filter(substr(HEROP_ID, 6, 7) %in% c("37"))

## View the first few values
head(county_data_filtered)
    HEROP_ID  FIPS SviSmryRnk SviTh1 SviTh2 SviTh3 SviTh4 EssnWrkP HghRskP
1 050US37083 37083     0.9870 0.9844 0.9007 0.9332 0.9360     57.3    25.9
2 050US37179 37179     0.2008 0.2221 0.3786 0.6701 0.0932     38.5    22.6
3 050US37129 37129     0.4289 0.5638 0.0824 0.5797 0.4989     40.5    15.5
4 050US37163 37163     0.9459 0.9300 0.8702 0.8721 0.8670     56.8    32.9
5 050US37031 37031     0.3888 0.4547 0.2558 0.4057 0.4432     46.8    16.7
6 050US37183 37183     0.3621 0.2526 0.3229 0.7989 0.4257     30.8    16.1
  HltCrP RetailP TotWrkE GiniCoeff MedInc  PciE PovP UnempP LngTermP MobileP
1   15.7    10.6   18108    0.4891  37991 26951 25.2    8.7     26.8    24.7
2   10.8    11.7  123262    0.4455  51241 45355  7.7    4.2     13.1     4.8
3   14.2    11.7  119132    0.4801  44816 46083 12.4    4.7     11.7     2.7
4   15.6    13.0   24859    0.4520  37968 26978 20.2    5.2     25.4    34.0
5   12.1    13.1   30470    0.4628  41628 42707 10.0    4.3     18.3    16.8
6   11.8     9.3  611492    0.4486  58520 52949  7.9    4.2     10.6     2.7
  RentalP TotUnits UnitDens VacantP FqhcAvTmDr FqhcCtTmDr FqhcTmDrP TotTracts
1    86.6    24849     34.3    18.4       4.29         16    100.00        16
2    51.5    86730    137.1     5.2      16.62         48    100.00        48
3    79.0   117236    609.7    12.7       9.00         54    100.00        54
4    78.5    25739     27.2    17.2       6.65         20    100.00        20
5    52.0    51615    101.7    39.6      31.04         20     64.52        31
6    79.2   481999    577.5     7.5       9.85        230    100.00       230
  OpRxRt BachelorsP EduHsP EduNoHsP EngProf GradSclP SomeCollegeP CrowdHsng
1   47.9       10.2   36.4     19.1   0.004      5.2         20.3      0.02
2   23.1       25.6   22.8      9.6   0.027     13.4         19.5      0.03
3   76.6       29.1   18.0      6.5   0.015     15.7         19.7      0.01
4   30.2       11.4   35.7     15.7   0.053      4.4         20.9      0.04
5   52.4       20.2   23.6      7.3   0.004     12.1         25.5      0.01
6   35.8       33.8   14.5      6.1   0.027     22.5         15.2      0.02
  DivrcdP FamSize HHSize HhldFA HhldFC HhldFS HhldMA HhldMC HhldMS HsdTot
1     8.6    2.95   2.32   19.1   10.0   11.8   13.5    1.1    5.3  20269
2     6.5    3.33   2.95    9.7    3.8    5.0    7.3    1.1    2.2  82231
3     7.9    2.85   2.20   19.7    3.8    9.3   13.8    0.7    3.8 102376
4     9.9    3.28   2.71   15.7    6.1    9.8   11.9    1.3    3.6  21321
5    10.4    2.69   2.16   17.0    2.7   10.0   13.2    0.7    5.4  31198
6     7.1    3.13   2.54   15.2    4.6    5.8   11.5    1.2    2.2 445636
  HsdTypCo HsdTypM HsdTypMC MrrdP NonRelFhhP NonRelNfhhP NvMrrdP OccupantP
1      6.2    38.8      9.5  44.4        6.8         1.1    40.7      81.6
2      4.5    65.7     30.8  61.1        6.4         1.9    29.5      94.8
3      8.1    43.4     15.0  49.4        3.9         4.0    38.3      87.3
4      6.1    48.2     14.8  48.1        9.6         2.6    35.4      82.8
5      5.1    52.8     14.8  58.8        4.3         1.6    25.5      60.4
6      6.7    51.0     23.6  54.3        4.9         3.3    35.5      92.5
  SepartedP TotPopHh WidwdP DisbP FemP MaleP MedAge  Ovr16 Ovr16P  Ovr18 Ovr18P
1       1.7    47017    4.5  20.9 51.4  48.6   43.8  39315   81.5  37968   78.7
2       1.4   242316    1.5   9.2 50.3  49.7   39.1 190283   77.7 181266   74.0
3       2.1   224956    2.3  11.8 52.3  47.7   40.1 193627   83.7 189042   81.8
4       2.0    57842    4.5  14.6 50.1  49.9   39.7  46618   78.7  44755   75.5
5       1.3    67436    4.0  17.6 51.1  48.9   50.1  58516   85.2  56945   82.9
6       1.4  1129867    1.7   9.0 51.0  49.0   37.2 914290   79.4 881746   76.6
   Ovr21 Ovr21P Ovr62P  Ovr65 Ovr65P SRatio SRatio18 SRatio65  TotPop Und18P
1  33485   75.6   26.3  10426   21.6   94.6     91.9     74.7   48219   21.3
2 150881   69.5   16.5  32401   13.2   98.7     96.4     81.5  244975   26.0
3 163335   76.8   22.4  43240   18.7   91.3     88.7     77.7  231214   18.2
4  38759   72.1   21.8  10636   18.0   99.7     97.7     77.7   59245   24.5
5  50406   80.2   32.9  18051   26.3   95.8     94.2     87.3   68652   17.1
6 758522   72.8   15.6 144054   12.5   96.0     93.6     78.0 1151009   23.4
  AmIndE AmIndP AsianE BlackE BlackP   HisE HisP OtherE OtherP PacIsE PacIsP
1   1424    3.0    446  25038   51.9   1558  3.2    776    1.6     24      0
2   1086    0.4  10002  27397   11.2  31683 12.9  14837    6.1     58      0
3    487    0.2   3204  26060   11.3  17675  7.6   8813    3.8     38      0
4   1288    2.2    289  14768   24.9  12600 21.3   8816   14.9      0      0
5    188    0.3    603   2998    4.4   3230  4.7   1427    2.1     15      0
6   3325    0.3  93193 221985   19.3 131104 11.4  55569    4.8    383      0
  TwoRaceE TwoRaceP WhiteE WhiteP DsmAs DsmBlk DsmHsp IntrAsWht IntrBlkWht
1     2574      5.3  17937   37.2  0.69   0.40   0.35      0.34       0.30
2    17450      7.1 174145   71.1  0.45   0.36   0.39      0.69       0.60
3    13627      5.9 178985   77.4  0.43   0.53   0.41      0.79       0.56
4     2545      4.3  31539   53.2  0.58   0.33   0.33      0.50       0.43
5     3871      5.6  59550   86.7  0.54   0.59   0.37      0.82       0.76
6    92986      8.1 683568   59.4  0.53   0.46   0.43      0.50       0.42
  IntrHspWht IsoAs IsoBlk IsoHsp TotVetPop VetP
1       0.43  0.04   0.60   0.07     37942  6.1
2       0.57  0.10   0.17   0.22    181172  6.4
3       0.66  0.03   0.29   0.15    188258  6.9
4       0.44  0.03   0.32   0.28     44570  7.1
5       0.81  0.03   0.13   0.10     56268 12.0
6       0.46  0.25   0.33   0.19    880615  5.7

6.7.4 Load the County Shapefile

Before we can map the data, we need to join the filtered dataset to a county shapefile. The shapefile is included in the OEPS data suite.

In your OEPS_Tutorial folder, open the data folder and locate county-2020-500k-shp.zip. If the file is still compressed, unzip it first. For best organization, create a folder named county-2020-500k-shp inside the data folder and extract the shapefile contents there.

Now, load the shapefile into R and display the geometry column to confirm that it imported correctly.

## Load in the county shapefile
county_geom <- st_read("data/county-2020-500k-shp/county-2020-500k.shp")
Reading layer `county-2020-500k' from data source 
  `/Users/maryniakolak/Code/opioid-environment-toolkit/data/county-2020-500k-shp/county-2020-500k.shp' 
  using driver `ESRI Shapefile'
Simple feature collection with 3234 features and 19 fields
Geometry type: MULTIPOLYGON
Dimension:     XY
Bounding box:  xmin: -179.1467 ymin: -14.5487 xmax: 179.7785 ymax: 71.38782
Geodetic CRS:  NAD83
## Create a quick plot
plot(county_geom$geometry)

6.7.5 Filter Shapefile to a Single State

Next, filter the shapefile to only include North Carolina counties.

## Filter to a single state
county_geom_filtered <- county_geom %>% filter(STATEFP %in% c("37"))

## Check plot
plot(county_geom_filtered$geometry)

6.7.6 Join the Data to the Shapefile

Now join the county_data_filtered dataframe to the filtered shapefile.

## Join county_data_filtered to county_geom_filtered
data_to_map_county <- county_geom_filtered %>%
  left_join(county_data_filtered, by = "HEROP_ID")

6.7.7 Create Maps

Finally, map PovP, the number of individuals earning below the poverty income threshold as a percentage of the total population and FqhcAvTmDr, the average driving time (minutes) across tracts in state to nearest Federally Qualified Health Center (FQHC).

## Map PovP
PovP <- tm_shape(data_to_map_county) +
  tm_polygons(
    fill = "PovP",
    fill.scale = tm_scale_intervals(style = "jenks"),
    fill.legend = tm_legend(title = ""),
    col = "grey85",
    lwd = 0.2) +
  tm_title("% of Individuals Earning Below Poverty Income Threshold")

## Map FqhcAvTmDr
FqhcAvTmDr <- tm_shape(data_to_map_county) +
  tm_polygons(
    fill = "FqhcAvTmDr",
    fill.scale = tm_scale_intervals(style = "jenks"),
    fill.legend = tm_legend(title = ""),
    col = "grey85",
    lwd = 0.2) +
  tm_title("Driving time (minutes) to nearest FQHC") 

## Arrange the maps
tmap_arrange(PovP, FqhcAvTmDr)

The PovP map shows us that poverty is unevenly distributed across the state. Poverty percentages appear higher in several Eastern and Southern counties, whereas many counties in the central Piedmont have lower poverty rates.

For the FqhcAvTmDr map, most counties appear to have relatively short driving times, especially in the central part of the state. Longer travel times seem more common in some coastal and mountain areas, where access may be more limited. There are no counties that exceed a driving time of 58 minutes.

6.8 Join OEPS data to outside data

Sometimes users will already have their own data and want to add OEPS variables for context. In this section, we will demonstrate how to join OEPS county-level data to an outside dataset using a shared geographic identifier.

The NC_cty_foodstamps_2025.csv contains information about the county-level prevalence of food stamps in North Carolina for 2025, which was originally pulled from CDC’s PLACES dataset.

## Read in non-OEPS data
foodstamps_data <- read.csv("data/NC_cty_foodstamps_2025.csv")

## Glimpse foodstamps_data
glimpse(foodstamps_data)
Rows: 100
Columns: 11
$ StateAbbr           <chr> "NC", "NC", "NC", "NC", "NC", "NC", "NC", "NC", "N…
$ StateDesc           <chr> "North Carolina", "North Carolina", "North Carolin…
$ CountyName          <chr> "Bertie", "Vance", "New Hanover", "Martin", "Edgec…
$ StateFIPS           <int> 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37…
$ CountyFIPS          <int> 15, 181, 129, 117, 65, 95, 159, 55, 193, 125, 45, …
$ TotalPopulation     <chr> "16,922", "42,301", "238,852", "21,447", "48,832",…
$ TotalPop18plus      <chr> "14,112", "32,113", "197,058", "17,125", "37,653",…
$ FOODSTAMP_CrudePrev <dbl> 24.9, 22.9, 9.0, 19.9, 24.8, 19.9, 14.1, 6.4, 13.6…
$ FOODSTAMP_Crude95ci <chr> "(20.8, 29.1)", "(18.8, 26.9)", "( 6.8, 11.6)", "(…
$ FOODSTAMP_AdjPrev   <dbl> 26.7, 24.2, 9.6, 21.6, 26.3, 20.8, 14.8, 7.0, 14.7…
$ FOODSTAMP_Adj95CI   <chr> "(22.1, 31.3)", "(19.9, 28.7)", "( 7.3, 12.5)", "(…

Often times, datasets will have a FIPS or GEOID column that can be used for joining to other datasets. Sometimes, these columns will not be the same format, in which case, joining will not work. Thus, you will need to re-format the columns before joining.

In foodstamps_data, we have a column called StateFIPS which stores the 2-digit state FIPS code and a CountyFIPS column which stores the 3-digit county FIPS code. However, CSV files often delete leading zeros, which is why you see some 1- or 2-digit values in CountyFIPS.

Let’s check again to see what identifier columns are in county_data_filtered.

## Glimpse county_data_filtered
glimpse(county_data_filtered)
Rows: 100
Columns: 100
$ HEROP_ID     <chr> "050US37083", "050US37179", "050US37129", "050US37163", "…
$ FIPS         <int> 37083, 37179, 37129, 37163, 37031, 37183, 37119, 37087, 3…
$ SviSmryRnk   <dbl> 0.9870, 0.2008, 0.4289, 0.9459, 0.3888, 0.3621, 0.6115, 0…
$ SviTh1       <dbl> 0.9844, 0.2221, 0.5638, 0.9300, 0.4547, 0.2526, 0.5787, 0…
$ SviTh2       <dbl> 0.9007, 0.3786, 0.0824, 0.8702, 0.2558, 0.3229, 0.4747, 0…
$ SviTh3       <dbl> 0.9332, 0.6701, 0.5797, 0.8721, 0.4057, 0.7989, 0.8915, 0…
$ SviTh4       <dbl> 0.9360, 0.0932, 0.4989, 0.8670, 0.4432, 0.4257, 0.5135, 0…
$ EssnWrkP     <dbl> 57.3, 38.5, 40.5, 56.8, 46.8, 30.8, 35.7, 49.7, 50.1, 46.…
$ HghRskP      <dbl> 25.9, 22.6, 15.5, 32.9, 16.7, 16.1, 16.3, 21.2, 24.6, 19.…
$ HltCrP       <dbl> 15.7, 10.8, 14.2, 15.6, 12.1, 11.8, 11.5, 15.1, 12.3, 16.…
$ RetailP      <dbl> 10.6, 11.7, 11.7, 13.0, 13.1, 9.3, 10.3, 11.1, 12.6, 11.7…
$ TotWrkE      <int> 18108, 123262, 119132, 24859, 30470, 611492, 616325, 2827…
$ GiniCoeff    <dbl> 0.4891, 0.4455, 0.4801, 0.4520, 0.4628, 0.4486, 0.4935, 0…
$ MedInc       <int> 37991, 51241, 44816, 37968, 41628, 58520, 51889, 40926, 4…
$ PciE         <int> 26951, 45355, 46083, 26978, 42707, 52949, 51490, 35580, 3…
$ PovP         <dbl> 25.2, 7.7, 12.4, 20.2, 10.0, 7.9, 10.4, 11.3, 10.7, 11.8,…
$ UnempP       <dbl> 8.7, 4.2, 4.7, 5.2, 4.3, 4.2, 4.2, 3.4, 6.0, 3.5, 4.7, 5.…
$ LngTermP     <dbl> 26.8, 13.1, 11.7, 25.4, 18.3, 10.6, 10.1, 23.0, 13.3, 17.…
$ MobileP      <dbl> 24.7, 4.8, 2.7, 34.0, 16.8, 2.7, 1.5, 16.7, 22.8, 12.1, 2…
$ RentalP      <dbl> 86.6, 51.5, 79.0, 78.5, 52.0, 79.2, 98.6, 51.9, 46.6, 84.…
$ TotUnits     <int> 24849, 86730, 117236, 25739, 51615, 481999, 491721, 35307…
$ UnitDens     <dbl> 34.3, 137.1, 609.7, 27.2, 101.7, 577.5, 939.1, 63.8, 35.6…
$ VacantP      <dbl> 18.4, 5.2, 12.7, 17.2, 39.6, 7.5, 7.4, 24.2, 22.1, 22.3, …
$ FqhcAvTmDr   <dbl> 4.29, 16.62, 9.00, 6.65, 31.04, 9.85, 11.08, 13.09, 19.40…
$ FqhcCtTmDr   <int> 16, 48, 54, 20, 20, 230, 305, 20, 12, 64, 15, 11, 22, 12,…
$ FqhcTmDrP    <dbl> 100.00, 100.00, 100.00, 100.00, 64.52, 100.00, 100.00, 95…
$ TotTracts    <int> 16, 48, 54, 20, 31, 230, 305, 21, 15, 65, 15, 12, 23, 12,…
$ OpRxRt       <dbl> 47.9, 23.1, 76.6, 30.2, 52.4, 35.8, 45.0, 52.7, 16.7, 73.…
$ BachelorsP   <dbl> 10.2, 25.6, 29.1, 11.4, 20.2, 33.8, 31.3, 17.9, 19.7, 26.…
$ EduHsP       <dbl> 36.4, 22.8, 18.0, 35.7, 23.6, 14.5, 16.4, 27.5, 27.7, 22.…
$ EduNoHsP     <dbl> 19.1, 9.6, 6.5, 15.7, 7.3, 6.1, 9.2, 8.4, 9.5, 7.7, 13.5,…
$ EngProf      <dbl> 0.004, 0.027, 0.015, 0.053, 0.004, 0.027, 0.055, 0.006, 0…
$ GradSclP     <dbl> 5.2, 13.4, 15.7, 4.4, 12.1, 22.5, 17.3, 11.5, 10.1, 18.2,…
$ SomeCollegeP <dbl> 20.3, 19.5, 19.7, 20.9, 25.5, 15.2, 17.6, 21.8, 21.7, 16.…
$ CrowdHsng    <dbl> 0.02, 0.03, 0.01, 0.04, 0.01, 0.02, 0.02, 0.01, 0.02, 0.0…
$ DivrcdP      <dbl> 8.6, 6.5, 7.9, 9.9, 10.4, 7.1, 7.7, 10.6, 11.7, 10.7, 12.…
$ FamSize      <dbl> 2.95, 3.33, 2.85, 3.28, 2.69, 3.13, 3.19, 2.85, 3.14, 3.2…
$ HHSize       <dbl> 2.32, 2.95, 2.20, 2.71, 2.16, 2.54, 2.45, 2.31, 2.58, 2.5…
$ HhldFA       <dbl> 19.1, 9.7, 19.7, 15.7, 17.0, 15.2, 18.7, 18.0, 14.6, 17.4…
$ HhldFC       <dbl> 10.0, 3.8, 3.8, 6.1, 2.7, 4.6, 6.0, 4.0, 3.9, 3.9, 6.3, 2…
$ HhldFS       <dbl> 11.8, 5.0, 9.3, 9.8, 10.0, 5.8, 5.7, 10.9, 8.6, 9.3, 11.9…
$ HhldMA       <dbl> 13.5, 7.3, 13.8, 11.9, 13.2, 11.5, 14.4, 12.6, 12.9, 13.5…
$ HhldMC       <dbl> 1.1, 1.1, 0.7, 1.3, 0.7, 1.2, 1.4, 0.7, 1.2, 1.3, 1.2, 0.…
$ HhldMS       <dbl> 5.3, 2.2, 3.8, 3.6, 5.4, 2.2, 2.2, 5.8, 4.6, 4.4, 5.3, 6.…
$ HsdTot       <dbl> 20269, 82231, 102376, 21321, 31198, 445636, 455494, 26772…
$ HsdTypCo     <dbl> 6.2, 4.5, 8.1, 6.1, 5.1, 6.7, 7.3, 6.2, 7.5, 6.4, 4.2, 3.…
$ HsdTypM      <dbl> 38.8, 65.7, 43.4, 48.2, 52.8, 51.0, 41.5, 49.9, 52.3, 46.…
$ HsdTypMC     <dbl> 9.5, 30.8, 15.0, 14.8, 14.8, 23.6, 17.9, 12.4, 20.3, 14.9…
$ MrrdP        <dbl> 44.4, 61.1, 49.4, 48.1, 58.8, 54.3, 47.4, 55.3, 52.3, 46.…
$ NonRelFhhP   <dbl> 6.8, 6.4, 3.9, 9.6, 4.3, 4.9, 6.2, 6.9, 6.4, 7.4, 8.3, 8.…
$ NonRelNfhhP  <dbl> 1.1, 1.9, 4.0, 2.6, 1.6, 3.3, 3.7, 2.6, 2.6, 7.0, 2.7, 3.…
$ NvMrrdP      <dbl> 40.7, 29.5, 38.3, 35.4, 25.5, 35.5, 41.8, 28.9, 29.9, 38.…
$ OccupantP    <dbl> 81.6, 94.8, 87.3, 82.8, 60.4, 92.5, 92.6, 75.8, 77.9, 77.…
$ SepartedP    <dbl> 1.7, 1.4, 2.1, 2.0, 1.3, 1.4, 1.6, 1.6, 2.3, 1.2, 2.6, 1.…
$ TotPopHh     <int> 47017, 242316, 224956, 57842, 67436, 1129867, 1115618, 61…
$ WidwdP       <dbl> 4.5, 1.5, 2.3, 4.5, 4.0, 1.7, 1.6, 3.6, 3.8, 3.1, 4.5, 5.…
$ DisbP        <dbl> 20.9, 9.2, 11.8, 14.6, 17.6, 9.0, 8.3, 17.3, 15.3, 13.8, …
$ FemP         <dbl> 51.4, 50.3, 52.3, 50.1, 51.1, 51.0, 51.7, 51.2, 49.4, 51.…
$ MaleP        <dbl> 48.6, 49.7, 47.7, 49.9, 48.9, 49.0, 48.3, 48.8, 50.6, 48.…
$ MedAge       <dbl> 43.8, 39.1, 40.1, 39.7, 50.1, 37.2, 35.4, 47.8, 41.8, 42.…
$ Ovr16        <dbl> 39315, 190283, 193627, 46618, 58516, 914290, 898405, 5279…
$ Ovr16P       <dbl> 81.5, 77.7, 83.7, 78.7, 85.2, 79.4, 79.4, 84.6, 80.4, 84.…
$ Ovr18        <dbl> 37968, 181266, 189042, 44755, 56945, 881746, 870043, 5129…
$ Ovr18P       <dbl> 78.7, 74.0, 81.8, 75.5, 82.9, 76.6, 76.9, 82.2, 77.5, 81.…
$ Ovr21        <dbl> 33485, 150881, 163335, 38759, 50406, 758522, 755280, 4549…
$ Ovr21P       <dbl> 75.6, 69.5, 76.8, 72.1, 80.2, 72.8, 73.3, 79.7, 73.8, 78.…
$ Ovr62P       <dbl> 26.3, 16.5, 22.4, 21.8, 32.9, 15.6, 14.7, 30.6, 22.4, 24.…
$ Ovr65        <dbl> 10426, 32401, 43240, 10636, 18051, 144054, 132281, 15827,…
$ Ovr65P       <dbl> 21.6, 13.2, 18.7, 18.0, 26.3, 12.5, 11.7, 25.4, 18.0, 20.…
$ SRatio       <dbl> 94.6, 98.7, 91.3, 99.7, 95.8, 96.0, 93.5, 95.3, 102.5, 93…
$ SRatio18     <dbl> 91.9, 96.4, 88.7, 97.7, 94.2, 93.6, 90.7, 91.9, 100.0, 91…
$ SRatio65     <dbl> 74.7, 81.5, 77.7, 77.7, 87.3, 78.0, 74.2, 83.7, 87.4, 78.…
$ TotPop       <int> 48219, 244975, 231214, 59245, 68652, 1151009, 1130906, 62…
$ Und18P       <dbl> 21.3, 26.0, 18.2, 24.5, 17.1, 23.4, 23.1, 17.8, 22.5, 18.…
$ AmIndE       <dbl> 1424, 1086, 487, 1288, 188, 3325, 4355, 361, 209, 565, 15…
$ AmIndP       <dbl> 3.0, 0.4, 0.2, 2.2, 0.3, 0.3, 0.4, 0.6, 0.3, 0.2, 3.1, 1.…
$ AsianE       <dbl> 446, 10002, 3204, 289, 603, 93193, 69579, 255, 437, 3144,…
$ BlackE       <dbl> 25038, 27397, 26060, 14768, 2998, 221985, 344569, 624, 78…
$ BlackP       <dbl> 51.9, 11.2, 11.3, 24.9, 4.4, 19.3, 30.5, 1.0, 12.4, 5.4, …
$ HisE         <dbl> 1558, 31683, 17675, 12600, 3230, 131104, 174580, 2953, 54…
$ HisP         <dbl> 3.2, 12.9, 7.6, 21.3, 4.7, 11.4, 15.4, 4.7, 8.5, 8.3, 5.5…
$ OtherE       <dbl> 776, 14837, 8813, 8816, 1427, 55569, 88166, 568, 3429, 58…
$ OtherP       <dbl> 1.6, 6.1, 3.8, 14.9, 2.1, 4.8, 7.8, 0.9, 5.4, 2.1, 3.1, 0…
$ PacIsE       <dbl> 24, 58, 38, 0, 15, 383, 434, 0, 41, 95, 22, 5, 39, 7, 0, …
$ PacIsP       <dbl> 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.1, 0.0, 0.0, 0.…
$ TwoRaceE     <dbl> 2574, 17450, 13627, 2545, 3871, 92986, 97368, 3271, 3432,…
$ TwoRaceP     <dbl> 5.3, 7.1, 5.9, 4.3, 5.6, 8.1, 8.6, 5.2, 5.4, 7.9, 4.5, 4.…
$ WhiteE       <dbl> 17937, 174145, 178985, 31539, 59550, 683568, 526435, 5735…
$ WhiteP       <dbl> 37.2, 71.1, 77.4, 53.2, 86.7, 59.4, 46.5, 91.9, 75.7, 83.…
$ DsmAs        <dbl> 0.69, 0.45, 0.43, 0.58, 0.54, 0.53, 0.45, 0.53, 0.45, 0.4…
$ DsmBlk       <dbl> 0.40, 0.36, 0.53, 0.33, 0.59, 0.46, 0.56, 0.56, 0.43, 0.5…
$ DsmHsp       <dbl> 0.35, 0.39, 0.41, 0.33, 0.37, 0.43, 0.54, 0.33, 0.43, 0.4…
$ IntrAsWht    <dbl> 0.34, 0.69, 0.79, 0.50, 0.82, 0.50, 0.42, 0.90, 0.79, 0.8…
$ IntrBlkWht   <dbl> 0.30, 0.60, 0.56, 0.43, 0.76, 0.42, 0.27, 0.88, 0.63, 0.7…
$ IntrHspWht   <dbl> 0.43, 0.57, 0.66, 0.44, 0.81, 0.46, 0.30, 0.88, 0.66, 0.7…
$ IsoAs        <dbl> 0.04, 0.10, 0.03, 0.03, 0.03, 0.25, 0.12, 0.01, 0.01, 0.0…
$ IsoBlk       <dbl> 0.60, 0.17, 0.29, 0.32, 0.13, 0.33, 0.46, 0.03, 0.22, 0.1…
$ IsoHsp       <dbl> 0.07, 0.22, 0.15, 0.28, 0.10, 0.19, 0.27, 0.08, 0.14, 0.1…
$ TotVetPop    <int> 37942, 181172, 188258, 44570, 56268, 880615, 868562, 5123…
$ VetP         <dbl> 6.1, 6.4, 6.9, 7.1, 12.0, 5.7, 5.3, 9.7, 9.7, 6.7, 6.7, 1…

In county_data_filtered, we have HEROP_ID, our unique OEPS identifier and FIPS which stores the 5-digit combined state and county FIPS codes.

It is obvious that the two datasets we want to join do not have a shared geographic identifier. But, we have all the information we need to create one!

Let’s keep the OEPS data as is, and instead alter foodstamps_data by creating a 5-digit FIPS column. To do so, we need to combine StateFIPS with CountyFIPS ensuring that CountyFIPS does not leave out leading zeros.

## Create a FIPS column
foodstamps_data <- foodstamps_data %>%
  mutate(FIPS = as.numeric(sprintf("%02d%03d", StateFIPS, CountyFIPS)))

## View the first few values
head(foodstamps_data)
  StateAbbr      StateDesc  CountyName StateFIPS CountyFIPS TotalPopulation
1        NC North Carolina      Bertie        37         15          16,922
2        NC North Carolina       Vance        37        181          42,301
3        NC North Carolina New Hanover        37        129         238,852
4        NC North Carolina      Martin        37        117          21,447
5        NC North Carolina   Edgecombe        37         65          48,832
6        NC North Carolina        Hyde        37         95           4,607
  TotalPop18plus FOODSTAMP_CrudePrev FOODSTAMP_Crude95ci FOODSTAMP_AdjPrev
1         14,112                24.9        (20.8, 29.1)              26.7
2         32,113                22.9        (18.8, 26.9)              24.2
3        197,058                 9.0        ( 6.8, 11.6)               9.6
4         17,125                19.9        (16.4, 23.7)              21.6
5         37,653                24.8        (20.5, 29.3)              26.3
6          3,922                19.9        (16.6, 23.3)              20.8
  FOODSTAMP_Adj95CI  FIPS
1      (22.1, 31.3) 37015
2      (19.9, 28.7) 37181
3      ( 7.3, 12.5) 37129
4      (17.9, 25.8) 37117
5      (21.8, 31.1) 37065
6      (17.4, 24.7) 37095

Now we have a shared geographic identifier, FIPS, and can join the two datasets.

## Join foodstamps_data to county_data_filtered
final_county_data <- county_data_filtered %>%
  left_join(foodstamps_data, by = "FIPS")

## Glimpse final_county_data
glimpse(final_county_data)
Rows: 100
Columns: 111
$ HEROP_ID            <chr> "050US37083", "050US37179", "050US37129", "050US37…
$ FIPS                <dbl> 37083, 37179, 37129, 37163, 37031, 37183, 37119, 3…
$ SviSmryRnk          <dbl> 0.9870, 0.2008, 0.4289, 0.9459, 0.3888, 0.3621, 0.…
$ SviTh1              <dbl> 0.9844, 0.2221, 0.5638, 0.9300, 0.4547, 0.2526, 0.…
$ SviTh2              <dbl> 0.9007, 0.3786, 0.0824, 0.8702, 0.2558, 0.3229, 0.…
$ SviTh3              <dbl> 0.9332, 0.6701, 0.5797, 0.8721, 0.4057, 0.7989, 0.…
$ SviTh4              <dbl> 0.9360, 0.0932, 0.4989, 0.8670, 0.4432, 0.4257, 0.…
$ EssnWrkP            <dbl> 57.3, 38.5, 40.5, 56.8, 46.8, 30.8, 35.7, 49.7, 50…
$ HghRskP             <dbl> 25.9, 22.6, 15.5, 32.9, 16.7, 16.1, 16.3, 21.2, 24…
$ HltCrP              <dbl> 15.7, 10.8, 14.2, 15.6, 12.1, 11.8, 11.5, 15.1, 12…
$ RetailP             <dbl> 10.6, 11.7, 11.7, 13.0, 13.1, 9.3, 10.3, 11.1, 12.…
$ TotWrkE             <int> 18108, 123262, 119132, 24859, 30470, 611492, 61632…
$ GiniCoeff           <dbl> 0.4891, 0.4455, 0.4801, 0.4520, 0.4628, 0.4486, 0.…
$ MedInc              <int> 37991, 51241, 44816, 37968, 41628, 58520, 51889, 4…
$ PciE                <int> 26951, 45355, 46083, 26978, 42707, 52949, 51490, 3…
$ PovP                <dbl> 25.2, 7.7, 12.4, 20.2, 10.0, 7.9, 10.4, 11.3, 10.7…
$ UnempP              <dbl> 8.7, 4.2, 4.7, 5.2, 4.3, 4.2, 4.2, 3.4, 6.0, 3.5, …
$ LngTermP            <dbl> 26.8, 13.1, 11.7, 25.4, 18.3, 10.6, 10.1, 23.0, 13…
$ MobileP             <dbl> 24.7, 4.8, 2.7, 34.0, 16.8, 2.7, 1.5, 16.7, 22.8, …
$ RentalP             <dbl> 86.6, 51.5, 79.0, 78.5, 52.0, 79.2, 98.6, 51.9, 46…
$ TotUnits            <int> 24849, 86730, 117236, 25739, 51615, 481999, 491721…
$ UnitDens            <dbl> 34.3, 137.1, 609.7, 27.2, 101.7, 577.5, 939.1, 63.…
$ VacantP             <dbl> 18.4, 5.2, 12.7, 17.2, 39.6, 7.5, 7.4, 24.2, 22.1,…
$ FqhcAvTmDr          <dbl> 4.29, 16.62, 9.00, 6.65, 31.04, 9.85, 11.08, 13.09…
$ FqhcCtTmDr          <int> 16, 48, 54, 20, 20, 230, 305, 20, 12, 64, 15, 11, …
$ FqhcTmDrP           <dbl> 100.00, 100.00, 100.00, 100.00, 64.52, 100.00, 100…
$ TotTracts           <int> 16, 48, 54, 20, 31, 230, 305, 21, 15, 65, 15, 12, …
$ OpRxRt              <dbl> 47.9, 23.1, 76.6, 30.2, 52.4, 35.8, 45.0, 52.7, 16…
$ BachelorsP          <dbl> 10.2, 25.6, 29.1, 11.4, 20.2, 33.8, 31.3, 17.9, 19…
$ EduHsP              <dbl> 36.4, 22.8, 18.0, 35.7, 23.6, 14.5, 16.4, 27.5, 27…
$ EduNoHsP            <dbl> 19.1, 9.6, 6.5, 15.7, 7.3, 6.1, 9.2, 8.4, 9.5, 7.7…
$ EngProf             <dbl> 0.004, 0.027, 0.015, 0.053, 0.004, 0.027, 0.055, 0…
$ GradSclP            <dbl> 5.2, 13.4, 15.7, 4.4, 12.1, 22.5, 17.3, 11.5, 10.1…
$ SomeCollegeP        <dbl> 20.3, 19.5, 19.7, 20.9, 25.5, 15.2, 17.6, 21.8, 21…
$ CrowdHsng           <dbl> 0.02, 0.03, 0.01, 0.04, 0.01, 0.02, 0.02, 0.01, 0.…
$ DivrcdP             <dbl> 8.6, 6.5, 7.9, 9.9, 10.4, 7.1, 7.7, 10.6, 11.7, 10…
$ FamSize             <dbl> 2.95, 3.33, 2.85, 3.28, 2.69, 3.13, 3.19, 2.85, 3.…
$ HHSize              <dbl> 2.32, 2.95, 2.20, 2.71, 2.16, 2.54, 2.45, 2.31, 2.…
$ HhldFA              <dbl> 19.1, 9.7, 19.7, 15.7, 17.0, 15.2, 18.7, 18.0, 14.…
$ HhldFC              <dbl> 10.0, 3.8, 3.8, 6.1, 2.7, 4.6, 6.0, 4.0, 3.9, 3.9,…
$ HhldFS              <dbl> 11.8, 5.0, 9.3, 9.8, 10.0, 5.8, 5.7, 10.9, 8.6, 9.…
$ HhldMA              <dbl> 13.5, 7.3, 13.8, 11.9, 13.2, 11.5, 14.4, 12.6, 12.…
$ HhldMC              <dbl> 1.1, 1.1, 0.7, 1.3, 0.7, 1.2, 1.4, 0.7, 1.2, 1.3, …
$ HhldMS              <dbl> 5.3, 2.2, 3.8, 3.6, 5.4, 2.2, 2.2, 5.8, 4.6, 4.4, …
$ HsdTot              <dbl> 20269, 82231, 102376, 21321, 31198, 445636, 455494…
$ HsdTypCo            <dbl> 6.2, 4.5, 8.1, 6.1, 5.1, 6.7, 7.3, 6.2, 7.5, 6.4, …
$ HsdTypM             <dbl> 38.8, 65.7, 43.4, 48.2, 52.8, 51.0, 41.5, 49.9, 52…
$ HsdTypMC            <dbl> 9.5, 30.8, 15.0, 14.8, 14.8, 23.6, 17.9, 12.4, 20.…
$ MrrdP               <dbl> 44.4, 61.1, 49.4, 48.1, 58.8, 54.3, 47.4, 55.3, 52…
$ NonRelFhhP          <dbl> 6.8, 6.4, 3.9, 9.6, 4.3, 4.9, 6.2, 6.9, 6.4, 7.4, …
$ NonRelNfhhP         <dbl> 1.1, 1.9, 4.0, 2.6, 1.6, 3.3, 3.7, 2.6, 2.6, 7.0, …
$ NvMrrdP             <dbl> 40.7, 29.5, 38.3, 35.4, 25.5, 35.5, 41.8, 28.9, 29…
$ OccupantP           <dbl> 81.6, 94.8, 87.3, 82.8, 60.4, 92.5, 92.6, 75.8, 77…
$ SepartedP           <dbl> 1.7, 1.4, 2.1, 2.0, 1.3, 1.4, 1.6, 1.6, 2.3, 1.2, …
$ TotPopHh            <int> 47017, 242316, 224956, 57842, 67436, 1129867, 1115…
$ WidwdP              <dbl> 4.5, 1.5, 2.3, 4.5, 4.0, 1.7, 1.6, 3.6, 3.8, 3.1, …
$ DisbP               <dbl> 20.9, 9.2, 11.8, 14.6, 17.6, 9.0, 8.3, 17.3, 15.3,…
$ FemP                <dbl> 51.4, 50.3, 52.3, 50.1, 51.1, 51.0, 51.7, 51.2, 49…
$ MaleP               <dbl> 48.6, 49.7, 47.7, 49.9, 48.9, 49.0, 48.3, 48.8, 50…
$ MedAge              <dbl> 43.8, 39.1, 40.1, 39.7, 50.1, 37.2, 35.4, 47.8, 41…
$ Ovr16               <dbl> 39315, 190283, 193627, 46618, 58516, 914290, 89840…
$ Ovr16P              <dbl> 81.5, 77.7, 83.7, 78.7, 85.2, 79.4, 79.4, 84.6, 80…
$ Ovr18               <dbl> 37968, 181266, 189042, 44755, 56945, 881746, 87004…
$ Ovr18P              <dbl> 78.7, 74.0, 81.8, 75.5, 82.9, 76.6, 76.9, 82.2, 77…
$ Ovr21               <dbl> 33485, 150881, 163335, 38759, 50406, 758522, 75528…
$ Ovr21P              <dbl> 75.6, 69.5, 76.8, 72.1, 80.2, 72.8, 73.3, 79.7, 73…
$ Ovr62P              <dbl> 26.3, 16.5, 22.4, 21.8, 32.9, 15.6, 14.7, 30.6, 22…
$ Ovr65               <dbl> 10426, 32401, 43240, 10636, 18051, 144054, 132281,…
$ Ovr65P              <dbl> 21.6, 13.2, 18.7, 18.0, 26.3, 12.5, 11.7, 25.4, 18…
$ SRatio              <dbl> 94.6, 98.7, 91.3, 99.7, 95.8, 96.0, 93.5, 95.3, 10…
$ SRatio18            <dbl> 91.9, 96.4, 88.7, 97.7, 94.2, 93.6, 90.7, 91.9, 10…
$ SRatio65            <dbl> 74.7, 81.5, 77.7, 77.7, 87.3, 78.0, 74.2, 83.7, 87…
$ TotPop              <int> 48219, 244975, 231214, 59245, 68652, 1151009, 1130…
$ Und18P              <dbl> 21.3, 26.0, 18.2, 24.5, 17.1, 23.4, 23.1, 17.8, 22…
$ AmIndE              <dbl> 1424, 1086, 487, 1288, 188, 3325, 4355, 361, 209, …
$ AmIndP              <dbl> 3.0, 0.4, 0.2, 2.2, 0.3, 0.3, 0.4, 0.6, 0.3, 0.2, …
$ AsianE              <dbl> 446, 10002, 3204, 289, 603, 93193, 69579, 255, 437…
$ BlackE              <dbl> 25038, 27397, 26060, 14768, 2998, 221985, 344569, …
$ BlackP              <dbl> 51.9, 11.2, 11.3, 24.9, 4.4, 19.3, 30.5, 1.0, 12.4…
$ HisE                <dbl> 1558, 31683, 17675, 12600, 3230, 131104, 174580, 2…
$ HisP                <dbl> 3.2, 12.9, 7.6, 21.3, 4.7, 11.4, 15.4, 4.7, 8.5, 8…
$ OtherE              <dbl> 776, 14837, 8813, 8816, 1427, 55569, 88166, 568, 3…
$ OtherP              <dbl> 1.6, 6.1, 3.8, 14.9, 2.1, 4.8, 7.8, 0.9, 5.4, 2.1,…
$ PacIsE              <dbl> 24, 58, 38, 0, 15, 383, 434, 0, 41, 95, 22, 5, 39,…
$ PacIsP              <dbl> 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.1, 0.0, …
$ TwoRaceE            <dbl> 2574, 17450, 13627, 2545, 3871, 92986, 97368, 3271…
$ TwoRaceP            <dbl> 5.3, 7.1, 5.9, 4.3, 5.6, 8.1, 8.6, 5.2, 5.4, 7.9, …
$ WhiteE              <dbl> 17937, 174145, 178985, 31539, 59550, 683568, 52643…
$ WhiteP              <dbl> 37.2, 71.1, 77.4, 53.2, 86.7, 59.4, 46.5, 91.9, 75…
$ DsmAs               <dbl> 0.69, 0.45, 0.43, 0.58, 0.54, 0.53, 0.45, 0.53, 0.…
$ DsmBlk              <dbl> 0.40, 0.36, 0.53, 0.33, 0.59, 0.46, 0.56, 0.56, 0.…
$ DsmHsp              <dbl> 0.35, 0.39, 0.41, 0.33, 0.37, 0.43, 0.54, 0.33, 0.…
$ IntrAsWht           <dbl> 0.34, 0.69, 0.79, 0.50, 0.82, 0.50, 0.42, 0.90, 0.…
$ IntrBlkWht          <dbl> 0.30, 0.60, 0.56, 0.43, 0.76, 0.42, 0.27, 0.88, 0.…
$ IntrHspWht          <dbl> 0.43, 0.57, 0.66, 0.44, 0.81, 0.46, 0.30, 0.88, 0.…
$ IsoAs               <dbl> 0.04, 0.10, 0.03, 0.03, 0.03, 0.25, 0.12, 0.01, 0.…
$ IsoBlk              <dbl> 0.60, 0.17, 0.29, 0.32, 0.13, 0.33, 0.46, 0.03, 0.…
$ IsoHsp              <dbl> 0.07, 0.22, 0.15, 0.28, 0.10, 0.19, 0.27, 0.08, 0.…
$ TotVetPop           <int> 37942, 181172, 188258, 44570, 56268, 880615, 86856…
$ VetP                <dbl> 6.1, 6.4, 6.9, 7.1, 12.0, 5.7, 5.3, 9.7, 9.7, 6.7,…
$ StateAbbr           <chr> "NC", "NC", "NC", "NC", "NC", "NC", "NC", "NC", "N…
$ StateDesc           <chr> "North Carolina", "North Carolina", "North Carolin…
$ CountyName          <chr> "Halifax", "Union", "New Hanover", "Sampson", "Car…
$ StateFIPS           <int> 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37…
$ CountyFIPS          <int> 83, 179, 129, 163, 31, 183, 119, 87, 141, 21, 47, …
$ TotalPopulation     <chr> "47,298", "256,452", "238,852", "59,601", "69,615"…
$ TotalPop18plus      <chr> "37,177", "191,719", "197,058", "44,966", "58,070"…
$ FOODSTAMP_CrudePrev <dbl> 27.8, 8.5, 9.0, 19.0, 7.8, 7.4, 10.8, 9.3, 10.3, 9…
$ FOODSTAMP_Crude95ci <chr> "(23.8, 31.9)", "( 6.6, 10.6)", "( 6.8, 11.6)", "(…
$ FOODSTAMP_AdjPrev   <dbl> 29.4, 8.7, 9.6, 19.8, 8.7, 7.4, 10.7, 10.2, 10.5, …
$ FOODSTAMP_Adj95CI   <chr> "(25.1, 34.0)", "( 6.9, 11.0)", "( 7.3, 12.5)", "(…

From here, you could use this dataset for further analysis and visualization.

6.9 Save Outputs

Finally, let’s save our outputs. We will save the final joined dataset as a CSV, and the county-level maps as images.

## Save final_county_data
write.csv(final_county_data, "final_county_data.csv")

## Save county-level maps
tmap_save(PovP, filename = "PovP_map.png")
tmap_save(FqhcAvTmDr, filename = "FqhcAvTmDr_map.png")

6.10 More Resources

To learn more about the OEPS ecosystem visit: https://oeps.healthyregions.org/

Visit the Geospatial Consortium & Community of Practice website to learn about past and upcoming workshops: https://gccp.healthyregions.org/

For more on the tmap package visit: https://r-tmap.github.io/tmap/reference/