6 Opioid Environment Data
6.1 Overview
The Opioid Environment Policy Scan Data Ecosystem, otherwise known as OEPS, is a data resource and warehouse with nationwide health, policy, and community data related to the opioid epidemic, available to download, analyze, and visualize at multiple spatial scales. It integrates social and structural determinants of health, which are grouped thematically. It currently includes hundreds variable constructs with the data available for download as CSV files which can be easily merged with geographic shapefiles (that are also provided).
This tutorial will show you how to access data from OEPS and directly integrate it within your own R environment. Objectives include:
- Navigating the most recent OEPS data suite
- Cleaning, merging, and mapping OEPS data in R
- Integrating additional data to OEPS measures
6.2 Environment Setup
To replicate the code and functions illustrated in this tutorial, you will need to have R and RStudio installed on your system. This tutorial assumes some basic familiarity with the R programming language, such as creating objects and running code from an R script.
6.2.1 Input/Output
We will download a data package from OEPS directly, and extract data from the state.csv file for our analyses, and later merge with geospatial data from the same package, state-2020-500k-shp, for mapping and exploration.
We’ll also work with the data at the county level from the same data package, county.csv and county-2020-500k-shp, and bring in some measures from an external dataset, NC_cty_foodstamps_2025.csv.
Our output will be a cleaned county-level CSV file, final-county-data.csv, to be used in further research & analyses, as well as several images of maps from our session.
6.2.2 Install Libraries
We will use the following packages in this tutorial:
sf: to read, manipulate, transform, and write spatial objectstidyverse: tidy data wrangling toolstmap: to create maps
If you have not already installed tidyverse, sf, and tmap, install them first using install.packages("PackageName"). Then load them with the library() function.
6.3 Download OEPS Data
First, go to the OEPS website at https://oeps.healthyregions.org/.
If this is your first time visiting the site, take a moment to explore the About OEPS, Methodology, Data Inventory and Data Standards pages to get familiar with site.
Next, under the Data Access tab, click Download. On this page, Click Aggregated Data Packages.
This tutorial uses data from the 2023 ACS release, which is a 5-year pooled estimate covering 2019-2023.
Under the DSuite2023 heading, download both of the following files:
- DSuite2023 data dictionary
- DSuite2023 data package (100mb+).
See screenshot below.
After downloading the files, create a new folder on your computer called OEPS_Tutorial and save the files there. If the data package downloads as a .zip file, open your Downloads folder, and extract the contents into your OEPS_Tutorial folder before continuing.
We also encourage you to pin this folder to Quick Access.
See the screenshot below.
6.5 Set Up Your R Project
Next, open RStudio and create a new R script called OEPS_Tutorial_RScript.R. Save it in your OEPS_Tutorial folder.
In the Files pane (usually on the bottom right-hand side), click the three dots on the right side, navigate to your OEPS_Tutorial folder, and click Open. Then click the gear icon and select Set As Working Directory.
If the Files pane is not open, click View in the main toolbar and then click Show Files.
At this stage, load the libraries referenced at the start of the tutorial.
## Load required packages
library(tidyverse)
library(sf)
library(tmap)6.6 Prepare the State-Level Data for Analysis and Visualization
6.6.1 Explore the State-Level Dataset
Now, we will read in the state-level CSV file and inspect its contents.
## Read state.csv
state_data <- read.csv("data/state.csv")
## Glimpse state_data
glimpse(state_data)Rows: 56
Columns: 83
$ HEROP_ID <chr> "040US35", "040US46", "040US06", "040US21", "040US01", "0…
$ FIPS <int> 35, 46, 6, 21, 1, 13, 5, 42, 29, 8, 49, 40, 47, 56, 36, 1…
$ FqhcAvTmDr <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, N…
$ FqhcCtTmDr <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, N…
$ FqhcTmDrP <lgl> NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, NA, N…
$ TotTracts <int> 612, 242, NA, 1376, NA, 2791, NA, 3445, 1654, NA, 716, 29…
$ A50_74HcvD <int> 92, 10, 1407, 218, 122, 205, 93, 328, 110, 289, 12, 457, …
$ AmInHcvD <int> NA, 15, 22, NA, NA, NA, NA, NA, NA, 11, NA, 57, NA, NA, N…
$ AsHcvD <dbl> NA, 0, 87, 0, 0, NA, NA, NA, NA, NA, 0, NA, NA, 0, 17, NA…
$ AvA50_74HcvD <dbl> 131, 6, 1703, 219, 132, 251, 115, 357, 163, 331, 43, 453,…
$ AvAmInHcvD <dbl> 7, 2, 32, NA, NA, 0, NA, 0, 0, 4, 0, 58, NA, NA, NA, 0, N…
$ AvBlkHcvD <dbl> NA, NA, 246, 34, 48, 96, 21, 97, 40, 36, NA, 44, 82, 0, 1…
$ AvFlHcvD <dbl> 40, 0, 605, 88, 46, 84, 42, 114, 60, 112, 20, 163, 139, N…
$ AvHcvD <dbl> 159, 33, 2085, 302, 161, 292, 135, 432, 197, 388, 64, 536…
$ AvHspHcvD <dbl> 88, 0, 645, NA, NA, NA, NA, 36, NA, 101, 2, 21, NA, NA, 1…
$ AvMlHcvD <dbl> 119, 25, 1479, 213, 115, 208, 94, 318, 137, 276, 43, 373,…
$ AvO75HcvD <dbl> NA, NA, 263, NA, NA, 10, NA, 22, 1, 20, NA, 22, 17, NA, 7…
$ AvU50HcvD <dbl> 1, NA, 108, 62, NA, 1, NA, 25, 4, 7, NA, 34, 66, NA, 14, …
$ MdMarijLaw <chr> "True", "True", "True", "False", "True", "False", "True",…
$ AnyPdmpDt <chr> "7/15/2004", "3/29/2010", "1/1/1939", "7/15/1998", "8/1/2…
$ AnyPdmpFr <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ AnyPdmphDt <chr> "7/1/2004", "3/1/2010", "1/1/1990", "7/1/1998", "11/1/200…
$ AnyPdmphFr <dbl> 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, 1, …
$ ElcPdmpDt <chr> "7/15/2004", "7/1/2010", "1/1/1998", "7/15/1998", "8/1/20…
$ ElcPdmpFr <dbl> 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 0.05, 1.0…
$ MsAcPdmpDt <chr> "9/28/2012", "", "10/2/2018", "7/20/2012", "3/9/2017", "1…
$ MsAcPdmpFr <dbl> 1, 0, 1, 1, 1, 1, 1, 1, 0, 1, 0, 1, 1, 1, 1, 1, 0, 0, 1, …
$ OpPdmpDt <chr> "8/1/2005", "3/1/2012", "9/1/2009", "7/1/1999", "4/1/2006…
$ OpPdmpFr <dbl> 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 1.00, 0.05, 1.0…
$ CrrctExp <int> 6359956, 783784, 112801895, 6670160, 12533999, 16334064, …
$ ExpnFedExp <dbl> 2026725000, 11769600, 28775203200, 5189196600, NA, NA, 23…
$ ExpnSttExp <dbl> 211870300, 960000, 3268018800, 578826600, NA, NA, 2771269…
$ HlthExp <int> 4692974, 316744, 76926071, 4620478, 10626825, 11019466, 2…
$ MedcdExp <dbl> 8173504300, 1167718100, 122791744500, 16172800500, 787630…
$ PlcFyrExp <int> 1214640, 369042, 35511690, 1636492, 2145358, 4925933, 116…
$ TradFedExp <dbl> 4739079800, 786090500, 50713624800, 8031034100, 611585800…
$ TradSttExp <dbl> 1195829200, 368898000, 40034897800, 2373743200, 176044310…
$ WlfrExp <int> 10049039, 1731185, 167181476, 16877514, 9596296, 17009307…
$ BachelorsP <dbl> 16.59, 21.52, 22.40, 15.91, 16.96, 20.73, 15.88, 20.57, 1…
$ EduHsP <dbl> 25.73, 29.48, 20.40, 32.71, 30.34, 26.89, 34.39, 33.21, 3…
$ EduNoHsP <dbl> 12.28, 6.99, 15.40, 11.47, 11.87, 11.05, 11.44, 8.09, 8.4…
$ EngProf <dbl> 91.34, 97.90, 82.72, 97.31, 97.64, 94.28, 96.75, 95.28, 9…
$ GradSclP <dbl> 13.61, 9.58, 14.10, 11.08, 10.79, 13.47, 9.23, 13.92, 12.…
$ SomeCollegeP <dbl> 19.47, 17.07, 18.18, 16.81, 17.43, 16.33, 17.82, 12.87, 1…
$ FemP <dbl> 50.33, 49.33, 50.04, 50.48, 51.46, 51.20, 50.67, 50.71, 5…
$ MaleP <dbl> 49.67, 50.67, 49.96, 49.52, 48.54, 48.80, 49.33, 49.29, 4…
$ Ovr16 <dbl> 1705686, 704344, 31545603, 3605426, 4056609, 8582918, 240…
$ Ovr16P <dbl> 80.66, 78.33, 80.39, 79.93, 80.26, 79.31, 79.46, 81.84, 8…
$ Ovr18 <dbl> 1648151, 679374, 30513773, 3487979, 3926252, 8280518, 232…
$ Ovr18P <dbl> 77.94, 75.55, 77.76, 77.33, 77.68, 76.51, 76.72, 79.41, 7…
$ Ovr21 <dbl> 1443310, 595134, 26488214, 3028981, 3389617, 7120517, 202…
$ Ovr21P <dbl> 68.25, 66.19, 67.50, 67.15, 67.06, 65.79, 66.69, 68.94, 6…
$ Ovr65 <dbl> 397156, 158103, 5994486, 767995, 884210, 1581263, 525644,…
$ Ovr65P <dbl> 18.78, 17.58, 15.28, 17.03, 17.49, 14.61, 17.33, 19.07, 1…
$ SRatio <dbl> 98.68, 102.71, 99.84, 98.11, 94.33, 95.32, 97.35, 97.20, …
$ SRatio18 <dbl> 97.27, 102.18, 98.16, 95.74, 91.27, 92.35, 94.85, 94.85, …
$ SRatio65 <dbl> 85.51, 89.35, 82.11, 81.51, 78.70, 78.77, 81.21, 80.64, 8…
$ TotPop <int> 2114768, 899194, 39242785, 4510725, 5054253, 10822590, 30…
$ AmIndE <dbl> 201346, 69514, 445219, 7610, 22491, 42058, 16893, 23921, …
$ AmIndP <dbl> 9.52, 7.73, 1.13, 0.17, 0.44, 0.39, 0.56, 0.18, 0.27, 1.0…
$ AsianE <dbl> 36065, 12554, 5997069, 68482, 71969, 472915, 47309, 48043…
$ AsianP <dbl> 1.71, 1.40, 15.28, 1.52, 1.42, 4.37, 1.56, 3.70, 2.07, 3.…
$ BlackE <dbl> 44709, 20149, 2173343, 355237, 1318507, 3391689, 452127, …
$ BlackP <dbl> 2.11, 2.24, 5.54, 7.88, 26.09, 31.34, 14.91, 10.73, 11.13…
$ HisE <dbl> 1018321, 41281, 15630830, 212163, 271640, 1158299, 265833…
$ HisP <dbl> 48.15, 4.59, 39.83, 4.70, 5.37, 10.70, 8.77, 8.38, 5.06, …
$ OtherE <dbl> 253794, 12547, 6820303, 67198, 107302, 448586, 91871, 445…
$ OtherP <dbl> 12.00, 1.40, 17.38, 1.49, 2.12, 4.14, 3.03, 3.43, 1.73, 5…
$ PacIsE <dbl> 2049, 606, 147827, 3737, 2564, 7338, 12040, 4614, 9681, 8…
$ PacIsP <dbl> 0.10, 0.07, 0.38, 0.08, 0.05, 0.07, 0.40, 0.04, 0.16, 0.1…
$ TwoRaceE <dbl> 442934, 50789, 6410245, 233880, 228050, 782473, 263525, 7…
$ TwoRaceP <dbl> 20.94, 5.65, 16.33, 5.18, 4.51, 7.23, 8.69, 6.12, 6.30, 1…
$ WhiteE <dbl> 1133871, 733035, 17248779, 3774581, 3303370, 5677531, 214…
$ WhiteP <dbl> 53.62, 81.52, 43.95, 83.68, 65.36, 52.46, 70.86, 75.80, 7…
$ DsmAs <dbl> 0.42, 0.24, 0.41, 0.52, 0.57, 0.47, 0.56, 0.47, 0.48, 0.3…
$ DsmBlk <dbl> 0.39, 0.28, 0.52, 0.45, 0.40, 0.34, 0.42, 0.55, 0.44, 0.4…
$ DsmHsp <dbl> 0.24, 0.19, 0.33, 0.40, 0.44, 0.34, 0.38, 0.40, 0.33, 0.2…
$ IntrAsWht <dbl> 0.36, 0.53, 0.47, 0.76, 0.62, 0.54, 0.68, 0.82, 0.83, 0.6…
$ IntrBlkWht <dbl> 0.38, 0.61, 0.45, 0.84, 0.54, 0.54, 0.66, 0.74, 0.84, 0.6…
$ IntrHspWht <dbl> 0.37, 0.74, 0.45, 0.86, 0.60, 0.59, 0.68, 0.78, 0.86, 0.6…
$ IsoAs <dbl> 0.03, 0.01, 0.13, 0.02, 0.02, 0.04, 0.02, 0.04, 0.02, 0.0…
$ IsoBlk <dbl> 0.03, 0.01, 0.07, 0.07, 0.37, 0.35, 0.21, 0.13, 0.07, 0.0…
$ IsoHsp <dbl> 0.53, 0.04, 0.40, 0.06, 0.08, 0.11, 0.10, 0.10, 0.06, 0.2…
The state-level dataset contains 56 rows, which represent the 50 states plus the District of Columbia and U.S. territories. It also contains 83 columns, each representing a different variable.
The glimpse() function shows the structure of the dataset, including each variable’s type and a preview of its values.
You can also view the loaded dataset by clicking on the name in the Environment tab.
6.6.2 Choose a Variable to Explore
For this tutorial, we will explore Hepatitis C (HCV) mortality as an example health outcome variable. The same workflow can be used to explore other health outcomes.
Open the HepatitisC_Rates.md in the metadata folder and read through the contents. The OEPS Hepatitis C (HCV) data was orginially sourced from HepVu. HepVu includes single-year data from 2018-2022. OEPS researchers cleaned and prepared this data for analysis by aggregating the multiple single year datasets into a single multi-year state-level dataset for 2018-2022.
Hepatitis C, also referred to as HCV, is relevant to opioid-environment research because injection drug use is one of the most common risk factors for HCV infection. People who use drugs are therefore at elevated risk for HCV. HCV can lead to chronic infection, serious liver disease, and/or death if left untreated. For these reasons, HCV is often used as a related health outcome in opioid research.
In the DSuite2023-data-dictionary.xlsx, in the Construct column, scroll to Hepatitis C Rates. You will see several variables related to HCV deaths by age, race, and ethnicity.
For simplicity, we will use AvHcvD, which is the average annual number of HCV deaths from 2018-2022. This variable is only available at the state-level in the OEPS database. While HepVu does include county-level HCV mortality, the data is a calculated estimation. For a complete description of the data, see HepVu’s Data Methods.
Return to RStudio and locate the AvHcvD variable in state_data. You will see that the row with FIPS code 35, which corresponds to New Mexico, has an average annual HCV death count of 159 people.
To check which FIPS code corresponds to which state or territory, visit this page from the U.S. Census Bureau.
6.6.3 Create a Population-Adjusted Measure
Next, we will prepare the data and generate descriptive statistics for the health outcome variable.
The AvHcvD variable is a raw count. Because states have different population sizes, we will standardize this measure by dividing AvHcvD by TotPop, the total population.
We will call this new variable PropHcvD. It represents the average annual proportion of the total state population that died from HCV.
We use total population as the denominator, rather than total HCV cases because of data availability constraints. Data for average annual HCV cases is only available as a 5-year pooled estimate covering 2013-2016. Whereas, this tutorial is using pooled mortality data from 2018-2022.
## Create a population-adjusted variable
state_data$PropHcvD <- (state_data$AvHcvD / state_data$TotPop)
## View the first few values
head(state_data$PropHcvD)[1] 7.518555e-05 3.669953e-05 5.313079e-05 6.695154e-05 3.185436e-05
[6] 2.698060e-05
6.6.4 Summarize the Variable
Next, use the summary() function to examine the distribution of PropHcvD. This output includes the mean, median, minimum, maximum, first quartile, third quartile, and the count of missing values.
## Summarize new variable
summary(state_data$PropHcvD) Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
0.000019 0.000032 0.000039 0.000046 0.000054 0.000134 5
Because these values are very small, they can be difficult to interpret directly. Additionally, the output shows five missing values, which we will address in a later section.
6.6.5 Re-express the Variable Per 100,000 People
To make the variable easier to interpret, multiply PropHcvD by 100,000. This gives the average annual HCV deaths per 100,000 people. We will call this new variable AvHcVD100K.
## Re-express the variable per 100,000 people
state_data$AvHcvD100K <- state_data$PropHcvD * 1000006.6.6 Identify Missing Values
Now we will identify the rows with missing values for AvHcvD100K.
## View the rows with missing AvHcvD100K data
state_data[is.na(state_data$AvHcvD100K),] HEROP_ID FIPS FqhcAvTmDr FqhcCtTmDr FqhcTmDrP TotTracts A50_74HcvD AmInHcvD
30 040US60 60 NA NA NA 8706 NA NA
35 040US72 72 NA NA NA 939 NA NA
48 040US78 78 NA NA NA 29 NA NA
54 040US66 66 NA NA NA 56 NA NA
55 040US69 69 NA NA NA 23 NA NA
AsHcvD AvA50_74HcvD AvAmInHcvD AvBlkHcvD AvFlHcvD AvHcvD AvHspHcvD AvMlHcvD
30 NA NA NA NA NA NA NA NA
35 NA NA NA NA NA NA NA NA
48 NA NA NA NA NA NA NA NA
54 NA NA NA NA NA NA NA NA
55 NA NA NA NA NA NA NA NA
AvO75HcvD AvU50HcvD MdMarijLaw AnyPdmpDt AnyPdmpFr AnyPdmphDt AnyPdmphFr
30 NA NA False NA NA
35 NA NA False NA NA
48 NA NA False NA NA
54 NA NA False NA NA
55 NA NA False NA NA
ElcPdmpDt ElcPdmpFr MsAcPdmpDt MsAcPdmpFr OpPdmpDt OpPdmpFr CrrctExp
30 NA NA NA NA
35 NA NA NA NA
48 NA NA NA NA
54 NA NA NA NA
55 NA NA NA NA
ExpnFedExp ExpnSttExp HlthExp MedcdExp PlcFyrExp TradFedExp TradSttExp
30 NA NA NA NA NA NA NA
35 NA NA NA NA NA NA NA
48 NA NA NA NA NA NA NA
54 NA NA NA NA NA NA NA
55 NA NA NA NA NA NA NA
WlfrExp BachelorsP EduHsP EduNoHsP EngProf GradSclP SomeCollegeP FemP MaleP
30 NA NA NA NA NA NA NA NA NA
35 NA NA NA NA NA NA NA NA NA
48 NA NA NA NA NA NA NA NA NA
54 NA NA NA NA NA NA NA NA NA
55 NA NA NA NA NA NA NA NA NA
Ovr16 Ovr16P Ovr18 Ovr18P Ovr21 Ovr21P Ovr65 Ovr65P SRatio SRatio18 SRatio65
30 NA NA NA NA NA NA NA NA NA NA NA
35 NA NA NA NA NA NA NA NA NA NA NA
48 NA NA NA NA NA NA NA NA NA NA NA
54 NA NA NA NA NA NA NA NA NA NA NA
55 NA NA NA NA NA NA NA NA NA NA NA
TotPop AmIndE AmIndP AsianE AsianP BlackE BlackP HisE HisP OtherE OtherP
30 NA NA NA NA NA NA NA NA NA NA NA
35 NA NA NA NA NA NA NA NA NA NA NA
48 NA NA NA NA NA NA NA NA NA NA NA
54 NA NA NA NA NA NA NA NA NA NA NA
55 NA NA NA NA NA NA NA NA NA NA NA
PacIsE PacIsP TwoRaceE TwoRaceP WhiteE WhiteP DsmAs DsmBlk DsmHsp IntrAsWht
30 NA NA NA NA NA NA NA NA NA NA
35 NA NA NA NA NA NA NA NA NA NA
48 NA NA NA NA NA NA NA NA NA NA
54 NA NA NA NA NA NA NA NA NA NA
55 NA NA NA NA NA NA NA NA NA NA
IntrBlkWht IntrHspWht IsoAs IsoBlk IsoHsp PropHcvD AvHcvD100K
30 NA NA NA NA NA NA NA
35 NA NA NA NA NA NA NA
48 NA NA NA NA NA NA NA
54 NA NA NA NA NA NA NA
55 NA NA NA NA NA NA NA
These missing values correspond to American Samoa (60), Guam (66), Northern Mariana Islands (69), Puerto Rico (72), and the U.S. Virgin Islands (78). For these locations, the dataset does not include HCV deaths or total population values.
6.6.7 Filter the Dataset
For the remainder of this tutorial, we will limit the analysis to the contiguous U.S., which will also make the mapping step easier. This excludes the territories listed above, as well as Alaska (02), and Hawaii (15).
## Filter to the contiguous U.S.
state_data_filtered <- state_data %>% filter(!FIPS %in% c(02, 15, 60, 66, 69, 72, 78))6.6.8 Summary Statistics
Now that AvHcvD100K has been calculated and processed, use the summary() function again to examine the distribution.
## Summarize new variable
summary(state_data_filtered$AvHcvD100K) Min. 1st Qu. Median Mean 3rd Qu. Max.
1.921 3.194 3.949 4.642 5.313 13.416
From this output, we can see the typical range of values across states and identify whether any states have unusually high or low rates. The summary also helps us confirm that the missing values have been removed from the filtered dataset.
The average annual HCV death rate ranges from 1.921 to 13.416 deaths per 100,000 people across states in the contiguous U.S. The median rate is 3.949 and the mean rate is 4.642, which is higher than the median, suggesting that the distribution is slightly right-skewed, with a few states having relatively high rates.
The distribution suggests that HCV mortality is not uniform across states. But it is important to note that summary statistics alone cannot tell us why rates differ.
6.6.9 Load the State Shapefile
Now that the data has been prepared, we can create a map of the outcome variable.
Before we can map the data, we need to join the state-level dataset to a state shapefile. The shapefile is included in the OEPS data suite.
In your OEPS_Tutorial folder, open the data folder and locate state-2020-500k-shp.zip. If the file is still compressed, unzip it first. For best organization, create a folder named state-2020-500k-shp inside the data folder and extract the shapefile contents there.
See the screenshot below.
Now load the shapefile into R and display the geometry column to confirm that it imported correctly.
## Load in the state shapefile
state_geom <- st_read("data/state-2020-500k-shp/state-2020-500k.shp")Reading layer `state-2020-500k' from data source
`/Users/maryniakolak/Code/opioid-environment-toolkit/data/state-2020-500k-shp/state-2020-500k.shp'
using driver `ESRI Shapefile'
Simple feature collection with 56 features and 16 fields
Geometry type: MULTIPOLYGON
Dimension: XY
Bounding box: xmin: -179.1467 ymin: -14.5487 xmax: 179.7785 ymax: 71.38782
Geodetic CRS: NAD83
## Create a quick plot
plot(state_geom$geometry)6.6.10 Filter to the Contiguous U.S.
Next, filter the shapefile to include only the contiguous U.S. This excludes Alaska, Hawaii, and the U.S. territories.
## Filter to the contiguous U.S.
state_geom_filtered <- state_geom %>% filter(!STATEFP %in% c("02", "15", "60", "66", "69", "72", "78"))
## Check plot
plot(state_geom_filtered$geometry)6.6.11 Join the Data to the Shapefile
Now join the state_data_filtered dataframe to the filtered shapefile. We join using HEROP_ID which serves as a unified geographic identifier for joins. To read more about this variable click here.
## Join state_data_filtered to state_geom_filtered
data_to_map <- state_geom_filtered %>%
left_join(state_data_filtered, by = "HEROP_ID")6.6.12 Create a Map
Finally, map AvHcvD100K, which represents the average annual HCV death rate per 100,000 people.
# Map
tm_shape(data_to_map) +
tm_polygons(
fill = "AvHcvD100K",
fill.scale = tm_scale_intervals(
values = c("white", "#FFF7BC", "#FEC44F", "#FE9929", "#D95F0E", "#993404"),
breaks = c(0, 0.0000001, 3.99999, 6.99999, 9.99999, 11.99999, 14.99999),
labels = c("0", "1–3", "4–6", "7–9", "10–12", "12-14"),
label.na = ""
),
fill.legend = tm_legend(title = "")) +
tm_title("Average Annual Hepatitis C (HCV) Deaths per 100,000 People")[plot mode] fit legend/component: Some legend items or map compoments do not
fit well, and are therefore rescaled.
ℹ Set the tmap option `component.autoscale = FALSE` to disable rescaling.
The map suggests that HCV death rates are not evenly distributed across the U.S. Instead, several Southern and Western states appear to have higher rates compared to Northeastern states. Oklahoma appears to have the highest rate of annual HCV deaths.
6.7 Prepare the County-Level Data for Analysis and Visualization
OEPS includes data at multiple geographic levels. In this section, we will demonstrate how the state-level workflow can be applied to county-level data. Additionally, we will show how to subset the data to a single state for closer examination.
6.7.1 Explore the County-Level Dataset
Now, we will read in the county-level CSV file and inspect its contents.
## Read county.csv
county_data <- read.csv("data/county.csv")
## Glimpse state_data
glimpse(county_data)Rows: 3,234
Columns: 100
$ HEROP_ID <chr> "050US13031", "050US13121", "050US13179", "050US13189", "…
$ FIPS <int> 13031, 13121, 13179, 13189, 13213, 13209, 13139, 13245, 1…
$ SviSmryRnk <dbl> 0.8374, 0.6599, 0.9007, 0.6567, 0.6847, 0.7547, 0.8136, 0…
$ SviTh1 <dbl> 0.9297, 0.5571, 0.8167, 0.7353, 0.7852, 0.8858, 0.7897, 0…
$ SviTh2 <dbl> 0.2326, 0.2278, 0.8431, 0.5002, 0.6061, 0.5756, 0.6640, 0…
$ SviTh3 <dbl> 0.7521, 0.9259, 0.9322, 0.8540, 0.5250, 0.7191, 0.7941, 0…
$ SviTh4 <dbl> 0.8781, 0.8501, 0.8600, 0.4076, 0.4909, 0.4518, 0.7416, 0…
$ EssnWrkP <dbl> 49.9, 27.6, 50.8, 54.9, 54.6, 57.4, 48.3, 53.4, 70.5, 61.…
$ HghRskP <dbl> 18.0, 10.5, 14.8, 25.2, 41.6, 30.2, 28.9, 18.0, 28.0, 27.…
$ HltCrP <dbl> 10.1, 10.2, 12.5, 15.5, 8.3, 14.6, 10.0, 16.3, 8.1, 14.0,…
$ RetailP <dbl> 12.3, 9.0, 14.0, 11.9, 12.0, 9.4, 11.1, 12.3, 12.8, 18.7,…
$ TotWrkE <int> 36964, 569609, 23579, 9334, 18970, 3565, 99845, 86611, 48…
$ GiniCoeff <dbl> 0.4849, 0.5296, 0.4089, 0.4515, 0.4337, 0.4652, 0.4493, 0…
$ MedInc <int> 35840, 61045, 36864, 36134, 41999, 41019, 41391, 35067, 3…
$ PciE <int> 30450, 61438, 27648, 28071, 30205, 27633, 37271, 30209, 2…
$ PovP <dbl> 22.5, 12.9, 14.5, 19.0, 13.0, 16.4, 12.7, 21.1, 22.5, 31.…
$ UnempP <dbl> 7.9, 5.4, 8.0, 3.7, 5.0, 6.1, 3.9, 7.9, 10.1, 5.9, 3.4, 6…
$ LngTermP <dbl> 14.2, 11.5, 12.3, 28.8, 19.1, 23.7, 13.6, 18.6, 22.4, 36.…
$ MobileP <dbl> 15.4, 0.5, 16.7, 27.8, 28.9, 37.9, 8.4, 6.5, 37.1, 44.1, …
$ RentalP <dbl> 111.0, 93.5, 139.0, 94.4, 68.5, 55.7, 89.4, 122.3, 50.0, …
$ TotUnits <int> 33607, 500404, 27228, 9466, 16163, 3776, 79233, 92652, 47…
$ UnitDens <dbl> 49.7, 950.1, 52.7, 36.8, 46.9, 15.7, 201.6, 285.7, 6.0, 1…
$ VacantP <dbl> 9.9, 8.5, 14.0, 12.3, 5.6, 20.0, 10.1, 19.0, 13.2, 42.3, …
$ FqhcAvTmDr <dbl> 10.51, 7.07, 11.77, 19.81, 12.37, 8.41, 10.97, 6.99, 18.7…
$ FqhcCtTmDr <int> 20, 326, 16, 6, 9, 3, 50, 55, 2, 2, 17, 4, 8, 4, 8, 9, 2,…
$ FqhcTmDrP <dbl> 100.00, 99.69, 94.12, 100.00, 90.00, 100.00, 100.00, 98.2…
$ TotTracts <int> 20, 327, 17, 6, 10, 3, 50, 56, 3, 2, 17, 4, 8, 4, 8, 9, 2…
$ OpRxRt <dbl> 72.4, 61.1, 27.5, 52.2, 28.3, 0.2, 79.0, 89.9, 58.6, 9.9,…
$ BachelorsP <dbl> 18.6, 33.6, 13.5, 11.2, 8.8, 10.8, 16.7, 14.7, 8.4, 5.7, …
$ EduHsP <dbl> 28.6, 15.4, 31.7, 40.4, 35.1, 37.3, 27.7, 31.8, 43.7, 39.…
$ EduNoHsP <dbl> 10.4, 6.4, 8.0, 14.5, 21.3, 14.7, 18.6, 12.1, 14.8, 26.1,…
$ EngProf <dbl> 0.007, 0.019, 0.011, 0.001, 0.026, 0.021, 0.080, 0.005, 0…
$ GradSclP <dbl> 12.2, 24.4, 6.4, 6.6, 4.9, 9.5, 10.0, 9.4, 3.0, 2.7, 8.7,…
$ SomeCollegeP <dbl> 21.5, 14.4, 28.6, 18.6, 21.9, 16.9, 19.3, 22.3, 24.9, 18.…
$ CrowdHsng <dbl> 0.01, 0.02, 0.02, 0.01, 0.03, 0.01, 0.05, 0.01, 0.03, 0.0…
$ DivrcdP <dbl> 9.0, 8.9, 9.3, 10.8, 10.5, 10.0, 8.6, 12.5, 13.4, 6.3, 8.…
$ FamSize <dbl> 3.05, 3.12, 3.39, 3.12, 3.00, 3.18, 3.37, 3.43, 3.71, 3.1…
$ HHSize <dbl> 2.48, 2.26, 2.75, 2.58, 2.62, 2.59, 2.89, 2.62, 2.94, 2.4…
$ HhldFA <dbl> 15.3, 21.0, 12.2, 16.0, 12.0, 14.7, 12.8, 19.5, 13.4, 20.…
$ HhldFC <dbl> 6.6, 6.1, 8.6, 12.8, 4.3, 7.8, 4.5, 12.0, 4.1, 6.0, 3.7, …
$ HhldFS <dbl> 6.1, 6.5, 4.4, 5.8, 6.6, 9.2, 6.7, 8.2, 6.1, 11.0, 5.5, 1…
$ HhldMA <dbl> 14.0, 17.4, 15.1, 10.5, 8.0, 14.2, 8.9, 15.5, 14.8, 13.0,…
$ HhldMC <dbl> 0.9, 1.1, 1.7, 1.3, 2.2, 1.5, 1.3, 1.3, 0.0, 0.0, 1.3, 2.…
$ HhldMS <dbl> 3.9, 3.0, 2.6, 3.3, 2.9, 5.8, 3.0, 4.2, 3.7, 6.1, 2.3, 6.…
$ HsdTot <dbl> 30276, 457791, 23406, 8298, 15258, 3021, 71263, 75023, 40…
$ HsdTypCo <dbl> 7.4, 6.6, 6.8, 4.3, 5.4, 4.7, 5.8, 5.4, 8.5, 3.0, 5.4, 3.…
$ HsdTypM <dbl> 37.0, 34.9, 45.8, 41.5, 55.8, 49.1, 55.6, 31.1, 41.6, 43.…
$ HsdTypMC <dbl> 14.9, 14.3, 18.9, 10.5, 21.2, 14.0, 21.7, 9.1, 14.8, 6.2,…
$ MrrdP <dbl> 37.2, 41.4, 47.3, 46.2, 56.7, 48.1, 55.3, 33.6, 34.5, 31.…
$ NonRelFhhP <dbl> 7.0, 6.8, 7.4, 7.6, 8.2, 8.2, 9.7, 9.3, 9.5, 12.7, 7.6, 9…
$ NonRelNfhhP <dbl> 10.3, 3.9, 2.0, 2.8, 1.8, 1.5, 3.1, 5.4, 3.4, 4.0, 2.3, 5…
$ NvMrrdP <dbl> 49.5, 46.4, 38.0, 37.8, 28.1, 37.7, 32.3, 49.2, 48.2, 54.…
$ OccupantP <dbl> 90.1, 91.5, 86.0, 87.7, 94.4, 80.0, 89.9, 81.0, 86.8, 57.…
$ SepartedP <dbl> 1.6, 1.3, 3.5, 2.2, 1.6, 0.9, 1.5, 2.3, 1.9, 2.1, 1.1, 3.…
$ TotPopHh <int> 75146, 1036112, 64336, 21421, 39993, 7827, 206163, 196186…
$ WidwdP <dbl> 2.7, 2.0, 1.9, 3.0, 3.1, 3.3, 2.2, 2.5, 2.0, 6.1, 2.3, 2.…
$ DisbP <dbl> 14.8, 10.2, 15.9, 12.4, 13.3, 14.0, 11.7, 19.1, 19.8, 23.…
$ FemP <dbl> 51.5, 51.4, 48.4, 50.7, 49.9, 48.4, 50.1, 51.6, 42.0, 47.…
$ MaleP <dbl> 48.5, 48.6, 51.6, 49.3, 50.1, 51.6, 49.9, 48.4, 58.0, 52.…
$ MedAge <dbl> 30.2, 36.2, 28.7, 39.1, 39.3, 38.2, 38.2, 35.0, 38.0, 47.…
$ Ovr16 <dbl> 66867, 869503, 49840, 16947, 31927, 7071, 163982, 163638,…
$ Ovr16P <dbl> 82.2, 81.4, 74.6, 78.1, 79.3, 81.5, 78.7, 79.4, 84.5, 86.…
$ Ovr18 <dbl> 64967, 842905, 48332, 16263, 30663, 6898, 158032, 158585,…
$ Ovr18P <dbl> 79.8, 78.9, 72.3, 75.0, 76.1, 79.5, 75.8, 77.0, 80.3, 84.…
$ Ovr21 <dbl> 50293, 726211, 41973, 14025, 26256, 5695, 135491, 137391,…
$ Ovr21P <dbl> 66.8, 74.7, 66.9, 70.8, 72.2, 71.6, 71.7, 71.9, 72.9, 81.…
$ Ovr62P <dbl> 14.8, 15.7, 12.4, 22.6, 19.1, 21.8, 19.7, 18.6, 18.8, 27.…
$ Ovr65 <dbl> 9836, 132189, 6558, 4004, 5997, 1512, 33285, 30664, 1872,…
$ Ovr65P <dbl> 12.1, 12.4, 9.8, 18.5, 14.9, 17.4, 16.0, 14.9, 14.7, 24.6…
$ SRatio <dbl> 94.3, 94.7, 106.5, 97.1, 100.5, 106.4, 99.6, 93.8, 138.0,…
$ SRatio18 <dbl> 94.1, 92.6, 107.3, 85.3, 98.4, 104.6, 97.8, 89.8, 147.8, …
$ SRatio65 <dbl> 83.5, 74.0, 81.5, 72.1, 82.1, 81.7, 84.0, 73.8, 104.6, 87…
$ TotPop <int> 81372, 1068507, 66826, 21687, 40282, 8675, 208395, 206040…
$ Und18P <dbl> 20.2, 21.1, 27.7, 25.0, 23.9, 20.5, 24.2, 23.0, 19.7, 15.…
$ AmIndE <dbl> 323, 2680, 310, 0, 337, 14, 896, 251, 75, 19, 151, 4, 10,…
$ AmIndP <dbl> 0.4, 0.3, 0.5, 0.0, 0.8, 0.2, 0.4, 0.1, 0.6, 0.2, 0.2, 0.…
$ AsianE <dbl> 1180, 81115, 1191, 96, 151, 53, 4351, 3687, 108, 47, 1679…
$ BlackE <dbl> 23003, 460766, 28654, 8911, 246, 1991, 14159, 115506, 318…
$ BlackP <dbl> 28.3, 43.1, 42.9, 41.1, 0.6, 23.0, 6.8, 56.1, 25.0, 71.4,…
$ HisE <dbl> 4105, 86110, 8298, 777, 6287, 618, 59586, 11627, 1967, 55…
$ HisP <dbl> 5.0, 8.1, 12.4, 3.6, 15.6, 7.1, 28.6, 5.6, 15.5, 0.6, 9.9…
$ OtherE <dbl> 1128, 33908, 2140, 253, 1556, 349, 17409, 5171, 199, 36, …
$ OtherP <dbl> 1.4, 3.2, 3.2, 1.2, 3.9, 4.0, 8.4, 2.5, 1.6, 0.4, 3.4, 1.…
$ PacIsE <dbl> 30, 299, 295, 0, 136, 0, 15, 372, 0, 0, 34, 1, 0, 0, 37, …
$ PacIsP <dbl> 0.0, 0.0, 0.4, 0.0, 0.3, 0.0, 0.0, 0.2, 0.0, 0.0, 0.0, 0.…
$ TwoRaceE <dbl> 4625, 69252, 7686, 495, 2700, 401, 31113, 12304, 1918, 87…
$ TwoRaceP <dbl> 5.7, 6.5, 11.5, 2.3, 6.7, 4.6, 14.9, 6.0, 15.1, 1.0, 6.8,…
$ WhiteE <dbl> 51083, 420487, 26550, 11932, 35156, 5867, 140452, 68749, …
$ WhiteP <dbl> 62.8, 39.4, 39.7, 55.0, 87.3, 67.6, 67.4, 33.4, 56.9, 26.…
$ DsmAs <dbl> 0.42, 0.47, 0.38, 0.75, 0.55, 0.28, 0.47, 0.45, 0.66, 0.4…
$ DsmBlk <dbl> 0.34, 0.71, 0.27, 0.34, 0.71, 0.45, 0.45, 0.45, 0.30, 0.2…
$ DsmHsp <dbl> 0.30, 0.44, 0.28, 0.37, 0.35, 0.13, 0.48, 0.44, 0.46, 0.4…
$ IntrAsWht <dbl> 0.65, 0.46, 0.34, 0.80, 0.80, 0.61, 0.57, 0.43, 0.50, 0.2…
$ IntrBlkWht <dbl> 0.53, 0.15, 0.34, 0.46, 0.77, 0.54, 0.50, 0.24, 0.53, 0.2…
$ IntrHspWht <dbl> 0.61, 0.38, 0.36, 0.56, 0.73, 0.66, 0.42, 0.32, 0.55, 0.2…
$ IsoAs <dbl> 0.03, 0.23, 0.03, 0.03, 0.01, 0.01, 0.05, 0.05, 0.02, 0.0…
$ IsoBlk <dbl> 0.37, 0.72, 0.46, 0.50, 0.03, 0.32, 0.11, 0.65, 0.35, 0.7…
$ IsoHsp <dbl> 0.07, 0.18, 0.16, 0.07, 0.23, 0.07, 0.45, 0.11, 0.32, 0.0…
$ TotVetPop <int> 64920, 841821, 40874, 16263, 30628, 6895, 157766, 151216,…
$ VetP <dbl> 5.8, 4.7, 19.7, 8.1, 5.0, 5.6, 6.2, 11.5, 8.4, 5.8, 6.7, …
The county-level dataset contains 3,234 rows representing individual counties and 100 columns representing different variables.
The glimpse() function shows the structure of the dataset, including each variable’s type and a preview of its values.
6.7.2 Choose Variables to Explore
Next, we will identify a few variables to explore.
Open DSuite2023-data-dictionary.xlsx. Explore the variables that have a year listed under the county column. These are variables where data exists at the county-level.
For this tutorial, let’s focus on PovP, the number of individuals earning below the poverty income threshold as a percentage of the total population. Open Economic_Characteristics.md in the metadata folder to read about this variable.
Additionally, let’s look at FqhcAvTmDr, the average driving time (minutes) across tracts in state to nearest Federally Qualified Health Center (FQHC). Open Access_FQHCs.md in the metadata folder to read about this variable.
## Summary statistics for PovP
summary(county_data$PovP) Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
1.70 10.00 13.30 14.26 17.50 52.80 99
The county-level poverty percentage ranges from 1.7 to 52.8%. The median is 13.3% and the mean is 14.26%, suggesting that the distribution is slightly right-skewed, with few counties having high percentages. There are 99 counties with no percent poverty data.
## Summary statistics for FqhcAvTmDr
summary(county_data$FqhcAvTmDr) Min. 1st Qu. Median Mean 3rd Qu. Max. NA's
0.00 8.52 14.13 20.85 27.55 89.82 168
The county-level average driving time to nearest FQHC ranges from 0 to 89.82 minutes. The median is 14.13 minutes and the mean is 20.85 minutes, which suggests that most counties are relatively close to an FQHC, but some are much further away. There are 168 counties with missing data, which corresponds to the worst access, where travel takes over 90 minutes in optimal conditions, or 180 minutes in normal conditions.
6.7.3 Filter the Dataset to a Single State
Next, we filter the county dataset to just counties in a single state using the HEROP_ID variable. In this tutorial, we will zoom in on North Carolina.
The HEROP_ID consists of four parts:
- The 3-digit summary level code for counties: 050
- The string: US
- The 2-digit state FIPS code
- The 3-digit county FIPS code
In the following code, we filter by the 6th and 7th positions of the HEROP_ID string which correspond to the 2-digit state FIPS code. North Carolina’s FIPS code is 37.
To check which FIPS code corresponds to which state or territory, visit this page from the U.S. Census Bureau.
## Filter to a single state
county_data_filtered <- county_data %>%
filter(substr(HEROP_ID, 6, 7) %in% c("37"))
## View the first few values
head(county_data_filtered) HEROP_ID FIPS SviSmryRnk SviTh1 SviTh2 SviTh3 SviTh4 EssnWrkP HghRskP
1 050US37083 37083 0.9870 0.9844 0.9007 0.9332 0.9360 57.3 25.9
2 050US37179 37179 0.2008 0.2221 0.3786 0.6701 0.0932 38.5 22.6
3 050US37129 37129 0.4289 0.5638 0.0824 0.5797 0.4989 40.5 15.5
4 050US37163 37163 0.9459 0.9300 0.8702 0.8721 0.8670 56.8 32.9
5 050US37031 37031 0.3888 0.4547 0.2558 0.4057 0.4432 46.8 16.7
6 050US37183 37183 0.3621 0.2526 0.3229 0.7989 0.4257 30.8 16.1
HltCrP RetailP TotWrkE GiniCoeff MedInc PciE PovP UnempP LngTermP MobileP
1 15.7 10.6 18108 0.4891 37991 26951 25.2 8.7 26.8 24.7
2 10.8 11.7 123262 0.4455 51241 45355 7.7 4.2 13.1 4.8
3 14.2 11.7 119132 0.4801 44816 46083 12.4 4.7 11.7 2.7
4 15.6 13.0 24859 0.4520 37968 26978 20.2 5.2 25.4 34.0
5 12.1 13.1 30470 0.4628 41628 42707 10.0 4.3 18.3 16.8
6 11.8 9.3 611492 0.4486 58520 52949 7.9 4.2 10.6 2.7
RentalP TotUnits UnitDens VacantP FqhcAvTmDr FqhcCtTmDr FqhcTmDrP TotTracts
1 86.6 24849 34.3 18.4 4.29 16 100.00 16
2 51.5 86730 137.1 5.2 16.62 48 100.00 48
3 79.0 117236 609.7 12.7 9.00 54 100.00 54
4 78.5 25739 27.2 17.2 6.65 20 100.00 20
5 52.0 51615 101.7 39.6 31.04 20 64.52 31
6 79.2 481999 577.5 7.5 9.85 230 100.00 230
OpRxRt BachelorsP EduHsP EduNoHsP EngProf GradSclP SomeCollegeP CrowdHsng
1 47.9 10.2 36.4 19.1 0.004 5.2 20.3 0.02
2 23.1 25.6 22.8 9.6 0.027 13.4 19.5 0.03
3 76.6 29.1 18.0 6.5 0.015 15.7 19.7 0.01
4 30.2 11.4 35.7 15.7 0.053 4.4 20.9 0.04
5 52.4 20.2 23.6 7.3 0.004 12.1 25.5 0.01
6 35.8 33.8 14.5 6.1 0.027 22.5 15.2 0.02
DivrcdP FamSize HHSize HhldFA HhldFC HhldFS HhldMA HhldMC HhldMS HsdTot
1 8.6 2.95 2.32 19.1 10.0 11.8 13.5 1.1 5.3 20269
2 6.5 3.33 2.95 9.7 3.8 5.0 7.3 1.1 2.2 82231
3 7.9 2.85 2.20 19.7 3.8 9.3 13.8 0.7 3.8 102376
4 9.9 3.28 2.71 15.7 6.1 9.8 11.9 1.3 3.6 21321
5 10.4 2.69 2.16 17.0 2.7 10.0 13.2 0.7 5.4 31198
6 7.1 3.13 2.54 15.2 4.6 5.8 11.5 1.2 2.2 445636
HsdTypCo HsdTypM HsdTypMC MrrdP NonRelFhhP NonRelNfhhP NvMrrdP OccupantP
1 6.2 38.8 9.5 44.4 6.8 1.1 40.7 81.6
2 4.5 65.7 30.8 61.1 6.4 1.9 29.5 94.8
3 8.1 43.4 15.0 49.4 3.9 4.0 38.3 87.3
4 6.1 48.2 14.8 48.1 9.6 2.6 35.4 82.8
5 5.1 52.8 14.8 58.8 4.3 1.6 25.5 60.4
6 6.7 51.0 23.6 54.3 4.9 3.3 35.5 92.5
SepartedP TotPopHh WidwdP DisbP FemP MaleP MedAge Ovr16 Ovr16P Ovr18 Ovr18P
1 1.7 47017 4.5 20.9 51.4 48.6 43.8 39315 81.5 37968 78.7
2 1.4 242316 1.5 9.2 50.3 49.7 39.1 190283 77.7 181266 74.0
3 2.1 224956 2.3 11.8 52.3 47.7 40.1 193627 83.7 189042 81.8
4 2.0 57842 4.5 14.6 50.1 49.9 39.7 46618 78.7 44755 75.5
5 1.3 67436 4.0 17.6 51.1 48.9 50.1 58516 85.2 56945 82.9
6 1.4 1129867 1.7 9.0 51.0 49.0 37.2 914290 79.4 881746 76.6
Ovr21 Ovr21P Ovr62P Ovr65 Ovr65P SRatio SRatio18 SRatio65 TotPop Und18P
1 33485 75.6 26.3 10426 21.6 94.6 91.9 74.7 48219 21.3
2 150881 69.5 16.5 32401 13.2 98.7 96.4 81.5 244975 26.0
3 163335 76.8 22.4 43240 18.7 91.3 88.7 77.7 231214 18.2
4 38759 72.1 21.8 10636 18.0 99.7 97.7 77.7 59245 24.5
5 50406 80.2 32.9 18051 26.3 95.8 94.2 87.3 68652 17.1
6 758522 72.8 15.6 144054 12.5 96.0 93.6 78.0 1151009 23.4
AmIndE AmIndP AsianE BlackE BlackP HisE HisP OtherE OtherP PacIsE PacIsP
1 1424 3.0 446 25038 51.9 1558 3.2 776 1.6 24 0
2 1086 0.4 10002 27397 11.2 31683 12.9 14837 6.1 58 0
3 487 0.2 3204 26060 11.3 17675 7.6 8813 3.8 38 0
4 1288 2.2 289 14768 24.9 12600 21.3 8816 14.9 0 0
5 188 0.3 603 2998 4.4 3230 4.7 1427 2.1 15 0
6 3325 0.3 93193 221985 19.3 131104 11.4 55569 4.8 383 0
TwoRaceE TwoRaceP WhiteE WhiteP DsmAs DsmBlk DsmHsp IntrAsWht IntrBlkWht
1 2574 5.3 17937 37.2 0.69 0.40 0.35 0.34 0.30
2 17450 7.1 174145 71.1 0.45 0.36 0.39 0.69 0.60
3 13627 5.9 178985 77.4 0.43 0.53 0.41 0.79 0.56
4 2545 4.3 31539 53.2 0.58 0.33 0.33 0.50 0.43
5 3871 5.6 59550 86.7 0.54 0.59 0.37 0.82 0.76
6 92986 8.1 683568 59.4 0.53 0.46 0.43 0.50 0.42
IntrHspWht IsoAs IsoBlk IsoHsp TotVetPop VetP
1 0.43 0.04 0.60 0.07 37942 6.1
2 0.57 0.10 0.17 0.22 181172 6.4
3 0.66 0.03 0.29 0.15 188258 6.9
4 0.44 0.03 0.32 0.28 44570 7.1
5 0.81 0.03 0.13 0.10 56268 12.0
6 0.46 0.25 0.33 0.19 880615 5.7
6.7.4 Load the County Shapefile
Before we can map the data, we need to join the filtered dataset to a county shapefile. The shapefile is included in the OEPS data suite.
In your OEPS_Tutorial folder, open the data folder and locate county-2020-500k-shp.zip. If the file is still compressed, unzip it first. For best organization, create a folder named county-2020-500k-shp inside the data folder and extract the shapefile contents there.
Now, load the shapefile into R and display the geometry column to confirm that it imported correctly.
## Load in the county shapefile
county_geom <- st_read("data/county-2020-500k-shp/county-2020-500k.shp")Reading layer `county-2020-500k' from data source
`/Users/maryniakolak/Code/opioid-environment-toolkit/data/county-2020-500k-shp/county-2020-500k.shp'
using driver `ESRI Shapefile'
Simple feature collection with 3234 features and 19 fields
Geometry type: MULTIPOLYGON
Dimension: XY
Bounding box: xmin: -179.1467 ymin: -14.5487 xmax: 179.7785 ymax: 71.38782
Geodetic CRS: NAD83
## Create a quick plot
plot(county_geom$geometry)6.7.5 Filter Shapefile to a Single State
Next, filter the shapefile to only include North Carolina counties.
## Filter to a single state
county_geom_filtered <- county_geom %>% filter(STATEFP %in% c("37"))
## Check plot
plot(county_geom_filtered$geometry)6.7.6 Join the Data to the Shapefile
Now join the county_data_filtered dataframe to the filtered shapefile.
## Join county_data_filtered to county_geom_filtered
data_to_map_county <- county_geom_filtered %>%
left_join(county_data_filtered, by = "HEROP_ID")6.7.7 Create Maps
Finally, map PovP, the number of individuals earning below the poverty income threshold as a percentage of the total population and FqhcAvTmDr, the average driving time (minutes) across tracts in state to nearest Federally Qualified Health Center (FQHC).
## Map PovP
PovP <- tm_shape(data_to_map_county) +
tm_polygons(
fill = "PovP",
fill.scale = tm_scale_intervals(style = "jenks"),
fill.legend = tm_legend(title = ""),
col = "grey85",
lwd = 0.2) +
tm_title("% of Individuals Earning Below Poverty Income Threshold")
## Map FqhcAvTmDr
FqhcAvTmDr <- tm_shape(data_to_map_county) +
tm_polygons(
fill = "FqhcAvTmDr",
fill.scale = tm_scale_intervals(style = "jenks"),
fill.legend = tm_legend(title = ""),
col = "grey85",
lwd = 0.2) +
tm_title("Driving time (minutes) to nearest FQHC")
## Arrange the maps
tmap_arrange(PovP, FqhcAvTmDr)The PovP map shows us that poverty is unevenly distributed across the state. Poverty percentages appear higher in several Eastern and Southern counties, whereas many counties in the central Piedmont have lower poverty rates.
For the FqhcAvTmDr map, most counties appear to have relatively short driving times, especially in the central part of the state. Longer travel times seem more common in some coastal and mountain areas, where access may be more limited. There are no counties that exceed a driving time of 58 minutes.
6.8 Join OEPS data to outside data
Sometimes users will already have their own data and want to add OEPS variables for context. In this section, we will demonstrate how to join OEPS county-level data to an outside dataset using a shared geographic identifier.
The NC_cty_foodstamps_2025.csv contains information about the county-level prevalence of food stamps in North Carolina for 2025, which was originally pulled from CDC’s PLACES dataset.
## Read in non-OEPS data
foodstamps_data <- read.csv("data/NC_cty_foodstamps_2025.csv")
## Glimpse foodstamps_data
glimpse(foodstamps_data)Rows: 100
Columns: 11
$ StateAbbr <chr> "NC", "NC", "NC", "NC", "NC", "NC", "NC", "NC", "N…
$ StateDesc <chr> "North Carolina", "North Carolina", "North Carolin…
$ CountyName <chr> "Bertie", "Vance", "New Hanover", "Martin", "Edgec…
$ StateFIPS <int> 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37…
$ CountyFIPS <int> 15, 181, 129, 117, 65, 95, 159, 55, 193, 125, 45, …
$ TotalPopulation <chr> "16,922", "42,301", "238,852", "21,447", "48,832",…
$ TotalPop18plus <chr> "14,112", "32,113", "197,058", "17,125", "37,653",…
$ FOODSTAMP_CrudePrev <dbl> 24.9, 22.9, 9.0, 19.9, 24.8, 19.9, 14.1, 6.4, 13.6…
$ FOODSTAMP_Crude95ci <chr> "(20.8, 29.1)", "(18.8, 26.9)", "( 6.8, 11.6)", "(…
$ FOODSTAMP_AdjPrev <dbl> 26.7, 24.2, 9.6, 21.6, 26.3, 20.8, 14.8, 7.0, 14.7…
$ FOODSTAMP_Adj95CI <chr> "(22.1, 31.3)", "(19.9, 28.7)", "( 7.3, 12.5)", "(…
Often times, datasets will have a FIPS or GEOID column that can be used for joining to other datasets. Sometimes, these columns will not be the same format, in which case, joining will not work. Thus, you will need to re-format the columns before joining.
In foodstamps_data, we have a column called StateFIPS which stores the 2-digit state FIPS code and a CountyFIPS column which stores the 3-digit county FIPS code. However, CSV files often delete leading zeros, which is why you see some 1- or 2-digit values in CountyFIPS.
Let’s check again to see what identifier columns are in county_data_filtered.
## Glimpse county_data_filtered
glimpse(county_data_filtered)Rows: 100
Columns: 100
$ HEROP_ID <chr> "050US37083", "050US37179", "050US37129", "050US37163", "…
$ FIPS <int> 37083, 37179, 37129, 37163, 37031, 37183, 37119, 37087, 3…
$ SviSmryRnk <dbl> 0.9870, 0.2008, 0.4289, 0.9459, 0.3888, 0.3621, 0.6115, 0…
$ SviTh1 <dbl> 0.9844, 0.2221, 0.5638, 0.9300, 0.4547, 0.2526, 0.5787, 0…
$ SviTh2 <dbl> 0.9007, 0.3786, 0.0824, 0.8702, 0.2558, 0.3229, 0.4747, 0…
$ SviTh3 <dbl> 0.9332, 0.6701, 0.5797, 0.8721, 0.4057, 0.7989, 0.8915, 0…
$ SviTh4 <dbl> 0.9360, 0.0932, 0.4989, 0.8670, 0.4432, 0.4257, 0.5135, 0…
$ EssnWrkP <dbl> 57.3, 38.5, 40.5, 56.8, 46.8, 30.8, 35.7, 49.7, 50.1, 46.…
$ HghRskP <dbl> 25.9, 22.6, 15.5, 32.9, 16.7, 16.1, 16.3, 21.2, 24.6, 19.…
$ HltCrP <dbl> 15.7, 10.8, 14.2, 15.6, 12.1, 11.8, 11.5, 15.1, 12.3, 16.…
$ RetailP <dbl> 10.6, 11.7, 11.7, 13.0, 13.1, 9.3, 10.3, 11.1, 12.6, 11.7…
$ TotWrkE <int> 18108, 123262, 119132, 24859, 30470, 611492, 616325, 2827…
$ GiniCoeff <dbl> 0.4891, 0.4455, 0.4801, 0.4520, 0.4628, 0.4486, 0.4935, 0…
$ MedInc <int> 37991, 51241, 44816, 37968, 41628, 58520, 51889, 40926, 4…
$ PciE <int> 26951, 45355, 46083, 26978, 42707, 52949, 51490, 35580, 3…
$ PovP <dbl> 25.2, 7.7, 12.4, 20.2, 10.0, 7.9, 10.4, 11.3, 10.7, 11.8,…
$ UnempP <dbl> 8.7, 4.2, 4.7, 5.2, 4.3, 4.2, 4.2, 3.4, 6.0, 3.5, 4.7, 5.…
$ LngTermP <dbl> 26.8, 13.1, 11.7, 25.4, 18.3, 10.6, 10.1, 23.0, 13.3, 17.…
$ MobileP <dbl> 24.7, 4.8, 2.7, 34.0, 16.8, 2.7, 1.5, 16.7, 22.8, 12.1, 2…
$ RentalP <dbl> 86.6, 51.5, 79.0, 78.5, 52.0, 79.2, 98.6, 51.9, 46.6, 84.…
$ TotUnits <int> 24849, 86730, 117236, 25739, 51615, 481999, 491721, 35307…
$ UnitDens <dbl> 34.3, 137.1, 609.7, 27.2, 101.7, 577.5, 939.1, 63.8, 35.6…
$ VacantP <dbl> 18.4, 5.2, 12.7, 17.2, 39.6, 7.5, 7.4, 24.2, 22.1, 22.3, …
$ FqhcAvTmDr <dbl> 4.29, 16.62, 9.00, 6.65, 31.04, 9.85, 11.08, 13.09, 19.40…
$ FqhcCtTmDr <int> 16, 48, 54, 20, 20, 230, 305, 20, 12, 64, 15, 11, 22, 12,…
$ FqhcTmDrP <dbl> 100.00, 100.00, 100.00, 100.00, 64.52, 100.00, 100.00, 95…
$ TotTracts <int> 16, 48, 54, 20, 31, 230, 305, 21, 15, 65, 15, 12, 23, 12,…
$ OpRxRt <dbl> 47.9, 23.1, 76.6, 30.2, 52.4, 35.8, 45.0, 52.7, 16.7, 73.…
$ BachelorsP <dbl> 10.2, 25.6, 29.1, 11.4, 20.2, 33.8, 31.3, 17.9, 19.7, 26.…
$ EduHsP <dbl> 36.4, 22.8, 18.0, 35.7, 23.6, 14.5, 16.4, 27.5, 27.7, 22.…
$ EduNoHsP <dbl> 19.1, 9.6, 6.5, 15.7, 7.3, 6.1, 9.2, 8.4, 9.5, 7.7, 13.5,…
$ EngProf <dbl> 0.004, 0.027, 0.015, 0.053, 0.004, 0.027, 0.055, 0.006, 0…
$ GradSclP <dbl> 5.2, 13.4, 15.7, 4.4, 12.1, 22.5, 17.3, 11.5, 10.1, 18.2,…
$ SomeCollegeP <dbl> 20.3, 19.5, 19.7, 20.9, 25.5, 15.2, 17.6, 21.8, 21.7, 16.…
$ CrowdHsng <dbl> 0.02, 0.03, 0.01, 0.04, 0.01, 0.02, 0.02, 0.01, 0.02, 0.0…
$ DivrcdP <dbl> 8.6, 6.5, 7.9, 9.9, 10.4, 7.1, 7.7, 10.6, 11.7, 10.7, 12.…
$ FamSize <dbl> 2.95, 3.33, 2.85, 3.28, 2.69, 3.13, 3.19, 2.85, 3.14, 3.2…
$ HHSize <dbl> 2.32, 2.95, 2.20, 2.71, 2.16, 2.54, 2.45, 2.31, 2.58, 2.5…
$ HhldFA <dbl> 19.1, 9.7, 19.7, 15.7, 17.0, 15.2, 18.7, 18.0, 14.6, 17.4…
$ HhldFC <dbl> 10.0, 3.8, 3.8, 6.1, 2.7, 4.6, 6.0, 4.0, 3.9, 3.9, 6.3, 2…
$ HhldFS <dbl> 11.8, 5.0, 9.3, 9.8, 10.0, 5.8, 5.7, 10.9, 8.6, 9.3, 11.9…
$ HhldMA <dbl> 13.5, 7.3, 13.8, 11.9, 13.2, 11.5, 14.4, 12.6, 12.9, 13.5…
$ HhldMC <dbl> 1.1, 1.1, 0.7, 1.3, 0.7, 1.2, 1.4, 0.7, 1.2, 1.3, 1.2, 0.…
$ HhldMS <dbl> 5.3, 2.2, 3.8, 3.6, 5.4, 2.2, 2.2, 5.8, 4.6, 4.4, 5.3, 6.…
$ HsdTot <dbl> 20269, 82231, 102376, 21321, 31198, 445636, 455494, 26772…
$ HsdTypCo <dbl> 6.2, 4.5, 8.1, 6.1, 5.1, 6.7, 7.3, 6.2, 7.5, 6.4, 4.2, 3.…
$ HsdTypM <dbl> 38.8, 65.7, 43.4, 48.2, 52.8, 51.0, 41.5, 49.9, 52.3, 46.…
$ HsdTypMC <dbl> 9.5, 30.8, 15.0, 14.8, 14.8, 23.6, 17.9, 12.4, 20.3, 14.9…
$ MrrdP <dbl> 44.4, 61.1, 49.4, 48.1, 58.8, 54.3, 47.4, 55.3, 52.3, 46.…
$ NonRelFhhP <dbl> 6.8, 6.4, 3.9, 9.6, 4.3, 4.9, 6.2, 6.9, 6.4, 7.4, 8.3, 8.…
$ NonRelNfhhP <dbl> 1.1, 1.9, 4.0, 2.6, 1.6, 3.3, 3.7, 2.6, 2.6, 7.0, 2.7, 3.…
$ NvMrrdP <dbl> 40.7, 29.5, 38.3, 35.4, 25.5, 35.5, 41.8, 28.9, 29.9, 38.…
$ OccupantP <dbl> 81.6, 94.8, 87.3, 82.8, 60.4, 92.5, 92.6, 75.8, 77.9, 77.…
$ SepartedP <dbl> 1.7, 1.4, 2.1, 2.0, 1.3, 1.4, 1.6, 1.6, 2.3, 1.2, 2.6, 1.…
$ TotPopHh <int> 47017, 242316, 224956, 57842, 67436, 1129867, 1115618, 61…
$ WidwdP <dbl> 4.5, 1.5, 2.3, 4.5, 4.0, 1.7, 1.6, 3.6, 3.8, 3.1, 4.5, 5.…
$ DisbP <dbl> 20.9, 9.2, 11.8, 14.6, 17.6, 9.0, 8.3, 17.3, 15.3, 13.8, …
$ FemP <dbl> 51.4, 50.3, 52.3, 50.1, 51.1, 51.0, 51.7, 51.2, 49.4, 51.…
$ MaleP <dbl> 48.6, 49.7, 47.7, 49.9, 48.9, 49.0, 48.3, 48.8, 50.6, 48.…
$ MedAge <dbl> 43.8, 39.1, 40.1, 39.7, 50.1, 37.2, 35.4, 47.8, 41.8, 42.…
$ Ovr16 <dbl> 39315, 190283, 193627, 46618, 58516, 914290, 898405, 5279…
$ Ovr16P <dbl> 81.5, 77.7, 83.7, 78.7, 85.2, 79.4, 79.4, 84.6, 80.4, 84.…
$ Ovr18 <dbl> 37968, 181266, 189042, 44755, 56945, 881746, 870043, 5129…
$ Ovr18P <dbl> 78.7, 74.0, 81.8, 75.5, 82.9, 76.6, 76.9, 82.2, 77.5, 81.…
$ Ovr21 <dbl> 33485, 150881, 163335, 38759, 50406, 758522, 755280, 4549…
$ Ovr21P <dbl> 75.6, 69.5, 76.8, 72.1, 80.2, 72.8, 73.3, 79.7, 73.8, 78.…
$ Ovr62P <dbl> 26.3, 16.5, 22.4, 21.8, 32.9, 15.6, 14.7, 30.6, 22.4, 24.…
$ Ovr65 <dbl> 10426, 32401, 43240, 10636, 18051, 144054, 132281, 15827,…
$ Ovr65P <dbl> 21.6, 13.2, 18.7, 18.0, 26.3, 12.5, 11.7, 25.4, 18.0, 20.…
$ SRatio <dbl> 94.6, 98.7, 91.3, 99.7, 95.8, 96.0, 93.5, 95.3, 102.5, 93…
$ SRatio18 <dbl> 91.9, 96.4, 88.7, 97.7, 94.2, 93.6, 90.7, 91.9, 100.0, 91…
$ SRatio65 <dbl> 74.7, 81.5, 77.7, 77.7, 87.3, 78.0, 74.2, 83.7, 87.4, 78.…
$ TotPop <int> 48219, 244975, 231214, 59245, 68652, 1151009, 1130906, 62…
$ Und18P <dbl> 21.3, 26.0, 18.2, 24.5, 17.1, 23.4, 23.1, 17.8, 22.5, 18.…
$ AmIndE <dbl> 1424, 1086, 487, 1288, 188, 3325, 4355, 361, 209, 565, 15…
$ AmIndP <dbl> 3.0, 0.4, 0.2, 2.2, 0.3, 0.3, 0.4, 0.6, 0.3, 0.2, 3.1, 1.…
$ AsianE <dbl> 446, 10002, 3204, 289, 603, 93193, 69579, 255, 437, 3144,…
$ BlackE <dbl> 25038, 27397, 26060, 14768, 2998, 221985, 344569, 624, 78…
$ BlackP <dbl> 51.9, 11.2, 11.3, 24.9, 4.4, 19.3, 30.5, 1.0, 12.4, 5.4, …
$ HisE <dbl> 1558, 31683, 17675, 12600, 3230, 131104, 174580, 2953, 54…
$ HisP <dbl> 3.2, 12.9, 7.6, 21.3, 4.7, 11.4, 15.4, 4.7, 8.5, 8.3, 5.5…
$ OtherE <dbl> 776, 14837, 8813, 8816, 1427, 55569, 88166, 568, 3429, 58…
$ OtherP <dbl> 1.6, 6.1, 3.8, 14.9, 2.1, 4.8, 7.8, 0.9, 5.4, 2.1, 3.1, 0…
$ PacIsE <dbl> 24, 58, 38, 0, 15, 383, 434, 0, 41, 95, 22, 5, 39, 7, 0, …
$ PacIsP <dbl> 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.1, 0.0, 0.0, 0.…
$ TwoRaceE <dbl> 2574, 17450, 13627, 2545, 3871, 92986, 97368, 3271, 3432,…
$ TwoRaceP <dbl> 5.3, 7.1, 5.9, 4.3, 5.6, 8.1, 8.6, 5.2, 5.4, 7.9, 4.5, 4.…
$ WhiteE <dbl> 17937, 174145, 178985, 31539, 59550, 683568, 526435, 5735…
$ WhiteP <dbl> 37.2, 71.1, 77.4, 53.2, 86.7, 59.4, 46.5, 91.9, 75.7, 83.…
$ DsmAs <dbl> 0.69, 0.45, 0.43, 0.58, 0.54, 0.53, 0.45, 0.53, 0.45, 0.4…
$ DsmBlk <dbl> 0.40, 0.36, 0.53, 0.33, 0.59, 0.46, 0.56, 0.56, 0.43, 0.5…
$ DsmHsp <dbl> 0.35, 0.39, 0.41, 0.33, 0.37, 0.43, 0.54, 0.33, 0.43, 0.4…
$ IntrAsWht <dbl> 0.34, 0.69, 0.79, 0.50, 0.82, 0.50, 0.42, 0.90, 0.79, 0.8…
$ IntrBlkWht <dbl> 0.30, 0.60, 0.56, 0.43, 0.76, 0.42, 0.27, 0.88, 0.63, 0.7…
$ IntrHspWht <dbl> 0.43, 0.57, 0.66, 0.44, 0.81, 0.46, 0.30, 0.88, 0.66, 0.7…
$ IsoAs <dbl> 0.04, 0.10, 0.03, 0.03, 0.03, 0.25, 0.12, 0.01, 0.01, 0.0…
$ IsoBlk <dbl> 0.60, 0.17, 0.29, 0.32, 0.13, 0.33, 0.46, 0.03, 0.22, 0.1…
$ IsoHsp <dbl> 0.07, 0.22, 0.15, 0.28, 0.10, 0.19, 0.27, 0.08, 0.14, 0.1…
$ TotVetPop <int> 37942, 181172, 188258, 44570, 56268, 880615, 868562, 5123…
$ VetP <dbl> 6.1, 6.4, 6.9, 7.1, 12.0, 5.7, 5.3, 9.7, 9.7, 6.7, 6.7, 1…
In county_data_filtered, we have HEROP_ID, our unique OEPS identifier and FIPS which stores the 5-digit combined state and county FIPS codes.
It is obvious that the two datasets we want to join do not have a shared geographic identifier. But, we have all the information we need to create one!
Let’s keep the OEPS data as is, and instead alter foodstamps_data by creating a 5-digit FIPS column. To do so, we need to combine StateFIPS with CountyFIPS ensuring that CountyFIPS does not leave out leading zeros.
## Create a FIPS column
foodstamps_data <- foodstamps_data %>%
mutate(FIPS = as.numeric(sprintf("%02d%03d", StateFIPS, CountyFIPS)))
## View the first few values
head(foodstamps_data) StateAbbr StateDesc CountyName StateFIPS CountyFIPS TotalPopulation
1 NC North Carolina Bertie 37 15 16,922
2 NC North Carolina Vance 37 181 42,301
3 NC North Carolina New Hanover 37 129 238,852
4 NC North Carolina Martin 37 117 21,447
5 NC North Carolina Edgecombe 37 65 48,832
6 NC North Carolina Hyde 37 95 4,607
TotalPop18plus FOODSTAMP_CrudePrev FOODSTAMP_Crude95ci FOODSTAMP_AdjPrev
1 14,112 24.9 (20.8, 29.1) 26.7
2 32,113 22.9 (18.8, 26.9) 24.2
3 197,058 9.0 ( 6.8, 11.6) 9.6
4 17,125 19.9 (16.4, 23.7) 21.6
5 37,653 24.8 (20.5, 29.3) 26.3
6 3,922 19.9 (16.6, 23.3) 20.8
FOODSTAMP_Adj95CI FIPS
1 (22.1, 31.3) 37015
2 (19.9, 28.7) 37181
3 ( 7.3, 12.5) 37129
4 (17.9, 25.8) 37117
5 (21.8, 31.1) 37065
6 (17.4, 24.7) 37095
Now we have a shared geographic identifier, FIPS, and can join the two datasets.
## Join foodstamps_data to county_data_filtered
final_county_data <- county_data_filtered %>%
left_join(foodstamps_data, by = "FIPS")
## Glimpse final_county_data
glimpse(final_county_data)Rows: 100
Columns: 111
$ HEROP_ID <chr> "050US37083", "050US37179", "050US37129", "050US37…
$ FIPS <dbl> 37083, 37179, 37129, 37163, 37031, 37183, 37119, 3…
$ SviSmryRnk <dbl> 0.9870, 0.2008, 0.4289, 0.9459, 0.3888, 0.3621, 0.…
$ SviTh1 <dbl> 0.9844, 0.2221, 0.5638, 0.9300, 0.4547, 0.2526, 0.…
$ SviTh2 <dbl> 0.9007, 0.3786, 0.0824, 0.8702, 0.2558, 0.3229, 0.…
$ SviTh3 <dbl> 0.9332, 0.6701, 0.5797, 0.8721, 0.4057, 0.7989, 0.…
$ SviTh4 <dbl> 0.9360, 0.0932, 0.4989, 0.8670, 0.4432, 0.4257, 0.…
$ EssnWrkP <dbl> 57.3, 38.5, 40.5, 56.8, 46.8, 30.8, 35.7, 49.7, 50…
$ HghRskP <dbl> 25.9, 22.6, 15.5, 32.9, 16.7, 16.1, 16.3, 21.2, 24…
$ HltCrP <dbl> 15.7, 10.8, 14.2, 15.6, 12.1, 11.8, 11.5, 15.1, 12…
$ RetailP <dbl> 10.6, 11.7, 11.7, 13.0, 13.1, 9.3, 10.3, 11.1, 12.…
$ TotWrkE <int> 18108, 123262, 119132, 24859, 30470, 611492, 61632…
$ GiniCoeff <dbl> 0.4891, 0.4455, 0.4801, 0.4520, 0.4628, 0.4486, 0.…
$ MedInc <int> 37991, 51241, 44816, 37968, 41628, 58520, 51889, 4…
$ PciE <int> 26951, 45355, 46083, 26978, 42707, 52949, 51490, 3…
$ PovP <dbl> 25.2, 7.7, 12.4, 20.2, 10.0, 7.9, 10.4, 11.3, 10.7…
$ UnempP <dbl> 8.7, 4.2, 4.7, 5.2, 4.3, 4.2, 4.2, 3.4, 6.0, 3.5, …
$ LngTermP <dbl> 26.8, 13.1, 11.7, 25.4, 18.3, 10.6, 10.1, 23.0, 13…
$ MobileP <dbl> 24.7, 4.8, 2.7, 34.0, 16.8, 2.7, 1.5, 16.7, 22.8, …
$ RentalP <dbl> 86.6, 51.5, 79.0, 78.5, 52.0, 79.2, 98.6, 51.9, 46…
$ TotUnits <int> 24849, 86730, 117236, 25739, 51615, 481999, 491721…
$ UnitDens <dbl> 34.3, 137.1, 609.7, 27.2, 101.7, 577.5, 939.1, 63.…
$ VacantP <dbl> 18.4, 5.2, 12.7, 17.2, 39.6, 7.5, 7.4, 24.2, 22.1,…
$ FqhcAvTmDr <dbl> 4.29, 16.62, 9.00, 6.65, 31.04, 9.85, 11.08, 13.09…
$ FqhcCtTmDr <int> 16, 48, 54, 20, 20, 230, 305, 20, 12, 64, 15, 11, …
$ FqhcTmDrP <dbl> 100.00, 100.00, 100.00, 100.00, 64.52, 100.00, 100…
$ TotTracts <int> 16, 48, 54, 20, 31, 230, 305, 21, 15, 65, 15, 12, …
$ OpRxRt <dbl> 47.9, 23.1, 76.6, 30.2, 52.4, 35.8, 45.0, 52.7, 16…
$ BachelorsP <dbl> 10.2, 25.6, 29.1, 11.4, 20.2, 33.8, 31.3, 17.9, 19…
$ EduHsP <dbl> 36.4, 22.8, 18.0, 35.7, 23.6, 14.5, 16.4, 27.5, 27…
$ EduNoHsP <dbl> 19.1, 9.6, 6.5, 15.7, 7.3, 6.1, 9.2, 8.4, 9.5, 7.7…
$ EngProf <dbl> 0.004, 0.027, 0.015, 0.053, 0.004, 0.027, 0.055, 0…
$ GradSclP <dbl> 5.2, 13.4, 15.7, 4.4, 12.1, 22.5, 17.3, 11.5, 10.1…
$ SomeCollegeP <dbl> 20.3, 19.5, 19.7, 20.9, 25.5, 15.2, 17.6, 21.8, 21…
$ CrowdHsng <dbl> 0.02, 0.03, 0.01, 0.04, 0.01, 0.02, 0.02, 0.01, 0.…
$ DivrcdP <dbl> 8.6, 6.5, 7.9, 9.9, 10.4, 7.1, 7.7, 10.6, 11.7, 10…
$ FamSize <dbl> 2.95, 3.33, 2.85, 3.28, 2.69, 3.13, 3.19, 2.85, 3.…
$ HHSize <dbl> 2.32, 2.95, 2.20, 2.71, 2.16, 2.54, 2.45, 2.31, 2.…
$ HhldFA <dbl> 19.1, 9.7, 19.7, 15.7, 17.0, 15.2, 18.7, 18.0, 14.…
$ HhldFC <dbl> 10.0, 3.8, 3.8, 6.1, 2.7, 4.6, 6.0, 4.0, 3.9, 3.9,…
$ HhldFS <dbl> 11.8, 5.0, 9.3, 9.8, 10.0, 5.8, 5.7, 10.9, 8.6, 9.…
$ HhldMA <dbl> 13.5, 7.3, 13.8, 11.9, 13.2, 11.5, 14.4, 12.6, 12.…
$ HhldMC <dbl> 1.1, 1.1, 0.7, 1.3, 0.7, 1.2, 1.4, 0.7, 1.2, 1.3, …
$ HhldMS <dbl> 5.3, 2.2, 3.8, 3.6, 5.4, 2.2, 2.2, 5.8, 4.6, 4.4, …
$ HsdTot <dbl> 20269, 82231, 102376, 21321, 31198, 445636, 455494…
$ HsdTypCo <dbl> 6.2, 4.5, 8.1, 6.1, 5.1, 6.7, 7.3, 6.2, 7.5, 6.4, …
$ HsdTypM <dbl> 38.8, 65.7, 43.4, 48.2, 52.8, 51.0, 41.5, 49.9, 52…
$ HsdTypMC <dbl> 9.5, 30.8, 15.0, 14.8, 14.8, 23.6, 17.9, 12.4, 20.…
$ MrrdP <dbl> 44.4, 61.1, 49.4, 48.1, 58.8, 54.3, 47.4, 55.3, 52…
$ NonRelFhhP <dbl> 6.8, 6.4, 3.9, 9.6, 4.3, 4.9, 6.2, 6.9, 6.4, 7.4, …
$ NonRelNfhhP <dbl> 1.1, 1.9, 4.0, 2.6, 1.6, 3.3, 3.7, 2.6, 2.6, 7.0, …
$ NvMrrdP <dbl> 40.7, 29.5, 38.3, 35.4, 25.5, 35.5, 41.8, 28.9, 29…
$ OccupantP <dbl> 81.6, 94.8, 87.3, 82.8, 60.4, 92.5, 92.6, 75.8, 77…
$ SepartedP <dbl> 1.7, 1.4, 2.1, 2.0, 1.3, 1.4, 1.6, 1.6, 2.3, 1.2, …
$ TotPopHh <int> 47017, 242316, 224956, 57842, 67436, 1129867, 1115…
$ WidwdP <dbl> 4.5, 1.5, 2.3, 4.5, 4.0, 1.7, 1.6, 3.6, 3.8, 3.1, …
$ DisbP <dbl> 20.9, 9.2, 11.8, 14.6, 17.6, 9.0, 8.3, 17.3, 15.3,…
$ FemP <dbl> 51.4, 50.3, 52.3, 50.1, 51.1, 51.0, 51.7, 51.2, 49…
$ MaleP <dbl> 48.6, 49.7, 47.7, 49.9, 48.9, 49.0, 48.3, 48.8, 50…
$ MedAge <dbl> 43.8, 39.1, 40.1, 39.7, 50.1, 37.2, 35.4, 47.8, 41…
$ Ovr16 <dbl> 39315, 190283, 193627, 46618, 58516, 914290, 89840…
$ Ovr16P <dbl> 81.5, 77.7, 83.7, 78.7, 85.2, 79.4, 79.4, 84.6, 80…
$ Ovr18 <dbl> 37968, 181266, 189042, 44755, 56945, 881746, 87004…
$ Ovr18P <dbl> 78.7, 74.0, 81.8, 75.5, 82.9, 76.6, 76.9, 82.2, 77…
$ Ovr21 <dbl> 33485, 150881, 163335, 38759, 50406, 758522, 75528…
$ Ovr21P <dbl> 75.6, 69.5, 76.8, 72.1, 80.2, 72.8, 73.3, 79.7, 73…
$ Ovr62P <dbl> 26.3, 16.5, 22.4, 21.8, 32.9, 15.6, 14.7, 30.6, 22…
$ Ovr65 <dbl> 10426, 32401, 43240, 10636, 18051, 144054, 132281,…
$ Ovr65P <dbl> 21.6, 13.2, 18.7, 18.0, 26.3, 12.5, 11.7, 25.4, 18…
$ SRatio <dbl> 94.6, 98.7, 91.3, 99.7, 95.8, 96.0, 93.5, 95.3, 10…
$ SRatio18 <dbl> 91.9, 96.4, 88.7, 97.7, 94.2, 93.6, 90.7, 91.9, 10…
$ SRatio65 <dbl> 74.7, 81.5, 77.7, 77.7, 87.3, 78.0, 74.2, 83.7, 87…
$ TotPop <int> 48219, 244975, 231214, 59245, 68652, 1151009, 1130…
$ Und18P <dbl> 21.3, 26.0, 18.2, 24.5, 17.1, 23.4, 23.1, 17.8, 22…
$ AmIndE <dbl> 1424, 1086, 487, 1288, 188, 3325, 4355, 361, 209, …
$ AmIndP <dbl> 3.0, 0.4, 0.2, 2.2, 0.3, 0.3, 0.4, 0.6, 0.3, 0.2, …
$ AsianE <dbl> 446, 10002, 3204, 289, 603, 93193, 69579, 255, 437…
$ BlackE <dbl> 25038, 27397, 26060, 14768, 2998, 221985, 344569, …
$ BlackP <dbl> 51.9, 11.2, 11.3, 24.9, 4.4, 19.3, 30.5, 1.0, 12.4…
$ HisE <dbl> 1558, 31683, 17675, 12600, 3230, 131104, 174580, 2…
$ HisP <dbl> 3.2, 12.9, 7.6, 21.3, 4.7, 11.4, 15.4, 4.7, 8.5, 8…
$ OtherE <dbl> 776, 14837, 8813, 8816, 1427, 55569, 88166, 568, 3…
$ OtherP <dbl> 1.6, 6.1, 3.8, 14.9, 2.1, 4.8, 7.8, 0.9, 5.4, 2.1,…
$ PacIsE <dbl> 24, 58, 38, 0, 15, 383, 434, 0, 41, 95, 22, 5, 39,…
$ PacIsP <dbl> 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.0, 0.1, 0.0, …
$ TwoRaceE <dbl> 2574, 17450, 13627, 2545, 3871, 92986, 97368, 3271…
$ TwoRaceP <dbl> 5.3, 7.1, 5.9, 4.3, 5.6, 8.1, 8.6, 5.2, 5.4, 7.9, …
$ WhiteE <dbl> 17937, 174145, 178985, 31539, 59550, 683568, 52643…
$ WhiteP <dbl> 37.2, 71.1, 77.4, 53.2, 86.7, 59.4, 46.5, 91.9, 75…
$ DsmAs <dbl> 0.69, 0.45, 0.43, 0.58, 0.54, 0.53, 0.45, 0.53, 0.…
$ DsmBlk <dbl> 0.40, 0.36, 0.53, 0.33, 0.59, 0.46, 0.56, 0.56, 0.…
$ DsmHsp <dbl> 0.35, 0.39, 0.41, 0.33, 0.37, 0.43, 0.54, 0.33, 0.…
$ IntrAsWht <dbl> 0.34, 0.69, 0.79, 0.50, 0.82, 0.50, 0.42, 0.90, 0.…
$ IntrBlkWht <dbl> 0.30, 0.60, 0.56, 0.43, 0.76, 0.42, 0.27, 0.88, 0.…
$ IntrHspWht <dbl> 0.43, 0.57, 0.66, 0.44, 0.81, 0.46, 0.30, 0.88, 0.…
$ IsoAs <dbl> 0.04, 0.10, 0.03, 0.03, 0.03, 0.25, 0.12, 0.01, 0.…
$ IsoBlk <dbl> 0.60, 0.17, 0.29, 0.32, 0.13, 0.33, 0.46, 0.03, 0.…
$ IsoHsp <dbl> 0.07, 0.22, 0.15, 0.28, 0.10, 0.19, 0.27, 0.08, 0.…
$ TotVetPop <int> 37942, 181172, 188258, 44570, 56268, 880615, 86856…
$ VetP <dbl> 6.1, 6.4, 6.9, 7.1, 12.0, 5.7, 5.3, 9.7, 9.7, 6.7,…
$ StateAbbr <chr> "NC", "NC", "NC", "NC", "NC", "NC", "NC", "NC", "N…
$ StateDesc <chr> "North Carolina", "North Carolina", "North Carolin…
$ CountyName <chr> "Halifax", "Union", "New Hanover", "Sampson", "Car…
$ StateFIPS <int> 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37, 37…
$ CountyFIPS <int> 83, 179, 129, 163, 31, 183, 119, 87, 141, 21, 47, …
$ TotalPopulation <chr> "47,298", "256,452", "238,852", "59,601", "69,615"…
$ TotalPop18plus <chr> "37,177", "191,719", "197,058", "44,966", "58,070"…
$ FOODSTAMP_CrudePrev <dbl> 27.8, 8.5, 9.0, 19.0, 7.8, 7.4, 10.8, 9.3, 10.3, 9…
$ FOODSTAMP_Crude95ci <chr> "(23.8, 31.9)", "( 6.6, 10.6)", "( 6.8, 11.6)", "(…
$ FOODSTAMP_AdjPrev <dbl> 29.4, 8.7, 9.6, 19.8, 8.7, 7.4, 10.7, 10.2, 10.5, …
$ FOODSTAMP_Adj95CI <chr> "(25.1, 34.0)", "( 6.9, 11.0)", "( 7.3, 12.5)", "(…
From here, you could use this dataset for further analysis and visualization.
6.9 Save Outputs
Finally, let’s save our outputs. We will save the final joined dataset as a CSV, and the county-level maps as images.
## Save final_county_data
write.csv(final_county_data, "final_county_data.csv")
## Save county-level maps
tmap_save(PovP, filename = "PovP_map.png")
tmap_save(FqhcAvTmDr, filename = "FqhcAvTmDr_map.png")6.10 More Resources
To learn more about the OEPS ecosystem visit: https://oeps.healthyregions.org/
Visit the Geospatial Consortium & Community of Practice website to learn about past and upcoming workshops: https://gccp.healthyregions.org/
For more on the tmap package visit: https://r-tmap.github.io/tmap/reference/