We are given an image which shows two flight information for two flights. Information provided is just departing state, destination, number of flights that on time and number of flights that are delayed.
I plan to use ocr_data to extract the image I screen shotted from assignment. It’s part of tesseract package, it reads in an image and extract the data from it.
Code Chunk I tried
library(knitr)library(dplyr)
Attaching package: 'dplyr'
The following objects are masked from 'package:stats':
filter, lag
The following objects are masked from 'package:base':
intersect, setdiff, setequal, union
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag() masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
Simple data extraction and outputting to FlightDelays.csv gave me a lots of untidy data. All the data has been extracted to one single column and second and third column is just garbage (probably image formatting.
I will be using select and saving the useful data to a new dataframe, and then formatting and tidying it using dplyr and tidyr. I’ll also try to use proper formatting that I read in chapter 4.