Assignment 5A: Airline Delays

Author

Mubin Ejaz

Approach

We are given an image which shows two flight information for two flights. Information provided is just departing state, destination, number of flights that on time and number of flights that are delayed.

I plan to use ocr_data to extract the image I screen shotted from assignment. It’s part of tesseract package, it reads in an image and extract the data from it.

Code Chunk I tried

library(knitr)
library(dplyr)

Attaching package: 'dplyr'
The following objects are masked from 'package:stats':

    filter, lag
The following objects are masked from 'package:base':

    intersect, setdiff, setequal, union
library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ forcats   1.0.1     ✔ readr     2.2.0
✔ ggplot2   4.0.3     ✔ stringr   1.6.0
✔ lubridate 1.9.5     ✔ tibble    3.3.1
✔ purrr     1.2.1     ✔ tidyr     1.3.2
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(tesseract)
df<-ocr_data("FlightDelays.png")
write_csv(df, "FlightDelays.csv")

Simple data extraction and outputting to FlightDelays.csv gave me a lots of untidy data. All the data has been extracted to one single column and second and third column is just garbage (probably image formatting.
I will be using select and saving the useful data to a new dataframe, and then formatting and tidying it using dplyr and tidyr. I’ll also try to use proper formatting that I read in chapter 4.

Some of the functions I’ll be trying:

pivot_wider()
pivot_longer()
str_extract()
str_detect()
select()
filter()
mutate()
str_replace() / str_replace_all()
as.numeric()
separate()
arrange()
group_by()
summarise()


I did have chat suggest some useful functions I’d be needing to manipulate text/csv file. I predict I’ll be having a lots of fun.

Eventually, after I’ve plotted the given tidy data, I’d conclude my observations.