Assignment 4- Part 2

# load data and see what's in it

library(tidyverse)
── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
✔ dplyr     1.2.1     ✔ readr     2.2.0
✔ forcats   1.0.1     ✔ stringr   1.6.0
✔ ggplot2   4.0.3     ✔ tibble    3.3.1
✔ lubridate 1.9.5     ✔ tidyr     1.3.2
✔ purrr     1.2.2     
── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
✖ dplyr::filter() masks stats::filter()
✖ dplyr::lag()    masks stats::lag()
ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(nycflights23)

library(gridExtra)

Attaching package: 'gridExtra'

The following object is masked from 'package:dplyr':

    combine
data(flights)

flights%>% as_tibble()
# A tibble: 435,352 × 19
    year month   day dep_time sched_dep_time dep_delay arr_time sched_arr_time
   <int> <int> <int>    <int>          <int>     <dbl>    <int>          <int>
 1  2023     1     1        1           2038       203      328              3
 2  2023     1     1       18           2300        78      228            135
 3  2023     1     1       31           2344        47      500            426
 4  2023     1     1       33           2140       173      238           2352
 5  2023     1     1       36           2048       228      223           2252
 6  2023     1     1      503            500         3      808            815
 7  2023     1     1      520            510        10      948            949
 8  2023     1     1      524            530        -6      645            710
 9  2023     1     1      537            520        17      926            818
10  2023     1     1      547            545         2      845            852
# ℹ 435,342 more rows
# ℹ 11 more variables: arr_delay <dbl>, carrier <chr>, flight <int>,
#   tailnum <chr>, origin <chr>, dest <chr>, air_time <dbl>, distance <dbl>,
#   hour <dbl>, minute <dbl>, time_hour <dttm>
# making sure the carrier acronyms become their actual names

flight_names <- flights %>%
  left_join(airlines, by = "carrier")

flight_names
# A tibble: 435,352 × 20
    year month   day dep_time sched_dep_time dep_delay arr_time sched_arr_time
   <int> <int> <int>    <int>          <int>     <dbl>    <int>          <int>
 1  2023     1     1        1           2038       203      328              3
 2  2023     1     1       18           2300        78      228            135
 3  2023     1     1       31           2344        47      500            426
 4  2023     1     1       33           2140       173      238           2352
 5  2023     1     1       36           2048       228      223           2252
 6  2023     1     1      503            500         3      808            815
 7  2023     1     1      520            510        10      948            949
 8  2023     1     1      524            530        -6      645            710
 9  2023     1     1      537            520        17      926            818
10  2023     1     1      547            545         2      845            852
# ℹ 435,342 more rows
# ℹ 12 more variables: arr_delay <dbl>, carrier <chr>, flight <int>,
#   tailnum <chr>, origin <chr>, dest <chr>, air_time <dbl>, distance <dbl>,
#   hour <dbl>, minute <dbl>, time_hour <dttm>, name <chr>
p1 <- flight_names %>%
  filter(year == 2023) %>%
  ggplot(aes(air_time, distance, color = name)) +
  geom_point() +
  labs(
    title = "Flight Distance Compared to Air Time",
    x = "Air Time",
    y = "Distance",
    color = "Carrier name",
    caption = "data source nycflights23" 
  ) +
  theme(legend.text = element_text(size = 7))


p1
Warning: Removed 12534 rows containing missing values or values outside the scale range
(`geom_point()`).

  1. Write a brief paragraph that describes the visualization you have created and at least one aspect of the plot that you would like to highlight. The paragraph should be around 150-250 words as a good estimate.

The visualization I created is meant to show the relationship between flight distance and air time. I would like to highlight how there’s a clear positive correlation between air time and distance of a flight. However, there are some distinct outliers between some of the points. For example, most of distance around 2500 had air times around the 400 area, but one of points had an air time of 600. That could have been influenced by several different factors that couldn’t be included in the graph. For example, if there were any delays. Despite the graph showing a positive correlation between air time and distance, it is evident that there are other factors involved determining a flight’s air time since the correlation isn’t exactly perfect. Additionally, I would like to highlight that not many flights are at the top of the graph, while more flights are combined near the bottom. This could be one of the results of airlines being domestic vs. international. Since this graph is based off nyc, it’s expected that most flights would be closer, so domestic, compared to international with longer air times and distances.