This report examines 2,600 food delivery orders placed throughout 2024. The data comes from Kaggle(https://www.kaggle.com/datasets/jayjoshi37/daily-food-delivery-orders-and-delivery-time?resource=download) and records each order separately, with details such as customer age, restaurant type, order value, delivery distance, delivery time, payment method and whether the order was delivered, delayed or cancelled. The purpose of this analysis is to uncover patterns in how and when customers order and what factors may affect delivery experience. The visualizations that follow look at the average order value by age group, order volume across months and days of the week, order outcomes by restaurant type and how delivery times change throughout the work week.
head(df, 5)
## order_id order_date customer_age restaurant_type order_value
## 1 1 2024-11-05 62 Indian 497.51
## 2 2 2024-08-20 35 Bakery 232.32
## 3 3 2024-02-28 34 Italian 540.82
## 4 4 2024-05-26 65 Cafe 1197.99
## 5 5 2024-09-21 40 Indian 947.03
## delivery_distance_km delivery_time_minutes payment_method
## 1 11.07 79 UPI
## 2 5.83 69 Wallet
## 3 3.61 70 Wallet
## 4 3.66 18 Card
## 5 12.08 57 UPI
## delivery_partner_rating order_status
## 1 3.9 Cancelled
## 2 2.7 Cancelled
## 3 3.4 Cancelled
## 4 4.6 Cancelled
## 5 4.9 Delayed
Pictured above are the first 5 lines of the data set which contains 10 variables. Each food order has a unique order_id along with the order_date it was placed in 2024. The customer_age variable gives the age of the person ordering which ranges from teenagers to customers who are 65+ years old. Restaurant_type identifies the cuisine the order came from which fall into one of six categories which are Bakery, Cafe, Chinese, Fast Food, Indian or Italian. The order_value variable is the total amount spent on the order, while delivery_distance_km measures how far the delivery drive was in kilometers and delivery_time_minutes records how many minutes the delivery took. The payment_method variale shows whether the customer paid by card, cash, digital wallet or UPI and delivery_partner_rating is the score the customer gave their driver on a 5 star scale. Finally, order_status indicates whether the order was delivered, delayed or cancelled. In the code, there are three more created variables which are age_group, which sorts customer_age into ranges, and month and dayoftheweek.
summary(df)
## order_id order_date customer_age restaurant_type
## Min. : 1.0 Length:2600 Min. :18.00 Length:2600
## 1st Qu.: 650.8 Class :character 1st Qu.:29.00 Class :character
## Median :1300.5 Mode :character Median :41.00 Mode :character
## Mean :1300.5 Mean :41.49
## 3rd Qu.:1950.2 3rd Qu.:54.00
## Max. :2600.0 Max. :65.00
## order_value delivery_distance_km delivery_time_minutes payment_method
## Min. : 150.9 Min. : 0.500 Min. :15.00 Length:2600
## 1st Qu.: 406.4 1st Qu.: 4.207 1st Qu.:32.00 Class :character
## Median : 667.6 Median : 7.965 Median :51.00 Mode :character
## Mean : 670.3 Mean : 7.887 Mean :51.75
## 3rd Qu.: 927.5 3rd Qu.:11.590 3rd Qu.:70.00
## Max. :1199.8 Max. :14.990 Max. :90.00
## delivery_partner_rating order_status
## Min. :2.50 Length:2600
## 1st Qu.:3.10 Class :character
## Median :3.80 Mode :character
## Mean :3.75
## 3rd Qu.:4.40
## Max. :5.00
The summary above gives a quick overview of the numeric variables in the data set. Customers range from 18 to 65 years old with an average of about 41. Order values run from $150.90 to $1,199.80 with a mean of $670.30. Delivery distances span from half a kilometer to just under 15 kilometers. Delivery time ranges from 15 to 90 minutes with an average of about 52. Delivery partner ratings fall between 2.5 and 5 stars averaging 3.75. One notable pattern is that the mean and median are almost identical for every single variable. This means that the values are spread very evenly with no extreme outliers pulling averages. This even spread will show up again and again in the visualizations.
Below are the visualizations and findings! Click through the tabs to see each.
avg_val_age <- df %>%
mutate(age_group = cut(customer_age, breaks = c(17, 25, 35, 45, 55, 65), labels = c("18-25", "26-35", "36-45","46-55", "56-65"))) %>%
group_by(age_group, restaurant_type) %>%
summarise(order_value = round(mean(order_value), 0), .groups = "drop")
ggplot(avg_val_age, aes(x = age_group, y = order_value, fill = restaurant_type)) +
geom_bar(stat = "identity", position = "dodge") +
labs(title = "Average Order Value by Age Group and Restaurant Type", x = "Age Group", y = "Average Order Value", fill = "Restaurant Type") +
theme_light() +
theme(plot.title = element_text(hjust = 0.5)) +
scale_fill_brewer(palette = "Paired") +
scale_y_continuous(labels = scales::dollar) +
geom_text(aes(label = scales::dollar(order_value)), position = position_dodge(width = 0.9), vjust = -0.5, size = 2)
Average order values are very consistent across age groups and restaurant types with every single bar falling between $620 and $735. The highest average in the chart is Bakery orders from 18 to 25 year olds at $735 and the lowest average is Indian orders from 46 to 55 year olds at $620. A different cuisine leads each age group. Bakery leads the 18 to 25 year olds, Fast Food for the 26 to 35 year olds, Italian for the 36 to 45 year olds, Cafe for the 46 to 55 year olds and Chinesse for the 56 to 65 year olds. This shows that no single restaurant type is consistenlty the most expensive. Overall the chart suggests that customer age and cuisine have a small effect on how much people spend per order.
monthly_counts <- df %>%
mutate(month = month(ymd(order_date))) %>%
group_by(month) %>%
summarise(n = n() )
ggplot(monthly_counts, aes(x = month, y = n)) +
geom_line() +
geom_point(shape = 1, size = 3, color = "red") +
geom_text(aes(label = n), vjust = -1.5, size = 3) +
scale_y_continuous(expand = expansion(mult = c(0.05, 0.12))) +
scale_x_continuous(breaks = 1:12, labels = month.abb) +
labs(title = "Food Delivery Orders By Month", x = "Month", y = "Order Count") +
theme_light() +
theme(plot.title = element_text(hjust = 0.5))
Monthly order volume stays within a fairly narrow range with a low of 189 orders in April to a high of 250 orders in October. The line moves very much in a zigzag pattern with sharp spikes in May, August and October with 239, 237 and 250 orders respectively. All were followed by drops the next month as well. The second half of the year was slightly busier with 1,339 orders from July to December compared to 1,261 from January through June. October and November together form the strongest 2 month run of the year. The y-axis starts near 190 rather than zero, so the swings are much more pronounced because of how focused in the graph is. The gap between the slowst and the busiest month is only about 60 orders, which in turn suggest that there is no strong seasonal pattern in food delivery demand.
days_df <- df %>%
mutate(month = month(ymd(order_date), label = TRUE),
dayoftheweek = as.character(wday(ymd(order_date), label = TRUE))) %>%
group_by(month, dayoftheweek) %>%
summarise(n = n(), .groups = "drop")
my_levels <- c("Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun")
days_df$dayoftheweek <- factor(days_df$dayoftheweek, levels = my_levels)
breaks <- c(seq(0, max(days_df$n), by = 4))
g <- ggplot(days_df, aes(x = month, y = dayoftheweek, fill = n)) +
geom_tile(color = "black") +
geom_text(aes(label = n)) +
coord_equal(ratio = 1) +
labs(title = "Heatmap of Food Delivery Orders by Day of the Week",
x = "Month",
y = "Day of the Week",
fill = "Order Count") +
theme_minimal() +
theme(plot.title = element_text(hjust = 0.5)) +
scale_y_discrete(limits = rev(levels(days_df$dayoftheweek))) +
scale_fill_continuous(low = "white", high = "red", breaks = breaks) +
guides(fill = guide_legend(reverse = TRUE, override.aes = list(colour = "black")))
ggplotly(g, tooltip = c("n", "month", "dayoftheweek")) %>%
style(hoverlabel = list(bgcolor = "white"))
The heat map breaks the 2,600 orders down by month and day of the week with the darker red tiles marking busier days while the whiter tiles mark less busy days. Most tiles fall between 25 and 35, but a few stand out. In particular Thursdays in October and Fridays in August at 47 orders each, Sundays in June at 45 and Wednesdays in November at 44. The lightest tiles are Wednesdays in April and Sundays in February, at only 20 orders each which follows along of the theme from the previous chart that April is the least busy month. The weekends seem to be some of the less busy days, which comes as a surprise because most people would associate food delivery with the weekend, especially Sunday. Overall, the colors are scattered throughout very evenly, which sticks with the theme of the data being very evenly spread out.
status_rest_df <- df %>%
group_by(restaurant_type, order_status) %>%
summarise(n = n(), .groups = "drop")
plot_ly(textposition = "inside", labels = ~order_status, values = ~n) %>%
add_pie(data = status_rest_df[status_rest_df$restaurant_type == "Bakery",],
name = "Bakery", title = "Bakery", sort = FALSE, domain = list(row = 0, column = 0)) %>%
add_pie(data = status_rest_df[status_rest_df$restaurant_type == "Cafe",],
name = "Cafe", title = "Cafe", sort = FALSE, domain = list(row = 0, column = 1)) %>%
add_pie(data = status_rest_df[status_rest_df$restaurant_type == "Chinese",],
name = "Chinese", title = "Chinese", sort = FALSE, domain = list(row = 0, column = 2)) %>%
add_pie(data = status_rest_df[status_rest_df$restaurant_type == "Fast Food",],
name = "Fast Food", title = "Fast Food", sort = FALSE, domain = list(row = 1, column = 0)) %>%
add_pie(data = status_rest_df[status_rest_df$restaurant_type == "Indian",],
name = "Indian", title = "Indian", sort = FALSE, domain = list(row = 1, column = 1)) %>%
add_pie(data = status_rest_df[status_rest_df$restaurant_type == "Italian",],
name = "Italian", title = "Italian", sort = FALSE, domain = list(row = 1, column = 2)) %>%
layout(title = "Order Status by Restaurant Type", showlegend = TRUE, grid = list(rows = 2, columns = 3))
Each pie splits a restaurant type’s order into delivered, delayed or cancelled and shockingly every single cuisine lands close to an even three way split. Indian has the best track record, with the highest delivered rate of 38.5% and the lowest cancellation rate with 30.9%. This delivery performance is followed by Cafe and Chinese who had delivery rates of 36.7% and 35% respectively. Italian performs the worst with the highest cancellation rate of 37.3% and the lowest delivery rate of 30.9%. Italian and Bakery are the only two types where cancellations outnumber successful deliveries. Delay rates are the most consistent across all graphs staying between 30.6% and 33.8% for every cuisine which shows that delays are not tied to a specific type of cuisine. The shocking discovery from the graphs is that almost 2/3 of all orders were either delayed or cancelled. This would cause almost certain business failure for a real company.
days_df <- df %>%
select(order_date, restaurant_type, delivery_time_minutes) %>%
mutate(dayoftheweek = weekdays(ymd(order_date), abbreviate = TRUE)) %>%
group_by(restaurant_type, dayoftheweek) %>%
summarise(avg_time = round(mean(delivery_time_minutes), 1), .groups = "keep") %>%
data.frame()
str(days_df)
## 'data.frame': 42 obs. of 3 variables:
## $ restaurant_type: chr "Bakery" "Bakery" "Bakery" "Bakery" ...
## $ dayoftheweek : chr "Fri" "Mon" "Sat" "Sun" ...
## $ avg_time : num 57.1 54.6 54.8 48.5 50.9 49.7 49.9 51 54.4 48 ...
days_df$restaurant_type <- as.factor(days_df$restaurant_type)
day_order <- factor(days_df$dayoftheweek, level = c("Mon", "Tue", "Wed", "Thu", "Fri", "Sat", "Sun"))
ggplot(days_df, aes(x = day_order, y = avg_time, group = restaurant_type)) +
geom_line(aes(color = restaurant_type), linewidth = 3) +
labs(title = "Avg Delivery Time by Day and by Resturant Type", x = "Days of the Weeks", "y = Avg Delivery Time in Minutes") +
theme_light() +
theme(plot.title = element_text(hjust = 0.5)) +
geom_point(shape = 21, size = 5, color = "black", fill = "white") +
scale_color_brewer(palette = "Paired", name = "Restaurant Type", guide = guide_legend(reverse = TRUE))
Average delivery times stay within a fairly tight window from about 47 to 59 minutes, but the lines constantly cross over each other. No restaurant type is consistently fast or slow. The largest spike happens on Thursdays when Fast Food and Chinese both jump to nearly 59 minutes, which is the slowest average in the entire chart. Indian also peaks that day to 56 minutes as well. This connects back into the heatmap where Thursday had the most orders out of any day, because a busier day would equal longer delivery times. Bakery follows its own pattern, peaking on Friday at 57.1 minutes before dropping to 48.5 on Sunday. Italian has some of the fastest delivery times early in the week, bottoming out near 47 minutes on Tuesday. Weekends are not noticeably slower than weekdays which does match the earlier finding that weekends were actually less busy. Overall, the difference between the fastest and the slowest averages is only around 12 minutes, so the day of the week and the type of cuisine have a small effect on delivery time.
This report explored 2,600 food delivery orders from 2024 through five different visualizations. Across every visualization, the most consistent finding was how evenly the data is spread. Average order values stayed between roughly $620 and $735 for almost every age group and cuisine, monthly order counts varied by only about 60 orders between the slowest and the busiest month and order outcomes were split close to evenly between delivered, delayed and cancelled. A few patters did stand out, October was the busiest month, weekdays saw slightly more orders than the weekends and Thursday combined the highest order volume with some of the slowest delivery times. Indian restaurants had the best delivery percentage while Italian restaurants had the highest cancellation rate. A shocking result was that roughly 2/3 of all orders were either cancelled or delayed which is a level of performance that a real world delivery platform would be unable to survive. Because this data set is synthetic, aka it is made up and is not real world data, the results are expected as this data set was created evenly. A real world data set would show stronger patterns in all of the visualizations that were ran throughout this report.