##
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ forcats 1.0.0 ✔ stringr 1.5.1
## ✔ lubridate 1.9.3 ✔ tibble 3.2.1
## ✔ purrr 1.0.2 ✔ tidyr 1.3.1
## ✔ readr 2.1.5
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
## Package 'qcc' version 2.7
##
## Type 'citation("qcc")' for citing this R package in publications.
In this part we will be reading the dataset. The chosen dataset is Restaurant Dataset, which contains real information about different restaurants and the idea that the author has is to develop a ML model to classify restaurants based on their cuisines.
df <- read.csv('Dataset.csv') #Reading of the dataset.
head(df)
## Restaurant.ID Restaurant.Name Country.Code City
## 1 6317637 Le Petit Souffle 162 Makati City
## 2 6304287 Izakaya Kikufuji 162 Makati City
## 3 6300002 Heat - Edsa Shangri-La 162 Mandaluyong City
## 4 6318506 Ooma 162 Mandaluyong City
## 5 6314302 Sambo Kojin 162 Mandaluyong City
## 6 18189371 Din Tai Fung 162 Mandaluyong City
## Address
## 1 Third Floor, Century City Mall, Kalayaan Avenue, Poblacion, Makati City
## 2 Little Tokyo, 2277 Chino Roces Avenue, Legaspi Village, Makati City
## 3 Edsa Shangri-La, 1 Garden Way, Ortigas, Mandaluyong City
## 4 Third Floor, Mega Fashion Hall, SM Megamall, Ortigas, Mandaluyong City
## 5 Third Floor, Mega Atrium, SM Megamall, Ortigas, Mandaluyong City
## 6 Ground Floor, Mega Fashion Hall, SM Megamall, Ortigas, Mandaluyong City
## Locality
## 1 Century City Mall, Poblacion, Makati City
## 2 Little Tokyo, Legaspi Village, Makati City
## 3 Edsa Shangri-La, Ortigas, Mandaluyong City
## 4 SM Megamall, Ortigas, Mandaluyong City
## 5 SM Megamall, Ortigas, Mandaluyong City
## 6 SM Megamall, Ortigas, Mandaluyong City
## Locality.Verbose Longitude
## 1 Century City Mall, Poblacion, Makati City, Makati City 121.0275
## 2 Little Tokyo, Legaspi Village, Makati City, Makati City 121.0141
## 3 Edsa Shangri-La, Ortigas, Mandaluyong City, Mandaluyong City 121.0568
## 4 SM Megamall, Ortigas, Mandaluyong City, Mandaluyong City 121.0565
## 5 SM Megamall, Ortigas, Mandaluyong City, Mandaluyong City 121.0575
## 6 SM Megamall, Ortigas, Mandaluyong City, Mandaluyong City 121.0563
## Latitude Cuisines Average.Cost.for.two
## 1 14.56544 French, Japanese, Desserts 1100
## 2 14.55371 Japanese 1200
## 3 14.58140 Seafood, Asian, Filipino, Indian 4000
## 4 14.58532 Japanese, Sushi 1500
## 5 14.58445 Japanese, Korean 1500
## 6 14.58376 Chinese 1000
## Currency Has.Table.booking Has.Online.delivery Is.delivering.now
## 1 Botswana Pula(P) Yes No No
## 2 Botswana Pula(P) Yes No No
## 3 Botswana Pula(P) Yes No No
## 4 Botswana Pula(P) No No No
## 5 Botswana Pula(P) Yes No No
## 6 Botswana Pula(P) No No No
## Switch.to.order.menu Price.range Aggregate.rating Rating.color Rating.text
## 1 No 3 4.8 Dark Green Excellent
## 2 No 3 4.5 Dark Green Excellent
## 3 No 4 4.4 Green Very Good
## 4 No 4 4.9 Dark Green Excellent
## 5 No 4 4.8 Dark Green Excellent
## 6 No 3 4.4 Green Very Good
## Votes
## 1 314
## 2 591
## 3 270
## 4 365
## 5 229
## 6 336
For this study I am going to set the following goals:
Clean the dataset to obtain clear food types and subtypes of foods
Make an exploratory analysis of the characteristics that define each of the food types
First of all we are going to be cleaning the dataset. As one column is in commas, we are going to split it in different columns:
df <- df %>%
separate(Cuisines, into = c("Cuisines_Type", "Cuisines_Subtype_L1", "Cuisines_Subtype_L2"), sep = ",")
## Warning: Expected 3 pieces. Additional pieces discarded in 864 rows [3, 8, 16, 20, 21,
## 22, 50, 258, 289, 465, 468, 470, 474, 476, 572, 573, 577, 585, 604, 605, ...].
## Warning: Expected 3 pieces. Missing pieces filled with `NA` in 6847 rows [2, 4, 5, 6, 7,
## 10, 11, 13, 14, 15, 17, 18, 23, 24, 25, 26, 27, 28, 29, 30, ...].
head(df)
## Restaurant.ID Restaurant.Name Country.Code City
## 1 6317637 Le Petit Souffle 162 Makati City
## 2 6304287 Izakaya Kikufuji 162 Makati City
## 3 6300002 Heat - Edsa Shangri-La 162 Mandaluyong City
## 4 6318506 Ooma 162 Mandaluyong City
## 5 6314302 Sambo Kojin 162 Mandaluyong City
## 6 18189371 Din Tai Fung 162 Mandaluyong City
## Address
## 1 Third Floor, Century City Mall, Kalayaan Avenue, Poblacion, Makati City
## 2 Little Tokyo, 2277 Chino Roces Avenue, Legaspi Village, Makati City
## 3 Edsa Shangri-La, 1 Garden Way, Ortigas, Mandaluyong City
## 4 Third Floor, Mega Fashion Hall, SM Megamall, Ortigas, Mandaluyong City
## 5 Third Floor, Mega Atrium, SM Megamall, Ortigas, Mandaluyong City
## 6 Ground Floor, Mega Fashion Hall, SM Megamall, Ortigas, Mandaluyong City
## Locality
## 1 Century City Mall, Poblacion, Makati City
## 2 Little Tokyo, Legaspi Village, Makati City
## 3 Edsa Shangri-La, Ortigas, Mandaluyong City
## 4 SM Megamall, Ortigas, Mandaluyong City
## 5 SM Megamall, Ortigas, Mandaluyong City
## 6 SM Megamall, Ortigas, Mandaluyong City
## Locality.Verbose Longitude
## 1 Century City Mall, Poblacion, Makati City, Makati City 121.0275
## 2 Little Tokyo, Legaspi Village, Makati City, Makati City 121.0141
## 3 Edsa Shangri-La, Ortigas, Mandaluyong City, Mandaluyong City 121.0568
## 4 SM Megamall, Ortigas, Mandaluyong City, Mandaluyong City 121.0565
## 5 SM Megamall, Ortigas, Mandaluyong City, Mandaluyong City 121.0575
## 6 SM Megamall, Ortigas, Mandaluyong City, Mandaluyong City 121.0563
## Latitude Cuisines_Type Cuisines_Subtype_L1 Cuisines_Subtype_L2
## 1 14.56544 French Japanese Desserts
## 2 14.55371 Japanese <NA> <NA>
## 3 14.58140 Seafood Asian Filipino
## 4 14.58532 Japanese Sushi <NA>
## 5 14.58445 Japanese Korean <NA>
## 6 14.58376 Chinese <NA> <NA>
## Average.Cost.for.two Currency Has.Table.booking Has.Online.delivery
## 1 1100 Botswana Pula(P) Yes No
## 2 1200 Botswana Pula(P) Yes No
## 3 4000 Botswana Pula(P) Yes No
## 4 1500 Botswana Pula(P) No No
## 5 1500 Botswana Pula(P) Yes No
## 6 1000 Botswana Pula(P) No No
## Is.delivering.now Switch.to.order.menu Price.range Aggregate.rating
## 1 No No 3 4.8
## 2 No No 3 4.5
## 3 No No 4 4.4
## 4 No No 4 4.9
## 5 No No 4 4.8
## 6 No No 3 4.4
## Rating.color Rating.text Votes
## 1 Dark Green Excellent 314
## 2 Dark Green Excellent 591
## 3 Green Very Good 270
## 4 Dark Green Excellent 365
## 5 Dark Green Excellent 229
## 6 Green Very Good 336
Now let’s explore if there are nulls:
colSums(is.na(df))
## Restaurant.ID Restaurant.Name Country.Code
## 0 0 0
## City Address Locality
## 0 0 0
## Locality.Verbose Longitude Latitude
## 0 0 0
## Cuisines_Type Cuisines_Subtype_L1 Cuisines_Subtype_L2
## 0 3403 6847
## Average.Cost.for.two Currency Has.Table.booking
## 0 0 0
## Has.Online.delivery Is.delivering.now Switch.to.order.menu
## 0 0 0
## Price.range Aggregate.rating Rating.color
## 0 0 0
## Rating.text Votes
## 0 0
We found that only subtypes are those that contain nulls. Then, no additional cleaning is required, as we are not going to be using the subtypes.
Let’s observe the count of cuisines in our dataset. If we focus only in the top 10, we will find that the preponderant cuisines are asian (Northern indian, Chinese, South Indian or Mithal). This suggests that the location of the study might be in an asian country, but it is going to be analyzed further in the following lines. Additionally, there are other cuisines, that are generally popular, such as Fast food, Bakery or Cafe.
As expected, most of the businesses in the analysis are from asia, specifically from india, as it is observed in the following visualization:
Now, in terms of rating, we can observe that most foods that were the most popular in terms of quantity, are not in the top 10 in terms of rating. Focusing again in the top 10, the cuisine types are the expected to be popular around the globe, such as american food, Italian, Burger or Cafe. Rating of Asian cuisines are below 2.5.
# Determine the top N cuisines
top_cuisines <- df %>%
count(Cuisines_Type, sort = TRUE) %>%
top_n(20, n) # for example, the top 20 cuisines
# Filter the dataframe for only the top cuisines
df_top_cuisines <- df %>%
filter(Cuisines_Type %in% top_cuisines$Cuisines_Type)
# Plot average rating for only the top cuisines
ggplot(df_top_cuisines, aes(x = reorder(Cuisines_Type, Aggregate.rating), y = Aggregate.rating)) +
geom_bar(stat = "summary", fun = "mean") +
coord_flip() +
labs(x = "Cuisine", y = "Average Rating", title = "Average Rating by Top 20 Cuisines")
If we go further analyzing the ratings, we will observe that there are some that contain outliers such as burgers, american and contitnental foods. Also those that are in the top 7 present their ratings distributed between 3.5 and 4.1, whereas the others are between 0 and 3.6.
# Identifying the top 20 cuisines based on count
top_cuisines <- df %>%
filter(!is.na(Cuisines_Type)) %>%
count(Cuisines_Type) %>%
top_n(20, n) %>%
pull(Cuisines_Type)
# Calculating median ratings for top cuisines
median_ratings <- df %>%
filter(Cuisines_Type %in% top_cuisines) %>%
group_by(Cuisines_Type) %>%
summarise(MedianRating = median(`Aggregate.rating`, na.rm = TRUE)) %>%
arrange(desc(MedianRating)) %>%
pull(Cuisines_Type)
# Filtering the data for the top 20 cuisines
df_top_cuisines <- df %>%
filter(Cuisines_Type %in% top_cuisines)
# Creating a boxplot for the aggregate ratings of the top 20 cuisines, ordered by median
ggplot(df_top_cuisines, aes(x = fct_reorder(Cuisines_Type, `Aggregate.rating`, .fun = median, .desc = TRUE), y = `Aggregate.rating`)) +
geom_boxplot() +
theme(axis.text.x = element_text(angle = 90, vjust = 0.5, hjust=1)) +
labs(x = "Cuisine Type", y = "Aggregate Rating", title = "Boxplot of Aggregate Ratings for Top 20 Cuisines Ordered by Median") +
theme(plot.title = element_text(hjust = 0.5))
Finally, let’s do a final exploration of our numeric variables. From the ratings, we could observe that ranges between 0 and 5. Price ranges are from 1 to 4, is an integer and is categorical.
db_num<-df %>% dplyr::select(where(is.numeric))
db_cat<-df %>% dplyr::select(where(is.factor))
par(mfrow=c(2,2))
for (i in 1:ncol(db_num)){
hist(db_num[[i]], main=paste("Plot ", colnames(db_num[i])), xlab = paste("Values Plot",i))
box(lty = "solid")
}