data <- read.csv("Zomato Restaurant Dataset.csv")
set.seed(123)
sample_data <- data[sample(1:nrow(data),
round(nrow(data) / 2)), ]
nrow(data)
## [1] 9551
nrow(sample_data)
## [1] 4776
This analysis uses the Zomato Restaurant Dataset. The dataset contains information about restaurants, including location, cuisine, price range, ratings, votes, table booking, and online delivery. The original dataset contained 9,551 observations and 21 variables. I used a random sample of 4,776 observations, which is approximately half of the original dataset.
hist(sample_data$Aggregate.rating,
main = "Distribution of Restaurant Ratings",
xlab = "Restaurant Rating",
col = "lightblue")
### EDA 1: Restaurant Ratings
The histogram shows how restaurant ratings are distributed. Most ratings appear to be between 2 and 4. The ratings range from 0 to 4.9.
barplot(table(sample_data$Price.range),
main = "Restaurant Price Range",
xlab = "Price Range",
ylab = "Number of Restaurants",
col = "lightgreen")
### EDA 2: Restaurant Price Range
The bar chart compares the number of restaurants in each price range. Price range 1 has the most restaurants, while price range 4 has the fewest restaurants in the sample.