This report provides a descriptive analysis of a dataset containing vintage-car listings. Each observation represents a vintage car included in the dataset. The dataset contains information about car names, descriptions, model years, countries of origin, total mileage, and listed prices. The purpose of this analysis is to examine how listed prices and mileage are distributed, explore whether mileage appears related to listed price, and compare vintage-car listings and prices across countries of origin. Five different visualization types are used to identify important patterns in the data.
The dataset contains 100 vintage-car listings and 6 variables.
The variables used in this analysis are:
Car_names: The name or model of the vintage carDescriptions: A written description of the carYear: The model year of the carCountry: The country of originTotal_distance: Mileage in kilometersPrice: Listed price in dollarshead(VintageCar_data)
The table below summarizes the primary numerical variables in the vintage-car dataset.
summary_stats <- data.frame(
Variable = c("Model Year", "Mileage (km)", "Listed Price ($)"),
Mean = c(
mean(VintageCar_data$Year, na.rm = TRUE),
mean(VintageCar_data$Total_distance, na.rm = TRUE),
mean(VintageCar_data$Price, na.rm = TRUE)
),
Median = c(
median(VintageCar_data$Year, na.rm = TRUE),
median(VintageCar_data$Total_distance, na.rm = TRUE),
median(VintageCar_data$Price, na.rm = TRUE)
),
Standard_Deviation = c(
sd(VintageCar_data$Year, na.rm = TRUE),
sd(VintageCar_data$Total_distance, na.rm = TRUE),
sd(VintageCar_data$Price, na.rm = TRUE)
),
Minimum = c(
min(VintageCar_data$Year, na.rm = TRUE),
min(VintageCar_data$Total_distance, na.rm = TRUE),
min(VintageCar_data$Price, na.rm = TRUE)
),
Maximum = c(
max(VintageCar_data$Year, na.rm = TRUE),
max(VintageCar_data$Total_distance, na.rm = TRUE),
max(VintageCar_data$Price, na.rm = TRUE)
)
)
summary_stats[, 2:6] <- round(summary_stats[, 2:6], 2)
knitr::kable(
summary_stats,
caption = "Descriptive Statistics for Key Numeric Variables"
)
| Variable | Mean | Median | Standard_Deviation | Minimum | Maximum |
|---|---|---|---|---|---|
| Model Year | 1968.92 | 1968.5 | 12.16 | 1951 | 1989 |
| Mileage (km) | 120400.61 | 127066.5 | 49728.92 | 24400 | 233990 |
| Listed Price ($) | 174493.46 | 185103.0 | 93126.74 | 18457 | 277672 |
The listed prices and mileage values vary substantially across the sample. This indicates that the dataset contains vehicles with different levels of use and a wide range of listed values. The descriptive statistics provide context for interpreting the visualizations that follow.
price_histogram <- ggplot(VintageCar_data, aes(x = Price)) +
geom_histogram(
binwidth = 25000,
fill = "pink",
color = "red",
na.rm = TRUE
) +
scale_x_continuous(labels = dollar) +
labs(
title = "Distribution of Vintage Car Prices",
subtitle = "How are listed prices distributed in the dataset?",
x = "Listed Price ($)",
y = "Number of Cars"
) +
theme_minimal()
price_histogram
This histogram shows the distribution of listed prices for the vintage cars in the dataset. The horizontal axis represents listed price in dollars, while the vertical axis shows the number of cars within each price range. The prices are spread across a wide range, from roughly $20,000 to nearly $300,000. The largest concentration of vehicles falls in the highest price range, approximately $260,000 to $290,000, with about 32 cars. The other price ranges each contain approximately 16 to 18 vehicles. This indicates that the dataset contains a substantial group of high-priced vintage cars, along with several similarly sized groups of lower- and middle-priced listings. Overall, the distribution is not centered around one typical price level; instead, it is clustered across several distinct price ranges.
mileage_price_plot <- ggplot(VintageCar_data, aes(x = Total_distance, y = Price)) +
geom_point(
color = "hotpink",
alpha = 0.65,
na.rm = TRUE
) +
geom_smooth(
method = "lm",
se = TRUE,
color = "purple",
na.rm = TRUE
) +
scale_x_continuous(labels = comma) +
scale_y_continuous(labels = dollar) +
labs(
title = "Vintage Car Price by Mileage",
subtitle = "Do higher-mileage cars tend to have lower listed prices?",
x = "Mileage (km)",
y = "Listed Price ($)"
) +
theme_minimal()
mileage_price_plot
## `geom_smooth()` using formula = 'y ~ x'
This scatterplot shows the relationship between a vintage car’s mileage and its listed price. Each point represents one vehicle. The downward-sloping purple regression line suggests that higher-mileage cars generally tend to have lower listed prices. Several cars with more than 160,000 kilometers are priced below $50,000 while a number of lower-mileage cars are listed above $200,000. However, the points are widely scattered, so mileage does not fully explain price. The gray confidence band is wide, particularly at higher mileage levels, which means the relationship should be interpreted cautiously because the sample is relatively small.
country_bar_chart <- ggplot(
VintageCar_data,
aes(x = reorder(Country, Country, FUN = length))
) +
geom_bar(
fill = "pink",
color = "hotpink",
na.rm = TRUE
) +
coord_flip() +
labs(
title = "Vintage Car Listings by Country of Origin",
subtitle = "Which countries are represented most often in the dataset?",
x = "Country of Origin",
y = "Number of Cars"
) +
theme_minimal()
country_bar_chart
This bar chart shows the number of vintage-car listings by country of origin. It answers the question, “Which countries are represented most often among the vintage-car listings in this dataset?” Germany is the most represented country, with approximately 49 listings. The United Kingdom is the second most represented country, with about 18 listings. Japan and Italy each have about 16 listings, while the United States is represented by only one listing. This shows that the dataset is heavily concentrated in German vehicles. Because Germany makes up such a large share of the sample and the United States has only one observation, results involving country, especially the price comparison in the boxplot should be interpreted with caution. The unequal country counts could affect how representative the results are for each country.
price_country_boxplot <- ggplot(
VintageCar_data,
aes(
x = reorder(Country, Price, FUN = median, na.rm = TRUE),
y = Price
)
) +
geom_boxplot(
fill = "magenta",
color = "purple",
na.rm = TRUE
) +
coord_flip() +
scale_y_continuous(labels = dollar) +
labs(
title = "Vintage Car Price Distribution by Country of Origin",
subtitle = "How do listed prices differ across countries of origin?",
x = "Country of Origin",
y = "Listed Price ($)"
) +
theme_minimal()
price_country_boxplot
This boxplot compares vintage-car listed prices across countries of origin. Germany has the clearest distribution because it has the most observations in the dataset, while the other countries have too few listings to create full boxplots. Germany’s listings show substantial price variation, ranging from relatively lower-priced vehicles to values above $250,000. The United Kingdom also has variation in price, including both lower- and higher-priced listings. Because Italy, the United States, and Japan have very small sample sizes, conclusions about their typical prices should be interpreted cautiously.
mileage_density_plot <- ggplot(
VintageCar_data,
aes(x = Total_distance)
) +
geom_density(
fill = "lightblue",
color = "navy",
alpha = 0.6,
na.rm = TRUE
) +
scale_x_continuous(labels = comma) +
scale_y_continuous(
labels = scales::label_number(accuracy = 0.000001)
) +
labs(
title = "Distribution of Vintage Car Mileage",
subtitle = "How is mileage distributed among the listed vintage cars?",
x = "Mileage (km)",
y = "Density"
) +
theme_minimal()
mileage_density_plot
This density plot shows how total mileage is distributed among the listed vintage cars. The horizontal axis shows mileage in kilometers, while the vertical axis shows density, which indicates where values are concentrated rather than the raw number of cars. The distribution has several distinct peaks, meaning it is multimodal rather than centered around one typical mileage value. The largest concentration appears around 110,000 to 145,000 kilometers, while smaller clusters appear at lower and higher mileage ranges. This suggests that the dataset includes several groups of vintage cars with different mileage levels.
This report provided a descriptive analysis of vintage-car listings using listed price, mileage, country of origin, and model year information. The five visualizations showed that listed prices vary widely, ranging from lower-priced vehicles near $20,000 to high-priced listings close to $300,000. The histogram also showed that the largest group of cars falls within the highest price range, while several other price ranges contain similar numbers of listings. The scatterplot suggested a negative relationship between mileage and listed price. In general, cars with higher mileage tended to have lower listed prices, although the points were widely spread around the regression line. This means that mileage is related to price in this sample but does not fully explain it. Factors such as the car’s year, country of origin, condition, model, and rarity may also influence its listed value. The country bar chart showed that Germany was by far the most represented country in the dataset, followed by the United Kingdom, Japan, and Italy. Because the number of listings differed substantially across countries, the price boxplot should be interpreted carefully. Germany had enough observations to show substantial price variation, while several other countries had too few observations to form a strong comparison. Finally, the mileage density plot showed that mileage is not concentrated around one single value. Instead, the dataset contains several mileage clusters, with the largest concentration appearing around the middle-to-higher mileage range. Overall, the analysis shows that vintage-car prices and mileage vary substantially across the dataset and that multiple vehicle characteristics are likely related to listed price.
Note that the echo = FALSE parameter was added to the
code chunk to prevent printing of the R code that generated the
plot.