Bianca Bosco s3999104
Social media is crucial for modern marketing, with platforms like Facebook playing a key role in brand building and customer engagement. This analysis uses a dataset from Moro et al. (2016), which investigates the performance metrics of Facebook posts by a renowned cosmetics brand in 2014. The dataset includes 500 posts with 19 attributes, focusing on pre-publication features and post-engagement metrics. The goal is to identify key predictors of post performance through statistical analysis, helping to optimise social media strategies for enhanced reach and interaction.
Important Variables:
Factors and Levels:
Scale of Numeric Variables:
To understand the key metrics of the dataset, we perform descriptive statistical analysis on the numerical variables. The important variables in our dataset include:
These variables are essential for optimising social media strategies, as they provide insights into content performance, user engagement, and overall effectiveness in reaching and interacting with the target audience.
The key numerical variables in the dataset show significant variability:
# Summary statistics for numerical variables
summary_stats <- facebook_data_scaled %>%
select(Total_Reach, Total_Impressions, Engaged_Users, Post_Consumers, Post_Consumptions,
Impressions_Liked_Page, Reach_Liked_Page, Engaged_Liked_Page, Comment, Like, Share, Total_Interactions) %>%
summary()
# Display summary statistics
knitr::kable(summary_stats, caption = "Summary Statistics for Numerical Variables")| Total_Reach | Total_Impressions | Engaged_Users | Post_Consumers | Post_Consumptions | Impressions_Liked_Page | Reach_Liked_Page | Engaged_Liked_Page | Comment | Like | Share | Total_Interactions | |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Min. :-0.60427 | Min. :-0.37965 | Min. :-0.9292 | Min. :-0.8983 | Min. :-0.70576 | Min. :-0.27215 | Min. :-0.8318 | Min. :-0.98501 | Min. :-0.35524 | Min. :-0.55222 | Min. :-0.6392 | Min. :-0.56060 | |
| 1st Qu.:-0.46874 | 1st Qu.:-0.31188 | 1st Qu.:-0.5344 | 1st Qu.:-0.5300 | 1st Qu.:-0.45497 | 1st Qu.:-0.21378 | 1st Qu.:-0.5751 | 1st Qu.:-0.51540 | 1st Qu.:-0.30824 | 1st Qu.:-0.37651 | 1st Qu.:-0.4047 | 1st Qu.:-0.37196 | |
| Median :-0.38290 | Median :-0.26928 | Median :-0.3005 | Median :-0.2815 | Median :-0.28138 | Median :-0.17702 | Median :-0.4108 | Median :-0.32251 | Median :-0.21423 | Median :-0.24088 | Median :-0.1937 | Median :-0.23310 | |
| Mean : 0.00000 | Mean : 0.00000 | Mean : 0.0000 | Mean : 0.0000 | Mean : 0.00000 | Mean : 0.00000 | Mean : 0.0000 | Mean : 0.00000 | Mean : 0.00000 | Mean : 0.00000 | Mean : 0.0000 | Mean : 0.00000 | |
| 3rd Qu.:-0.03418 | 3rd Qu.:-0.09533 | 3rd Qu.: 0.1369 | 3rd Qu.: 0.1862 | 3rd Qu.: 0.02644 | 3rd Qu.:-0.02952 | 3rd Qu.: 0.1788 | 3rd Qu.: 0.07221 | 3rd Qu.:-0.02621 | 3rd Qu.: 0.02729 | 3rd Qu.: 0.1227 | 3rd Qu.: 0.04462 | |
| Max. : 7.29379 | Max. :14.00550 | Max. :10.6561 | Max. :11.8889 | Max. : 9.14151 | Max. :18.15954 | Max. : 5.8199 | Max. : 6.12336 | Max. :17.13058 | Max. :15.39047 | Max. :17.8809 | Max. :16.03457 |
The box plot shows the distribution of Total Reach for different Post Types on Facebook.
# Boxplot for Total Reach by Post Type
p <- ggplot(facebook_data_scaled, aes(x = Post_Type, y = Total_Reach)) +
geom_boxplot() +
theme_minimal() +
labs(title = "Total Reach by Post Type", x = "Post Type", y = "Total Reach")
# Convert ggplot object to plotly object
interactive_boxplot <- ggplotly(p)
interactive_boxplotThe graph indicates:
Skewed Distribution: The histogram of Total Reach is highly skewed to the right, indicating that most posts have a relatively low reach.
Majority Low Reach: A significant majority of the posts have a Total Reach close to zero, with the highest frequency in the first bin.
Sparse High Reach: There are only a few posts with a Total Reach greater than 2, indicating that very few posts achieve high reach.
# Histogram for Total Reach
p_hist <- ggplot(facebook_data_scaled, aes(x = Total_Reach)) +
geom_histogram(binwidth = 0.5, fill = "skyblue", color = "black") +
theme_minimal() +
labs(title = "Distribution of Total Reach", x = "Total Reach", y = "Frequency")
# Convert ggplot object to plotly object
interactive_histogram <- ggplotly(p_hist)
interactive_histogramHighest Mean Reach: Video posts have the highest mean Total Reach (295.86), followed by Status posts (217.04), indicating they generally perform better in terms of reach compared to other post types.
Wide Range in Photo Posts: Photo posts have a broad range of Total Reach, from 0 to 6334, and a high standard deviation (407.57), suggesting that the performance of photo posts varies widely.
Consistent Performance for Links: Link posts have the lowest median Total Reach (52.5) and the smallest standard deviation (95.72), indicating more consistent but generally lower performance in terms of reach.
facebook_data %>% group_by(Post_Type) %>% summarise(Min = min(Total_Interactions, na.rm = TRUE),Q1 = quantile(Total_Interactions, probs = .25, na.rm = TRUE), Median = median(Total_Interactions, na.rm = TRUE), Q3 = quantile(Total_Interactions, probs = .75, na.rm = TRUE), Max = max(Total_Interactions, na.rm = TRUE), Mean = mean(Total_Interactions, na.rm = TRUE),SD = sd(Total_Interactions, na.rm = TRUE), n = n(),
Missing = sum(is.na(Total_Interactions))) -> table1
knitr::kable(table1)| Post_Type | Min | Q1 | Median | Q3 | Max | Mean | SD | n | Missing |
|---|---|---|---|---|---|---|---|---|---|
| Link | 6 | 32.75 | 52.5 | 125.0 | 420 | 89.04545 | 95.72056 | 22 | 0 |
| Photo | 0 | 72.00 | 124.0 | 226.0 | 6334 | 218.80523 | 407.56850 | 421 | 0 |
| Status | 17 | 106.00 | 186.0 | 265.0 | 1009 | 217.04444 | 178.47994 | 45 | 0 |
| Video | 81 | 144.00 | 271.0 | 440.5 | 550 | 295.85714 | 183.99224 | 7 | 0 |
\[H_A: \text{Observed frequencies do not match expected frequencies}\]
The chi-squared test statistic is calculated as:
\[\chi^2 = \sum \frac{(O_i - E_i)^2}{E_i}\]
Where: - \(O_i\) is the observed frequency - \(E_i\) is the expected frequency
The chi-squared test resulted in a chi-squared statistic of 957.92 and a p-value of less than 2.2e-16, indicating that we reject the null hypothesis and conclude that the distribution of post types is not equal.
# Observed frequencies
observed <- table(facebook_data$Post_Type)
expected <- c(0.25, 0.25, 0.25, 0.25) * sum(observed)
chi_square_test <- chisq.test(observed, p = expected / sum(expected))
chi_square_test##
## Chi-squared test for given probabilities
##
## data: observed
## X-squared = 957.92, df = 3, p-value < 2.2e-16
For the confidence intervals, we calculated the 95% confidence intervals for the mean reach of each post type. The general formula for the confidence interval is:
\[CI = \bar{x} \pm t_{\alpha/2, n-1} \frac{s}{\sqrt{n}}\]
Where: \(\bar{x}\) is the sample mean, \(t_{\alpha/2, n-1}\) is the t-value for a 95% confidence interval, \(s\) is the sample standard deviation, \(n\) is the sample size
Video Posts Have the Highest Reach: Among all post types, video posts have the highest mean reach by a significant margin. This suggests that video content is highly effective in engaging a larger audience on Facebook. Brands and social media managers should consider prioritising video content to maximise their reach.
Photo Posts Perform Better Than Link and Status Posts: Photo posts have a higher mean reach compared to link and status posts, although not as high as video posts. This indicates that visual content, in general, tends to perform better in terms of reach. Incorporating more photo content into the posting strategy could improve engagement.
Variability in Reach: The confidence intervals for each post type show the range within which the true mean reach is expected to lie with 95% confidence. Video posts have a very wide confidence interval, indicating high variability in their performance. This could be due to a few viral videos skewing the average. On the other hand, link, photo, and status posts have narrower confidence intervals, suggesting more consistent performance.
#Calculate confidence intervals for the mean reach of each post type
conf_intervals <- facebook_data %>%
group_by(Post_Type) %>%
summarise(
Mean_Reach = mean(Total_Reach, na.rm = TRUE),
CI_Lower = Mean_Reach - qt(0.975, n() - 1) * sd(Total_Reach, na.rm = TRUE) / sqrt(n()),
CI_Upper = Mean_Reach + qt(0.975, n() - 1) * sd(Total_Reach, na.rm = TRUE) / sqrt(n())
)
knitr::kable(conf_intervals, caption = "95% Confidence Intervals for Mean Reach by Post Type")| Post_Type | Mean_Reach | CI_Lower | CI_Upper |
|---|---|---|---|
| Link | 18544.59 | 9056.623 | 28032.56 |
| Photo | 13275.39 | 11074.128 | 15476.65 |
| Status | 13078.89 | 11511.202 | 14646.58 |
| Video | 51205.71 | 6049.944 | 96361.48 |
# Fit linear regression model
lm_model <- lm(Total_Reach ~ Post_Type, data = facebook_data)
# Summarise the model
summary(lm_model)##
## Call:
## lm(formula = Total_Reach ~ Post_Type, data = facebook_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -37662 -10154 -8189 -579 167205
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 18545 4781 3.879 0.000119 ***
## Post_TypePhoto -5269 4904 -1.074 0.283134
## Post_TypeStatus -5466 5833 -0.937 0.349229
## Post_TypeVideo 32661 9730 3.357 0.000850 ***
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 22420 on 491 degrees of freedom
## Multiple R-squared: 0.04044, Adjusted R-squared: 0.03457
## F-statistic: 6.897 on 3 and 491 DF, p-value: 0.000148
# Visualise the relationship
ggplot(facebook_data, aes(x = Post_Type, y = Total_Reach)) +
geom_boxplot() +
geom_jitter(width = 0.2, alpha = 0.3) +
stat_summary(fun = mean, geom = "point", shape = 20, size = 4, color = "red") +
labs(title = "Total Reach by Post Type", x = "Post Type", y = "Total Reach")Residual Standard Error: The residual standard error has decreased from 22420 to 1.097, indicating that the transformed model fits the data better by reducing the error.
R-squared and Adjusted R-squared: The R-squared value has increased from 0.04044 to 0.0671, and the adjusted R-squared value has also increased from 0.03457 to 0.0614. Although these values are still relatively low, the increase suggests an improved fit.
Overall Model Significance: The F-statistic has improved from 6.897 to 11.77, with a corresponding p-value of 1.854×10^-7, indicating that the overall model is statistically significant.
# Apply logarithmic transformation
facebook_data$log_Total_Reach <- log(facebook_data$Total_Reach + 1)
# Fit the regression model with transformed data
regression_model_transformed <- lm(log_Total_Reach ~ Post_Type, data = facebook_data)
# Display the summary of the transformed regression model
summary(regression_model_transformed)##
## Call:
## lm(formula = log_Total_Reach ~ Post_Type, data = facebook_data)
##
## Residuals:
## Min 1Q Median 3Q Max
## -3.2381 -0.6289 -0.2157 0.4788 3.3888
##
## Coefficients:
## Estimate Std. Error t value Pr(>|t|)
## (Intercept) 9.1815 0.2339 39.259 < 2e-16 ***
## Post_TypePhoto -0.4670 0.2399 -1.947 0.05216 .
## Post_TypeStatus 0.2256 0.2854 0.791 0.42955
## Post_TypeVideo 1.3045 0.4760 2.740 0.00636 **
## ---
## Signif. codes: 0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
##
## Residual standard error: 1.097 on 491 degrees of freedom
## Multiple R-squared: 0.0671, Adjusted R-squared: 0.0614
## F-statistic: 11.77 on 3 and 491 DF, p-value: 1.854e-07
# Plot diagnostic plots for the transformed regression model
par(mfrow = c(2, 2))
plot(regression_model_transformed)Major Findings
Strengths:
Limitations:
The analysis addressed the problem statement by identifying key predictors of post performance, specifically highlighting the effectiveness of video content. The hypothesis testing confirmed that different post types have significantly different engagement levels. Investing in high-quality video content can significantly enhance a brand’s reach and engagement on Facebook, making it a key component of effective social media strategies. The main takeaway is clear: Prioritise video content to maximise engagement and visibility on Facebook
Class Modules 1-9
Understanding How Your Videos Perform on Facebook Www.facebook.com. Retrieved June 13, 2024, from https://www.facebook.com/formedia/blog/understanding-how-your-videos-perform-on-facebook#:~:text=Metrics%20are%20available%20at%20the
Teves, C. (2022, April 21). Facebook Analytics: how to analyze your data. Sprout Social. https://sproutsocial.com/insights/facebook-analytics/
ChatGPT https://chatgpt.com/
DataFlair Team. (2017, June 30). Introduction to Hypothesis Testing in R - Learn every concept from Scratch! - DataFlair. DataFlair. https://data-flair.training/blogs/hypothesis-testing-in-r/
Calculating Confidence Intervals — R Tutorial. (n.d.). Www.cyclismo.org. https://www.cyclismo.org/tutorial/R/confidence.html
Frost, J. (2018). How To Interpret R-squared in Regression Analysis. Statistics by Jim; Statistics By Jim. https://statisticsbyjim.com/regression/interpret-r-squared-regression/
Keyhole. (2023, November 24). Facebook Analytics Guide 2024: How To Analyze & Use Your FB Data. Keyhole. https://keyhole.co/blog/facebook-analytics/
Dopson, E. (2021, March 13). Videos vs. Images: Which Drives More Engagement in Facebook Ads? | Databox Blog. Databox. https://databox.com/videos-vs-images-in-facebook-ads