Purpose and study design

This project demonstrates how to present statistical evidence about a portfolio manager’s services. It examines client satisfaction, gold holdings and USD strength, returns compared with a bank, and changes in client stock-portfolio values.

All data are synthetic and were created with ChatGPT-assisted R code. There are no real clients, private records, actual bank results, or historical market observations. The findings illustrate an analysis; they cannot establish that Ledy is an effective real-world investment manager.

Two datasets are used. Dataset 1 contains 120 different clients: 60 using Ledy and 60 using a fictional bank. Each provider has 30 domestic and 30 international clients. Dataset 2 contains the same 60 Ledy clients, measured before and after 12 months and matched by Client_ID. The two files overlap intentionally; they do not represent 180 unique people or two independent replications.

Returns are measured over the same 12-month horizon, net of management fees and before personal taxes. The simulation assumes comparable risk categories and no deposits or withdrawals. The before-and-after stock values are consistent with each Ledy client’s return in Dataset 1.

For the instructor’s question about ounces of gold and the value of the U.S. dollar, USD value is operationalized as a fictional dollar-strength index with base 100. Gold is the number of troy ounces held. The index is not the dollar price of gold, an actual observed exchange rate, or the USD value of the gold holding. Each row represents an independent fictional client-market scenario, rather than a shared-date financial time series. This is the operational definition used in this project. It should be confirmed with the instructor because “value of the U.S. dollar” can also mean an exchange rate. The correlation describes the scenarios and is not a direct measure of manager effectiveness.

Variables

Dataset Variable Type and meaning
1 and 2 Client_ID Dataset 1: L001–L060 and B001–B060. Dataset 2: L001–L060 only, linking the repeated measurements.
1 Provider Nominal: Ledy or Bank. Different clients belong to each group.
1 Client_Type Nominal: Domestic or International, relative to the manager’s home market.
1 Satisfied Nominal: Yes or No, measured after 12 months.
1 Net_Return_Pct Numeric: 12-month net stock-portfolio return in percent. A stored value of 8 means 8%.
1 Gold_Ounces Numeric: gold held, in troy ounces, in each simulated scenario.
1 USD_Index Numeric: fictional USD strength index, base 100. Higher means a stronger dollar in the scenario.
2 Stock_Before_USD Numeric: stock-portfolio value in USD before management.
2 Stock_After_USD Numeric: value in USD after 12 months, net of fees.

Data generation

The reproducible generator is supplied in Generate_Synthetic_Data.R, using seed 52212026. It creates the two R data frames used to build the supplied Excel files. A single dataset realization was kept; the seed was not changed to obtain significant p-values.

The simulation specifies mean net returns of 8.5% for Ledy and 6.0% for Bank, both with a standard deviation of 4.5 percentage points. These are design inputs, not conclusions about actual services. Satisfaction probabilities are .80/.60 for Ledy’s domestic/international clients and .70/.60 for Bank’s. Gold is generated around 12 ounces (SD 3); USD strength is generated as 100 - 0.9 * (Gold_Ounces - 12) + error, with error SD 4. This negative relation is intentionally part of the synthetic scenario. Before values are generated around $100,000 (SD $12,000), and after values are calculated using the corresponding return. Numeric precision and generation code are retained for reproducibility.

The sample sizes were chosen for the course exercise, not from a formal power analysis. The chi-square analysis has expected counts of at least 7; the other analyses have 120 observations or 60 matched pairs. Independence across clients is a simulation assumption, not something established by a normality test.

All four tests use alpha = .05 and two-sided alternatives, following the course templates. The tests are educational and exploratory; the p-values are not adjusted for multiple testing. A significant before-after test is interpreted as an increase only when the observed change is positive.

Import and data checks

The code imports the supplied Excel datasets when both files are available. If neither file is in the rendering folder, it uses exact embedded copies of those same synthetic datasets. This portability correction does not generate new observations or change the analyses. The code verifies sample sizes, missing values, unique IDs, matching pairs, and agreement between the two datasets. No records are removed.

library(readxl)
library(effsize)
# Fix: handle rendering a downloaded RMarkdown outside the project folder.
# Import the original Excel files when available; otherwise use exact embedded copies.
dataset_files <- c("Dataset1_Between_Subjects.xlsx", "Dataset2_Within_Subjects.xlsx")
if (all(file.exists(dataset_files))) {
  between <- as.data.frame(read_excel(dataset_files[1], sheet = "Synthetic_Data"))
  within <- as.data.frame(read_excel(dataset_files[2], sheet = "Synthetic_Data"))
} else if (!any(file.exists(dataset_files))) {
  between <- between_backup
  within <- within_backup
} else {
  stop("Only one Excel dataset was found. Keep both original Excel files together.")
}
between$Provider <- factor(between$Provider, levels = c("Ledy", "Bank"))
between$Client_Type <- factor(between$Client_Type, levels = c("Domestic", "International"))
between$Satisfied <- factor(between$Satisfied, levels = c("Yes", "No"))
stopifnot(nrow(between) == 120, nrow(within) == 60,
          !anyDuplicated(between$Client_ID), !anyDuplicated(within$Client_ID),
          all(complete.cases(between)), all(complete.cases(within)))
ledy <- subset(between, Provider == "Ledy")
# Match IDs explicitly; do not rely on row order across files.
within <- within[match(ledy$Client_ID, within$Client_ID), ]
stopifnot(identical(within$Client_ID, ledy$Client_ID))
reconstructed_return <- (within$Stock_After_USD / within$Stock_Before_USD - 1) * 100
stopifnot(max(abs(reconstructed_return - ledy$Net_Return_Pct)) < .0001)
print(data.frame(Dataset = c("Between-subjects", "Within-subjects"),
                 Records = c(nrow(between), nrow(within)), Missing = c(sum(is.na(between)), sum(is.na(within)))))
##            Dataset Records Missing
## 1 Between-subjects     120       0
## 2  Within-subjects      60       0
describe <- function(x) c(N = length(x), Mean = mean(x), Median = median(x), SD = sd(x))
p_text <- function(p) if (p < .001) "p < .001" else if (p > .05) "p > .05" else sprintf("p = %.3f", p)
d_size <- function(d) {
 a <- abs(d)
 if (a < .2) "negligible" else if (a < .5) "small" else if (a < .8) "medium" else if (a < 1.2) "large" else "very large"
}

1. Client type and satisfaction

Research question: Is there a relationship between client type (domestic or international) and satisfaction with Ledy’s services (yes or no)?

Variables: Client_Type and Satisfied, both categorical. This analysis includes only the 60 Ledy clients; Bank responses are not mixed into an assessment of Ledy’s satisfaction.

Hypotheses: H0: Client type and satisfaction are independent among Ledy’s clients. H1: They are associated.

Statistical test: Pearson chi-square test of independence, without continuity correction. Each client contributes one response, and every expected cell count must be at least 5. Cramer’s V measures the strength of association. For this 2-by-2 table, reference magnitudes are .10 small, .30 medium, and .50 large.

# Analysis 1: Chi-square test of independence, Ledy clients only.
observed <- table(Client_Type = ledy$Client_Type, Satisfied = ledy$Satisfied)
print(observed)
##                Satisfied
## Client_Type     Yes No
##   Domestic       24  6
##   International  22  8
satisfaction_pct <- round(100 * prop.table(observed, margin = 1), 2)
print(satisfaction_pct)
##                Satisfied
## Client_Type       Yes    No
##   Domestic      80.00 20.00
##   International 73.33 26.67
barplot(t(observed), beside = TRUE, col = c("steelblue", "grey70"),
        legend.text = TRUE, args.legend = list(x = "topright", bty = "n"),
        ylim = c(0, max(observed) + 8),
        xlab = "Client type", ylab = "Number of clients",
        main = "Satisfaction with Ledy's services")

chi_result <- chisq.test(observed, correct = FALSE)
print(chi_result)
## 
##  Pearson's Chi-squared test
## 
## data:  observed
## X-squared = 0.37267, df = 1, p-value = 0.5416
print(chi_result$expected)
##                Satisfied
## Client_Type     Yes No
##   Domestic       23  7
##   International  23  7
stopifnot(all(chi_result$expected >= 5))
cramers_v <- sqrt(as.numeric(chi_result$statistic) /
                  (sum(observed) * min(nrow(observed)-1, ncol(observed)-1)))
print(cramers_v)
## [1] 0.07881104
# Expected counts are 23 and 7 in each row, meeting the course criterion.
# Pearson chi-square is reported without Yates's continuity correction.
# p = .5416; fail to reject independence. V = .079 is negligible.

Results: A Chi-Square Test of Independence was conducted to determine if there was an association between client type (domestic or international) and satisfaction with Ledy’s services (yes or no). The results showed that there was not a statistically significant association between the two variables, chi-square(1) = 0.37, p > .05. The association was negligible (Cramer’s V = 0.08).

Interpretation: Satisfaction was reported by 24 of 30 domestic clients (80%) and 22 of 30 international clients (73.3%). The observed difference is not statistically significant (p = 0.5416). Fail to reject H0. This does not prove equal satisfaction rates, and the independence test does not test whether the overall satisfaction rate is sufficiently high.

2. Gold holdings and USD strength

Research question: Is there a correlation between ounces of gold held and the value of the U.S. dollar, defined here as the synthetic USD strength index?

Variables: Gold_Ounces and USD_Index, both numerical, across all 120 independent synthetic scenarios.

Hypotheses: H0: The population Pearson correlation is zero. H1: The population Pearson correlation is not zero.

Statistical test: Pearson correlation. The scatterplot is checked for linearity, and histograms, Q-Q plots, outlier checks, and Shapiro-Wilk tests are used to assess the numerical variables. The course rule selects Pearson when both variables pass the normality check; otherwise it would select Spearman. A Shapiro-Wilk p-value above .05 does not prove normality. Here the approximately linear scatterplot, Q-Q plots, and approximately jointly normal construction also support Pearson. The coefficient itself is the effect size.

# Analysis 2: Gold holdings and USD strength in 120 independent synthetic scenarios.
print(rbind(Gold_Ounces = describe(between$Gold_Ounces),
            USD_Index = describe(between$USD_Index)))
##               N     Mean   Median       SD
## Gold_Ounces 120 11.93420 11.87200 2.973844
## USD_Index   120 99.67086 99.86535 4.786532
plot(between$Gold_Ounces, between$USD_Index, pch = 19,
     col = "steelblue", xlab = "Gold holdings (troy ounces)",
     ylab = "Synthetic USD strength index (base 100)",
     main = "Gold holdings and USD strength")
abline(lm(USD_Index ~ Gold_Ounces, data = between), col = "firebrick", lwd = 2)

# The scatterplot shows an approximately linear negative relationship.
par(mfrow = c(2,2))
hist(between$Gold_Ounces, breaks = 12, col = "skyblue", border = "white",
     main = "Gold holdings", xlab = "Troy ounces")
hist(between$USD_Index, breaks = 12, col = "grey75", border = "white",
     main = "USD strength", xlab = "Index points")
qqnorm(between$Gold_Ounces, main = "Gold Q-Q plot"); qqline(between$Gold_Ounces)
qqnorm(between$USD_Index, main = "USD Q-Q plot"); qqline(between$USD_Index)

par(mfrow = c(1,1))
print(list(Gold_outliers = boxplot.stats(between$Gold_Ounces)$out,
           USD_outliers = boxplot.stats(between$USD_Index)$out))
## $Gold_outliers
## numeric(0)
## 
## $USD_outliers
## numeric(0)
gold_normality <- shapiro.test(between$Gold_Ounces)
usd_normality <- shapiro.test(between$USD_Index)
print(gold_normality); print(usd_normality)
## 
##  Shapiro-Wilk normality test
## 
## data:  between$Gold_Ounces
## W = 0.98751, p-value = 0.3405
## 
##  Shapiro-Wilk normality test
## 
## data:  between$USD_Index
## W = 0.98814, p-value = 0.3838
# Both variables pass the course normality check, with no univariate boxplot outliers.
# Gold p = .341 and USD p = .384. Pearson correlation is appropriate here.
correlation <- cor.test(between$Gold_Ounces, between$USD_Index,
                        method = "pearson", alternative = "two.sided")
print(correlation)
## 
##  Pearson's product-moment correlation
## 
## data:  between$Gold_Ounces and between$USD_Index
## t = -7.9228, df = 118, p-value = 1.441e-12
## alternative hypothesis: true correlation is not equal to 0
## 95 percent confidence interval:
##  -0.6950937 -0.4584500
## sample estimates:
##        cor 
## -0.5892692
# r = -.589, p < .001: a moderate negative relationship using the course thresholds.

Results: A Pearson correlation was conducted to test the relationship between gold holdings (M = 11.93, SD = 2.97 troy ounces) and USD strength (M = 99.67, SD = 4.79 index points). There was a statistically significant relationship between the two variables, r(118) = -0.59, p < .001. The relationship was negative and moderate. As gold holdings increased, the USD strength index decreased.

Interpretation: Reject H0. Scenarios with more gold generally have lower USD strength. The 95% confidence interval for r is [-0.70, -0.46]. This association was built into the generator. It does not show that buying gold weakens the dollar, and it is not evidence of Ledy’s skill.

3. Ledy versus bank investing services

Research question: Is there a difference in mean 12-month net returns between Ledy’s services and bank investing services?

Variables: Provider (Ledy or Bank) and Net_Return_Pct (numeric). Each group contains 60 different clients.

Hypotheses: H0: The two population mean net returns are equal. H1: The population mean net returns differ.

Statistical test: Independent two-sample t-test, using equal variances as instructed in the course template. Normality is checked separately within each group. The generator uses the same return SD for both providers, and the observed SDs are similar. This is a stated equal-variance assumption rather than a conclusion from Shapiro-Wilk. Cohen’s d is calculated using the pooled within-group SD. The comparison order is Ledy minus Bank.

# Analysis 3: Independent t-test comparing 12-month net returns.
Ledy_Return <- between$Net_Return_Pct[between$Provider == "Ledy"]
Bank_Return <- between$Net_Return_Pct[between$Provider == "Bank"]
print(rbind(Ledy = describe(Ledy_Return), Bank = describe(Bank_Return)))
##       N     Mean Median       SD
## Ledy 60 8.697627 8.6630 4.780462
## Bank 60 6.313375 6.4177 4.630998
par(mfrow = c(1,2))
hist(Ledy_Return, breaks = 10, col = "skyblue", border = "white",
     main = "Ledy", xlab = "12-month net return (%)")
hist(Bank_Return, breaks = 10, col = "grey75", border = "white",
     main = "Bank", xlab = "12-month net return (%)")

par(mfrow = c(1,1))
boxplot(Net_Return_Pct ~ Provider, data = between,
        col = c("skyblue", "grey75"), ylab = "12-month net return (%)",
        main = "Returns by service provider")

print(list(Ledy_outliers = boxplot.stats(Ledy_Return)$out,
           Bank_outliers = boxplot.stats(Bank_Return)$out))
## $Ledy_outliers
## [1] -4.9595
## 
## $Bank_outliers
## numeric(0)
print(shapiro.test(Ledy_Return)); print(shapiro.test(Bank_Return))
## 
##  Shapiro-Wilk normality test
## 
## data:  Ledy_Return
## W = 0.98811, p-value = 0.8266
## 
##  Shapiro-Wilk normality test
## 
## data:  Bank_Return
## W = 0.98175, p-value = 0.507
# Both distributions are approximately bell-shaped. Ledy has one potential outlier;
# Bank has none. All synthetic observations are retained.
# Ledy p = .827 and Bank p = .507; neither fails the normality check.
# Equal variances follow the course template and the simulation's common SD.
independent <- t.test(Net_Return_Pct ~ Provider, data = between,
                      var.equal = TRUE, alternative = "two.sided")
print(independent)
## 
##  Two Sample t-test
## 
## data:  Net_Return_Pct by Provider
## t = 2.7748, df = 118, p-value = 0.006424
## alternative hypothesis: true difference in means between group Ledy and group Bank is not equal to 0
## 95 percent confidence interval:
##  0.6826964 4.0858070
## sample estimates:
## mean in group Ledy mean in group Bank 
##           8.697627           6.313375
pooled_sd <- sqrt(((length(Ledy_Return)-1)*var(Ledy_Return) +
                   (length(Bank_Return)-1)*var(Bank_Return)) /
                  (length(Ledy_Return)+length(Bank_Return)-2))
independent_d <- (mean(Ledy_Return)-mean(Bank_Return))/pooled_sd
print(independent_d)
## [1] 0.506606
# The group order is Ledy minus Bank. The difference is in percentage points.
# p = .00642; reject equal means. Cohen's d = .51 is medium.

Results: An Independent T-Test was conducted to determine if there was a difference in 12-month net stock-portfolio returns between clients using Ledy’s services and clients using bank investing services. Ledy returns (M = 8.70%, SD = 4.78 percentage points) were significantly different from Bank returns (M = 6.31%, SD = 4.63 percentage points), t(118) = 2.77, p = 0.006. The effect size was medium, Cohen’s d = 0.51.

Interpretation: Reject H0. Ledy’s mean net return is 2.38 percentage points higher than Bank’s, with a 95% confidence interval of [0.68, 4.09] percentage points. This supports a higher return within the simulated scenario. The generator assumed different provider means, so it cannot independently demonstrate real-world superiority. Actual comparisons would also require risk, fees, client selection, and comparable market exposure to be considered.

4. Stock value before and after management

Research question: Is there an increase in client stock-portfolio value after 12 months of management by Ledy?

Variables: Stock_Before_USD and Stock_After_USD for the same 60 clients, matched by Client_ID. The analysis uses After - Before.

Hypotheses: H0: The population mean paired change is zero. H1: The population mean paired change is not zero. A two-sided test follows the course example; a significant positive mean change supports the proposed increase.

Statistical test: Dependent (paired) t-test. Normality is assessed on the paired differences. The three potential tail outliers are retained because they are valid generated observations. The paired effect size uses effsize::cohen.d(After, Before, paired = TRUE), matching the course’s standardization rather than substituting d_z. Because paired observations are correlated, a small p-value can coexist with a medium effect under this convention; the statistic and effect size answer different questions.

# Analysis 4: Dependent t-test, the same 60 Ledy clients before and after 12 months.
Before <- within$Stock_Before_USD
After <- within$Stock_After_USD
Differences <- After - Before
print(rbind(Before = describe(Before), After = describe(After),
            Change = describe(Differences)))
##         N       Mean     Median        SD
## Before 60  98394.298  97331.595 11117.335
## After  60 106997.446 107239.175 13400.794
## Change 60   8603.148   8267.855  4974.117
par(mfrow = c(1,2))
hist(Differences, breaks = 12, col = "skyblue", border = "white",
     main = "Change in stock value", xlab = "After - Before (USD)")
boxplot(Differences, col = "skyblue", main = "Paired differences",
        ylab = "After - Before (USD)")

par(mfrow = c(1,1))
print(boxplot.stats(Differences)$out)
## [1] 22115.93 -4643.68 22053.14
print(shapiro.test(Differences))
## 
##  Shapiro-Wilk normality test
## 
## data:  Differences
## W = 0.97723, p-value = 0.3228
# Normality applies to differences, not the before and after columns separately.
# The histogram is approximately bell-shaped with a few tail observations.
# Three potential outliers are retained. Shapiro-Wilk p = .323.
matplot(c(0,12), t(cbind(Before,After)), type = "l", lty = 1,
        col = adjustcolor("steelblue", alpha.f = .35), xaxt = "n",
        xlab = "Measurement", ylab = "Stock-portfolio value (USD)",
        main = "Matched clients before and after management")
axis(1, at = c(0,12), labels = c("Before", "After 12 months"))

paired_result <- t.test(After, Before, paired = TRUE, alternative = "two.sided")
print(paired_result)
## 
##  Paired t-test
## 
## data:  After and Before
## t = 13.397, df = 59, p-value < 2.2e-16
## alternative hypothesis: true mean difference is not equal to 0
## 95 percent confidence interval:
##  7318.197 9888.099
## sample estimates:
## mean difference 
##        8603.148
paired_effect <- effsize::cohen.d(After, Before, paired = TRUE)
print(paired_effect)
## 
## Cohen's d
## 
## d estimate: 0.6261824 (medium)
## 95 percent confidence interval:
##     lower     upper 
## 0.5249583 0.7274065
# A two-sided test follows the course example. A significant positive change
# supports an increase in this simulated sample. This direction was not selected
# after testing a one-sided alternative.
# The paired Cohen's d uses the effsize convention taught in the course.
# It is not d_z, which uses only the SD of the differences.

Results: A Dependent T-Test was conducted to determine if there was a difference in stock-portfolio value before versus after 12 months of management by Ledy. Before values (M = $98394.30, SD = $11117.34) were significantly different from After values (M = $106997.45, SD = $13400.79), t(59) = 13.40, p < .001. The mean change was positive. The effect size was medium, Cohen’s d = 0.63.

Interpretation: Reject H0. The average change is an increase of $8,603.15, with a 95% confidence interval from $7,318.20 to $9,888.10. Clients with gains: 59; clients with losses: 1. An average increase does not mean every client gained. No deposits or withdrawals occur in this simulation. A real before-after increase alone would not separate manager performance from general market movement.

Overall findings and limits

Numerical results at a glance

Analysis Sample Statistic p-value Effect size Decision at .05
Client type and satisfaction 60 Ledy clients Chi-square(1) = 0.37 .5416 Cramer’s V = .079, negligible Fail to reject H0
Gold and USD strength 120 scenarios r(118) = -.59 1.441 × 10^-12 r = -.589, moderate negative Reject H0
Ledy versus Bank returns 60 per group t(118) = 2.77 .006424 Cohen’s d = .507, medium Reject H0
Stock value before/after 60 matched clients t(59) = 13.40 1.510 × 10^-19 Paired Cohen’s d = .626, medium Reject H0

The p-values above give numerical detail; the results paragraphs follow the course reporting format. A nonsignificant result does not establish equivalence. All inferential statements refer to the simulated data-generating scenario.

Question Test Main finding in the synthetic data
Client type and satisfaction Chi-square independence No statistically significant association; overall satisfaction is a separate question.
Gold holdings and dollar strength Pearson correlation A moderate negative association, intentionally included in the simulation.
Ledy versus Bank Independent t-test Ledy’s mean net return is higher in this scenario.
Before versus after Dependent t-test Mean stock-portfolio value increases over 12 months.

This project shows how Ledy could organize an evidence-based client report: define the outcomes, check assumptions, present the test results, and explain the limits. The synthetic returns and stock values are favorable to Ledy, while the satisfaction analysis does not establish a client-type difference. The gold relationship provides context rather than direct evidence of effectiveness.

These findings must not be presented as an actual performance record. They depend on the simulation parameters. Dataset 2 uses clients already included in Dataset 1, so the two return-related tests are related analyses rather than independent confirmations. Real evidence would require verified client records, suitable benchmarks, comparable risk, and a defensible study design.

Course references

Reproducibility

Run Final_Project_Analysis.R with both Excel files in the working directory, or open this RMarkdown file and Knit. Required packages are readxl, effsize, knitr, and rmarkdown. Generate_Synthetic_Data.R recreates the two source data frames without overwriting the submitted Excel files.

sessionInfo()
## R version 4.6.1 (2026-06-24 ucrt)
## Platform: x86_64-w64-mingw32/x64
## Running under: Windows 11 x64 (build 26200)
## 
## Matrix products: default
##   LAPACK version 3.12.1
## 
## locale:
## [1] LC_COLLATE=Spanish_Ecuador.utf8  LC_CTYPE=Spanish_Ecuador.utf8   
## [3] LC_MONETARY=Spanish_Ecuador.utf8 LC_NUMERIC=C                    
## [5] LC_TIME=Spanish_Ecuador.utf8    
## 
## time zone: America/Chicago
## tzcode source: internal
## 
## attached base packages:
## [1] stats     graphics  grDevices utils     datasets  methods   base     
## 
## other attached packages:
## [1] effsize_0.8.1 readxl_1.5.0 
## 
## loaded via a namespace (and not attached):
##  [1] digest_0.6.39    R6_2.6.1         fastmap_1.2.0    cellranger_1.1.0
##  [5] xfun_0.60        cachem_1.1.0     knitr_1.51       htmltools_0.5.9 
##  [9] rmarkdown_2.31   lifecycle_1.0.5  cli_3.6.6        sass_0.4.10     
## [13] jquerylib_0.1.4  compiler_4.6.1   tools_4.6.1      evaluate_1.0.5  
## [17] bslib_0.12.0     yaml_2.3.12      rlang_1.3.0      jsonlite_2.0.0