Analisis data merupakan proses penting untuk memperoleh informasi yang lebih terstruktur dari suatu dataset. Pada analisis ini digunakan dataset Cleaning Data 1.csv untuk melihat struktur data, membersihkan variabel yang tidak diperlukan, memeriksa data kosong, mengubah variabel kategorik menjadi faktor, serta melakukan encoding agar data dapat digunakan pada tahap analisis berikutnya.
Analisis ini bertujuan untuk memahami struktur dataset, melakukan pembersihan data, memeriksa kualitas data, melakukan factor encoding dan one-hot encoding, membuat visualisasi sederhana, serta mengekspor dataset hasil pengolahan.
Dataset dibaca dari lokasi
D:\\Downloads\\Cleaning Data 1.csv dengan pemisah titik
koma (;), sesuai dengan sintaks R yang diberikan.
file_csv <- "D:/Downloads/Cleaning Data 1.csv"
data <- read.csv(
file_csv,
header = TRUE,
sep = ";",
stringsAsFactors = FALSE,
check.names = FALSE
)
head(data)ukuran_data <- data.frame(
Keterangan = c("Jumlah baris", "Jumlah kolom"),
Nilai = c(nrow(data), ncol(data))
)
kable(ukuran_data, caption = "Ukuran dataset sebelum pembersihan")| Keterangan | Nilai |
|---|---|
| Jumlah baris | 1000 |
| Jumlah kolom | 8 |
## [1] "Transaction.ID" "Item" "Quantity" "Price.Per.Unit"
## [5] "Total.Spent" "Payment.Method" "Location" ""
## 'data.frame': 1000 obs. of 8 variables:
## $ Transaction.ID: chr "TXN_2176024" "TXN_6327139" "TXN_5488764" "TXN_9530003" ...
## $ Item : chr "Coffee" "Sandwich" "Coffee" "Salad" ...
## $ Quantity : int 5 4 4 1 5 3 3 1 3 1 ...
## $ Price.Per.Unit: chr "2.0" "4.0" "2.0" "5.0" ...
## $ Total.Spent : chr "10.0" "16.0" "8.0" "5.0" ...
## $ Payment.Method: chr "Digital Wallet" "Credit Card" "Digital Wallet" "Digital Wallet" ...
## $ Location : chr "In-store" "Takeaway" "Takeaway" "Takeaway" ...
## $ : logi NA NA NA NA NA NA ...
| Transaction.ID | Item | Quantity | Price.Per.Unit | Total.Spent | Payment.Method | Location | |
|---|---|---|---|---|---|---|---|
| TXN_2176024 | Coffee | 5 | 2.0 | 10.0 | Digital Wallet | In-store | NA |
| TXN_6327139 | Sandwich | 4 | 4.0 | 16.0 | Credit Card | Takeaway | NA |
| TXN_5488764 | Coffee | 4 | 2.0 | 8.0 | Digital Wallet | Takeaway | NA |
| TXN_9530003 | Salad | 1 | 5.0 | 5.0 | Digital Wallet | Takeaway | NA |
| TXN_3753993 | Sandwich | 5 | 4.0 | 20.0 | Credit Card | Takeaway | NA |
Berdasarkan sintaks R yang diberikan, beberapa variabel dihapus
karena tidak digunakan dalam proses pengolahan berikutnya, yaitu
X, Payment.Method, Total.Spent,
Price.Per.Unit, dan Transaction.ID.
variabel_hapus <- intersect(
c("X", "Payment.Method", "Total.Spent", "Price.Per.Unit", "Transaction.ID"),
names(data)
)
if (length(variabel_hapus) > 0) {
data <- data[, !names(data) %in% variabel_hapus, drop = FALSE]
}
kable(
data.frame(Variabel_Dihapus = variabel_hapus),
caption = "Variabel yang dihapus dari dataset"
)| Variabel_Dihapus |
|---|
| Payment.Method |
| Total.Spent |
| Price.Per.Unit |
| Transaction.ID |
## 'data.frame': 1000 obs. of 4 variables:
## $ Item : chr "Coffee" "Sandwich" "Coffee" "Salad" ...
## $ Quantity: int 5 4 4 1 5 3 3 1 3 1 ...
## $ Location: chr "In-store" "Takeaway" "Takeaway" "Takeaway" ...
## $ : logi NA NA NA NA NA NA ...
Pemeriksaan missing value dilakukan pada setiap variabel untuk mengetahui jumlah data yang kosong sebelum proses encoding.
missing_data <- data.frame(
Variabel = names(data),
Jumlah_NA = colSums(is.na(data)),
Persentase_NA = round(colMeans(is.na(data)) * 100, 2)
)
kable(
missing_data,
caption = "Jumlah dan persentase missing value"
)| Variabel | Jumlah_NA | Persentase_NA | |
|---|---|---|---|
| Item | Item | 0 | 0 |
| Quantity | Quantity | 0 | 0 |
| Location | Location | 0 | 0 |
| 1000 | 100 |
missing_plot <- missing_data %>%
filter(Jumlah_NA > 0)
if (nrow(missing_plot) > 0) {
ggplot(missing_plot, aes(x = reorder(Variabel, Jumlah_NA), y = Jumlah_NA)) +
geom_col() +
coord_flip() +
labs(
title = "Jumlah Missing Value per Variabel",
x = "Variabel",
y = "Jumlah missing value"
) +
tema_laporan
} else {
ggplot(data.frame(x = 1, y = 1), aes(x, y)) +
geom_text(aes(label = "Tidak terdapat missing value"), size = 6) +
theme_void() +
labs(title = "Pemeriksaan Missing Value")
}Sesuai sintaks awal, frekuensi Item,
Location, dan Quantity diperiksa menggunakan
table().
##
## Cake Coffee Cookie Juice Salad Sandwich Smoothie Tea
## 138 118 110 114 126 123 136 135
##
## In-store Takeaway
## 300 700
##
## 1 2 3 4 5
## 195 203 179 204 219
Variabel Item, Location, dan
Quantity diubah menjadi tipe factor agar dapat
diperlakukan sebagai variabel kategorik.
variabel_factor <- intersect(c("Item", "Location", "Quantity"), names(data))
data[variabel_factor] <- lapply(data[variabel_factor], factor)
str(data[variabel_factor])## 'data.frame': 1000 obs. of 3 variables:
## $ Item : Factor w/ 8 levels "Cake","Coffee",..: 2 6 2 5 6 7 5 5 4 2 ...
## $ Location: Factor w/ 2 levels "In-store","Takeaway": 1 2 2 2 2 1 2 2 1 2 ...
## $ Quantity: Factor w/ 5 levels "1","2","3","4",..: 5 4 4 1 5 3 3 1 3 1 ...
## [1] "Cake" "Coffee" "Cookie" "Juice" "Salad" "Sandwich" "Smoothie"
## [8] "Tea"
## [1] "In-store" "Takeaway"
## [1] "1" "2" "3" "4" "5"
if ("Item" %in% names(data)) {
ggplot(data, aes(x = Item)) +
geom_bar() +
labs(
title = "Distribusi Item",
x = "Item",
y = "Frekuensi"
) +
tema_laporan +
theme(axis.text.x = element_text(angle = 45, hjust = 1))
}One-hot encoding digunakan untuk mengubah kategori menjadi
beberapa variabel indikator. Proses ini mengikuti penggunaan package
fastDummies pada sintaks R yang diberikan.
data_onehot <- data
kolom_onehot <- intersect(c("Item", "Location"), names(data_onehot))
if (length(kolom_onehot) > 0) {
data_onehot <- fastDummies::dummy_cols(
data_onehot,
select_columns = kolom_onehot,
remove_first_dummy = FALSE,
remove_selected_columns = TRUE
)
}
kable(
head(data_onehot, 5),
caption = "Hasil one-hot encoding pada Item dan Location"
)| Quantity | Item_Cake | Item_Coffee | Item_Cookie | Item_Juice | Item_Salad | Item_Sandwich | Item_Smoothie | Item_Tea | Location_In-store | Location_Takeaway | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 5 | NA | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 |
| 4 | NA | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 1 |
| 4 | NA | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| 1 | NA | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 1 |
| 5 | NA | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 1 |
perbandingan_dimensi <- data.frame(
Tahap = c("Sebelum encoding", "Sesudah one-hot encoding"),
Baris = c(nrow(data), nrow(data_onehot)),
Kolom = c(ncol(data), ncol(data_onehot))
)
kable(
perbandingan_dimensi,
caption = "Perbandingan dimensi dataset"
)| Tahap | Baris | Kolom |
|---|---|---|
| Sebelum encoding | 1000 | 4 |
| Sesudah one-hot encoding | 1000 | 12 |
if ("Location" %in% names(data)) {
location_dummy <- model.matrix(~ Location - 1, data = data)
kable(
head(as.data.frame(location_dummy), 5),
caption = "Matriks dummy untuk Location"
)
}| LocationIn-store | LocationTakeaway |
|---|---|
| 1 | 0 |
| 0 | 1 |
| 0 | 1 |
| 0 | 1 |
| 0 | 1 |
if ("Item" %in% names(data)) {
item_dummy <- model.matrix(~ Item - 1, data = data)
kable(
head(as.data.frame(item_dummy), 5),
caption = "Matriks dummy untuk Item"
)
}| ItemCake | ItemCoffee | ItemCookie | ItemJuice | ItemSalad | ItemSandwich | ItemSmoothie | ItemTea |
|---|---|---|---|---|---|---|---|
| 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 |
| 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 |
if ("Quantity" %in% names(data)) {
quantity_dummy <- model.matrix(~ Quantity - 1, data = data)
kable(
head(as.data.frame(quantity_dummy), 5),
caption = "Matriks dummy untuk Quantity"
)
}| Quantity1 | Quantity2 | Quantity3 | Quantity4 | Quantity5 |
|---|---|---|---|---|
| 0 | 0 | 0 | 0 | 1 |
| 0 | 0 | 0 | 1 | 0 |
| 0 | 0 | 0 | 1 | 0 |
| 1 | 0 | 0 | 0 | 0 |
| 0 | 0 | 0 | 0 | 1 |
missing_after <- data.frame(
Variabel = names(data_onehot),
Jumlah_NA = colSums(is.na(data_onehot)),
Persentase_NA = round(colMeans(is.na(data_onehot)) * 100, 2)
)
kable(
missing_after,
caption = "Missing value setelah one-hot encoding"
)| Variabel | Jumlah_NA | Persentase_NA | |
|---|---|---|---|
| Quantity | Quantity | 0 | 0 |
| 1000 | 100 | ||
| Item_Cake | Item_Cake | 0 | 0 |
| Item_Coffee | Item_Coffee | 0 | 0 |
| Item_Cookie | Item_Cookie | 0 | 0 |
| Item_Juice | Item_Juice | 0 | 0 |
| Item_Salad | Item_Salad | 0 | 0 |
| Item_Sandwich | Item_Sandwich | 0 | 0 |
| Item_Smoothie | Item_Smoothie | 0 | 0 |
| Item_Tea | Item_Tea | 0 | 0 |
| Location_In-store | Location_In-store | 0 | 0 |
| Location_Takeaway | Location_Takeaway | 0 | 0 |
## 'data.frame': 1000 obs. of 12 variables:
## $ Quantity : Factor w/ 5 levels "1","2","3","4",..: 5 4 4 1 5 3 3 1 3 1 ...
## $ : logi NA NA NA NA NA NA ...
## $ Item_Cake : int 0 0 0 0 0 0 0 0 0 0 ...
## $ Item_Coffee : int 1 0 1 0 0 0 0 0 0 1 ...
## $ Item_Cookie : int 0 0 0 0 0 0 0 0 0 0 ...
## $ Item_Juice : int 0 0 0 0 0 0 0 0 1 0 ...
## $ Item_Salad : int 0 0 0 1 0 0 1 1 0 0 ...
## $ Item_Sandwich : int 0 1 0 0 1 0 0 0 0 0 ...
## $ Item_Smoothie : int 0 0 0 0 0 1 0 0 0 0 ...
## $ Item_Tea : int 0 0 0 0 0 0 0 0 0 0 ...
## $ Location_In-store: int 1 0 0 0 0 1 0 0 1 0 ...
## $ Location_Takeaway: int 0 1 1 1 1 0 1 1 0 1 ...
| Quantity | Item_Cake | Item_Coffee | Item_Cookie | Item_Juice | Item_Salad | Item_Sandwich | Item_Smoothie | Item_Tea | Location_In-store | Location_Takeaway | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| 5 | NA | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 1 | 0 |
| 4 | NA | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 1 |
| 4 | NA | 0 | 1 | 0 | 0 | 0 | 0 | 0 | 0 | 0 | 1 |
| 1 | NA | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 0 | 1 |
| 5 | NA | 0 | 0 | 0 | 0 | 0 | 1 | 0 | 0 | 0 | 1 |
Dataset hasil one-hot encoding diekspor dalam format CSV ke
folder D:\\Downloads dengan nama
data_encoded.csv.
file_output <- "D:/Downloads/data_encoded.csv"
write.csv(
data_onehot,
file_output,
row.names = FALSE,
na = ""
)
cat("Dataset berhasil diekspor ke:\n", file_output, "\n")## Dataset berhasil diekspor ke:
## D:/Downloads/data_encoded.csv
Analisis dilakukan dengan membaca dataset Cleaning Data
1.csv, menghapus variabel yang tidak digunakan, memeriksa
struktur dan missing value, mengubah variabel kategorik menjadi
factor, serta melakukan one-hot encoding pada
variabel Item dan Location. Dataset hasil
pengolahan kemudian diperiksa kembali dan diekspor sebagai
data_encoded.csv.
Hasil akhir berupa dataset yang telah disiapkan dalam bentuk yang lebih sesuai untuk digunakan pada tahap analisis atau pemodelan berikutnya.