FİNAL PROJESİ
SPACESHİP TİTANİC
Kozmik bir gizemi çözmek için veri bilimi becerilerinize ihtiyaç duyulan 2912 yılına hoş geldiniz. Dört ışık yılı öteden bir sinyal aldık ve işler pek iyi görünmüyor.
Uzay Gemisi Titanik, bir ay önce fırlatılan yıldızlararası bir yolcu gemisiydi. Gemide neredeyse 13.000 yolcu bulunan gemi, güneş sistemimizden göçmenleri yakın yıldızların yörüngesinde bulunan üç yeni yaşanabilir dış gezegene taşımak üzere ilk yolculuğuna çıktı.
Dikkatsiz Uzay Gemisi Titanic, ilk varış noktası olan kavurucu 55 Cancri E’ye giderken Alpha Centauri’yi dönerken, bir toz bulutunun içine gizlenmiş bir uzay-zaman anormalliğiyle çarpıştı. Ne yazık ki 1000 yıl öncesindeki adaşı ile benzer bir kaderle karşılaştı. Gemi sağlam kalmasına rağmen yolcuların neredeyse yarısı alternatif bir boyuta taşındı!
DATALARI YÜKLEME
Dataları indirdikten sonra gereken paket yüklemelerini yapalım.
“tidyverse” ve “explore” paketleri bize verileri görselleştirmede ve analiz yapmada yardımcı olacaktır.
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr 1.1.4 ✔ purrr 1.0.2
## ✔ forcats 1.0.0 ✔ stringr 1.5.1
## ✔ ggplot2 3.4.4 ✔ tibble 3.2.1
## ✔ lubridate 1.9.3 ✔ tidyr 1.3.0
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag() masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
Gereken paket yüklemelerini yaptıktan sonra bize verilen data bilgilerini kontrol edelim.
## # A tibble: 14 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId chr 0 0 8693 NA NA NA
## 2 HomePlanet chr 201 2.3 4 NA NA NA
## 3 CryoSleep lgl 217 2.5 3 0 0.36 1
## 4 Cabin chr 199 2.3 6561 NA NA NA
## 5 Destination chr 182 2.1 4 NA NA NA
## 6 Age dbl 179 2.1 81 0 28.8 79
## 7 VIP lgl 203 2.3 3 0 0.02 1
## 8 RoomService dbl 181 2.1 1274 0 225. 14327
## 9 FoodCourt dbl 183 2.1 1508 0 458. 29813
## 10 ShoppingMall dbl 208 2.4 1116 0 174. 23492
## 11 Spa dbl 183 2.1 1328 0 311. 22408
## 12 VRDeck dbl 188 2.2 1307 0 305. 24133
## 13 Name chr 200 2.3 8474 NA NA NA
## 14 Transported lgl 0 0 2 0 0.5 1
## # A tibble: 13 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId chr 0 0 4277 NA NA NA
## 2 HomePlanet chr 87 2 4 NA NA NA
## 3 CryoSleep lgl 93 2.2 3 0 0.37 1
## 4 Cabin chr 100 2.3 3266 NA NA NA
## 5 Destination chr 92 2.2 4 NA NA NA
## 6 Age dbl 91 2.1 80 0 28.7 79
## 7 VIP lgl 93 2.2 3 0 0.02 1
## 8 RoomService dbl 82 1.9 843 0 219. 11567
## 9 FoodCourt dbl 106 2.5 903 0 439. 25273
## 10 ShoppingMall dbl 98 2.3 716 0 177. 8292
## 11 Spa dbl 101 2.4 834 0 303. 19844
## 12 VRDeck dbl 80 1.9 797 0 311. 22272
## 13 Name chr 94 2.2 4177 NA NA NA
Verilen data bilgilerinin özelliklerine baktığımızda bilinmeyen eksik değerlerin fazlalığını görüyoruz. Biz bu dataları daha temiz daha okunabilirlik sağlamak için bir kaç işlemde bulunacağız.
İlk olarak PASSENGERID sütunu ve CABİN sütununa baktığımızda bize her bir bilgide birden fazla değer gösteriyor. Biz bu bilgileri farklı sütunlara ayıralım.
SÜTUNLARI AYIRMA
PASSENGERID
PassengerId sütununu “ailenum ve ailesıra” olmak üzere ikiye ayıralım.
CABİN
Daha sonra Cabin sütununu üçe ayıralım. Bu 3 sütunumuzun isimleri ” deck, num, side” olarak belirleyelim.
train[c('deck',"num", "side")]<- str_split_fixed(train$Cabin, "/",3)
test[c('deck',"num", "side")]<- str_split_fixed(test$Cabin, "/",3)Bu kodları kullandıktan sonra Cabin sütunu bir işimize yaramayacak bunun için veri setlerimizde silelim.
BİLİNMEYEN DEĞERLERİ TEMİZLEME VE DOLDURMA
Sütunları ayırma işlemini bitirdik. Şimdi bilinmeyen eksik değerleri temizleyelim ve yerlerini NA değeriyle dolduralım.
Bu işlem genel bir temizleme işlemi ve doldurma işlemi yaptı. Biz daha iyi bir sonuca ulaşmak istiyorsak bütün verilen bilgileri detaylı bir şekilde temizlemeliyiz.
İlk öncelikle verilere baktığımızda sayısal ve sözel ifadeler görüyoruz. Bu sayısal ve sözel ifadeler için farklı işlemlerde bulunacağız. İlk önce sözel ifadeler olan karakter(chr) ve logical(lgl) tip veriler için işleme başlayalım.
HOMEPLANET
CRYOSLEEP
DESTİNATİON
VIP
DECK
NUM
SİDE
Daha sonra bu eklediğimiz NA değeri sütun boşluklarında gözükmeyecektir. Bu NA değerini “NA” olarak değiştirmemiz gerekiyor.
HOMEPLANET
levels(train$HomePlanet)[is.na(levels(train$HomePlanet))] <- "NA"
levels(test$HomePlanet)[is.na(levels(test$HomePlanet))] <- "NA"CRYOSLEEP
levels(train$CryoSleep)[is.na(levels(train$CryoSleep))] <- "NA"
levels(test$CryoSleep)[is.na(levels(test$CryoSleep))] <- "NA"DESTİNATİON
levels(train$Destination)[is.na(levels(train$Destination))] <- "NA"
levels(test$Destination)[is.na(levels(test$Destination))] <- "NA"VIP
levels(train$VIP)[is.na(levels(train$VIP))] <- "NA"
levels(test$VIP)[is.na(levels(test$VIP))] <- "NA"DECK
levels(train$deck)[is.na(levels(train$deck))] <- "NA"
levels(test$deck)[is.na(levels(test$deck))] <- "NA"NUM
levels(train$num)[is.na(levels(train$num))] <- "NA"
levels(test$num)[is.na(levels(test$num))] <- "NA"SIDE
levels(train$side)[is.na(levels(train$side))] <- "NA"
levels(test$side)[is.na(levels(test$side))] <- "NA"## [1] "55 Cancri e" "PSO J318.5-22" "TRAPPIST-1e" "NA"
Şimdi double(dbl) yani sayısal değerler için işlem yapalım. Burada yapacağımız işlem eksik değerler yerine ortalamalarını koymaktır. Burada kullanacağımız kod da her iki data için DESTİNATİON verilerine göre ortalamalarını almasını sağlamaktır.
AGE
train <- train %>%
group_by(Destination) %>%
mutate(Age = replace(Age, is.na(Age), mean(Age, na.rm = TRUE)))test <- test %>%
group_by(Destination) %>%
mutate(Age = replace(Age, is.na(Age), mean(Age, na.rm = TRUE)))ROOMSERVİCE
train <- train %>%
group_by(Destination) %>%
mutate(RoomService = replace(RoomService, is.na(RoomService), mean(RoomService, na.rm = TRUE)))test <- test %>%
group_by(Destination) %>%
mutate(RoomService = replace(RoomService, is.na(RoomService), mean(RoomService, na.rm = TRUE)))FOODCOURT
train <- train %>%
group_by(Destination) %>%
mutate(FoodCourt = replace(FoodCourt, is.na(FoodCourt), mean(FoodCourt, na.rm = TRUE)))test <- test %>%
group_by(Destination) %>%
mutate(FoodCourt = replace(FoodCourt, is.na(FoodCourt), mean(FoodCourt, na.rm = TRUE)))SHOPPİNGMALL
train <- train %>%
group_by(Destination) %>%
mutate(ShoppingMall = replace(ShoppingMall, is.na(ShoppingMall), mean(ShoppingMall, na.rm = TRUE)))test <- test %>%
group_by(Destination) %>%
mutate(ShoppingMall = replace(ShoppingMall, is.na(ShoppingMall), mean(ShoppingMall, na.rm = TRUE)))SPA
train <- train %>%
group_by(Destination) %>%
mutate(Spa = replace(Spa, is.na(Spa), mean(Spa, na.rm = TRUE)))test <- test %>%
group_by(Destination) %>%
mutate(Spa = replace(Spa, is.na(Spa), mean(Spa, na.rm = TRUE)))VRDECK
train <- train %>%
group_by(Destination) %>%
mutate(VRDeck = replace(VRDeck, is.na(VRDeck), mean(VRDeck, na.rm = TRUE)))test <- test %>%
group_by(Destination) %>%
mutate(VRDeck = replace(VRDeck, is.na(VRDeck), mean(VRDeck, na.rm = TRUE)))Böylelikle bilinmeyene eksik değerleri temzilemiş bulunmaktayız. Şimdi dataları kontrol edelim.
## # A tibble: 18 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId chr 0 0 8693 NA NA NA
## 2 HomePlanet fct 0 0 4 NA NA NA
## 3 CryoSleep fct 0 0 3 NA NA NA
## 4 Destination fct 0 0 4 NA NA NA
## 5 Age dbl 0 0 84 0 28.8 79
## 6 VIP fct 0 0 3 NA NA NA
## 7 RoomService dbl 0 0 1277 0 225. 14327
## 8 FoodCourt dbl 0 0 1511 0 458. 29813
## 9 ShoppingMall dbl 0 0 1119 0 174. 23492
## 10 Spa dbl 0 0 1331 0 311. 22408
## 11 VRDeck dbl 0 0 1310 0 305. 24133
## 12 Name chr 200 2.3 8474 NA NA NA
## 13 Transported lgl 0 0 2 0 0.5 1
## 14 ailenum chr 0 0 6217 NA NA NA
## 15 ailesıra chr 0 0 8 NA NA NA
## 16 deck fct 0 0 9 NA NA NA
## 17 num fct 0 0 1818 NA NA NA
## 18 side fct 0 0 3 NA NA NA
## # A tibble: 17 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId chr 0 0 4277 NA NA NA
## 2 HomePlanet fct 0 0 4 NA NA NA
## 3 CryoSleep fct 0 0 3 NA NA NA
## 4 Destination fct 0 0 4 NA NA NA
## 5 Age dbl 0 0 83 0 28.7 79
## 6 VIP fct 0 0 3 NA NA NA
## 7 RoomService dbl 0 0 846 0 219. 11567
## 8 FoodCourt dbl 0 0 905 0 439. 25273
## 9 ShoppingMall dbl 0 0 719 0 177. 8292
## 10 Spa dbl 0 0 837 0 303. 19844
## 11 VRDeck dbl 0 0 800 0 311. 22272
## 12 Name chr 94 2.2 4177 NA NA NA
## 13 ailenum chr 0 0 3063 NA NA NA
## 14 ailesıra chr 0 0 8 NA NA NA
## 15 deck fct 0 0 9 NA NA NA
## 16 num fct 0 0 1506 NA NA NA
## 17 side fct 0 0 3 NA NA NA
Kontrol sağladığımızda verilerin temizlendiğini görüyoruz. “name” veri bilgisi herhangi bir özellik sağlamıyor ve işimize yaramıyacaktır. Bunun için datalardan kaldıralım.
Veri setlerini incelediğimizde “ailenum” veri bilgisinde bulunun tekrarlanmaları ve okunabilirliği düzeltmeliyiz. Bunun için aşağıdaki kodu kullanalım. Her iki data için yapalım.
TRAİN
trainailenum <- ifelse(duplicated(train$ailenum) | duplicated(train$ailenum, fromLast = TRUE), 1, 0)TEST
Bütün eksik değerleri ve okunuabilirliği kolaylaştırcak işlemleri yaptık. Şimdi dataları kontrol edelim.
TRAİN describe_all
## # A tibble: 15 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId chr 0 0 8693 NA NA NA
## 2 HomePlanet fct 0 0 4 NA NA NA
## 3 CryoSleep fct 0 0 3 NA NA NA
## 4 Destination fct 0 0 4 NA NA NA
## 5 Age dbl 0 0 84 0 28.8 79
## 6 VIP fct 0 0 3 NA NA NA
## 7 RoomService dbl 0 0 1277 0 225. 14327
## 8 FoodCourt dbl 0 0 1511 0 458. 29813
## 9 ShoppingMall dbl 0 0 1119 0 174. 23492
## 10 Spa dbl 0 0 1331 0 311. 22408
## 11 VRDeck dbl 0 0 1310 0 305. 24133
## 12 Transported lgl 0 0 2 0 0.5 1
## 13 ailesıra chr 0 0 8 NA NA NA
## 14 deck fct 0 0 9 NA NA NA
## 15 side fct 0 0 3 NA NA NA
TEST describe_all
## # A tibble: 14 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId chr 0 0 4277 NA NA NA
## 2 HomePlanet fct 0 0 4 NA NA NA
## 3 CryoSleep fct 0 0 3 NA NA NA
## 4 Destination fct 0 0 4 NA NA NA
## 5 Age dbl 0 0 83 0 28.7 79
## 6 VIP fct 0 0 3 NA NA NA
## 7 RoomService dbl 0 0 846 0 219. 11567
## 8 FoodCourt dbl 0 0 905 0 439. 25273
## 9 ShoppingMall dbl 0 0 719 0 177. 8292
## 10 Spa dbl 0 0 837 0 303. 19844
## 11 VRDeck dbl 0 0 800 0 311. 22272
## 12 ailesıra chr 0 0 8 NA NA NA
## 13 deck fct 0 0 9 NA NA NA
## 14 side fct 0 0 3 NA NA NA
Görüldüğü üzere tertemiz şekilde datalarımız bulunuyor. Böylelikle daha iyi tahmin sonuçlarına ulaşabileceğiz.
Şimdi dataların veri profillerinin raporunu oluşturalım ve bir kaç yorumda bulunalım.
VERİ PROFİL RAPORLARI
TRAİN
##
|
| | 0%
|
|. | 2%
|
|.. | 5% [global_options]
|
|... | 7%
|
|.... | 10% [introduce]
|
|.... | 12%
|
|..... | 14% [plot_intro]
|
|...... | 17%
|
|....... | 19% [data_structure]
|
|........ | 21%
|
|......... | 24% [missing_profile]
|
|.......... | 26%
|
|........... | 29% [univariate_distribution_header]
|
|........... | 31%
|
|............ | 33% [plot_histogram]
|
|............. | 36%
|
|.............. | 38% [plot_density]
|
|............... | 40%
|
|................ | 43% [plot_frequency_bar]
|
|................. | 45%
|
|.................. | 48% [plot_response_bar]
|
|.................. | 50%
|
|................... | 52% [plot_with_bar]
|
|.................... | 55%
|
|..................... | 57% [plot_normal_qq]
|
|...................... | 60%
|
|....................... | 62% [plot_response_qq]
|
|........................ | 64%
|
|......................... | 67% [plot_by_qq]
|
|.......................... | 69%
|
|.......................... | 71% [correlation_analysis]
|
|........................... | 74%
|
|............................ | 76% [principal_component_analysis]
|
|............................. | 79%
|
|.............................. | 81% [bivariate_distribution_header]
|
|............................... | 83%
|
|................................ | 86% [plot_response_boxplot]
|
|................................. | 88%
|
|................................. | 90% [plot_by_boxplot]
|
|.................................. | 93%
|
|................................... | 95% [plot_response_scatterplot]
|
|.................................... | 98%
|
|.....................................| 100% [plot_by_scatterplot]
## "C:/Program Files/RStudio/resources/app/bin/quarto/bin/tools/pandoc" +RTS -K512m -RTS "C:\Users\Win10\Desktop\finalpro\report.knit.md" --to html4 --from markdown+autolink_bare_uris+tex_math_single_backslash --output pandoc37a824497682.html --lua-filter "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmarkdown\lua\pagebreak.lua" --lua-filter "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmarkdown\lua\latex-div.lua" --embed-resources --standalone --variable bs3=TRUE --section-divs --table-of-contents --toc-depth 6 --template "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmd\h\default.html" --no-highlight --variable highlightjs=1 --variable theme=yeti --mathjax --variable "mathjax-url=https://mathjax.rstudio.com/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML" --include-in-header "C:\Users\Win10\AppData\Local\Temp\Rtmp2FgHWT\rmarkdown-str37a86b6266bc.html"
TEST
##
|
| | 0%
|
|. | 2%
|
|.. | 5% [global_options]
|
|... | 7%
|
|.... | 10% [introduce]
|
|.... | 12%
|
|..... | 14% [plot_intro]
|
|...... | 17%
|
|....... | 19% [data_structure]
|
|........ | 21%
|
|......... | 24% [missing_profile]
|
|.......... | 26%
|
|........... | 29% [univariate_distribution_header]
|
|........... | 31%
|
|............ | 33% [plot_histogram]
|
|............. | 36%
|
|.............. | 38% [plot_density]
|
|............... | 40%
|
|................ | 43% [plot_frequency_bar]
|
|................. | 45%
|
|.................. | 48% [plot_response_bar]
|
|.................. | 50%
|
|................... | 52% [plot_with_bar]
|
|.................... | 55%
|
|..................... | 57% [plot_normal_qq]
|
|...................... | 60%
|
|....................... | 62% [plot_response_qq]
|
|........................ | 64%
|
|......................... | 67% [plot_by_qq]
|
|.......................... | 69%
|
|.......................... | 71% [correlation_analysis]
|
|........................... | 74%
|
|............................ | 76% [principal_component_analysis]
|
|............................. | 79%
|
|.............................. | 81% [bivariate_distribution_header]
|
|............................... | 83%
|
|................................ | 86% [plot_response_boxplot]
|
|................................. | 88%
|
|................................. | 90% [plot_by_boxplot]
|
|.................................. | 93%
|
|................................... | 95% [plot_response_scatterplot]
|
|.................................... | 98%
|
|.....................................| 100% [plot_by_scatterplot]
## "C:/Program Files/RStudio/resources/app/bin/quarto/bin/tools/pandoc" +RTS -K512m -RTS "C:\Users\Win10\Desktop\finalpro\report.knit.md" --to html4 --from markdown+autolink_bare_uris+tex_math_single_backslash --output pandoc37a851f768b6.html --lua-filter "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmarkdown\lua\pagebreak.lua" --lua-filter "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmarkdown\lua\latex-div.lua" --embed-resources --standalone --variable bs3=TRUE --section-divs --table-of-contents --toc-depth 6 --template "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmd\h\default.html" --no-highlight --variable highlightjs=1 --variable theme=yeti --mathjax --variable "mathjax-url=https://mathjax.rstudio.com/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML" --include-in-header "C:\Users\Win10\AppData\Local\Temp\Rtmp2FgHWT\rmarkdown-str37a866ea44c9.html"
Şimdi bir histogram grafiği oluşturalım.
HİSTOGRAM GRAFİĞİ OLUŞTURMA
ÖRNEK1
test datasından “ShoppingMall” ve “Destination” verilerini göz önüne alalım. Bunun için aşağıdaki kodları kullanabiliriz.
ggplot(test, aes(x = ShoppingMall)) +
geom_histogram(fill = "white", color = "black") +
facet_grid(Destination ~ .)YORUM
Yaptığımız histogram grafiğini incelediğimide neredeyse hiçbir yolcunun hiç harcama yapmadığı görülüyor. 0-200 kişi TRAPPIST-1e gezegenine gidenlerden harcama yapma olasılığı oluşmuş.
TAHMİN MODELLERİ OLUŞTURMA
Öncelikle modelleri oluşturmadan önce yeni veri setleri ve alt veri setleri oluşturmalıyız.
Daha sonra daha iyi görselleştirme, analiz yapma ve okunabilirliği yapmak için “caTools” paketini yüklemeliyiz.
Pakaeti yükledikten sonra kuracağımız modellerde tekrarlanabilirliği sürekli farklı bir şekilde tahminde bulunması için “set.seed” fonksiyonunu kullanalım.
Şimdi alt veri setleri oluşturmamız gerekiyor. İlk önce “sample.split” fonksiyonu ile oluşturduğumuz yeni veri setinden ne kadar oranda veriler çekmek istediğimizi yazmalıyız.
Daha sonra bu oluşturduğumuz oran ile alt veri setleri oluşturalım.
Böylelikle bütün gereken bilgileri oluşturduktan sonra tahmin modellerini oluşturabiliriz.
SVM MODELİ
SVM MODELİ NEDİR?
Support Vector Machine (SVM), sınıflandırma ve regresyon problemleri için kullanılan bir makine öğrenimi algoritmasıdır. SVM, özellikle sınıflandırma problemlerinde etkili olan bir algoritmadır ve doğrusal olarak ayrılabilir veya doğrusal olarak ayrılamayan veri setlerinde kullanılabilir.
SVM’nin temel amacı, veri setindeki sınıfları en iyi şekilde ayıran bir karar sınırı (hyperplane) oluşturmaktır. Bu sınır, iki sınıf arasındaki maksimum marjı (uzaklığı) maksimize etmek için belirlenir. SVM, özellikle düşük boyutlu veri setlerinde ve özellikle yüksek boyutlu uzaylarda iyi performans gösterir.
library(e1071)
fit_svm <- svm(Transported ~ ., data = training_set,
type = 'C-classification',
kernel = 'linear')
preds <- predict(fit_svm, newdata = testing_set, type = "raw") %>%
data.frame()## y_pred
## y_true 0 1
## 0 688 175
## 1 164 712
## [1] 0.8050604
SVM MODEL SONUCU
SVM RADİAL(KERNEL) MODELİ
SVM RADİAL(KERNEL) MODELİ NEDİR?
Support Vector Machine (SVM), sınıflandırma ve regresyon problemleri için kullanılan bir makine öğrenimi algoritmasıdır. SVM’nin çekirdek (kernel) fonksiyonları, özellikle doğrusal olarak ayrılamayan veri setleri üzerinde etkili olmasını sağlar. Radyal bazlı fonksiyon çekirdeği (Radial Basis Function Kernel), bu çekirdek fonksiyonlarından biridir.
SVM RADİAL SONUCU
DECİSİON TREES MODELİ
DECİSİON TREES NEDİR?
Decision Trees (Karar Ağaçları), sınıflandırma ve regresyon problemleri için kullanılan bir makine öğrenimi algoritmasıdır. Veri kümesindeki özelliklere dayanarak bir dizi karar yapısı oluşturur ve bu yapıyı kullanarak veri noktalarını sınıflandırır veya tahminler yapar. Decision Trees, ağaç yapısındaki bir dizi karar düğümünden oluşur ve her düğüm, bir özellik testi yaparak veriyi iki veya daha fazla alt kümeye böler.
training_set$Transported <- as.factor(training_set$Transported )
testing_set$Transported <- as.factor(testing_set$Transported)
train_set$Transported <- as.factor(train_set$Transported )## Call:
## rpart::rpart(formula = Transported ~ ., data = training_set)
## n= 6954
##
## CP nsplit rel error xerror xstd
## 1 0.42873696 0 1.0000000 1.0162225 0.01207812
## 2 0.02887215 1 0.5712630 0.5712630 0.01088848
## 3 0.01289108 4 0.4846466 0.4846466 0.01032566
## 4 0.01129780 6 0.4588644 0.4776941 0.01027460
## 5 0.01013905 7 0.4475666 0.4669757 0.01019404
## 6 0.01000000 8 0.4374276 0.4658169 0.01018520
##
## Variable importance
## CryoSleep Spa VRDeck RoomService FoodCourt ShoppingMall
## 39 16 15 10 9 4
## HomePlanet deck
## 3 2
##
## Node number 1: 6954 observations, complexity param=0.428737
## predicted class=TRUE expected loss=0.4964049 P(node) =1
## class counts: 3452 3502
## probabilities: 0.496 0.504
## left son=2 (4524 obs) right son=3 (2430 obs)
## Primary splits:
## CryoSleep splits as LRL, improve=723.5736, (0 missing)
## RoomService < 0.5 to the right, improve=425.3976, (0 missing)
## Spa < 0.5 to the right, improve=395.9205, (0 missing)
## VRDeck < 0.5 to the right, improve=362.9468, (0 missing)
## ShoppingMall < 0.5 to the right, improve=209.5968, (0 missing)
## Surrogate splits:
## Spa < 0.5 to the right, agree=0.720, adj=0.199, (0 split)
## VRDeck < 0.5 to the right, agree=0.705, adj=0.157, (0 split)
## FoodCourt < 0.5 to the right, agree=0.703, adj=0.149, (0 split)
## RoomService < 0.5 to the right, agree=0.696, adj=0.131, (0 split)
## ShoppingMall < 0.5 to the right, agree=0.683, adj=0.093, (0 split)
##
## Node number 2: 4524 observations, complexity param=0.02887215
## predicted class=FALSE expected loss=0.3364279 P(node) =0.6505608
## class counts: 3002 1522
## probabilities: 0.664 0.336
## left son=4 (1163 obs) right son=5 (3361 obs)
## Primary splits:
## RoomService < 346.5 to the right, improve=94.69975, (0 missing)
## Spa < 291.5 to the right, improve=86.28603, (0 missing)
## Age < 12.5 to the right, improve=85.70985, (0 missing)
## FoodCourt < 1331 to the left, improve=77.60358, (0 missing)
## VRDeck < 417.5 to the right, improve=65.50853, (0 missing)
## Surrogate splits:
## HomePlanet splits as RRLR, agree=0.782, adj=0.153, (0 split)
## deck splits as RRRLRRRLR, agree=0.746, adj=0.010, (0 split)
## Age < 78.5 to the right, agree=0.743, adj=0.001, (0 split)
##
## Node number 3: 2430 observations
## predicted class=TRUE expected loss=0.1851852 P(node) =0.3494392
## class counts: 450 1980
## probabilities: 0.185 0.815
##
## Node number 4: 1163 observations
## predicted class=FALSE expected loss=0.1625107 P(node) =0.1672419
## class counts: 974 189
## probabilities: 0.837 0.163
##
## Node number 5: 3361 observations, complexity param=0.02887215
## predicted class=FALSE expected loss=0.3966082 P(node) =0.483319
## class counts: 2028 1333
## probabilities: 0.603 0.397
## left son=10 (815 obs) right son=11 (2546 obs)
## Primary splits:
## Spa < 523 to the right, improve=124.74930, (0 missing)
## VRDeck < 407 to the right, improve=107.54040, (0 missing)
## Age < 12.5 to the right, improve= 60.01738, (0 missing)
## FoodCourt < 2507.5 to the left, improve= 48.57491, (0 missing)
## ShoppingMall < 622.5 to the left, improve= 42.28737, (0 missing)
## Surrogate splits:
## VRDeck < 12683.5 to the right, agree=0.758, adj=0.002, (0 split)
## Age < 65.5 to the right, agree=0.758, adj=0.001, (0 split)
##
## Node number 10: 815 observations
## predicted class=FALSE expected loss=0.1558282 P(node) =0.1171987
## class counts: 688 127
## probabilities: 0.844 0.156
##
## Node number 11: 2546 observations, complexity param=0.02887215
## predicted class=FALSE expected loss=0.4736842 P(node) =0.3661202
## class counts: 1340 1206
## probabilities: 0.526 0.474
## left son=22 (777 obs) right son=23 (1769 obs)
## Primary splits:
## VRDeck < 355 to the right, improve=142.39180, (0 missing)
## FoodCourt < 2063.5 to the left, improve= 62.90754, (0 missing)
## HomePlanet splits as LRRL, improve= 37.56650, (0 missing)
## Age < 12.5 to the right, improve= 33.03081, (0 missing)
## ShoppingMall < 1205 to the left, improve= 29.11563, (0 missing)
## Surrogate splits:
## deck splits as RRLRRRR-R, agree=0.705, adj=0.032, (0 split)
## HomePlanet splits as RLRR, agree=0.700, adj=0.015, (0 split)
## FoodCourt < 6922.5 to the right, agree=0.698, adj=0.012, (0 split)
## VIP splits as RLR, agree=0.698, adj=0.009, (0 split)
## Age < 64.5 to the right, agree=0.697, adj=0.006, (0 split)
##
## Node number 22: 777 observations, complexity param=0.01013905
## predicted class=FALSE expected loss=0.2213642 P(node) =0.1117343
## class counts: 605 172
## probabilities: 0.779 0.221
## left son=44 (692 obs) right son=45 (85 obs)
## Primary splits:
## FoodCourt < 2866 to the left, improve=44.810930, (0 missing)
## deck splits as LRRRLLL-L, improve=22.050440, (0 missing)
## HomePlanet splits as LRLR, improve=15.305190, (0 missing)
## ailesıra splits as LRLRRRLL, improve= 8.888805, (0 missing)
## VRDeck < 1678.5 to the right, improve= 8.710888, (0 missing)
## Surrogate splits:
## ailesıra splits as LLLLRLLL, agree=0.893, adj=0.024, (0 split)
## RoomService < 324 to the left, agree=0.892, adj=0.012, (0 split)
##
## Node number 23: 1769 observations, complexity param=0.01289108
## predicted class=TRUE expected loss=0.415489 P(node) =0.254386
## class counts: 735 1034
## probabilities: 0.415 0.585
## left son=46 (1528 obs) right son=47 (241 obs)
## Primary splits:
## HomePlanet splits as LRLL, improve=47.25637, (0 missing)
## FoodCourt < 1738.5 to the left, improve=44.18785, (0 missing)
## deck splits as RRRLLLL-L, improve=43.43835, (0 missing)
## Spa < 205.5 to the right, improve=22.04282, (0 missing)
## ShoppingMall < 1481.5 to the left, improve=18.38204, (0 missing)
## Surrogate splits:
## deck splits as RRRLLLL-L, agree=0.979, adj=0.842, (0 split)
## FoodCourt < 1809.5 to the left, agree=0.921, adj=0.419, (0 split)
## ShoppingMall < 8081 to the left, agree=0.865, adj=0.012, (0 split)
## Spa < 518 to the left, agree=0.865, adj=0.008, (0 split)
##
## Node number 44: 692 observations
## predicted class=FALSE expected loss=0.1618497 P(node) =0.09951107
## class counts: 580 112
## probabilities: 0.838 0.162
##
## Node number 45: 85 observations
## predicted class=TRUE expected loss=0.2941176 P(node) =0.01222318
## class counts: 25 60
## probabilities: 0.294 0.706
##
## Node number 46: 1528 observations, complexity param=0.01289108
## predicted class=TRUE expected loss=0.4613874 P(node) =0.2197297
## class counts: 705 823
## probabilities: 0.461 0.539
## left son=92 (255 obs) right son=93 (1273 obs)
## Primary splits:
## Spa < 111 to the right, improve=27.805020, (0 missing)
## ShoppingMall < 1205 to the left, improve=20.602200, (0 missing)
## VRDeck < 11.5 to the right, improve=15.930710, (0 missing)
## Age < 4.5 to the right, improve=13.375610, (0 missing)
## RoomService < 103.5 to the right, improve= 9.157256, (0 missing)
## Surrogate splits:
## Age < 74.5 to the right, agree=0.834, adj=0.008, (0 split)
##
## Node number 47: 241 observations
## predicted class=TRUE expected loss=0.1244813 P(node) =0.03465631
## class counts: 30 211
## probabilities: 0.124 0.876
##
## Node number 92: 255 observations
## predicted class=FALSE expected loss=0.3254902 P(node) =0.03666954
## class counts: 172 83
## probabilities: 0.675 0.325
##
## Node number 93: 1273 observations, complexity param=0.0112978
## predicted class=TRUE expected loss=0.418696 P(node) =0.1830601
## class counts: 533 740
## probabilities: 0.419 0.581
## left son=186 (125 obs) right son=187 (1148 obs)
## Primary splits:
## VRDeck < 129 to the right, improve=15.611210, (0 missing)
## ShoppingMall < 1205 to the left, improve=15.326260, (0 missing)
## RoomService < 104 to the right, improve=10.677220, (0 missing)
## Age < 3.5 to the right, improve= 7.897389, (0 missing)
## FoodCourt < 1738.5 to the left, improve= 6.552242, (0 missing)
##
## Node number 186: 125 observations
## predicted class=FALSE expected loss=0.344 P(node) =0.01797527
## class counts: 82 43
## probabilities: 0.656 0.344
##
## Node number 187: 1148 observations
## predicted class=TRUE expected loss=0.3928571 P(node) =0.1650848
## class counts: 451 697
## probabilities: 0.393 0.607
## y_pred
## y_true 0 1
## 0 623 240
## 1 123 753
## [1] 0.7912593
## Call:
## rpart::rpart(formula = Transported ~ ., data = train_set)
## n= 8693
##
## CP nsplit rel error xerror xstd
## 1 0.43244496 0 1.0000000 1.0000000 0.010803454
## 2 0.03429896 1 0.5675550 0.5675550 0.009719864
## 3 0.01000000 4 0.4646582 0.4730012 0.009158661
##
## Variable importance
## CryoSleep Spa VRDeck RoomService FoodCourt ShoppingMall
## 44 17 15 11 7 4
## HomePlanet deck
## 2 1
##
## Node number 1: 8693 observations, complexity param=0.432445
## predicted class=TRUE expected loss=0.4963764 P(node) =1
## class counts: 4315 4378
## probabilities: 0.496 0.504
## left son=2 (5656 obs) right son=3 (3037 obs)
## Primary splits:
## CryoSleep splits as LRL, improve=920.2004, (0 missing)
## RoomService < 0.5 to the right, improve=523.3930, (0 missing)
## Spa < 0.5 to the right, improve=504.8747, (0 missing)
## VRDeck < 0.5 to the right, improve=462.3201, (0 missing)
## ShoppingMall < 0.5 to the right, improve=282.7649, (0 missing)
## Surrogate splits:
## Spa < 0.5 to the right, agree=0.722, adj=0.204, (0 split)
## FoodCourt < 0.5 to the right, agree=0.706, adj=0.157, (0 split)
## VRDeck < 0.5 to the right, agree=0.703, adj=0.150, (0 split)
## RoomService < 0.5 to the right, agree=0.692, adj=0.119, (0 split)
## ShoppingMall < 0.5 to the right, agree=0.685, adj=0.097, (0 split)
##
## Node number 2: 5656 observations, complexity param=0.03429896
## predicted class=FALSE expected loss=0.3350424 P(node) =0.6506384
## class counts: 3761 1895
## probabilities: 0.665 0.335
## left son=4 (1432 obs) right son=5 (4224 obs)
## Primary splits:
## RoomService < 346.5 to the right, improve=121.3964, (0 missing)
## Spa < 266.5 to the right, improve=112.9389, (0 missing)
## Age < 12.5 to the right, improve=109.4055, (0 missing)
## FoodCourt < 1331 to the left, improve= 98.1198, (0 missing)
## VRDeck < 587.5 to the right, improve= 75.0317, (0 missing)
## Surrogate splits:
## HomePlanet splits as RRLR, agree=0.785, adj=0.151, (0 split)
## deck splits as RRRLRRRRR, agree=0.749, adj=0.007, (0 split)
## Age < 78.5 to the right, agree=0.747, adj=0.001, (0 split)
##
## Node number 3: 3037 observations
## predicted class=TRUE expected loss=0.1824169 P(node) =0.3493616
## class counts: 554 2483
## probabilities: 0.182 0.818
##
## Node number 4: 1432 observations
## predicted class=FALSE expected loss=0.1571229 P(node) =0.1647302
## class counts: 1207 225
## probabilities: 0.843 0.157
##
## Node number 5: 4224 observations, complexity param=0.03429896
## predicted class=FALSE expected loss=0.3953598 P(node) =0.4859082
## class counts: 2554 1670
## probabilities: 0.605 0.395
## left son=10 (1459 obs) right son=11 (2765 obs)
## Primary splits:
## Spa < 205 to the right, improve=166.33280, (0 missing)
## VRDeck < 417.5 to the right, improve=127.52200, (0 missing)
## Age < 12.5 to the right, improve= 76.46367, (0 missing)
## FoodCourt < 2507.5 to the left, improve= 63.32833, (0 missing)
## ShoppingMall < 627 to the left, improve= 59.33765, (0 missing)
## Surrogate splits:
## HomePlanet splits as RLRR, agree=0.686, adj=0.092, (0 split)
## deck splits as LLLRRRRLR, agree=0.674, adj=0.055, (0 split)
## FoodCourt < 3197.5 to the right, agree=0.662, adj=0.022, (0 split)
## VRDeck < 2052 to the right, agree=0.659, adj=0.014, (0 split)
## Age < 65.5 to the right, agree=0.656, adj=0.003, (0 split)
##
## Node number 10: 1459 observations
## predicted class=FALSE expected loss=0.2021933 P(node) =0.1678362
## class counts: 1164 295
## probabilities: 0.798 0.202
##
## Node number 11: 2765 observations, complexity param=0.03429896
## predicted class=FALSE expected loss=0.4972875 P(node) =0.318072
## class counts: 1390 1375
## probabilities: 0.503 0.497
## left son=22 (807 obs) right son=23 (1958 obs)
## Primary splits:
## VRDeck < 355 to the right, improve=180.83390, (0 missing)
## FoodCourt < 2069.5 to the left, improve= 60.83808, (0 missing)
## HomePlanet splits as LRRL, improve= 46.66727, (0 missing)
## ShoppingMall < 1540.5 to the left, improve= 37.98636, (0 missing)
## Age < 7.5 to the right, improve= 34.98638, (0 missing)
## Surrogate splits:
## deck splits as RRLRRRRRR, agree=0.712, adj=0.014, (0 split)
## FoodCourt < 6922.5 to the right, agree=0.710, adj=0.007, (0 split)
## Age < 68.5 to the right, agree=0.709, adj=0.002, (0 split)
## RoomService < 343 to the right, agree=0.708, adj=0.001, (0 split)
##
## Node number 22: 807 observations
## predicted class=FALSE expected loss=0.2156134 P(node) =0.09283331
## class counts: 633 174
## probabilities: 0.784 0.216
##
## Node number 23: 1958 observations
## predicted class=TRUE expected loss=0.386619 P(node) =0.2252387
## class counts: 757 1201
## probabilities: 0.387 0.613