1910701526

SPACESCHIP TİTANİC KAGGLE COMPETİTİON

Kaggle, ünlü Titanic yarışmasına benzer şekilde Spaceship Titanic adında eğlenceli bir yarışma başlattı. Bu yarışma, veri bilimine yeni başlayanlar için makine öğrenimi hakkında bilgi edinmenin, Kaggle’ı keşfetmenin ve topluluğun diğer üyeleriyle tanışmanın bir yoludur. Bu makale Spaceship Titanic yarışmasını analiz etmekte ve RandomForestClassifier kullanarak test seti için anlamlı içgörülerin nasıl elde edileceğini ve “temel gerçeğin” yaklaşık %80 doğrulukla nasıl tahmin edileceğini açıklamaktadır.

Problem tanımı

Uzay gemisi Titanik bir uzay-zaman anomalisiyle çarpıştı ve yolcuların yarısı farklı bir boyuta taşındı. Bu kayıp yolcuları bulmak için geminin hasarlı bilgisayar sisteminden geri yüklenen kayıtları kullanmamız gerekiyor.

Veriler hakkında

Train dosyası (spaceship_titanic_train.csv) — Kişisel kayıtlar,makine öğrenimi modeli oluşturmak için kullanılacak yolcuların bilgilerini içerir. Test dosyası (spaceship_titanic_test.csv) — Metin, yolcuların üçtebirinin kişisel kayıtlarını içerdiğini, ancak Taşınabilir değeri içermediğini belirtmektedir. Bu verilerin, modelin görünmeyen veriler üzerinde ne kadar iyi performans gösterdiğini ölçmek için kullanılacağı söylenmektedir.

Bu problem çözmek için Rstudio ve RMarkdown kullanacağız.

library(rmarkdown)

Veri Okuma

Veri setinin yapısını kontrol edeceğiz.Önce özelliklere bakacağız, sonra türleri kontrol edeceğiz.

library(readr)
train <- read_csv("train.csv")
library(readr)
test <- read_csv("test.csv")

Train veri setinde 13 bağımsız değişken ve 1 hedef değişken bulunuyor(Transported). Aynı şekilde test veri setinde de sütunlar mevcut. Test veri kümesinde, train veri kümesiyle benzer özelliklere sahip olduğumuz için tahminlerimizi train verileriyle oluşturulan modele dayanarak yapacağız.

paged_table(train)
paged_table(test)

Yukarıdaki değişkenleri açıklamasına vereceğiz.

*PassengerId:Her yolcu için benzersiz bir kimlik olan gggg_pp formatındaki kimlikler,gggg’nin yolcunun seyahat ettiği grubu ve pp’nin grup içindeki numarasını gösterir. Bu grup genellikle aile üyelerinden oluşur, fakat her zaman olmayabilir.

*HomePlanet:Yolcunun ayrıldığı gezegen daimi ikamet gezegeni olarak bilinir.

*CryoSleep:Yolcunun yolculuk süresinde dondurucu uykuya alınıp alınmayacağını gösterir. Yolcular cabinlere kapatılır.

*Cabin :Yolcunun kaldığı cabin numarası.deck/num/side biçimini alır, burada yan İskele için P veya Sancak için S olabilir.

*Destination:Yolcunun ineceği gezegen.

*Age:Yolcunun yaşı.

*VIP:Yolcunun yolculuk sırasında özel VIP hizmeti için ödeme yapıp yapmadığı.

*Roomservice,Foodcourt,ShoppingMall,Spa,VRDeck:Yolcunun Uzay Gemisi Titanic’in faturalandırılan her bir lüks özelliği için tutar ödediği belirtiliyor.

*Name:Yolcunun adı ve soyadı.

*Transported:Gözlemlediğimiz şey, yolcunun başka bir boyuta geçip geçmediğini belirlemeye çalışmak. Bunun hedefi sütundur.

unique(train$HomePlanet)
## [1] "Europa" "Earth"  "Mars"   NA
unique(test$HomePlanet)
## [1] "Earth"  "Europa" "Mars"   NA
unique(train$CryoSleep)
## [1] FALSE  TRUE    NA
unique(test$CryoSleep)
## [1]  TRUE FALSE    NA
unique(train$Destination)
## [1] "TRAPPIST-1e"   "PSO J318.5-22" "55 Cancri e"   NA
unique(test$Destination)
## [1] "TRAPPIST-1e"   "55 Cancri e"   "PSO J318.5-22" NA

Tidyverse

Tidyverse, veri analizi ve görselleştirme için kullanılan bir paket serisidir. Temiz, düzenli ve etkili bir veri analizi süreci sağlar.

library(tidyverse)
library(explore)

train ve test veri çerçevesinde boş değerleri kontrol eder ve boş olan hücrelere NA değeri atanır.

train[train == ""] <- NA
test[test == ""] <- NA
train %>% describe_all()
## # A tibble: 14 × 8
##    variable     type     na na_pct unique   min   mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl>  <dbl> <dbl>
##  1 PassengerId  chr       0    0     8693    NA  NA       NA
##  2 HomePlanet   chr     201    2.3      4    NA  NA       NA
##  3 CryoSleep    lgl     217    2.5      3     0   0.36     1
##  4 Cabin        chr     199    2.3   6561    NA  NA       NA
##  5 Destination  chr     182    2.1      4    NA  NA       NA
##  6 Age          dbl     179    2.1     81     0  28.8     79
##  7 VIP          lgl     203    2.3      3     0   0.02     1
##  8 RoomService  dbl     181    2.1   1274     0 225.   14327
##  9 FoodCourt    dbl     183    2.1   1508     0 458.   29813
## 10 ShoppingMall dbl     208    2.4   1116     0 174.   23492
## 11 Spa          dbl     183    2.1   1328     0 311.   22408
## 12 VRDeck       dbl     188    2.2   1307     0 305.   24133
## 13 Name         chr     200    2.3   8474    NA  NA       NA
## 14 Transported  lgl       0    0        2     0   0.5      1
test %>% describe_all()
## # A tibble: 13 × 8
##    variable     type     na na_pct unique   min   mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl>  <dbl> <dbl>
##  1 PassengerId  chr       0    0     4277    NA  NA       NA
##  2 HomePlanet   chr      87    2        4    NA  NA       NA
##  3 CryoSleep    lgl      93    2.2      3     0   0.37     1
##  4 Cabin        chr     100    2.3   3266    NA  NA       NA
##  5 Destination  chr      92    2.2      4    NA  NA       NA
##  6 Age          dbl      91    2.1     80     0  28.7     79
##  7 VIP          lgl      93    2.2      3     0   0.02     1
##  8 RoomService  dbl      82    1.9    843     0 219.   11567
##  9 FoodCourt    dbl     106    2.5    903     0 439.   25273
## 10 ShoppingMall dbl      98    2.3    716     0 177.    8292
## 11 Spa          dbl     101    2.4    834     0 303.   19844
## 12 VRDeck       dbl      80    1.9    797     0 311.   22272
## 13 Name         chr      94    2.2   4177    NA  NA       NA

“train ve test” veri çerçevesindeki “Cabin” sütununu karakterine göre bölerek, “Deck”, “Num” ve “Side” adlı yeni sütunları oluşturur. Sonra İskele için P veya Sancak için S kolayca göstereceğiz.

train[c('Deck', 'Num', 'Side')] <-str_split_fixed(train$Cabin,'/', 3)
test[c('Deck', 'Num', 'Side')] <-str_split_fixed(test$Cabin,'/', 3)

Train ve test bulunan PassengerId sütununu içinde iki tane bilgi (gggg_pp) vardı ve bu iki bilgi çıkartacağız,(gggg)ailenum ve (pp)ailesıra.

train[c('ailenum','ailesıra')] <- str_split_fixed(train$PassengerId,"_",2)
test[c('ailenum','ailesıra')] <- str_split_fixed(test$PassengerId,"_",2)

Train ve test, PassengerId sütunundaki ayırılan hücreyi başıya alalım.

train <- train[,c(18,19,1:17)]
test <- test[,c(17,18,1:16)]

Şimdi train ve test, Cabinden üç parçaya ayırılan hücreyi (Deck ,Num ,Side) Cabinin yanına alalım.

train <- train[,c(1:6,17,18,19,7:16)]
test <- test[,c(1:6,16,17,18,7:15)]

Train ve test veri kümesinin Cabinin bilgisine aldık,artık ihtiyacımız kadığı için temizliğeceğiz.

train <- train %>% select(- Cabin)
test <- test %>% select(- Cabin)

addNA() fonksiyonu, belirtilen vektöre veya faktöre NA değerini eklemek için kullanılır.Örneğin, eğer “train ve test” veri sütunun başlangıçta NA değerleri içermiyorsa, bu kod parçası her hücreye bir NA değeri ekleyecektir. Eğer “train ve test” veri sütunun zaten NA değerleri içeriyorsa, bu kod değişikliğe neden olmayacaktır.

train$HomePlanet<-addNA(train$HomePlanet)
test$HomePlanet<-addNA(test$HomePlanet)
train$CryoSleep<-addNA(train$CryoSleep)
test$CryoSleep<-addNA(test$CryoSleep)
train$Deck<-addNA(train$Deck)
test$Deck<-addNA(test$Deck)
train$Side <-addNA(train$Side)
test$Side <- addNA(test$Side)
train$Destination<-addNA(train$Destination)
test$Destination<-addNA(test$Destination)
train$VIP <- addNA(train$VIP)
test$VIP <- addNA(test$VIP)

Şimdi “train” ve “test” veri çerçevelerinin “HomePlanet” ve “Destination” sütunlarına göre gruplanmasını sağlar. Daha sonra, eksik değerleri (NA) “Age” sütununda ortalama değerlerle değiştiririz.

library(dplyr)
library(tidyr)
train <- train %>%
  group_by(HomePlanet, Destination) %>%
  mutate_at(vars(Age), ~replace_na(., mean(., na.rm = TRUE)))
test <- test %>%
  group_by(HomePlanet, Destination) %>%
  mutate_at(vars(Age), ~replace_na(., mean(., na.rm = TRUE)))

Train ve test, isim(Name) sütunun ihtiyacımız yok o yüzden sileceğiz.

train <- train %>% select(- Name)
test <- test %>% select(- Name)

Şimdi Bu dönüşüm işlemi, sütunlardaki eksik veya boş değerleri belirli bir değerle (burada 0 ile) doldurmayı amaçlacımızdır. Bu, veri analizi veya modelleme sürecinde eksik değerlerin doğru şekilde ele alınması ve sorun yaşanmaması için önemlidir.

train <- train %>% 
  mutate(RoomService=coalesce(RoomService, 0),
         FoodCourt=coalesce(FoodCourt, 0),
         ShoppingMall=coalesce(ShoppingMall, 0),
         Spa=coalesce(Spa,0),
         VRDeck=coalesce(VRDeck,0))
test <- test %>% 
  mutate(RoomService=coalesce(RoomService, 0),
         FoodCourt=coalesce(FoodCourt, 0),
         ShoppingMall=coalesce(ShoppingMall, 0),
         Spa=coalesce(Spa,0),
         VRDeck=coalesce(VRDeck,0))

Artık train ve test setinin içinde na kısmın değişkenler sıfırdır.

“Train ve test” veri çerçevesine “aile” adlı bir sütun eklenir. Bu sütun, “ailenum” sütunundaki yinelenen değerlerin varlığını gösteren bir etiket alır, 1 yinelenen değerler için ve 0 yinelenmeyen değerler için işaretlenir.

train$aile <- ifelse(duplicated(train$ailenum) | duplicated(train$ailenum, fromLast = TRUE),1,0 )
test$aile <- ifelse(duplicated(test$ailenum) | duplicated(test$ailenum, fromLast = TRUE),1,0 )
head(train[,c("PassengerId","ailenum","aile")],20 )
## # A tibble: 20 × 3
##    PassengerId ailenum  aile
##    <chr>       <chr>   <dbl>
##  1 0001_01     0001        0
##  2 0002_01     0002        0
##  3 0003_01     0003        1
##  4 0003_02     0003        1
##  5 0004_01     0004        0
##  6 0005_01     0005        0
##  7 0006_01     0006        1
##  8 0006_02     0006        1
##  9 0007_01     0007        0
## 10 0008_01     0008        1
## 11 0008_02     0008        1
## 12 0008_03     0008        1
## 13 0009_01     0009        0
## 14 0010_01     0010        0
## 15 0011_01     0011        0
## 16 0012_01     0012        0
## 17 0014_01     0014        0
## 18 0015_01     0015        0
## 19 0016_01     0016        0
## 20 0017_01     0017        1
head(test[,c("PassengerId","ailenum","aile")],20 )
## # A tibble: 20 × 3
##    PassengerId ailenum  aile
##    <chr>       <chr>   <dbl>
##  1 0013_01     0013        0
##  2 0018_01     0018        0
##  3 0019_01     0019        0
##  4 0021_01     0021        0
##  5 0023_01     0023        0
##  6 0027_01     0027        0
##  7 0029_01     0029        0
##  8 0032_01     0032        1
##  9 0032_02     0032        1
## 10 0033_01     0033        0
## 11 0037_01     0037        0
## 12 0040_01     0040        1
## 13 0040_02     0040        1
## 14 0042_01     0042        0
## 15 0046_01     0046        1
## 16 0046_02     0046        1
## 17 0046_03     0046        1
## 18 0047_01     0047        1
## 19 0047_02     0047        1
## 20 0047_03     0047        1

Artık bunlar (ailenum,ailesıra ve Num ) ihtiyacımız yok onlara sileceğiz.

train <- train %>% select(- c(ailenum,ailesıra ,Num ))
test <- test %>% select(- c(ailenum,ailesıra ,Num ))
most_frequent_hp <- train %>%
  filter(!is.na(HomePlanet)) %>%
  group_by(Destination, HomePlanet) %>%
  summarize(count = n()) %>%
  arrange(Destination, desc(count)) %>%
  slice(1) %>%
  ungroup()
## `summarise()` has grouped output by 'Destination'. You can override using the
## `.groups` argument.
most_frequent_hp <- test %>%
  filter(!is.na(HomePlanet)) %>%
  group_by(Destination, HomePlanet) %>%
  summarize(count = n()) %>%
  arrange(Destination, desc(count)) %>%
  slice(1) %>%
  ungroup()
## `summarise()` has grouped output by 'Destination'. You can override using the
## `.groups` argument.
most_frequent_hp
## # A tibble: 4 × 3
##   Destination   HomePlanet count
##   <fct>         <fct>      <int>
## 1 55 Cancri e   Europa       424
## 2 PSO J318.5-22 Earth        353
## 3 TRAPPIST-1e   Earth       1571
## 4 <NA>          Earth         45
train$HomePlanet <- as.character(train$HomePlanet)
train$Destination<- as.character(train$Destination)
test$HomePlanet <- as.character(test$HomePlanet)
test$Destination<- as.character(test$Destination)
train <- train %>%
  mutate(HomePlanet = ifelse(is.na(HomePlanet) & Destination == "55 cancri e", "Europa",                                     ifelse(is.na(HomePlanet), "Earth", HomePlanet)))
test <- test %>%
  mutate(HomePlanet = ifelse(is.na(HomePlanet) & Destination == "55 cancri e", "Europa",                                     ifelse(is.na(HomePlanet), "Earth", HomePlanet)))
train <- transform (train, HomePlanet  = replace (HomePlanet, is.na (HomePlanet), "Earth"))
test <- transform (test, HomePlanet  = replace (HomePlanet, is.na (HomePlanet), "Earth"))
most_frequent_Destinations <- train %>%
  filter(!is.na(Destination)) %>%
  group_by(HomePlanet, Destination) %>%
  summarize(count = n()) %>%
  arrange(Destination, desc(count)) %>%
  slice(1) %>%
  ungroup()
## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Destinations <- test %>%
  filter(!is.na(Destination)) %>%
  group_by(HomePlanet, Destination) %>%
  summarize(count = n()) %>%
  arrange(Destination, desc(count)) %>%
  slice(1) %>%
  ungroup()
## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Destinations
## # A tibble: 3 × 3
##   HomePlanet Destination count
##   <chr>      <chr>       <int>
## 1 Earth      55 Cancri e   316
## 2 Europa     55 Cancri e   424
## 3 Mars       55 Cancri e   101
train<- transform (train, Destination = replace(Destination, is.na(Destination), "TRAPPIST-1e"))
test<- transform (test, Destination = replace(Destination, is.na(Destination), "TRAPPIST-1e"))
train$HomePlanet <- as.factor (train$HomePlanet)
train $destination <- as.factor(train$Destination)
test$HomePlanet <- as.factor (test$HomePlanet)
test$destination <- as.factor(test$Destination)
train <- train %>%
  group_by(HomePlanet, Destination) %>%
  mutate_at(vars(Age), ~replace_na(., mean (., na.rm =TRUE)))
test <- test %>%
  group_by(HomePlanet, Destination) %>%
  mutate_at(vars(Age), ~replace_na(., mean (., na.rm =TRUE)))
train$expense <- train$RoomService + train$FoodCourt + train$ShoppingMall + train$Spa + train$VRDeck
test$expense <- test$RoomService + test$FoodCourt + test$ShoppingMall + test$Spa + test$VRDeck
train <- transform(train, CryoSleep = replace(CryoSleep,is.na(CryoSleep) & expense>0 & Age>12, "FALSE"))
test <- transform(test, CryoSleep = replace(CryoSleep,is.na(CryoSleep) & expense>0 & Age>12, "FALSE"))
summary(train)
##  PassengerId         HomePlanet   CryoSleep         Deck      Side     
##  Length:8693        Earth :4803   FALSE:5439   F      :2794     : 199  
##  Class :character   Europa:2131   TRUE :3037   G      :2559   P :4206  
##  Mode  :character   Mars  :1759   NA   : 217   E      : 876   S :4288  
##                                                B      : 779   NA:   0  
##                                                C      : 747            
##                                                D      : 478            
##                                                (Other): 460            
##  Destination             Age           VIP        RoomService   
##  Length:8693        Min.   : 0.00   FALSE:8291   Min.   :    0  
##  Class :character   1st Qu.:20.00   TRUE : 199   1st Qu.:    0  
##  Mode  :character   Median :27.00   NA   : 203   Median :    0  
##                     Mean   :28.83                Mean   :  220  
##                     3rd Qu.:37.00                3rd Qu.:   41  
##                     Max.   :79.00                Max.   :14327  
##                                                                 
##    FoodCourt        ShoppingMall          Spa              VRDeck       
##  Min.   :    0.0   Min.   :    0.0   Min.   :    0.0   Min.   :    0.0  
##  1st Qu.:    0.0   1st Qu.:    0.0   1st Qu.:    0.0   1st Qu.:    0.0  
##  Median :    0.0   Median :    0.0   Median :    0.0   Median :    0.0  
##  Mean   :  448.4   Mean   :  169.6   Mean   :  304.6   Mean   :  298.3  
##  3rd Qu.:   61.0   3rd Qu.:   22.0   3rd Qu.:   53.0   3rd Qu.:   40.0  
##  Max.   :29813.0   Max.   :23492.0   Max.   :22408.0   Max.   :24133.0  
##                                                                         
##  Transported          aile               destination      expense     
##  Mode :logical   Min.   :0.0000   55 Cancri e  :1800   Min.   :    0  
##  FALSE:4315      1st Qu.:0.0000   PSO J318.5-22: 796   1st Qu.:    0  
##  TRUE :4378      Median :0.0000   TRAPPIST-1e  :6097   Median :  716  
##                  Mean   :0.4473                        Mean   : 1441  
##                  3rd Qu.:1.0000                        3rd Qu.: 1441  
##                  Max.   :1.0000                        Max.   :35987  
## 
summary(test)
##  PassengerId         HomePlanet   CryoSleep         Deck      Side     
##  Length:4277        Earth :2350   FALSE:2640   F      :1445     : 100  
##  Class :character   Europa:1002   TRUE :1544   G      :1222   P :2084  
##  Mode  :character   Mars  : 925   NA   :  93   E      : 447   S :2093  
##                                                B      : 362   NA:   0  
##                                                C      : 355            
##                                                D      : 242            
##                                                (Other): 204            
##  Destination             Age           VIP        RoomService     
##  Length:4277        Min.   : 0.00   FALSE:4110   Min.   :    0.0  
##  Class :character   1st Qu.:20.00   TRUE :  74   1st Qu.:    0.0  
##  Mode  :character   Median :26.22   NA   :  93   Median :    0.0  
##                     Mean   :28.66                Mean   :  215.1  
##                     3rd Qu.:37.00                3rd Qu.:   48.0  
##                     Max.   :79.00                Max.   :11567.0  
##                                                                   
##    FoodCourt        ShoppingMall         Spa              VRDeck       
##  Min.   :    0.0   Min.   :   0.0   Min.   :    0.0   Min.   :    0.0  
##  1st Qu.:    0.0   1st Qu.:   0.0   1st Qu.:    0.0   1st Qu.:    0.0  
##  Median :    0.0   Median :   0.0   Median :    0.0   Median :    0.0  
##  Mean   :  428.6   Mean   : 173.2   Mean   :  295.9   Mean   :  304.9  
##  3rd Qu.:   66.0   3rd Qu.:  27.0   3rd Qu.:   43.0   3rd Qu.:   31.0  
##  Max.   :25273.0   Max.   :8292.0   Max.   :19844.0   Max.   :22272.0  
##                                                                        
##       aile               destination      expense     
##  Min.   :0.0000   55 Cancri e  : 841   Min.   :    0  
##  1st Qu.:0.0000   PSO J318.5-22: 388   1st Qu.:    0  
##  Median :0.0000   TRAPPIST-1e  :3048   Median :  714  
##  Mean   :0.4529                        Mean   : 1418  
##  3rd Qu.:1.0000                        3rd Qu.: 1444  
##  Max.   :1.0000                        Max.   :33666  
## 
train <- transform(train, CryoSleep = replace(CryoSleep, is.na(CryoSleep) & expense==0 & Age>12, "TRUE"))
test <- transform(test, CryoSleep = replace(CryoSleep, is.na(CryoSleep) & expense==0 & Age>12, "TRUE"))
train$CryoSleep <- as.factor(train$CryoSleep)
test$CryoSleep <- as.factor(test$CryoSleep)
most_frequent_Deck <- train %>%
  filter(!is.na(Deck)) %>%
  group_by(HomePlanet,  Deck) %>%
  summarize (count = n()) %>%
  arrange(HomePlanet, desc(count)) %>%
  slice (1) %>%
  ungroup ()
## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Deck <- test %>%
  filter(!is.na(Deck)) %>%
  group_by(HomePlanet,  Deck) %>%
  summarize (count = n()) %>%
  arrange(HomePlanet, desc(count)) %>%
  slice (1) %>%
  ungroup ()
## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Deck
## # A tibble: 3 × 3
##   HomePlanet Deck  count
##   <fct>      <fct> <int>
## 1 Earth      G      1222
## 2 Europa     B       358
## 3 Mars       F       603
train$HomePlanet  <- as.character(train$HomePlanet)
train$Deck <- as.character(train$Deck)
test$HomePlanet  <- as.character(test$HomePlanet)
test$Deck <- as.character(test$Deck)
train <- train %>%
  mutate(Deck = ifelse(is.na(Deck) & HomePlanet == "Earth", "G",
                       ifelse(is.na(Deck) & HomePlanet == "Europa", "B" ,
                              ifelse(is.na(Deck) & HomePlanet == "Mars", "F", Deck))))
test <- test %>%
  mutate(Deck = ifelse(is.na(Deck) & HomePlanet == "Earth", "G",
                       ifelse(is.na(Deck) & HomePlanet == "Europa", "B" ,
                              ifelse(is.na(Deck) & HomePlanet == "Mars", "F", Deck))))
most_frequent_Side <- train %>%
  filter(!is.na(Side)) %>%
  group_by(HomePlanet,  Side) %>%
  summarize (count = n()) %>%
  arrange(HomePlanet, desc(count)) %>%
  slice (1) %>%
  ungroup ()
## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Side <- test %>%
  filter(!is.na(Side)) %>%
  group_by(HomePlanet,  Side) %>%
  summarize (count = n()) %>%
  arrange(HomePlanet, desc(count)) %>%
  slice (1) %>%
  ungroup ()
## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Side
## # A tibble: 3 × 3
##   HomePlanet Side  count
##   <chr>      <fct> <int>
## 1 Earth      P      1147
## 2 Europa     P       495
## 3 Mars       S       463
train$Side <- as.character(train$Side)
test$Side <- as.character(test$Side)
train <- train %>%
  mutate(Side = ifelse(is.na(Side) & HomePlanet == "Earth", "P",
                       ifelse(is.na(Side) & HomePlanet == "Europa", "S" ,
                              ifelse(is.na(Side) & HomePlanet == "Mars", "P", Side))))
test <- test %>%
  mutate(Side = ifelse(is.na(Side) & HomePlanet == "Earth", "P",
                       ifelse(is.na(Side) & HomePlanet == "Europa", "S" ,
                              ifelse(is.na(Side) & HomePlanet == "Mars", "P", Side))))
train <- train %>% mutate_if(is.character,as.factor)
test <- test %>% mutate_if(is.character,as.factor)

Train ve test veri setimizin yapısına bakalım.

train %>% describe_all()
## # A tibble: 17 × 8
##    variable     type     na na_pct unique   min    mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl>   <dbl> <dbl>
##  1 PassengerId  fct       0      0   8693    NA   NA       NA
##  2 HomePlanet   fct       0      0      3    NA   NA       NA
##  3 CryoSleep    fct       0      0      3    NA   NA       NA
##  4 Deck         fct       0      0      8    NA   NA       NA
##  5 Side         fct       0      0      3    NA   NA       NA
##  6 Destination  fct       0      0      3    NA   NA       NA
##  7 Age          dbl       0      0     91     0   28.8     79
##  8 VIP          fct       0      0      3    NA   NA       NA
##  9 RoomService  dbl       0      0   1273     0  220.   14327
## 10 FoodCourt    dbl       0      0   1507     0  448.   29813
## 11 ShoppingMall dbl       0      0   1115     0  170.   23492
## 12 Spa          dbl       0      0   1327     0  305.   22408
## 13 VRDeck       dbl       0      0   1306     0  298.   24133
## 14 Transported  lgl       0      0      2     0    0.5      1
## 15 aile         dbl       0      0      2     0    0.45     1
## 16 destination  fct       0      0      3    NA   NA       NA
## 17 expense      dbl       0      0   2336     0 1441.   35987
test %>% describe_all()
## # A tibble: 16 × 8
##    variable     type     na na_pct unique   min    mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl>   <dbl> <dbl>
##  1 PassengerId  fct       0      0   4277    NA   NA       NA
##  2 HomePlanet   fct       0      0      3    NA   NA       NA
##  3 CryoSleep    fct       0      0      3    NA   NA       NA
##  4 Deck         fct       0      0      8    NA   NA       NA
##  5 Side         fct       0      0      3    NA   NA       NA
##  6 Destination  fct       0      0      3    NA   NA       NA
##  7 Age          dbl       0      0     91     0   28.7     79
##  8 VIP          fct       0      0      3    NA   NA       NA
##  9 RoomService  dbl       0      0    842     0  215.   11567
## 10 FoodCourt    dbl       0      0    902     0  429.   25273
## 11 ShoppingMall dbl       0      0    715     0  173.    8292
## 12 Spa          dbl       0      0    833     0  296.   19844
## 13 VRDeck       dbl       0      0    796     0  305.   22272
## 14 aile         dbl       0      0      2     0    0.45     1
## 15 destination  fct       0      0      3    NA   NA       NA
## 16 expense      dbl       0      0   1437     0 1418.   33666
cat("Eğitilmiş veri kümesinin şekli şöyledir: ", nrow(train), " satırlar ve ", ncol(train), " columns\n")
## Eğitilmiş veri kümesinin şekli şöyledir:  8693  satırlar ve  17  columns
cat("Eğitilmiş veri kümesinin şekli şöyledir: ", nrow(test), " satırlar ve ", ncol(test), " columns\n")
## Eğitilmiş veri kümesinin şekli şöyledir:  4277  satırlar ve  16  columns

Train veri kümesinde 8693 satır ve 17 sütun ve test veri kümesinde 4277 satır ve 16 sütun var.

Şimdi sayısal değişkenleri histogramlara görselleştirelim.

Age

hist(train$Age)

hist(test$Age)

Train ve test, yaş değişkeninde aykırı değerler vardır ve dağılım oldukça normaldir.

RoomService

hist(train$RoomService)

hist(test$RoomService)

RoomService’deki verilerin çoğunluğu sola doğru dağılmıştır ve aykırı değerleri fazladır.Bunları düzelteceğiz.

Spa

hist(train$Spa)

hist(test$Spa)

RoomService’inkine benzer bir dağılım var. Çok fazla aykırı değer içerir ve normal olarak dağıtılmaz.

VRDeck

hist(train$VRDeck)

hist(test$VRDeck)

RoomService, FoodCourt, ShoppingMall, Spa, VRDeck, yolcunun Spaceship Titanic’in birçok lüks olanağının her birinde faturalandırdığı tutardır, bu yüzden VRDeck, FoodCourt ve ShoppingMall’un benzer bir dağılımı olup olmadığını görelim.

FoodCourt

hist(train$FoodCourt)

hist(test$FoodCourt)

ShoppingMall

hist(train$ShoppingMall)

hist(test$ShoppingMall)

VRDeck, FoodCourt ve ShoppingMall’un benzer bir dağılıma sahip olduğunu görebiliriz. Hepsi normal olarak dağılmamıştır ve hepsinin aykırı değerleri vardır.

Aile

hist(train$aile)

hist(test$aile)

Expense

hist(train$expense)

hist(test$expense)

D <- train[,2:17] %>% mutate(across(everything(), ~as.integer(.)))
kor<- cor(D)
library(corrplot)
corrplot.mixed(kor)

Logistic Regresyon

Logistik regresyon, istatistiksel bir modeldir ve sınıflandırma problemlerinde kullanılır. İki kategorili bağımlı değişkenleri tahmin etmek için bağımsız değişkenlerin değerlerini kullanır. Log-odds oranlarını temel alır ve bağımlı değişkenin olasılığını tahmin etmeyi amaçlar. Model eğitiminde regresyon katsayıları tahmin edilir ve bu katsayılarla yeni verilere dayalı olarak sınıflandırma yapılır. Logistik regresyon, sağlam istatistiksel özelliklere sahiptir ve birçok farklı alanda başarıyla kullanılır.

train_set <- train[2:17]
test_set <- test[2:16]
library(caTools)
set.seed(123)
split = sample.split(train_set$Transported,SplitRatio = 0.75)
training_set = subset(train_set, split == TRUE)
testing_set = subset(train_set, split == FALSE)
logistic = glm(formula = Transported ~ . ,family = binomial, data = training_set)
## Warning: glm.fit: des probabilités ont été ajustées numériquement à 0 ou 1
prob_pred = predict(logistic, type = 'response', newdata = testing_set[-13])
y_pred = ifelse(prob_pred > 0.5,1,0)
y_true <- ifelse(testing_set[13] == TRUE,1,0)
cm = table(y_true, y_pred)
cm
##       y_pred
## y_true   0   1
##      0 819 260
##      1 188 906
(819+906)/(819+906+260+188)
## [1] 0.7938334
logistic_son = glm(formula = Transported ~ . ,family = binomial, data = train_set)
## Warning: glm.fit: des probabilités ont été ajustées numériquement à 0 ou 1
prob_pred = predict(logistic_son, type = 'response', newdata = test_set)
y_pred = ifelse(prob_pred > 0.5, TRUE,FALSE)
Transported <- as.character(y_pred)
PassengerId <- test$PassengerId
Transported <- as.vector(Transported)
submission <- cbind(PassengerId,Transported)
submission <- as.data.frame(submission)
library(stringr)
write.csv(submission,"sub_logis.csv", row.names = FALSE,quote = FALSE)

SVM

SVM, sınıflandırma ve regresyon problemleri için kullanılan bir makine öğrenimi algoritmasıdır. Ana amacı, sınıfları ayıran bir hiper düzlemi bulmaktır. Bu düzlemi belirleyen noktalara “destek vektör” denir. SVM, sınıflar arasındaki marjı (mesafe) maksimize eden en iyi hiper düzlemi bulmaya yardımcı olur. Ayrıca, non-lineer sınıflandırma problemlerini çözebilmek için kernel fonksiyonlarını kullanır. SVM, genelleme yeteneğini artırmak amacıyla sınıflar arasındaki marjı maksimize eder.

library(e1071)
fit_svm <- svm(Transported ~ ., data = training_set, type = 'C-classification', kernel = 'linear')
preds <- predict(fit_svm, newdata = testing_set, type = "raw") %>% data.frame()

Kısaca, e1071 paketi kullanılarak bir destek vektör makinesi (SVM) modeli eğitilir ve bu model kullanılarak yeni veriye sınıflandırma tahminleri yapılır.

y_pred = ifelse(preds$. == TRUE , 1, 0)
cm = table(y_true, y_pred)
cm
##       y_pred
## y_true   0   1
##      0 832 247
##      1 177 917
(832+917)/(832+917+247+177)
## [1] 0.804878
svm_son <- svm(Transported ~ ., data = train_set, type = 'C-classification', kernel = 'linear')
preds <- predict(svm_son, newdata = test_set, type = "response") %>% data.frame()
y_pred = preds$.
Transported <- as.character(y_pred)
PassengerId <- test$PassengerId
Transported <- as.vector(Transported)
submission <- cbind(PassengerId, Transported)
submission <- as.data.frame(submission)
submission$Transported <- str_to_title(submission$Transported)
write.csv(submission, "sub_svm.csv",row.names = FALSE, quote = FALSE)

Dicision Trees

Karar ağaçları, sınıflandırma ve regresyon problemlerini çözmek için kullanılan bir makine öğrenimi algoritmasıdır. Karar ağaçlarında, veri kümesi üzerinde yapılan testlere göre ağaç dallara ayrılır ve sonuçlar yaprak düğümlerinde temsil edilir. Karar ağaçları, veri analizi, tanımlayıcı analiz ve tahmin problemlerinde kullanılır ve veri hazırlama sürecini kolaylaştıran özellik seçimi yapabilir. Ayrıca, Random Forest veya Gradient Boosting gibi gelişmiş teknikler de kullanılabilir.

library(rpart)
library(rpart.plot)
library(randomForest)
library(caret)
training_set$Transported <- as.factor(training_set$Transported)
testing_set$Transported <- as.factor(testing_set$Transported)
train_set$Transported <- as.factor(train_set$Transported)
fit_tree <- rpart(Transported ~ . , data = training_set)
summary(fit_tree)
## Call:
## rpart(formula = Transported ~ ., data = training_set)
##   n= 6520 
## 
##           CP nsplit rel error    xerror       xstd
## 1 0.47126082      0 1.0000000 1.0000000 0.01247595
## 2 0.02487639      1 0.5287392 0.5287392 0.01097792
## 3 0.01205192      3 0.4789864 0.4808405 0.01063625
## 4 0.01127936      4 0.4669345 0.4749691 0.01059132
## 5 0.01112485      6 0.4443758 0.4694067 0.01054812
## 6 0.01000000      7 0.4332509 0.4629172 0.01049692
## 
## Variable importance
##      expense    CryoSleep          Spa    FoodCourt       VRDeck  RoomService 
##           25           20           14           14           12           11 
## ShoppingMall         Deck   HomePlanet 
##            3            1            1 
## 
## Node number 1: 6520 observations,    complexity param=0.4712608
##   predicted class=TRUE   expected loss=0.496319  P(node) =1
##     class counts:  3236  3284
##    probabilities: 0.496 0.504 
##   left son=2 (3795 obs) right son=3 (2725 obs)
##   Primary splits:
##       expense     < 0.5     to the right, improve=760.2351, (0 missing)
##       CryoSleep   splits as  LRL,         improve=683.0583, (0 missing)
##       RoomService < 0.5     to the right, improve=408.8215, (0 missing)
##       Spa         < 0.5     to the right, improve=372.3306, (0 missing)
##       VRDeck      < 0.5     to the right, improve=353.3889, (0 missing)
##   Surrogate splits:
##       CryoSleep   splits as  LRL,         agree=0.932, adj=0.837, (0 split)
##       Spa         < 0.5     to the right, agree=0.784, adj=0.484, (0 split)
##       FoodCourt   < 0.5     to the right, agree=0.770, adj=0.451, (0 split)
##       VRDeck      < 0.5     to the right, agree=0.768, adj=0.444, (0 split)
##       RoomService < 0.5     to the right, agree=0.757, adj=0.419, (0 split)
## 
## Node number 2: 3795 observations,    complexity param=0.02487639
##   predicted class=FALSE  expected loss=0.2990777  P(node) =0.5820552
##     class counts:  2660  1135
##    probabilities: 0.701 0.299 
##   left son=4 (3243 obs) right son=5 (552 obs)
##   Primary splits:
##       FoodCourt    < 1331    to the left,  improve=91.50674, (0 missing)
##       ShoppingMall < 627.5   to the left,  improve=72.08933, (0 missing)
##       RoomService  < 365.5   to the right, improve=70.65071, (0 missing)
##       Spa          < 257.5   to the right, improve=53.26111, (0 missing)
##       VRDeck       < 721     to the right, improve=36.44331, (0 missing)
##   Surrogate splits:
##       expense    < 5981    to the left,  agree=0.885, adj=0.210, (0 split)
##       Deck       splits as  RRRLLLLL,    agree=0.884, adj=0.199, (0 split)
##       HomePlanet splits as  LRL,         agree=0.878, adj=0.159, (0 split)
##       Spa        < 8955.5  to the left,  agree=0.856, adj=0.009, (0 split)
##       VRDeck     < 11692   to the left,  agree=0.856, adj=0.009, (0 split)
## 
## Node number 3: 2725 observations
##   predicted class=TRUE   expected loss=0.2113761  P(node) =0.4179448
##     class counts:   576  2149
##    probabilities: 0.211 0.789 
## 
## Node number 4: 3243 observations,    complexity param=0.01127936
##   predicted class=FALSE  expected loss=0.2537774  P(node) =0.4973926
##     class counts:  2420   823
##    probabilities: 0.746 0.254 
##   left son=8 (2577 obs) right son=9 (666 obs)
##   Primary splits:
##       ShoppingMall < 541.5   to the left,  improve=90.77444, (0 missing)
##       RoomService  < 365.5   to the right, improve=42.55270, (0 missing)
##       Spa          < 240.5   to the right, improve=40.86813, (0 missing)
##       VRDeck       < 114     to the right, improve=34.19891, (0 missing)
##       expense      < 2867.5  to the right, improve=22.24986, (0 missing)
##   Surrogate splits:
##       expense < 18644   to the left,  agree=0.795, adj=0.003, (0 split)
## 
## Node number 5: 552 observations,    complexity param=0.02487639
##   predicted class=TRUE   expected loss=0.4347826  P(node) =0.08466258
##     class counts:   240   312
##    probabilities: 0.435 0.565 
##   left son=10 (123 obs) right son=11 (429 obs)
##   Primary splits:
##       Spa       < 1372.5  to the right, improve=57.714490, (0 missing)
##       VRDeck    < 1063.5  to the right, improve=46.364550, (0 missing)
##       expense   < 5395    to the right, improve=17.936380, (0 missing)
##       Side      splits as  LLR,         improve= 7.616904, (0 missing)
##       FoodCourt < 2513    to the left,  improve= 7.542383, (0 missing)
##   Surrogate splits:
##       expense     < 12647   to the right, agree=0.790, adj=0.057, (0 split)
##       Age         < 13.5    to the left,  agree=0.779, adj=0.008, (0 split)
##       RoomService < 3895.5  to the right, agree=0.779, adj=0.008, (0 split)
## 
## Node number 8: 2577 observations
##   predicted class=FALSE  expected loss=0.193636  P(node) =0.3952454
##     class counts:  2078   499
##    probabilities: 0.806 0.194 
## 
## Node number 9: 666 observations,    complexity param=0.01127936
##   predicted class=FALSE  expected loss=0.4864865  P(node) =0.1021472
##     class counts:   342   324
##    probabilities: 0.514 0.486 
##   left son=18 (157 obs) right son=19 (509 obs)
##   Primary splits:
##       RoomService  < 310     to the right, improve=31.364140, (0 missing)
##       Spa          < 200     to the right, improve=19.676920, (0 missing)
##       ShoppingMall < 1586.5  to the left,  improve=13.953460, (0 missing)
##       VRDeck       < 120.5   to the right, improve=13.338010, (0 missing)
##       HomePlanet   splits as  RLL,         improve= 6.864426, (0 missing)
##   Surrogate splits:
##       FoodCourt < 1102    to the right, agree=0.766, adj=0.006, (0 split)
## 
## Node number 10: 123 observations
##   predicted class=FALSE  expected loss=0.1382114  P(node) =0.01886503
##     class counts:   106    17
##    probabilities: 0.862 0.138 
## 
## Node number 11: 429 observations,    complexity param=0.01205192
##   predicted class=TRUE   expected loss=0.3123543  P(node) =0.06579755
##     class counts:   134   295
##    probabilities: 0.312 0.688 
##   left son=22 (143 obs) right son=23 (286 obs)
##   Primary splits:
##       VRDeck      < 611     to the right, improve=45.037300, (0 missing)
##       Spa         < 225     to the right, improve= 9.768013, (0 missing)
##       FoodCourt   < 3119.5  to the left,  improve= 9.524139, (0 missing)
##       Side        splits as  LLR,         improve= 7.968308, (0 missing)
##       RoomService < 1719.5  to the right, improve= 7.244213, (0 missing)
##   Surrogate splits:
##       expense   < 6032    to the right, agree=0.702, adj=0.105, (0 split)
##       Age       < 53.5    to the right, agree=0.674, adj=0.021, (0 split)
##       FoodCourt < 12128.5 to the right, agree=0.671, adj=0.014, (0 split)
## 
## Node number 18: 157 observations
##   predicted class=FALSE  expected loss=0.2101911  P(node) =0.02407975
##     class counts:   124    33
##    probabilities: 0.790 0.210 
## 
## Node number 19: 509 observations,    complexity param=0.01112485
##   predicted class=TRUE   expected loss=0.4282908  P(node) =0.07806748
##     class counts:   218   291
##    probabilities: 0.428 0.572 
##   left son=38 (66 obs) right son=39 (443 obs)
##   Primary splits:
##       Spa          < 209     to the right, improve=17.99311, (0 missing)
##       VRDeck       < 33.5    to the right, improve=17.76420, (0 missing)
##       ShoppingMall < 1540.5  to the left,  improve=10.97812, (0 missing)
##       Deck         splits as  LRLRLRR-,    improve= 7.32794, (0 missing)
##       expense      < 4099.5  to the right, improve= 5.09474, (0 missing)
##   Surrogate splits:
##       expense      < 4161    to the right, agree=0.898, adj=0.212, (0 split)
##       Deck         splits as  LLLRRRR-,    agree=0.896, adj=0.197, (0 split)
##       HomePlanet   splits as  RLR,         agree=0.884, adj=0.106, (0 split)
##       ShoppingMall < 7126    to the right, agree=0.876, adj=0.045, (0 split)
##       VRDeck       < 1206.5  to the right, agree=0.876, adj=0.045, (0 split)
## 
## Node number 22: 143 observations
##   predicted class=FALSE  expected loss=0.3636364  P(node) =0.02193252
##     class counts:    91    52
##    probabilities: 0.636 0.364 
## 
## Node number 23: 286 observations
##   predicted class=TRUE   expected loss=0.1503497  P(node) =0.04386503
##     class counts:    43   243
##    probabilities: 0.150 0.850 
## 
## Node number 38: 66 observations
##   predicted class=FALSE  expected loss=0.2272727  P(node) =0.0101227
##     class counts:    51    15
##    probabilities: 0.773 0.227 
## 
## Node number 39: 443 observations
##   predicted class=TRUE   expected loss=0.3769752  P(node) =0.06794479
##     class counts:   167   276
##    probabilities: 0.377 0.623
rpart.plot(fit_tree)

Ağacımız şöyle bir ifade veriyor: Diyor ki eğer Transported deki CryoSleep aldılarsa ( FALSE ve ya NA) no ya denk geliyor ama CryoSleep almadılarsa (TRUE) eğer RoomService daha fazla para ediyorsa onlar Transported olmamış ama daha az ediyorsa fakat Spa ya daha fazla verirse onlar Transported olmuş. Diyelim ki RoomService 366 daha üzeri ver Spa ya 205 daha üzeri ver ve VRDeck ye 248 daha fazla üzeri ver ama diyor ki FoodCourt en azından sonuç 0 dan fazla uygularsa , onlar da Transported olmuş

preds = predict(fit_tree, newdata = testing_set[-13],type = "class")
y_pred = ifelse(preds == TRUE, 1,0)
cm = table(y_true, y_pred)
cm
##       y_pred
## y_true   0   1
##      0 798 281
##      1 196 898
(798+898)/(798+898+281+196)
## [1] 0.7804878
fit_tree <- rpart(Transported ~ ., data = train_set)
summary(fit_tree)
## Call:
## rpart(formula = Transported ~ ., data = train_set)
##   n= 8693 
## 
##           CP nsplit rel error    xerror        xstd
## 1 0.47045191      0 1.0000000 1.0000000 0.010803454
## 2 0.02143685      1 0.5295481 0.5295481 0.009511274
## 3 0.01993048      3 0.4866744 0.5015064 0.009342999
## 4 0.01008111      4 0.4667439 0.4769409 0.009184965
## 5 0.01000000      7 0.4347625 0.4602549 0.009071689
## 
## Variable importance
##      expense    CryoSleep          Spa    FoodCourt       VRDeck  RoomService 
##           25           19           14           14           12           11 
## ShoppingMall   HomePlanet         Deck 
##            3            1            1 
## 
## Node number 1: 8693 observations,    complexity param=0.4704519
##   predicted class=TRUE   expected loss=0.4963764  P(node) =1
##     class counts:  4315  4378
##    probabilities: 0.496 0.504 
##   left son=2 (5040 obs) right son=3 (3653 obs)
##   Primary splits:
##       expense     < 0.5    to the right, improve=1008.1870, (0 missing)
##       CryoSleep   splits as  LRL,        improve= 920.2004, (0 missing)
##       RoomService < 0.5    to the right, improve= 526.1339, (0 missing)
##       Spa         < 0.5    to the right, improve= 514.4709, (0 missing)
##       VRDeck      < 0.5    to the right, improve= 479.5694, (0 missing)
##   Surrogate splits:
##       CryoSleep   splits as  LRL,        agree=0.929, adj=0.831, (0 split)
##       Spa         < 0.5    to the right, agree=0.787, adj=0.492, (0 split)
##       FoodCourt   < 0.5    to the right, agree=0.772, adj=0.456, (0 split)
##       VRDeck      < 0.5    to the right, agree=0.766, adj=0.444, (0 split)
##       RoomService < 0.5    to the right, agree=0.758, adj=0.424, (0 split)
## 
## Node number 2: 5040 observations,    complexity param=0.02143685
##   predicted class=FALSE  expected loss=0.2986111  P(node) =0.5797768
##     class counts:  3535  1505
##    probabilities: 0.701 0.299 
##   left son=4 (3827 obs) right son=5 (1213 obs)
##   Primary splits:
##       FoodCourt    < 668.5  to the left,  improve=136.56630, (0 missing)
##       ShoppingMall < 627.5  to the left,  improve= 95.37971, (0 missing)
##       RoomService  < 346.5  to the right, improve= 80.08985, (0 missing)
##       Spa          < 452.5  to the right, improve= 73.42969, (0 missing)
##       VRDeck       < 613.5  to the right, improve= 48.08949, (0 missing)
##   Surrogate splits:
##       HomePlanet splits as  LRL,        agree=0.845, adj=0.358, (0 split)
##       Deck       splits as  RRRLLLLR,   agree=0.834, adj=0.309, (0 split)
##       expense    < 3589.5 to the left,  agree=0.823, adj=0.263, (0 split)
##       Spa        < 4710.5 to the left,  agree=0.763, adj=0.016, (0 split)
##       VRDeck     < 4254.5 to the left,  agree=0.762, adj=0.012, (0 split)
## 
## Node number 3: 3653 observations
##   predicted class=TRUE   expected loss=0.2135231  P(node) =0.4202232
##     class counts:   780  2873
##    probabilities: 0.214 0.786 
## 
## Node number 4: 3827 observations,    complexity param=0.01008111
##   predicted class=FALSE  expected loss=0.2330807  P(node) =0.4402393
##     class counts:  2935   892
##    probabilities: 0.767 0.233 
##   left son=8 (2961 obs) right son=9 (866 obs)
##   Primary splits:
##       ShoppingMall < 541    to the left,  improve=143.35850, (0 missing)
##       Spa          < 445.5  to the right, improve= 42.68628, (0 missing)
##       RoomService  < 343    to the right, improve= 33.14410, (0 missing)
##       VRDeck       < 114    to the right, improve= 29.58027, (0 missing)
##       expense      < 2716   to the right, improve= 16.47007, (0 missing)
##   Surrogate splits:
##       expense < 18816  to the left,  agree=0.774, adj=0.002, (0 split)
## 
## Node number 5: 1213 observations,    complexity param=0.02143685
##   predicted class=TRUE   expected loss=0.4946414  P(node) =0.1395376
##     class counts:   600   613
##    probabilities: 0.495 0.505 
##   left son=10 (230 obs) right son=11 (983 obs)
##   Primary splits:
##       Spa       < 1372.5 to the right, improve=81.65183, (0 missing)
##       VRDeck    < 1114.5 to the right, improve=62.27701, (0 missing)
##       FoodCourt < 2507.5 to the left,  improve=27.03659, (0 missing)
##       Side      splits as  LLR,        improve=17.25608, (0 missing)
##       expense   < 6806   to the right, improve=16.91079, (0 missing)
##   Surrogate splits:
##       expense     < 12957  to the right, agree=0.821, adj=0.057, (0 split)
##       Age         < 67.5   to the right, agree=0.812, adj=0.009, (0 split)
##       RoomService < 5257.5 to the right, agree=0.811, adj=0.004, (0 split)
## 
## Node number 8: 2961 observations
##   predicted class=FALSE  expected loss=0.1590679  P(node) =0.3406189
##     class counts:  2490   471
##    probabilities: 0.841 0.159 
## 
## Node number 9: 866 observations,    complexity param=0.01008111
##   predicted class=FALSE  expected loss=0.4861432  P(node) =0.09962038
##     class counts:   445   421
##    probabilities: 0.514 0.486 
##   left son=18 (201 obs) right son=19 (665 obs)
##   Primary splits:
##       RoomService  < 310    to the right, improve=36.00767, (0 missing)
##       Spa          < 200    to the right, improve=28.52457, (0 missing)
##       VRDeck       < 120.5  to the right, improve=15.13357, (0 missing)
##       ShoppingMall < 1968.5 to the left,  improve=14.39760, (0 missing)
##       HomePlanet   splits as  RLL,        improve=11.22426, (0 missing)
##   Surrogate splits:
##       FoodCourt < 591.5  to the right, agree=0.769, adj=0.005, (0 split)
##       expense   < 2343.5 to the right, agree=0.769, adj=0.005, (0 split)
## 
## Node number 10: 230 observations
##   predicted class=FALSE  expected loss=0.126087  P(node) =0.02645807
##     class counts:   201    29
##    probabilities: 0.874 0.126 
## 
## Node number 11: 983 observations,    complexity param=0.01993048
##   predicted class=TRUE   expected loss=0.4059003  P(node) =0.1130795
##     class counts:   399   584
##    probabilities: 0.406 0.594 
##   left son=22 (136 obs) right son=23 (847 obs)
##   Primary splits:
##       VRDeck      < 1685.5 to the right, improve=53.136330, (0 missing)
##       FoodCourt   < 3125   to the left,  improve=36.358710, (0 missing)
##       RoomService < 364    to the right, improve=14.375770, (0 missing)
##       Side        splits as  LLR,        improve=14.088100, (0 missing)
##       Deck        splits as  LRRRLLLL,   improve= 9.618287, (0 missing)
##   Surrogate splits:
##       expense < 12248  to the right, agree=0.87, adj=0.059, (0 split)
## 
## Node number 18: 201 observations
##   predicted class=FALSE  expected loss=0.2238806  P(node) =0.02312205
##     class counts:   156    45
##    probabilities: 0.776 0.224 
## 
## Node number 19: 665 observations,    complexity param=0.01008111
##   predicted class=TRUE   expected loss=0.4345865  P(node) =0.07649833
##     class counts:   289   376
##    probabilities: 0.435 0.565 
##   left son=38 (87 obs) right son=39 (578 obs)
##   Primary splits:
##       Spa          < 201    to the right, improve=25.731350, (0 missing)
##       VRDeck       < 128.5  to the right, improve=19.130020, (0 missing)
##       ShoppingMall < 1815.5 to the left,  improve=10.453460, (0 missing)
##       Deck         splits as  LRLRLRR-,   improve= 5.398272, (0 missing)
##       expense      < 2958.5 to the right, improve= 4.372934, (0 missing)
##   Surrogate splits:
##       expense      < 3966.5 to the right, agree=0.898, adj=0.218, (0 split)
##       Deck         splits as  LLLRRRR-,   agree=0.893, adj=0.184, (0 split)
##       HomePlanet   splits as  RLR,        agree=0.889, adj=0.149, (0 split)
##       VRDeck       < 1206.5 to the right, agree=0.880, adj=0.080, (0 split)
##       ShoppingMall < 7126   to the right, agree=0.874, adj=0.034, (0 split)
## 
## Node number 22: 136 observations
##   predicted class=FALSE  expected loss=0.1838235  P(node) =0.01564477
##     class counts:   111    25
##    probabilities: 0.816 0.184 
## 
## Node number 23: 847 observations
##   predicted class=TRUE   expected loss=0.3400236  P(node) =0.09743472
##     class counts:   288   559
##    probabilities: 0.340 0.660 
## 
## Node number 38: 87 observations
##   predicted class=FALSE  expected loss=0.2068966  P(node) =0.01000805
##     class counts:    69    18
##    probabilities: 0.793 0.207 
## 
## Node number 39: 578 observations
##   predicted class=TRUE   expected loss=0.3806228  P(node) =0.06649028
##     class counts:   220   358
##    probabilities: 0.381 0.619
rpart.plot(fit_tree)

Ağacımız şöyle bir ifade veriyor:

Eğer CryoSleep almışlarsa onlar Transported olmuş ama RoomService fazla almışsa ölmüştür. Ve diyor ki RoomService daha az para harcamışlar , yine Spa ya daha fazla para harcamış ama onlar da daha fazla harcamışsa bile VRDeck eğer 355 daha fazla para harcamışsa ölmüştür. Bu üçü ( RoomService, Spa ve VRDeck ) eğer 347 den fazla harcamamışsa , 205 den fazla harcamamışsa, 355 den fazla harcamamışsa uyku anmışlarsa bile ölmüştür.

preds = predict(fit_tree, newdata = test_set,type = "class")
y_pred = ifelse(preds == TRUE,TRUE,FALSE)
Transported <- as.character(y_pred)
PassengerId <- test$PassengerId
Transported <- as.vector(Transported)
submission <- cbind(PassengerId,Transported)
submission <- as.data.frame(submission)
submission$Transported <- str_to_title(submission$Transported)
write.csv(submission,"subm_dt.csv", row.names = FALSE,quote = FALSE)

Naive Bayes

Naive Bayes, sınıflandırma problemlerinde kullanılan bir makine öğrenme algoritmasıdır. Bu algoritma, Bayes teoremini kullanarak sınıflandırma yapar. Naive Bayes, özellikle metin sınıflandırma gibi doğal dil işleme problemlerinde ve spam filtrelemede yaygın olarak kullanılır. Bu algoritma basit ve hızlıdır ancak bağımsızlık varsayımının gerçek durumu yansıtmadığı durumlarda yanlı sonuçlara yol açabilir.

library(e1071)
fit_nb <- naiveBayes(Transported ~ . , data = training_set)
preds <-predict(fit_nb, newdata = testing_set[-13], type = "raw") %>% data.frame()
y_pred = ifelse(preds$TRUE. > 0.5,1,0)
cm = table(y_true, y_pred)
cm
##       y_pred
## y_true    0    1
##      0  487  592
##      1   79 1015
(487+1015)/(487+1015+592+79)
## [1] 0.6912103
nb_son = naiveBayes(Transported ~ . , data = train_set)
preds <-predict(nb_son, newdata = test_set, type = "raw") %>% data.frame()
y_pred = ifelse(preds$TRUE. > 0.5, TRUE,FALSE)
Transported <- as.character(y_pred)
PassengerId <- test$PassengerId
Transported <- as.vector(Transported)
submission <- cbind(PassengerId,Transported)
submission <- as.data.frame(submission)
write.csv(submission,"sub_nb.csv", row.names = FALSE,quote = FALSE)