1910701526
SPACESCHIP TİTANİC KAGGLE COMPETİTİON
Kaggle, ünlü Titanic yarışmasına benzer şekilde Spaceship Titanic adında eğlenceli bir yarışma başlattı. Bu yarışma, veri bilimine yeni başlayanlar için makine öğrenimi hakkında bilgi edinmenin, Kaggle’ı keşfetmenin ve topluluğun diğer üyeleriyle tanışmanın bir yoludur. Bu makale Spaceship Titanic yarışmasını analiz etmekte ve RandomForestClassifier kullanarak test seti için anlamlı içgörülerin nasıl elde edileceğini ve “temel gerçeğin” yaklaşık %80 doğrulukla nasıl tahmin edileceğini açıklamaktadır.
Problem tanımı
Uzay gemisi Titanik bir uzay-zaman anomalisiyle çarpıştı ve yolcuların yarısı farklı bir boyuta taşındı. Bu kayıp yolcuları bulmak için geminin hasarlı bilgisayar sisteminden geri yüklenen kayıtları kullanmamız gerekiyor.
Veriler hakkında
Train dosyası (spaceship_titanic_train.csv) — Kişisel kayıtlar,makine öğrenimi modeli oluşturmak için kullanılacak yolcuların bilgilerini içerir. Test dosyası (spaceship_titanic_test.csv) — Metin, yolcuların üçtebirinin kişisel kayıtlarını içerdiğini, ancak Taşınabilir değeri içermediğini belirtmektedir. Bu verilerin, modelin görünmeyen veriler üzerinde ne kadar iyi performans gösterdiğini ölçmek için kullanılacağı söylenmektedir.
Bu problem çözmek için Rstudio ve RMarkdown kullanacağız.
Veri Okuma
Veri setinin yapısını kontrol edeceğiz.Önce özelliklere bakacağız, sonra türleri kontrol edeceğiz.
Train veri setinde 13 bağımsız değişken ve 1 hedef değişken bulunuyor(Transported). Aynı şekilde test veri setinde de sütunlar mevcut. Test veri kümesinde, train veri kümesiyle benzer özelliklere sahip olduğumuz için tahminlerimizi train verileriyle oluşturulan modele dayanarak yapacağız.
Yukarıdaki değişkenleri açıklamasına vereceğiz.
*PassengerId:Her yolcu için benzersiz bir kimlik olan gggg_pp formatındaki kimlikler,gggg’nin yolcunun seyahat ettiği grubu ve pp’nin grup içindeki numarasını gösterir. Bu grup genellikle aile üyelerinden oluşur, fakat her zaman olmayabilir.
*HomePlanet:Yolcunun ayrıldığı gezegen daimi ikamet gezegeni olarak bilinir.
*CryoSleep:Yolcunun yolculuk süresinde dondurucu uykuya alınıp alınmayacağını gösterir. Yolcular cabinlere kapatılır.
*Cabin :Yolcunun kaldığı cabin numarası.deck/num/side biçimini alır, burada yan İskele için P veya Sancak için S olabilir.
*Destination:Yolcunun ineceği gezegen.
*Age:Yolcunun yaşı.
*VIP:Yolcunun yolculuk sırasında özel VIP hizmeti için ödeme yapıp yapmadığı.
*Roomservice,Foodcourt,ShoppingMall,Spa,VRDeck:Yolcunun Uzay Gemisi Titanic’in faturalandırılan her bir lüks özelliği için tutar ödediği belirtiliyor.
*Name:Yolcunun adı ve soyadı.
*Transported:Gözlemlediğimiz şey, yolcunun başka bir boyuta geçip geçmediğini belirlemeye çalışmak. Bunun hedefi sütundur.
## [1] "Europa" "Earth" "Mars" NA
## [1] "Earth" "Europa" "Mars" NA
## [1] FALSE TRUE NA
## [1] TRUE FALSE NA
## [1] "TRAPPIST-1e" "PSO J318.5-22" "55 Cancri e" NA
## [1] "TRAPPIST-1e" "55 Cancri e" "PSO J318.5-22" NA
Tidyverse
Tidyverse, veri analizi ve görselleştirme için kullanılan bir paket serisidir. Temiz, düzenli ve etkili bir veri analizi süreci sağlar.
train ve test veri çerçevesinde boş değerleri kontrol eder ve boş olan hücrelere NA değeri atanır.
## # A tibble: 14 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId chr 0 0 8693 NA NA NA
## 2 HomePlanet chr 201 2.3 4 NA NA NA
## 3 CryoSleep lgl 217 2.5 3 0 0.36 1
## 4 Cabin chr 199 2.3 6561 NA NA NA
## 5 Destination chr 182 2.1 4 NA NA NA
## 6 Age dbl 179 2.1 81 0 28.8 79
## 7 VIP lgl 203 2.3 3 0 0.02 1
## 8 RoomService dbl 181 2.1 1274 0 225. 14327
## 9 FoodCourt dbl 183 2.1 1508 0 458. 29813
## 10 ShoppingMall dbl 208 2.4 1116 0 174. 23492
## 11 Spa dbl 183 2.1 1328 0 311. 22408
## 12 VRDeck dbl 188 2.2 1307 0 305. 24133
## 13 Name chr 200 2.3 8474 NA NA NA
## 14 Transported lgl 0 0 2 0 0.5 1
## # A tibble: 13 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId chr 0 0 4277 NA NA NA
## 2 HomePlanet chr 87 2 4 NA NA NA
## 3 CryoSleep lgl 93 2.2 3 0 0.37 1
## 4 Cabin chr 100 2.3 3266 NA NA NA
## 5 Destination chr 92 2.2 4 NA NA NA
## 6 Age dbl 91 2.1 80 0 28.7 79
## 7 VIP lgl 93 2.2 3 0 0.02 1
## 8 RoomService dbl 82 1.9 843 0 219. 11567
## 9 FoodCourt dbl 106 2.5 903 0 439. 25273
## 10 ShoppingMall dbl 98 2.3 716 0 177. 8292
## 11 Spa dbl 101 2.4 834 0 303. 19844
## 12 VRDeck dbl 80 1.9 797 0 311. 22272
## 13 Name chr 94 2.2 4177 NA NA NA
“train ve test” veri çerçevesindeki “Cabin” sütununu karakterine göre bölerek, “Deck”, “Num” ve “Side” adlı yeni sütunları oluşturur. Sonra İskele için P veya Sancak için S kolayca göstereceğiz.
Train ve test bulunan PassengerId sütununu içinde iki tane bilgi (gggg_pp) vardı ve bu iki bilgi çıkartacağız,(gggg)ailenum ve (pp)ailesıra.
Train ve test, PassengerId sütunundaki ayırılan hücreyi başıya alalım.
Şimdi train ve test, Cabinden üç parçaya ayırılan hücreyi (Deck ,Num ,Side) Cabinin yanına alalım.
Train ve test veri kümesinin Cabinin bilgisine aldık,artık ihtiyacımız kadığı için temizliğeceğiz.
addNA() fonksiyonu, belirtilen vektöre veya faktöre NA değerini eklemek için kullanılır.Örneğin, eğer “train ve test” veri sütunun başlangıçta NA değerleri içermiyorsa, bu kod parçası her hücreye bir NA değeri ekleyecektir. Eğer “train ve test” veri sütunun zaten NA değerleri içeriyorsa, bu kod değişikliğe neden olmayacaktır.
Şimdi “train” ve “test” veri çerçevelerinin “HomePlanet” ve “Destination” sütunlarına göre gruplanmasını sağlar. Daha sonra, eksik değerleri (NA) “Age” sütununda ortalama değerlerle değiştiririz.
library(tidyr)
train <- train %>%
group_by(HomePlanet, Destination) %>%
mutate_at(vars(Age), ~replace_na(., mean(., na.rm = TRUE)))test <- test %>%
group_by(HomePlanet, Destination) %>%
mutate_at(vars(Age), ~replace_na(., mean(., na.rm = TRUE)))Train ve test, isim(Name) sütunun ihtiyacımız yok o yüzden sileceğiz.
Şimdi Bu dönüşüm işlemi, sütunlardaki eksik veya boş değerleri belirli bir değerle (burada 0 ile) doldurmayı amaçlacımızdır. Bu, veri analizi veya modelleme sürecinde eksik değerlerin doğru şekilde ele alınması ve sorun yaşanmaması için önemlidir.
train <- train %>%
mutate(RoomService=coalesce(RoomService, 0),
FoodCourt=coalesce(FoodCourt, 0),
ShoppingMall=coalesce(ShoppingMall, 0),
Spa=coalesce(Spa,0),
VRDeck=coalesce(VRDeck,0))
test <- test %>%
mutate(RoomService=coalesce(RoomService, 0),
FoodCourt=coalesce(FoodCourt, 0),
ShoppingMall=coalesce(ShoppingMall, 0),
Spa=coalesce(Spa,0),
VRDeck=coalesce(VRDeck,0))Artık train ve test setinin içinde na kısmın değişkenler sıfırdır.
“Train ve test” veri çerçevesine “aile” adlı bir sütun eklenir. Bu sütun, “ailenum” sütunundaki yinelenen değerlerin varlığını gösteren bir etiket alır, 1 yinelenen değerler için ve 0 yinelenmeyen değerler için işaretlenir.
train$aile <- ifelse(duplicated(train$ailenum) | duplicated(train$ailenum, fromLast = TRUE),1,0 )
test$aile <- ifelse(duplicated(test$ailenum) | duplicated(test$ailenum, fromLast = TRUE),1,0 )## # A tibble: 20 × 3
## PassengerId ailenum aile
## <chr> <chr> <dbl>
## 1 0001_01 0001 0
## 2 0002_01 0002 0
## 3 0003_01 0003 1
## 4 0003_02 0003 1
## 5 0004_01 0004 0
## 6 0005_01 0005 0
## 7 0006_01 0006 1
## 8 0006_02 0006 1
## 9 0007_01 0007 0
## 10 0008_01 0008 1
## 11 0008_02 0008 1
## 12 0008_03 0008 1
## 13 0009_01 0009 0
## 14 0010_01 0010 0
## 15 0011_01 0011 0
## 16 0012_01 0012 0
## 17 0014_01 0014 0
## 18 0015_01 0015 0
## 19 0016_01 0016 0
## 20 0017_01 0017 1
## # A tibble: 20 × 3
## PassengerId ailenum aile
## <chr> <chr> <dbl>
## 1 0013_01 0013 0
## 2 0018_01 0018 0
## 3 0019_01 0019 0
## 4 0021_01 0021 0
## 5 0023_01 0023 0
## 6 0027_01 0027 0
## 7 0029_01 0029 0
## 8 0032_01 0032 1
## 9 0032_02 0032 1
## 10 0033_01 0033 0
## 11 0037_01 0037 0
## 12 0040_01 0040 1
## 13 0040_02 0040 1
## 14 0042_01 0042 0
## 15 0046_01 0046 1
## 16 0046_02 0046 1
## 17 0046_03 0046 1
## 18 0047_01 0047 1
## 19 0047_02 0047 1
## 20 0047_03 0047 1
Artık bunlar (ailenum,ailesıra ve Num ) ihtiyacımız yok onlara sileceğiz.
train <- train %>% select(- c(ailenum,ailesıra ,Num ))
test <- test %>% select(- c(ailenum,ailesıra ,Num ))most_frequent_hp <- train %>%
filter(!is.na(HomePlanet)) %>%
group_by(Destination, HomePlanet) %>%
summarize(count = n()) %>%
arrange(Destination, desc(count)) %>%
slice(1) %>%
ungroup()## `summarise()` has grouped output by 'Destination'. You can override using the
## `.groups` argument.
most_frequent_hp <- test %>%
filter(!is.na(HomePlanet)) %>%
group_by(Destination, HomePlanet) %>%
summarize(count = n()) %>%
arrange(Destination, desc(count)) %>%
slice(1) %>%
ungroup()## `summarise()` has grouped output by 'Destination'. You can override using the
## `.groups` argument.
## # A tibble: 4 × 3
## Destination HomePlanet count
## <fct> <fct> <int>
## 1 55 Cancri e Europa 424
## 2 PSO J318.5-22 Earth 353
## 3 TRAPPIST-1e Earth 1571
## 4 <NA> Earth 45
train$HomePlanet <- as.character(train$HomePlanet)
train$Destination<- as.character(train$Destination)
test$HomePlanet <- as.character(test$HomePlanet)
test$Destination<- as.character(test$Destination)train <- train %>%
mutate(HomePlanet = ifelse(is.na(HomePlanet) & Destination == "55 cancri e", "Europa", ifelse(is.na(HomePlanet), "Earth", HomePlanet)))test <- test %>%
mutate(HomePlanet = ifelse(is.na(HomePlanet) & Destination == "55 cancri e", "Europa", ifelse(is.na(HomePlanet), "Earth", HomePlanet)))train <- transform (train, HomePlanet = replace (HomePlanet, is.na (HomePlanet), "Earth"))
test <- transform (test, HomePlanet = replace (HomePlanet, is.na (HomePlanet), "Earth"))most_frequent_Destinations <- train %>%
filter(!is.na(Destination)) %>%
group_by(HomePlanet, Destination) %>%
summarize(count = n()) %>%
arrange(Destination, desc(count)) %>%
slice(1) %>%
ungroup()## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Destinations <- test %>%
filter(!is.na(Destination)) %>%
group_by(HomePlanet, Destination) %>%
summarize(count = n()) %>%
arrange(Destination, desc(count)) %>%
slice(1) %>%
ungroup()## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
## # A tibble: 3 × 3
## HomePlanet Destination count
## <chr> <chr> <int>
## 1 Earth 55 Cancri e 316
## 2 Europa 55 Cancri e 424
## 3 Mars 55 Cancri e 101
train<- transform (train, Destination = replace(Destination, is.na(Destination), "TRAPPIST-1e"))
test<- transform (test, Destination = replace(Destination, is.na(Destination), "TRAPPIST-1e"))train$HomePlanet <- as.factor (train$HomePlanet)
train $destination <- as.factor(train$Destination)
test$HomePlanet <- as.factor (test$HomePlanet)
test$destination <- as.factor(test$Destination)train <- train %>%
group_by(HomePlanet, Destination) %>%
mutate_at(vars(Age), ~replace_na(., mean (., na.rm =TRUE)))
test <- test %>%
group_by(HomePlanet, Destination) %>%
mutate_at(vars(Age), ~replace_na(., mean (., na.rm =TRUE)))train$expense <- train$RoomService + train$FoodCourt + train$ShoppingMall + train$Spa + train$VRDeck
test$expense <- test$RoomService + test$FoodCourt + test$ShoppingMall + test$Spa + test$VRDecktrain <- transform(train, CryoSleep = replace(CryoSleep,is.na(CryoSleep) & expense>0 & Age>12, "FALSE"))
test <- transform(test, CryoSleep = replace(CryoSleep,is.na(CryoSleep) & expense>0 & Age>12, "FALSE"))## PassengerId HomePlanet CryoSleep Deck Side
## Length:8693 Earth :4803 FALSE:5439 F :2794 : 199
## Class :character Europa:2131 TRUE :3037 G :2559 P :4206
## Mode :character Mars :1759 NA : 217 E : 876 S :4288
## B : 779 NA: 0
## C : 747
## D : 478
## (Other): 460
## Destination Age VIP RoomService
## Length:8693 Min. : 0.00 FALSE:8291 Min. : 0
## Class :character 1st Qu.:20.00 TRUE : 199 1st Qu.: 0
## Mode :character Median :27.00 NA : 203 Median : 0
## Mean :28.83 Mean : 220
## 3rd Qu.:37.00 3rd Qu.: 41
## Max. :79.00 Max. :14327
##
## FoodCourt ShoppingMall Spa VRDeck
## Min. : 0.0 Min. : 0.0 Min. : 0.0 Min. : 0.0
## 1st Qu.: 0.0 1st Qu.: 0.0 1st Qu.: 0.0 1st Qu.: 0.0
## Median : 0.0 Median : 0.0 Median : 0.0 Median : 0.0
## Mean : 448.4 Mean : 169.6 Mean : 304.6 Mean : 298.3
## 3rd Qu.: 61.0 3rd Qu.: 22.0 3rd Qu.: 53.0 3rd Qu.: 40.0
## Max. :29813.0 Max. :23492.0 Max. :22408.0 Max. :24133.0
##
## Transported aile destination expense
## Mode :logical Min. :0.0000 55 Cancri e :1800 Min. : 0
## FALSE:4315 1st Qu.:0.0000 PSO J318.5-22: 796 1st Qu.: 0
## TRUE :4378 Median :0.0000 TRAPPIST-1e :6097 Median : 716
## Mean :0.4473 Mean : 1441
## 3rd Qu.:1.0000 3rd Qu.: 1441
## Max. :1.0000 Max. :35987
##
## PassengerId HomePlanet CryoSleep Deck Side
## Length:4277 Earth :2350 FALSE:2640 F :1445 : 100
## Class :character Europa:1002 TRUE :1544 G :1222 P :2084
## Mode :character Mars : 925 NA : 93 E : 447 S :2093
## B : 362 NA: 0
## C : 355
## D : 242
## (Other): 204
## Destination Age VIP RoomService
## Length:4277 Min. : 0.00 FALSE:4110 Min. : 0.0
## Class :character 1st Qu.:20.00 TRUE : 74 1st Qu.: 0.0
## Mode :character Median :26.22 NA : 93 Median : 0.0
## Mean :28.66 Mean : 215.1
## 3rd Qu.:37.00 3rd Qu.: 48.0
## Max. :79.00 Max. :11567.0
##
## FoodCourt ShoppingMall Spa VRDeck
## Min. : 0.0 Min. : 0.0 Min. : 0.0 Min. : 0.0
## 1st Qu.: 0.0 1st Qu.: 0.0 1st Qu.: 0.0 1st Qu.: 0.0
## Median : 0.0 Median : 0.0 Median : 0.0 Median : 0.0
## Mean : 428.6 Mean : 173.2 Mean : 295.9 Mean : 304.9
## 3rd Qu.: 66.0 3rd Qu.: 27.0 3rd Qu.: 43.0 3rd Qu.: 31.0
## Max. :25273.0 Max. :8292.0 Max. :19844.0 Max. :22272.0
##
## aile destination expense
## Min. :0.0000 55 Cancri e : 841 Min. : 0
## 1st Qu.:0.0000 PSO J318.5-22: 388 1st Qu.: 0
## Median :0.0000 TRAPPIST-1e :3048 Median : 714
## Mean :0.4529 Mean : 1418
## 3rd Qu.:1.0000 3rd Qu.: 1444
## Max. :1.0000 Max. :33666
##
train <- transform(train, CryoSleep = replace(CryoSleep, is.na(CryoSleep) & expense==0 & Age>12, "TRUE"))
test <- transform(test, CryoSleep = replace(CryoSleep, is.na(CryoSleep) & expense==0 & Age>12, "TRUE"))most_frequent_Deck <- train %>%
filter(!is.na(Deck)) %>%
group_by(HomePlanet, Deck) %>%
summarize (count = n()) %>%
arrange(HomePlanet, desc(count)) %>%
slice (1) %>%
ungroup ()## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Deck <- test %>%
filter(!is.na(Deck)) %>%
group_by(HomePlanet, Deck) %>%
summarize (count = n()) %>%
arrange(HomePlanet, desc(count)) %>%
slice (1) %>%
ungroup ()## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
## # A tibble: 3 × 3
## HomePlanet Deck count
## <fct> <fct> <int>
## 1 Earth G 1222
## 2 Europa B 358
## 3 Mars F 603
train$HomePlanet <- as.character(train$HomePlanet)
train$Deck <- as.character(train$Deck)
test$HomePlanet <- as.character(test$HomePlanet)
test$Deck <- as.character(test$Deck)train <- train %>%
mutate(Deck = ifelse(is.na(Deck) & HomePlanet == "Earth", "G",
ifelse(is.na(Deck) & HomePlanet == "Europa", "B" ,
ifelse(is.na(Deck) & HomePlanet == "Mars", "F", Deck))))
test <- test %>%
mutate(Deck = ifelse(is.na(Deck) & HomePlanet == "Earth", "G",
ifelse(is.na(Deck) & HomePlanet == "Europa", "B" ,
ifelse(is.na(Deck) & HomePlanet == "Mars", "F", Deck))))most_frequent_Side <- train %>%
filter(!is.na(Side)) %>%
group_by(HomePlanet, Side) %>%
summarize (count = n()) %>%
arrange(HomePlanet, desc(count)) %>%
slice (1) %>%
ungroup ()## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
most_frequent_Side <- test %>%
filter(!is.na(Side)) %>%
group_by(HomePlanet, Side) %>%
summarize (count = n()) %>%
arrange(HomePlanet, desc(count)) %>%
slice (1) %>%
ungroup ()## `summarise()` has grouped output by 'HomePlanet'. You can override using the
## `.groups` argument.
## # A tibble: 3 × 3
## HomePlanet Side count
## <chr> <fct> <int>
## 1 Earth P 1147
## 2 Europa P 495
## 3 Mars S 463
train <- train %>%
mutate(Side = ifelse(is.na(Side) & HomePlanet == "Earth", "P",
ifelse(is.na(Side) & HomePlanet == "Europa", "S" ,
ifelse(is.na(Side) & HomePlanet == "Mars", "P", Side))))
test <- test %>%
mutate(Side = ifelse(is.na(Side) & HomePlanet == "Earth", "P",
ifelse(is.na(Side) & HomePlanet == "Europa", "S" ,
ifelse(is.na(Side) & HomePlanet == "Mars", "P", Side))))train <- train %>% mutate_if(is.character,as.factor)
test <- test %>% mutate_if(is.character,as.factor)Train ve test veri setimizin yapısına bakalım.
## # A tibble: 17 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId fct 0 0 8693 NA NA NA
## 2 HomePlanet fct 0 0 3 NA NA NA
## 3 CryoSleep fct 0 0 3 NA NA NA
## 4 Deck fct 0 0 8 NA NA NA
## 5 Side fct 0 0 3 NA NA NA
## 6 Destination fct 0 0 3 NA NA NA
## 7 Age dbl 0 0 91 0 28.8 79
## 8 VIP fct 0 0 3 NA NA NA
## 9 RoomService dbl 0 0 1273 0 220. 14327
## 10 FoodCourt dbl 0 0 1507 0 448. 29813
## 11 ShoppingMall dbl 0 0 1115 0 170. 23492
## 12 Spa dbl 0 0 1327 0 305. 22408
## 13 VRDeck dbl 0 0 1306 0 298. 24133
## 14 Transported lgl 0 0 2 0 0.5 1
## 15 aile dbl 0 0 2 0 0.45 1
## 16 destination fct 0 0 3 NA NA NA
## 17 expense dbl 0 0 2336 0 1441. 35987
## # A tibble: 16 × 8
## variable type na na_pct unique min mean max
## <chr> <chr> <int> <dbl> <int> <dbl> <dbl> <dbl>
## 1 PassengerId fct 0 0 4277 NA NA NA
## 2 HomePlanet fct 0 0 3 NA NA NA
## 3 CryoSleep fct 0 0 3 NA NA NA
## 4 Deck fct 0 0 8 NA NA NA
## 5 Side fct 0 0 3 NA NA NA
## 6 Destination fct 0 0 3 NA NA NA
## 7 Age dbl 0 0 91 0 28.7 79
## 8 VIP fct 0 0 3 NA NA NA
## 9 RoomService dbl 0 0 842 0 215. 11567
## 10 FoodCourt dbl 0 0 902 0 429. 25273
## 11 ShoppingMall dbl 0 0 715 0 173. 8292
## 12 Spa dbl 0 0 833 0 296. 19844
## 13 VRDeck dbl 0 0 796 0 305. 22272
## 14 aile dbl 0 0 2 0 0.45 1
## 15 destination fct 0 0 3 NA NA NA
## 16 expense dbl 0 0 1437 0 1418. 33666
cat("Eğitilmiş veri kümesinin şekli şöyledir: ", nrow(train), " satırlar ve ", ncol(train), " columns\n")## Eğitilmiş veri kümesinin şekli şöyledir: 8693 satırlar ve 17 columns
cat("Eğitilmiş veri kümesinin şekli şöyledir: ", nrow(test), " satırlar ve ", ncol(test), " columns\n")## Eğitilmiş veri kümesinin şekli şöyledir: 4277 satırlar ve 16 columns
Train veri kümesinde 8693 satır ve 17 sütun ve test veri kümesinde 4277 satır ve 16 sütun var.
Şimdi sayısal değişkenleri histogramlara görselleştirelim.
Age
Train ve test, yaş değişkeninde aykırı değerler vardır ve dağılım oldukça normaldir.
RoomService
RoomService’deki verilerin çoğunluğu sola doğru dağılmıştır ve aykırı değerleri fazladır.Bunları düzelteceğiz.
Spa
RoomService’inkine benzer bir dağılım var. Çok fazla aykırı değer içerir ve normal olarak dağıtılmaz.
VRDeck
RoomService, FoodCourt, ShoppingMall, Spa, VRDeck, yolcunun Spaceship Titanic’in birçok lüks olanağının her birinde faturalandırdığı tutardır, bu yüzden VRDeck, FoodCourt ve ShoppingMall’un benzer bir dağılımı olup olmadığını görelim.
ShoppingMall
VRDeck, FoodCourt ve ShoppingMall’un benzer bir dağılıma sahip olduğunu görebiliriz. Hepsi normal olarak dağılmamıştır ve hepsinin aykırı değerleri vardır.
Expense
Logistic Regresyon
Logistik regresyon, istatistiksel bir modeldir ve sınıflandırma problemlerinde kullanılır. İki kategorili bağımlı değişkenleri tahmin etmek için bağımsız değişkenlerin değerlerini kullanır. Log-odds oranlarını temel alır ve bağımlı değişkenin olasılığını tahmin etmeyi amaçlar. Model eğitiminde regresyon katsayıları tahmin edilir ve bu katsayılarla yeni verilere dayalı olarak sınıflandırma yapılır. Logistik regresyon, sağlam istatistiksel özelliklere sahiptir ve birçok farklı alanda başarıyla kullanılır.
library(caTools)
set.seed(123)
split = sample.split(train_set$Transported,SplitRatio = 0.75)
training_set = subset(train_set, split == TRUE)
testing_set = subset(train_set, split == FALSE)## Warning: glm.fit: des probabilités ont été ajustées numériquement à 0 ou 1
## y_pred
## y_true 0 1
## 0 819 260
## 1 188 906
## [1] 0.7938334
## Warning: glm.fit: des probabilités ont été ajustées numériquement à 0 ou 1
SVM
SVM, sınıflandırma ve regresyon problemleri için kullanılan bir makine öğrenimi algoritmasıdır. Ana amacı, sınıfları ayıran bir hiper düzlemi bulmaktır. Bu düzlemi belirleyen noktalara “destek vektör” denir. SVM, sınıflar arasındaki marjı (mesafe) maksimize eden en iyi hiper düzlemi bulmaya yardımcı olur. Ayrıca, non-lineer sınıflandırma problemlerini çözebilmek için kernel fonksiyonlarını kullanır. SVM, genelleme yeteneğini artırmak amacıyla sınıflar arasındaki marjı maksimize eder.
library(e1071)
fit_svm <- svm(Transported ~ ., data = training_set, type = 'C-classification', kernel = 'linear')
preds <- predict(fit_svm, newdata = testing_set, type = "raw") %>% data.frame()Kısaca, e1071 paketi kullanılarak bir destek vektör makinesi (SVM) modeli eğitilir ve bu model kullanılarak yeni veriye sınıflandırma tahminleri yapılır.
## y_pred
## y_true 0 1
## 0 832 247
## 1 177 917
## [1] 0.804878
Dicision Trees
Karar ağaçları, sınıflandırma ve regresyon problemlerini çözmek için kullanılan bir makine öğrenimi algoritmasıdır. Karar ağaçlarında, veri kümesi üzerinde yapılan testlere göre ağaç dallara ayrılır ve sonuçlar yaprak düğümlerinde temsil edilir. Karar ağaçları, veri analizi, tanımlayıcı analiz ve tahmin problemlerinde kullanılır ve veri hazırlama sürecini kolaylaştıran özellik seçimi yapabilir. Ayrıca, Random Forest veya Gradient Boosting gibi gelişmiş teknikler de kullanılabilir.
training_set$Transported <- as.factor(training_set$Transported)
testing_set$Transported <- as.factor(testing_set$Transported)
train_set$Transported <- as.factor(train_set$Transported)## Call:
## rpart(formula = Transported ~ ., data = training_set)
## n= 6520
##
## CP nsplit rel error xerror xstd
## 1 0.47126082 0 1.0000000 1.0000000 0.01247595
## 2 0.02487639 1 0.5287392 0.5287392 0.01097792
## 3 0.01205192 3 0.4789864 0.4808405 0.01063625
## 4 0.01127936 4 0.4669345 0.4749691 0.01059132
## 5 0.01112485 6 0.4443758 0.4694067 0.01054812
## 6 0.01000000 7 0.4332509 0.4629172 0.01049692
##
## Variable importance
## expense CryoSleep Spa FoodCourt VRDeck RoomService
## 25 20 14 14 12 11
## ShoppingMall Deck HomePlanet
## 3 1 1
##
## Node number 1: 6520 observations, complexity param=0.4712608
## predicted class=TRUE expected loss=0.496319 P(node) =1
## class counts: 3236 3284
## probabilities: 0.496 0.504
## left son=2 (3795 obs) right son=3 (2725 obs)
## Primary splits:
## expense < 0.5 to the right, improve=760.2351, (0 missing)
## CryoSleep splits as LRL, improve=683.0583, (0 missing)
## RoomService < 0.5 to the right, improve=408.8215, (0 missing)
## Spa < 0.5 to the right, improve=372.3306, (0 missing)
## VRDeck < 0.5 to the right, improve=353.3889, (0 missing)
## Surrogate splits:
## CryoSleep splits as LRL, agree=0.932, adj=0.837, (0 split)
## Spa < 0.5 to the right, agree=0.784, adj=0.484, (0 split)
## FoodCourt < 0.5 to the right, agree=0.770, adj=0.451, (0 split)
## VRDeck < 0.5 to the right, agree=0.768, adj=0.444, (0 split)
## RoomService < 0.5 to the right, agree=0.757, adj=0.419, (0 split)
##
## Node number 2: 3795 observations, complexity param=0.02487639
## predicted class=FALSE expected loss=0.2990777 P(node) =0.5820552
## class counts: 2660 1135
## probabilities: 0.701 0.299
## left son=4 (3243 obs) right son=5 (552 obs)
## Primary splits:
## FoodCourt < 1331 to the left, improve=91.50674, (0 missing)
## ShoppingMall < 627.5 to the left, improve=72.08933, (0 missing)
## RoomService < 365.5 to the right, improve=70.65071, (0 missing)
## Spa < 257.5 to the right, improve=53.26111, (0 missing)
## VRDeck < 721 to the right, improve=36.44331, (0 missing)
## Surrogate splits:
## expense < 5981 to the left, agree=0.885, adj=0.210, (0 split)
## Deck splits as RRRLLLLL, agree=0.884, adj=0.199, (0 split)
## HomePlanet splits as LRL, agree=0.878, adj=0.159, (0 split)
## Spa < 8955.5 to the left, agree=0.856, adj=0.009, (0 split)
## VRDeck < 11692 to the left, agree=0.856, adj=0.009, (0 split)
##
## Node number 3: 2725 observations
## predicted class=TRUE expected loss=0.2113761 P(node) =0.4179448
## class counts: 576 2149
## probabilities: 0.211 0.789
##
## Node number 4: 3243 observations, complexity param=0.01127936
## predicted class=FALSE expected loss=0.2537774 P(node) =0.4973926
## class counts: 2420 823
## probabilities: 0.746 0.254
## left son=8 (2577 obs) right son=9 (666 obs)
## Primary splits:
## ShoppingMall < 541.5 to the left, improve=90.77444, (0 missing)
## RoomService < 365.5 to the right, improve=42.55270, (0 missing)
## Spa < 240.5 to the right, improve=40.86813, (0 missing)
## VRDeck < 114 to the right, improve=34.19891, (0 missing)
## expense < 2867.5 to the right, improve=22.24986, (0 missing)
## Surrogate splits:
## expense < 18644 to the left, agree=0.795, adj=0.003, (0 split)
##
## Node number 5: 552 observations, complexity param=0.02487639
## predicted class=TRUE expected loss=0.4347826 P(node) =0.08466258
## class counts: 240 312
## probabilities: 0.435 0.565
## left son=10 (123 obs) right son=11 (429 obs)
## Primary splits:
## Spa < 1372.5 to the right, improve=57.714490, (0 missing)
## VRDeck < 1063.5 to the right, improve=46.364550, (0 missing)
## expense < 5395 to the right, improve=17.936380, (0 missing)
## Side splits as LLR, improve= 7.616904, (0 missing)
## FoodCourt < 2513 to the left, improve= 7.542383, (0 missing)
## Surrogate splits:
## expense < 12647 to the right, agree=0.790, adj=0.057, (0 split)
## Age < 13.5 to the left, agree=0.779, adj=0.008, (0 split)
## RoomService < 3895.5 to the right, agree=0.779, adj=0.008, (0 split)
##
## Node number 8: 2577 observations
## predicted class=FALSE expected loss=0.193636 P(node) =0.3952454
## class counts: 2078 499
## probabilities: 0.806 0.194
##
## Node number 9: 666 observations, complexity param=0.01127936
## predicted class=FALSE expected loss=0.4864865 P(node) =0.1021472
## class counts: 342 324
## probabilities: 0.514 0.486
## left son=18 (157 obs) right son=19 (509 obs)
## Primary splits:
## RoomService < 310 to the right, improve=31.364140, (0 missing)
## Spa < 200 to the right, improve=19.676920, (0 missing)
## ShoppingMall < 1586.5 to the left, improve=13.953460, (0 missing)
## VRDeck < 120.5 to the right, improve=13.338010, (0 missing)
## HomePlanet splits as RLL, improve= 6.864426, (0 missing)
## Surrogate splits:
## FoodCourt < 1102 to the right, agree=0.766, adj=0.006, (0 split)
##
## Node number 10: 123 observations
## predicted class=FALSE expected loss=0.1382114 P(node) =0.01886503
## class counts: 106 17
## probabilities: 0.862 0.138
##
## Node number 11: 429 observations, complexity param=0.01205192
## predicted class=TRUE expected loss=0.3123543 P(node) =0.06579755
## class counts: 134 295
## probabilities: 0.312 0.688
## left son=22 (143 obs) right son=23 (286 obs)
## Primary splits:
## VRDeck < 611 to the right, improve=45.037300, (0 missing)
## Spa < 225 to the right, improve= 9.768013, (0 missing)
## FoodCourt < 3119.5 to the left, improve= 9.524139, (0 missing)
## Side splits as LLR, improve= 7.968308, (0 missing)
## RoomService < 1719.5 to the right, improve= 7.244213, (0 missing)
## Surrogate splits:
## expense < 6032 to the right, agree=0.702, adj=0.105, (0 split)
## Age < 53.5 to the right, agree=0.674, adj=0.021, (0 split)
## FoodCourt < 12128.5 to the right, agree=0.671, adj=0.014, (0 split)
##
## Node number 18: 157 observations
## predicted class=FALSE expected loss=0.2101911 P(node) =0.02407975
## class counts: 124 33
## probabilities: 0.790 0.210
##
## Node number 19: 509 observations, complexity param=0.01112485
## predicted class=TRUE expected loss=0.4282908 P(node) =0.07806748
## class counts: 218 291
## probabilities: 0.428 0.572
## left son=38 (66 obs) right son=39 (443 obs)
## Primary splits:
## Spa < 209 to the right, improve=17.99311, (0 missing)
## VRDeck < 33.5 to the right, improve=17.76420, (0 missing)
## ShoppingMall < 1540.5 to the left, improve=10.97812, (0 missing)
## Deck splits as LRLRLRR-, improve= 7.32794, (0 missing)
## expense < 4099.5 to the right, improve= 5.09474, (0 missing)
## Surrogate splits:
## expense < 4161 to the right, agree=0.898, adj=0.212, (0 split)
## Deck splits as LLLRRRR-, agree=0.896, adj=0.197, (0 split)
## HomePlanet splits as RLR, agree=0.884, adj=0.106, (0 split)
## ShoppingMall < 7126 to the right, agree=0.876, adj=0.045, (0 split)
## VRDeck < 1206.5 to the right, agree=0.876, adj=0.045, (0 split)
##
## Node number 22: 143 observations
## predicted class=FALSE expected loss=0.3636364 P(node) =0.02193252
## class counts: 91 52
## probabilities: 0.636 0.364
##
## Node number 23: 286 observations
## predicted class=TRUE expected loss=0.1503497 P(node) =0.04386503
## class counts: 43 243
## probabilities: 0.150 0.850
##
## Node number 38: 66 observations
## predicted class=FALSE expected loss=0.2272727 P(node) =0.0101227
## class counts: 51 15
## probabilities: 0.773 0.227
##
## Node number 39: 443 observations
## predicted class=TRUE expected loss=0.3769752 P(node) =0.06794479
## class counts: 167 276
## probabilities: 0.377 0.623
Ağacımız şöyle bir ifade veriyor: Diyor ki eğer Transported deki CryoSleep aldılarsa ( FALSE ve ya NA) no ya denk geliyor ama CryoSleep almadılarsa (TRUE) eğer RoomService daha fazla para ediyorsa onlar Transported olmamış ama daha az ediyorsa fakat Spa ya daha fazla verirse onlar Transported olmuş. Diyelim ki RoomService 366 daha üzeri ver Spa ya 205 daha üzeri ver ve VRDeck ye 248 daha fazla üzeri ver ama diyor ki FoodCourt en azından sonuç 0 dan fazla uygularsa , onlar da Transported olmuş
## y_pred
## y_true 0 1
## 0 798 281
## 1 196 898
## [1] 0.7804878
## Call:
## rpart(formula = Transported ~ ., data = train_set)
## n= 8693
##
## CP nsplit rel error xerror xstd
## 1 0.47045191 0 1.0000000 1.0000000 0.010803454
## 2 0.02143685 1 0.5295481 0.5295481 0.009511274
## 3 0.01993048 3 0.4866744 0.5015064 0.009342999
## 4 0.01008111 4 0.4667439 0.4769409 0.009184965
## 5 0.01000000 7 0.4347625 0.4602549 0.009071689
##
## Variable importance
## expense CryoSleep Spa FoodCourt VRDeck RoomService
## 25 19 14 14 12 11
## ShoppingMall HomePlanet Deck
## 3 1 1
##
## Node number 1: 8693 observations, complexity param=0.4704519
## predicted class=TRUE expected loss=0.4963764 P(node) =1
## class counts: 4315 4378
## probabilities: 0.496 0.504
## left son=2 (5040 obs) right son=3 (3653 obs)
## Primary splits:
## expense < 0.5 to the right, improve=1008.1870, (0 missing)
## CryoSleep splits as LRL, improve= 920.2004, (0 missing)
## RoomService < 0.5 to the right, improve= 526.1339, (0 missing)
## Spa < 0.5 to the right, improve= 514.4709, (0 missing)
## VRDeck < 0.5 to the right, improve= 479.5694, (0 missing)
## Surrogate splits:
## CryoSleep splits as LRL, agree=0.929, adj=0.831, (0 split)
## Spa < 0.5 to the right, agree=0.787, adj=0.492, (0 split)
## FoodCourt < 0.5 to the right, agree=0.772, adj=0.456, (0 split)
## VRDeck < 0.5 to the right, agree=0.766, adj=0.444, (0 split)
## RoomService < 0.5 to the right, agree=0.758, adj=0.424, (0 split)
##
## Node number 2: 5040 observations, complexity param=0.02143685
## predicted class=FALSE expected loss=0.2986111 P(node) =0.5797768
## class counts: 3535 1505
## probabilities: 0.701 0.299
## left son=4 (3827 obs) right son=5 (1213 obs)
## Primary splits:
## FoodCourt < 668.5 to the left, improve=136.56630, (0 missing)
## ShoppingMall < 627.5 to the left, improve= 95.37971, (0 missing)
## RoomService < 346.5 to the right, improve= 80.08985, (0 missing)
## Spa < 452.5 to the right, improve= 73.42969, (0 missing)
## VRDeck < 613.5 to the right, improve= 48.08949, (0 missing)
## Surrogate splits:
## HomePlanet splits as LRL, agree=0.845, adj=0.358, (0 split)
## Deck splits as RRRLLLLR, agree=0.834, adj=0.309, (0 split)
## expense < 3589.5 to the left, agree=0.823, adj=0.263, (0 split)
## Spa < 4710.5 to the left, agree=0.763, adj=0.016, (0 split)
## VRDeck < 4254.5 to the left, agree=0.762, adj=0.012, (0 split)
##
## Node number 3: 3653 observations
## predicted class=TRUE expected loss=0.2135231 P(node) =0.4202232
## class counts: 780 2873
## probabilities: 0.214 0.786
##
## Node number 4: 3827 observations, complexity param=0.01008111
## predicted class=FALSE expected loss=0.2330807 P(node) =0.4402393
## class counts: 2935 892
## probabilities: 0.767 0.233
## left son=8 (2961 obs) right son=9 (866 obs)
## Primary splits:
## ShoppingMall < 541 to the left, improve=143.35850, (0 missing)
## Spa < 445.5 to the right, improve= 42.68628, (0 missing)
## RoomService < 343 to the right, improve= 33.14410, (0 missing)
## VRDeck < 114 to the right, improve= 29.58027, (0 missing)
## expense < 2716 to the right, improve= 16.47007, (0 missing)
## Surrogate splits:
## expense < 18816 to the left, agree=0.774, adj=0.002, (0 split)
##
## Node number 5: 1213 observations, complexity param=0.02143685
## predicted class=TRUE expected loss=0.4946414 P(node) =0.1395376
## class counts: 600 613
## probabilities: 0.495 0.505
## left son=10 (230 obs) right son=11 (983 obs)
## Primary splits:
## Spa < 1372.5 to the right, improve=81.65183, (0 missing)
## VRDeck < 1114.5 to the right, improve=62.27701, (0 missing)
## FoodCourt < 2507.5 to the left, improve=27.03659, (0 missing)
## Side splits as LLR, improve=17.25608, (0 missing)
## expense < 6806 to the right, improve=16.91079, (0 missing)
## Surrogate splits:
## expense < 12957 to the right, agree=0.821, adj=0.057, (0 split)
## Age < 67.5 to the right, agree=0.812, adj=0.009, (0 split)
## RoomService < 5257.5 to the right, agree=0.811, adj=0.004, (0 split)
##
## Node number 8: 2961 observations
## predicted class=FALSE expected loss=0.1590679 P(node) =0.3406189
## class counts: 2490 471
## probabilities: 0.841 0.159
##
## Node number 9: 866 observations, complexity param=0.01008111
## predicted class=FALSE expected loss=0.4861432 P(node) =0.09962038
## class counts: 445 421
## probabilities: 0.514 0.486
## left son=18 (201 obs) right son=19 (665 obs)
## Primary splits:
## RoomService < 310 to the right, improve=36.00767, (0 missing)
## Spa < 200 to the right, improve=28.52457, (0 missing)
## VRDeck < 120.5 to the right, improve=15.13357, (0 missing)
## ShoppingMall < 1968.5 to the left, improve=14.39760, (0 missing)
## HomePlanet splits as RLL, improve=11.22426, (0 missing)
## Surrogate splits:
## FoodCourt < 591.5 to the right, agree=0.769, adj=0.005, (0 split)
## expense < 2343.5 to the right, agree=0.769, adj=0.005, (0 split)
##
## Node number 10: 230 observations
## predicted class=FALSE expected loss=0.126087 P(node) =0.02645807
## class counts: 201 29
## probabilities: 0.874 0.126
##
## Node number 11: 983 observations, complexity param=0.01993048
## predicted class=TRUE expected loss=0.4059003 P(node) =0.1130795
## class counts: 399 584
## probabilities: 0.406 0.594
## left son=22 (136 obs) right son=23 (847 obs)
## Primary splits:
## VRDeck < 1685.5 to the right, improve=53.136330, (0 missing)
## FoodCourt < 3125 to the left, improve=36.358710, (0 missing)
## RoomService < 364 to the right, improve=14.375770, (0 missing)
## Side splits as LLR, improve=14.088100, (0 missing)
## Deck splits as LRRRLLLL, improve= 9.618287, (0 missing)
## Surrogate splits:
## expense < 12248 to the right, agree=0.87, adj=0.059, (0 split)
##
## Node number 18: 201 observations
## predicted class=FALSE expected loss=0.2238806 P(node) =0.02312205
## class counts: 156 45
## probabilities: 0.776 0.224
##
## Node number 19: 665 observations, complexity param=0.01008111
## predicted class=TRUE expected loss=0.4345865 P(node) =0.07649833
## class counts: 289 376
## probabilities: 0.435 0.565
## left son=38 (87 obs) right son=39 (578 obs)
## Primary splits:
## Spa < 201 to the right, improve=25.731350, (0 missing)
## VRDeck < 128.5 to the right, improve=19.130020, (0 missing)
## ShoppingMall < 1815.5 to the left, improve=10.453460, (0 missing)
## Deck splits as LRLRLRR-, improve= 5.398272, (0 missing)
## expense < 2958.5 to the right, improve= 4.372934, (0 missing)
## Surrogate splits:
## expense < 3966.5 to the right, agree=0.898, adj=0.218, (0 split)
## Deck splits as LLLRRRR-, agree=0.893, adj=0.184, (0 split)
## HomePlanet splits as RLR, agree=0.889, adj=0.149, (0 split)
## VRDeck < 1206.5 to the right, agree=0.880, adj=0.080, (0 split)
## ShoppingMall < 7126 to the right, agree=0.874, adj=0.034, (0 split)
##
## Node number 22: 136 observations
## predicted class=FALSE expected loss=0.1838235 P(node) =0.01564477
## class counts: 111 25
## probabilities: 0.816 0.184
##
## Node number 23: 847 observations
## predicted class=TRUE expected loss=0.3400236 P(node) =0.09743472
## class counts: 288 559
## probabilities: 0.340 0.660
##
## Node number 38: 87 observations
## predicted class=FALSE expected loss=0.2068966 P(node) =0.01000805
## class counts: 69 18
## probabilities: 0.793 0.207
##
## Node number 39: 578 observations
## predicted class=TRUE expected loss=0.3806228 P(node) =0.06649028
## class counts: 220 358
## probabilities: 0.381 0.619
Ağacımız şöyle bir ifade veriyor:
Eğer CryoSleep almışlarsa onlar Transported olmuş ama RoomService fazla almışsa ölmüştür. Ve diyor ki RoomService daha az para harcamışlar , yine Spa ya daha fazla para harcamış ama onlar da daha fazla harcamışsa bile VRDeck eğer 355 daha fazla para harcamışsa ölmüştür. Bu üçü ( RoomService, Spa ve VRDeck ) eğer 347 den fazla harcamamışsa , 205 den fazla harcamamışsa, 355 den fazla harcamamışsa uyku anmışlarsa bile ölmüştür.
Naive Bayes
Naive Bayes, sınıflandırma problemlerinde kullanılan bir makine öğrenme algoritmasıdır. Bu algoritma, Bayes teoremini kullanarak sınıflandırma yapar. Naive Bayes, özellikle metin sınıflandırma gibi doğal dil işleme problemlerinde ve spam filtrelemede yaygın olarak kullanılır. Bu algoritma basit ve hızlıdır ancak bağımsızlık varsayımının gerçek durumu yansıtmadığı durumlarda yanlı sonuçlara yol açabilir.
library(e1071)
fit_nb <- naiveBayes(Transported ~ . , data = training_set)
preds <-predict(fit_nb, newdata = testing_set[-13], type = "raw") %>% data.frame()## y_pred
## y_true 0 1
## 0 487 592
## 1 79 1015
## [1] 0.6912103