FİNAL PROJESİ

SPACESHİP TİTANİC

Kozmik bir gizemi çözmek için veri bilimi becerilerinize ihtiyaç duyulan 2912 yılına hoş geldiniz. Dört ışık yılı öteden bir sinyal aldık ve işler pek iyi görünmüyor.

Uzay Gemisi Titanik, bir ay önce fırlatılan yıldızlararası bir yolcu gemisiydi. Gemide neredeyse 13.000 yolcu bulunan gemi, güneş sistemimizden göçmenleri yakın yıldızların yörüngesinde bulunan üç yeni yaşanabilir dış gezegene taşımak üzere ilk yolculuğuna çıktı.

Dikkatsiz Uzay Gemisi Titanic, ilk varış noktası olan kavurucu 55 Cancri E’ye giderken Alpha Centauri’yi dönerken, bir toz bulutunun içine gizlenmiş bir uzay-zaman anormalliğiyle çarpıştı. Ne yazık ki 1000 yıl öncesindeki adaşı ile benzer bir kaderle karşılaştı. Gemi sağlam kalmasına rağmen yolcuların neredeyse yarısı alternatif bir boyuta taşındı!

DATALARI YÜKLEME

library(readr)
train <- read_csv("train.csv")
test <- read_csv("test.csv")

Dataları indirdikten sonra gereken paket yüklemelerini yapalım.

“tidyverse” ve “explore” paketleri bize verileri görselleştirmede ve analiz yapmada yardımcı olacaktır.

library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.1.4     ✔ purrr     1.0.2
## ✔ forcats   1.0.0     ✔ stringr   1.5.1
## ✔ ggplot2   3.4.4     ✔ tibble    3.2.1
## ✔ lubridate 1.9.3     ✔ tidyr     1.3.0
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::filter() masks stats::filter()
## ✖ dplyr::lag()    masks stats::lag()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors
library(explore)

Gereken paket yüklemelerini yaptıktan sonra bize verilen data bilgilerini kontrol edelim.

train %>% describe_all()
## # A tibble: 14 × 8
##    variable     type     na na_pct unique   min   mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl>  <dbl> <dbl>
##  1 PassengerId  chr       0    0     8693    NA  NA       NA
##  2 HomePlanet   chr     201    2.3      4    NA  NA       NA
##  3 CryoSleep    lgl     217    2.5      3     0   0.36     1
##  4 Cabin        chr     199    2.3   6561    NA  NA       NA
##  5 Destination  chr     182    2.1      4    NA  NA       NA
##  6 Age          dbl     179    2.1     81     0  28.8     79
##  7 VIP          lgl     203    2.3      3     0   0.02     1
##  8 RoomService  dbl     181    2.1   1274     0 225.   14327
##  9 FoodCourt    dbl     183    2.1   1508     0 458.   29813
## 10 ShoppingMall dbl     208    2.4   1116     0 174.   23492
## 11 Spa          dbl     183    2.1   1328     0 311.   22408
## 12 VRDeck       dbl     188    2.2   1307     0 305.   24133
## 13 Name         chr     200    2.3   8474    NA  NA       NA
## 14 Transported  lgl       0    0        2     0   0.5      1
test %>% describe_all()
## # A tibble: 13 × 8
##    variable     type     na na_pct unique   min   mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl>  <dbl> <dbl>
##  1 PassengerId  chr       0    0     4277    NA  NA       NA
##  2 HomePlanet   chr      87    2        4    NA  NA       NA
##  3 CryoSleep    lgl      93    2.2      3     0   0.37     1
##  4 Cabin        chr     100    2.3   3266    NA  NA       NA
##  5 Destination  chr      92    2.2      4    NA  NA       NA
##  6 Age          dbl      91    2.1     80     0  28.7     79
##  7 VIP          lgl      93    2.2      3     0   0.02     1
##  8 RoomService  dbl      82    1.9    843     0 219.   11567
##  9 FoodCourt    dbl     106    2.5    903     0 439.   25273
## 10 ShoppingMall dbl      98    2.3    716     0 177.    8292
## 11 Spa          dbl     101    2.4    834     0 303.   19844
## 12 VRDeck       dbl      80    1.9    797     0 311.   22272
## 13 Name         chr      94    2.2   4177    NA  NA       NA

Verilen data bilgilerinin özelliklerine baktığımızda bilinmeyen eksik değerlerin fazlalığını görüyoruz. Biz bu dataları daha temiz daha okunabilirlik sağlamak için bir kaç işlemde bulunacağız.

İlk olarak PASSENGERID sütunu ve CABİN sütununa baktığımızda bize her bir bilgide birden fazla değer gösteriyor. Biz bu bilgileri farklı sütunlara ayıralım.

SÜTUNLARI AYIRMA

PASSENGERID

PassengerId sütununu “ailenum ve ailesıra” olmak üzere ikiye ayıralım.

train[c('ailenum','ailesıra')]<- str_split_fixed(train$PassengerId, "_",2)
test[c('ailenum','ailesıra')]<- str_split_fixed(test$PassengerId, "_",2)

CABİN

Daha sonra Cabin sütununu üçe ayıralım. Bu 3 sütunumuzun isimleri ” deck, num, side” olarak belirleyelim.

train[c('deck',"num", "side")]<- str_split_fixed(train$Cabin, "/",3)
test[c('deck',"num", "side")]<- str_split_fixed(test$Cabin, "/",3)

Bu kodları kullandıktan sonra Cabin sütunu bir işimize yaramayacak bunun için veri setlerimizde silelim.

train <- train %>% select(-Cabin)
test <- test %>% select(-Cabin)

BİLİNMEYEN DEĞERLERİ TEMİZLEME VE DOLDURMA

Sütunları ayırma işlemini bitirdik. Şimdi bilinmeyen eksik değerleri temizleyelim ve yerlerini NA değeriyle dolduralım.

train[train == ''] <- NA
test[test == ''] <- NA

Bu işlem genel bir temizleme işlemi ve doldurma işlemi yaptı. Biz daha iyi bir sonuca ulaşmak istiyorsak bütün verilen bilgileri detaylı bir şekilde temizlemeliyiz.

İlk öncelikle verilere baktığımızda sayısal ve sözel ifadeler görüyoruz. Bu sayısal ve sözel ifadeler için farklı işlemlerde bulunacağız. İlk önce sözel ifadeler olan karakter(chr) ve logical(lgl) tip veriler için işleme başlayalım.

HOMEPLANET

train $HomePlanet<-addNA(train$HomePlanet)
test $HomePlanet<-addNA(test$HomePlanet)

CRYOSLEEP

train $CryoSleep<-addNA(train$CryoSleep)
test $CryoSleep<-addNA(test$CryoSleep)

DESTİNATİON

train $Destination<-addNA(train$Destination)
test $Destination<-addNA(test$Destination)

VIP

train $VIP<-addNA(train$VIP)
test $VIP<-addNA(test$VIP)

DECK

train $deck<-addNA(train$deck)
test $deck<-addNA(test$deck)

NUM

train $num<-addNA(train$num)
test $num<-addNA(test$num)

SİDE

train $side<-addNA(train$side)
test $side<-addNA(test$side)

Daha sonra bu eklediğimiz NA değeri sütun boşluklarında gözükmeyecektir. Bu NA değerini “NA” olarak değiştirmemiz gerekiyor.

HOMEPLANET

levels(train$HomePlanet)[is.na(levels(train$HomePlanet))] <- "NA"
levels(test$HomePlanet)[is.na(levels(test$HomePlanet))] <- "NA"

CRYOSLEEP

levels(train$CryoSleep)[is.na(levels(train$CryoSleep))] <- "NA"
levels(test$CryoSleep)[is.na(levels(test$CryoSleep))] <- "NA"

DESTİNATİON

levels(train$Destination)[is.na(levels(train$Destination))] <- "NA"
levels(test$Destination)[is.na(levels(test$Destination))] <- "NA"

VIP

levels(train$VIP)[is.na(levels(train$VIP))] <- "NA"
levels(test$VIP)[is.na(levels(test$VIP))] <- "NA"

DECK

levels(train$deck)[is.na(levels(train$deck))] <- "NA"
levels(test$deck)[is.na(levels(test$deck))] <- "NA"

NUM

levels(train$num)[is.na(levels(train$num))] <- "NA"
levels(test$num)[is.na(levels(test$num))] <- "NA"

SIDE

levels(train$side)[is.na(levels(train$side))] <- "NA"
levels(test$side)[is.na(levels(test$side))] <- "NA"
levels(train$Destination)
## [1] "55 Cancri e"   "PSO J318.5-22" "TRAPPIST-1e"   "NA"

Şimdi double(dbl) yani sayısal değerler için işlem yapalım. Burada yapacağımız işlem eksik değerler yerine ortalamalarını koymaktır. Burada kullanacağımız kod da her iki data için DESTİNATİON verilerine göre ortalamalarını almasını sağlamaktır.

AGE

train <- train %>%
  group_by(Destination) %>%
  mutate(Age = replace(Age, is.na(Age), mean(Age, na.rm = TRUE)))
test <- test %>%
  group_by(Destination) %>%
  mutate(Age = replace(Age, is.na(Age), mean(Age, na.rm = TRUE)))

ROOMSERVİCE

train <- train %>%
  group_by(Destination) %>%
  mutate(RoomService = replace(RoomService, is.na(RoomService), mean(RoomService, na.rm = TRUE)))
test <- test %>%
  group_by(Destination) %>%
  mutate(RoomService = replace(RoomService, is.na(RoomService), mean(RoomService, na.rm = TRUE)))

FOODCOURT

train <- train %>%
  group_by(Destination) %>%
  mutate(FoodCourt = replace(FoodCourt, is.na(FoodCourt), mean(FoodCourt, na.rm = TRUE)))
test <- test %>%
  group_by(Destination) %>%
  mutate(FoodCourt = replace(FoodCourt, is.na(FoodCourt), mean(FoodCourt, na.rm = TRUE)))

SHOPPİNGMALL

train <- train %>%
  group_by(Destination) %>%
  mutate(ShoppingMall = replace(ShoppingMall, is.na(ShoppingMall), mean(ShoppingMall, na.rm = TRUE)))
test <- test %>%
  group_by(Destination) %>%
  mutate(ShoppingMall = replace(ShoppingMall, is.na(ShoppingMall), mean(ShoppingMall, na.rm = TRUE)))

SPA

train <- train %>%
  group_by(Destination) %>%
  mutate(Spa = replace(Spa, is.na(Spa), mean(Spa, na.rm = TRUE)))
test <- test %>%
  group_by(Destination) %>%
  mutate(Spa = replace(Spa, is.na(Spa), mean(Spa, na.rm = TRUE)))

VRDECK

train <- train %>%
  group_by(Destination) %>%
  mutate(VRDeck = replace(VRDeck, is.na(VRDeck), mean(VRDeck, na.rm = TRUE)))
test <- test %>%
  group_by(Destination) %>%
  mutate(VRDeck = replace(VRDeck, is.na(VRDeck), mean(VRDeck, na.rm = TRUE)))

Böylelikle bilinmeyene eksik değerleri temzilemiş bulunmaktayız. Şimdi dataları kontrol edelim.

train%>%describe_all()
## # A tibble: 18 × 8
##    variable     type     na na_pct unique   min  mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl> <dbl> <dbl>
##  1 PassengerId  chr       0    0     8693    NA  NA      NA
##  2 HomePlanet   fct       0    0        4    NA  NA      NA
##  3 CryoSleep    fct       0    0        3    NA  NA      NA
##  4 Destination  fct       0    0        4    NA  NA      NA
##  5 Age          dbl       0    0       84     0  28.8    79
##  6 VIP          fct       0    0        3    NA  NA      NA
##  7 RoomService  dbl       0    0     1277     0 225.  14327
##  8 FoodCourt    dbl       0    0     1511     0 458.  29813
##  9 ShoppingMall dbl       0    0     1119     0 174.  23492
## 10 Spa          dbl       0    0     1331     0 311.  22408
## 11 VRDeck       dbl       0    0     1310     0 305.  24133
## 12 Name         chr     200    2.3   8474    NA  NA      NA
## 13 Transported  lgl       0    0        2     0   0.5     1
## 14 ailenum      chr       0    0     6217    NA  NA      NA
## 15 ailesıra     chr       0    0        8    NA  NA      NA
## 16 deck         fct       0    0        9    NA  NA      NA
## 17 num          fct       0    0     1818    NA  NA      NA
## 18 side         fct       0    0        3    NA  NA      NA
test%>%describe_all()
## # A tibble: 17 × 8
##    variable     type     na na_pct unique   min  mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl> <dbl> <dbl>
##  1 PassengerId  chr       0    0     4277    NA  NA      NA
##  2 HomePlanet   fct       0    0        4    NA  NA      NA
##  3 CryoSleep    fct       0    0        3    NA  NA      NA
##  4 Destination  fct       0    0        4    NA  NA      NA
##  5 Age          dbl       0    0       83     0  28.7    79
##  6 VIP          fct       0    0        3    NA  NA      NA
##  7 RoomService  dbl       0    0      846     0 219.  11567
##  8 FoodCourt    dbl       0    0      905     0 439.  25273
##  9 ShoppingMall dbl       0    0      719     0 177.   8292
## 10 Spa          dbl       0    0      837     0 303.  19844
## 11 VRDeck       dbl       0    0      800     0 311.  22272
## 12 Name         chr      94    2.2   4177    NA  NA      NA
## 13 ailenum      chr       0    0     3063    NA  NA      NA
## 14 ailesıra     chr       0    0        8    NA  NA      NA
## 15 deck         fct       0    0        9    NA  NA      NA
## 16 num          fct       0    0     1506    NA  NA      NA
## 17 side         fct       0    0        3    NA  NA      NA

Kontrol sağladığımızda verilerin temizlendiğini görüyoruz. “name” veri bilgisi herhangi bir özellik sağlamıyor ve işimize yaramıyacaktır. Bunun için datalardan kaldıralım.

train <- train %>% select(-Name)
test <- test %>% select(-Name)

Veri setlerini incelediğimizde “ailenum” veri bilgisinde bulunun tekrarlanmaları ve okunabilirliği düzeltmeliyiz. Bunun için aşağıdaki kodu kullanalım. Her iki data için yapalım.

TRAİN

trainailenum <- ifelse(duplicated(train$ailenum) | duplicated(train$ailenum, fromLast = TRUE), 1, 0)

TEST

testailenum <- ifelse(duplicated(test$ailenum) | duplicated(test$ailenum, fromLast = TRUE), 1, 0)
train <- train %>% select(-ailenum)
test <- test %>% select(-ailenum)
train <- train %>% select(-num)
test <- test %>% select(-num)

Bütün eksik değerleri ve okunuabilirliği kolaylaştırcak işlemleri yaptık. Şimdi dataları kontrol edelim.

TRAİN describe_all

train%>%describe_all()
## # A tibble: 15 × 8
##    variable     type     na na_pct unique   min  mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl> <dbl> <dbl>
##  1 PassengerId  chr       0      0   8693    NA  NA      NA
##  2 HomePlanet   fct       0      0      4    NA  NA      NA
##  3 CryoSleep    fct       0      0      3    NA  NA      NA
##  4 Destination  fct       0      0      4    NA  NA      NA
##  5 Age          dbl       0      0     84     0  28.8    79
##  6 VIP          fct       0      0      3    NA  NA      NA
##  7 RoomService  dbl       0      0   1277     0 225.  14327
##  8 FoodCourt    dbl       0      0   1511     0 458.  29813
##  9 ShoppingMall dbl       0      0   1119     0 174.  23492
## 10 Spa          dbl       0      0   1331     0 311.  22408
## 11 VRDeck       dbl       0      0   1310     0 305.  24133
## 12 Transported  lgl       0      0      2     0   0.5     1
## 13 ailesıra     chr       0      0      8    NA  NA      NA
## 14 deck         fct       0      0      9    NA  NA      NA
## 15 side         fct       0      0      3    NA  NA      NA

TEST describe_all

test%>%describe_all()
## # A tibble: 14 × 8
##    variable     type     na na_pct unique   min  mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl> <dbl> <dbl>
##  1 PassengerId  chr       0      0   4277    NA  NA      NA
##  2 HomePlanet   fct       0      0      4    NA  NA      NA
##  3 CryoSleep    fct       0      0      3    NA  NA      NA
##  4 Destination  fct       0      0      4    NA  NA      NA
##  5 Age          dbl       0      0     83     0  28.7    79
##  6 VIP          fct       0      0      3    NA  NA      NA
##  7 RoomService  dbl       0      0    846     0 219.  11567
##  8 FoodCourt    dbl       0      0    905     0 439.  25273
##  9 ShoppingMall dbl       0      0    719     0 177.   8292
## 10 Spa          dbl       0      0    837     0 303.  19844
## 11 VRDeck       dbl       0      0    800     0 311.  22272
## 12 ailesıra     chr       0      0      8    NA  NA      NA
## 13 deck         fct       0      0      9    NA  NA      NA
## 14 side         fct       0      0      3    NA  NA      NA

Görüldüğü üzere tertemiz şekilde datalarımız bulunuyor. Böylelikle daha iyi tahmin sonuçlarına ulaşabileceğiz.

Şimdi dataların veri profillerinin raporunu oluşturalım ve bir kaç yorumda bulunalım.

VERİ PROFİL RAPORLARI

TRAİN

library(DataExplorer)
create_report(train)
## 
  |                                           
  |                                     |   0%
  |                                           
  |.                                    |   2%                                 
  |                                           
  |..                                   |   5% [global_options]                
  |                                           
  |...                                  |   7%                                 
  |                                           
  |....                                 |  10% [introduce]                     
  |                                           
  |....                                 |  12%                                 
  |                                           
  |.....                                |  14% [plot_intro]                    
  |                                           
  |......                               |  17%                                 
  |                                           
  |.......                              |  19% [data_structure]                
  |                                           
  |........                             |  21%                                 
  |                                           
  |.........                            |  24% [missing_profile]               
  |                                           
  |..........                           |  26%                                 
  |                                           
  |...........                          |  29% [univariate_distribution_header]
  |                                           
  |...........                          |  31%                                 
  |                                           
  |............                         |  33% [plot_histogram]                
  |                                           
  |.............                        |  36%                                 
  |                                           
  |..............                       |  38% [plot_density]                  
  |                                           
  |...............                      |  40%                                 
  |                                           
  |................                     |  43% [plot_frequency_bar]            
  |                                           
  |.................                    |  45%                                 
  |                                           
  |..................                   |  48% [plot_response_bar]             
  |                                           
  |..................                   |  50%                                 
  |                                           
  |...................                  |  52% [plot_with_bar]                 
  |                                           
  |....................                 |  55%                                 
  |                                           
  |.....................                |  57% [plot_normal_qq]                
  |                                           
  |......................               |  60%                                 
  |                                           
  |.......................              |  62% [plot_response_qq]              
  |                                           
  |........................             |  64%                                 
  |                                           
  |.........................            |  67% [plot_by_qq]                    
  |                                           
  |..........................           |  69%                                 
  |                                           
  |..........................           |  71% [correlation_analysis]          
  |                                           
  |...........................          |  74%                                 
  |                                           
  |............................         |  76% [principal_component_analysis]  
  |                                           
  |.............................        |  79%                                 
  |                                           
  |..............................       |  81% [bivariate_distribution_header] 
  |                                           
  |...............................      |  83%                                 
  |                                           
  |................................     |  86% [plot_response_boxplot]         
  |                                           
  |.................................    |  88%                                 
  |                                           
  |.................................    |  90% [plot_by_boxplot]               
  |                                           
  |..................................   |  93%                                 
  |                                           
  |...................................  |  95% [plot_response_scatterplot]     
  |                                           
  |.................................... |  98%                                 
  |                                           
  |.....................................| 100% [plot_by_scatterplot]           
                                                                                                                           
## "C:/Program Files/RStudio/resources/app/bin/quarto/bin/tools/pandoc" +RTS -K512m -RTS "C:\Users\Win10\Desktop\finalpro\report.knit.md" --to html4 --from markdown+autolink_bare_uris+tex_math_single_backslash --output pandoc37a824497682.html --lua-filter "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmarkdown\lua\pagebreak.lua" --lua-filter "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmarkdown\lua\latex-div.lua" --embed-resources --standalone --variable bs3=TRUE --section-divs --table-of-contents --toc-depth 6 --template "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmd\h\default.html" --no-highlight --variable highlightjs=1 --variable theme=yeti --mathjax --variable "mathjax-url=https://mathjax.rstudio.com/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML" --include-in-header "C:\Users\Win10\AppData\Local\Temp\Rtmp2FgHWT\rmarkdown-str37a86b6266bc.html"

TEST

library(DataExplorer)
create_report(test)
## 
  |                                           
  |                                     |   0%
  |                                           
  |.                                    |   2%                                 
  |                                           
  |..                                   |   5% [global_options]                
  |                                           
  |...                                  |   7%                                 
  |                                           
  |....                                 |  10% [introduce]                     
  |                                           
  |....                                 |  12%                                 
  |                                           
  |.....                                |  14% [plot_intro]                    
  |                                           
  |......                               |  17%                                 
  |                                           
  |.......                              |  19% [data_structure]                
  |                                           
  |........                             |  21%                                 
  |                                           
  |.........                            |  24% [missing_profile]               
  |                                           
  |..........                           |  26%                                 
  |                                           
  |...........                          |  29% [univariate_distribution_header]
  |                                           
  |...........                          |  31%                                 
  |                                           
  |............                         |  33% [plot_histogram]                
  |                                           
  |.............                        |  36%                                 
  |                                           
  |..............                       |  38% [plot_density]                  
  |                                           
  |...............                      |  40%                                 
  |                                           
  |................                     |  43% [plot_frequency_bar]            
  |                                           
  |.................                    |  45%                                 
  |                                           
  |..................                   |  48% [plot_response_bar]             
  |                                           
  |..................                   |  50%                                 
  |                                           
  |...................                  |  52% [plot_with_bar]                 
  |                                           
  |....................                 |  55%                                 
  |                                           
  |.....................                |  57% [plot_normal_qq]                
  |                                           
  |......................               |  60%                                 
  |                                           
  |.......................              |  62% [plot_response_qq]              
  |                                           
  |........................             |  64%                                 
  |                                           
  |.........................            |  67% [plot_by_qq]                    
  |                                           
  |..........................           |  69%                                 
  |                                           
  |..........................           |  71% [correlation_analysis]          
  |                                           
  |...........................          |  74%                                 
  |                                           
  |............................         |  76% [principal_component_analysis]  
  |                                           
  |.............................        |  79%                                 
  |                                           
  |..............................       |  81% [bivariate_distribution_header] 
  |                                           
  |...............................      |  83%                                 
  |                                           
  |................................     |  86% [plot_response_boxplot]         
  |                                           
  |.................................    |  88%                                 
  |                                           
  |.................................    |  90% [plot_by_boxplot]               
  |                                           
  |..................................   |  93%                                 
  |                                           
  |...................................  |  95% [plot_response_scatterplot]     
  |                                           
  |.................................... |  98%                                 
  |                                           
  |.....................................| 100% [plot_by_scatterplot]           
                                                                                                                           
## "C:/Program Files/RStudio/resources/app/bin/quarto/bin/tools/pandoc" +RTS -K512m -RTS "C:\Users\Win10\Desktop\finalpro\report.knit.md" --to html4 --from markdown+autolink_bare_uris+tex_math_single_backslash --output pandoc37a851f768b6.html --lua-filter "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmarkdown\lua\pagebreak.lua" --lua-filter "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmarkdown\lua\latex-div.lua" --embed-resources --standalone --variable bs3=TRUE --section-divs --table-of-contents --toc-depth 6 --template "C:\Users\Win10\AppData\Local\R\win-library\4.3\rmarkdown\rmd\h\default.html" --no-highlight --variable highlightjs=1 --variable theme=yeti --mathjax --variable "mathjax-url=https://mathjax.rstudio.com/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML" --include-in-header "C:\Users\Win10\AppData\Local\Temp\Rtmp2FgHWT\rmarkdown-str37a866ea44c9.html"

Şimdi bir histogram grafiği oluşturalım.

HİSTOGRAM GRAFİĞİ OLUŞTURMA

ÖRNEK1

test datasından “ShoppingMall” ve “Destination” verilerini göz önüne alalım. Bunun için aşağıdaki kodları kullanabiliriz.

ggplot(test, aes(x = ShoppingMall)) + 
  geom_histogram(fill = "white", color = "black") +
  facet_grid(Destination ~ .)

YORUM

Yaptığımız histogram grafiğini incelediğimide neredeyse hiçbir yolcunun hiç harcama yapmadığı görülüyor. 0-200 kişi TRAPPIST-1e gezegenine gidenlerden harcama yapma olasılığı oluşmuş.

ÖRNEK2

İkinci bir örnek olarak da train datasından “FoodCourt” ve “HomePlanet” veri bilgilerini alalım.

ggplot(train, aes(x = FoodCourt)) + 
  geom_histogram(stat = "count", fill = "white", color = "black") +
  facet_grid(HomePlanet ~ .)

Tahmin modellerini oluşturmaya başlayabiliriz.

TAHMİN MODELLERİ OLUŞTURMA

Öncelikle modelleri oluşturmadan önce yeni veri setleri ve alt veri setleri oluşturmalıyız.

train_set <- train [2:15]
test_set <- test[2:14]

Daha sonra daha iyi görselleştirme, analiz yapma ve okunabilirliği yapmak için “caTools” paketini yüklemeliyiz.

library(caTools)

Pakaeti yükledikten sonra kuracağımız modellerde tekrarlanabilirliği sürekli farklı bir şekilde tahminde bulunması için “set.seed” fonksiyonunu kullanalım.

set.seed(1576)

Şimdi alt veri setleri oluşturmamız gerekiyor. İlk önce “sample.split” fonksiyonu ile oluşturduğumuz yeni veri setinden ne kadar oranda veriler çekmek istediğimizi yazmalıyız.

split = sample.split(train_set$Transported, SplitRatio = 0.80)

Daha sonra bu oluşturduğumuz oran ile alt veri setleri oluşturalım.

training_set = subset(train_set, split == TRUE)
testing_set = subset(train_set, split == FALSE)

Böylelikle bütün gereken bilgileri oluşturduktan sonra tahmin modellerini oluşturabiliriz.

SVM MODELİ

SVM MODELİ NEDİR?

Support Vector Machine (SVM), sınıflandırma ve regresyon problemleri için kullanılan bir makine öğrenimi algoritmasıdır. SVM, özellikle sınıflandırma problemlerinde etkili olan bir algoritmadır ve doğrusal olarak ayrılabilir veya doğrusal olarak ayrılamayan veri setlerinde kullanılabilir.

SVM’nin temel amacı, veri setindeki sınıfları en iyi şekilde ayıran bir karar sınırı (hyperplane) oluşturmaktır. Bu sınır, iki sınıf arasındaki maksimum marjı (uzaklığı) maksimize etmek için belirlenir. SVM, özellikle düşük boyutlu veri setlerinde ve özellikle yüksek boyutlu uzaylarda iyi performans gösterir.

library(e1071)
fit_svm <- svm(Transported ~ ., data = training_set,
               type = 'C-classification',
               kernel = 'linear')

preds <- predict(fit_svm, newdata = testing_set, type = "raw") %>%
  data.frame()
y_pred = ifelse(preds$. == TRUE, 1, 0)
y_true <- ifelse(testing_set[11] == TRUE, 1,0)
cm = table(y_true, y_pred)
cm
##       y_pred
## y_true   0   1
##      0 688 175
##      1 164 712
(688+712) / (688+712+175+164)
## [1] 0.8050604
svm_son = svm(Transported ~ ., data = train_set,
              type = 'C-classification',
              kernel = 'linear')
preds <- predict(svm_son, newdata = test_set, type = "raw") %>%
  data.frame()
y_pred = preds$.
Transported <- as.character(y_pred)
PassengerId <- test$PassengerId
Transported <- as.vector(Transported)
svmsonucu <- cbind(PassengerId, Transported)
svmsonucu <- as.data.frame(svmsonucu)
svmsonucu$Transported <- str_to_title(svmsonucu$Transported)
write.csv(svmsonucu, "svmsonucu.csv", row.names = FALSE, quote = FALSE)

SVM MODEL SONUCU

SVM RADİAL(KERNEL) MODELİ

SVM RADİAL(KERNEL) MODELİ NEDİR?

Support Vector Machine (SVM), sınıflandırma ve regresyon problemleri için kullanılan bir makine öğrenimi algoritmasıdır. SVM’nin çekirdek (kernel) fonksiyonları, özellikle doğrusal olarak ayrılamayan veri setleri üzerinde etkili olmasını sağlar. Radyal bazlı fonksiyon çekirdeği (Radial Basis Function Kernel), bu çekirdek fonksiyonlarından biridir.

svm_ker_son = svm(Transported ~ ., data = train_set,
                  type = 'C-classification',
                  kernel = 'radial')
preds <- predict(svm_ker_son, newdata = test_set, type = "raw") %>%
  data.frame()
y_pred = preds$.
Transported <- as.character(y_pred)
PassengerId <- test$PassengerId
Transported <- as.vector(Transported)
radialsonucu <- cbind(PassengerId, Transported)
radialsonucu <- as.data.frame(radialsonucu)
radialsonucu$Transported <- str_to_title(radialsonucu$Transported)
write.csv(radialsonucu, "radialsonucu.csv", row.names = FALSE , quote = FALSE)

SVM RADİAL SONUCU

DECİSİON TREES MODELİ

DECİSİON TREES NEDİR?

Decision Trees (Karar Ağaçları), sınıflandırma ve regresyon problemleri için kullanılan bir makine öğrenimi algoritmasıdır. Veri kümesindeki özelliklere dayanarak bir dizi karar yapısı oluşturur ve bu yapıyı kullanarak veri noktalarını sınıflandırır veya tahminler yapar. Decision Trees, ağaç yapısındaki bir dizi karar düğümünden oluşur ve her düğüm, bir özellik testi yaparak veriyi iki veya daha fazla alt kümeye böler.

library(rpart)
library(rpart.plot)
library(randomForest)
library(caret)
training_set$Transported <- as.factor(training_set$Transported )
testing_set$Transported <- as.factor(testing_set$Transported)
train_set$Transported <- as.factor(train_set$Transported )
fit_tree <- rpart :: rpart(Transported ~ ., data = training_set)
summary(fit_tree)
## Call:
## rpart::rpart(formula = Transported ~ ., data = training_set)
##   n= 6954 
## 
##           CP nsplit rel error    xerror       xstd
## 1 0.42873696      0 1.0000000 1.0162225 0.01207812
## 2 0.02887215      1 0.5712630 0.5712630 0.01088848
## 3 0.01289108      4 0.4846466 0.4846466 0.01032566
## 4 0.01129780      6 0.4588644 0.4776941 0.01027460
## 5 0.01013905      7 0.4475666 0.4669757 0.01019404
## 6 0.01000000      8 0.4374276 0.4658169 0.01018520
## 
## Variable importance
##    CryoSleep          Spa       VRDeck  RoomService    FoodCourt ShoppingMall 
##           39           16           15           10            9            4 
##   HomePlanet         deck 
##            3            2 
## 
## Node number 1: 6954 observations,    complexity param=0.428737
##   predicted class=TRUE   expected loss=0.4964049  P(node) =1
##     class counts:  3452  3502
##    probabilities: 0.496 0.504 
##   left son=2 (4524 obs) right son=3 (2430 obs)
##   Primary splits:
##       CryoSleep    splits as  LRL,         improve=723.5736, (0 missing)
##       RoomService  < 0.5     to the right, improve=425.3976, (0 missing)
##       Spa          < 0.5     to the right, improve=395.9205, (0 missing)
##       VRDeck       < 0.5     to the right, improve=362.9468, (0 missing)
##       ShoppingMall < 0.5     to the right, improve=209.5968, (0 missing)
##   Surrogate splits:
##       Spa          < 0.5     to the right, agree=0.720, adj=0.199, (0 split)
##       VRDeck       < 0.5     to the right, agree=0.705, adj=0.157, (0 split)
##       FoodCourt    < 0.5     to the right, agree=0.703, adj=0.149, (0 split)
##       RoomService  < 0.5     to the right, agree=0.696, adj=0.131, (0 split)
##       ShoppingMall < 0.5     to the right, agree=0.683, adj=0.093, (0 split)
## 
## Node number 2: 4524 observations,    complexity param=0.02887215
##   predicted class=FALSE  expected loss=0.3364279  P(node) =0.6505608
##     class counts:  3002  1522
##    probabilities: 0.664 0.336 
##   left son=4 (1163 obs) right son=5 (3361 obs)
##   Primary splits:
##       RoomService < 346.5   to the right, improve=94.69975, (0 missing)
##       Spa         < 291.5   to the right, improve=86.28603, (0 missing)
##       Age         < 12.5    to the right, improve=85.70985, (0 missing)
##       FoodCourt   < 1331    to the left,  improve=77.60358, (0 missing)
##       VRDeck      < 417.5   to the right, improve=65.50853, (0 missing)
##   Surrogate splits:
##       HomePlanet splits as  RRLR,        agree=0.782, adj=0.153, (0 split)
##       deck       splits as  RRRLRRRLR,   agree=0.746, adj=0.010, (0 split)
##       Age        < 78.5    to the right, agree=0.743, adj=0.001, (0 split)
## 
## Node number 3: 2430 observations
##   predicted class=TRUE   expected loss=0.1851852  P(node) =0.3494392
##     class counts:   450  1980
##    probabilities: 0.185 0.815 
## 
## Node number 4: 1163 observations
##   predicted class=FALSE  expected loss=0.1625107  P(node) =0.1672419
##     class counts:   974   189
##    probabilities: 0.837 0.163 
## 
## Node number 5: 3361 observations,    complexity param=0.02887215
##   predicted class=FALSE  expected loss=0.3966082  P(node) =0.483319
##     class counts:  2028  1333
##    probabilities: 0.603 0.397 
##   left son=10 (815 obs) right son=11 (2546 obs)
##   Primary splits:
##       Spa          < 523     to the right, improve=124.74930, (0 missing)
##       VRDeck       < 407     to the right, improve=107.54040, (0 missing)
##       Age          < 12.5    to the right, improve= 60.01738, (0 missing)
##       FoodCourt    < 2507.5  to the left,  improve= 48.57491, (0 missing)
##       ShoppingMall < 622.5   to the left,  improve= 42.28737, (0 missing)
##   Surrogate splits:
##       VRDeck < 12683.5 to the right, agree=0.758, adj=0.002, (0 split)
##       Age    < 65.5    to the right, agree=0.758, adj=0.001, (0 split)
## 
## Node number 10: 815 observations
##   predicted class=FALSE  expected loss=0.1558282  P(node) =0.1171987
##     class counts:   688   127
##    probabilities: 0.844 0.156 
## 
## Node number 11: 2546 observations,    complexity param=0.02887215
##   predicted class=FALSE  expected loss=0.4736842  P(node) =0.3661202
##     class counts:  1340  1206
##    probabilities: 0.526 0.474 
##   left son=22 (777 obs) right son=23 (1769 obs)
##   Primary splits:
##       VRDeck       < 355     to the right, improve=142.39180, (0 missing)
##       FoodCourt    < 2063.5  to the left,  improve= 62.90754, (0 missing)
##       HomePlanet   splits as  LRRL,        improve= 37.56650, (0 missing)
##       Age          < 12.5    to the right, improve= 33.03081, (0 missing)
##       ShoppingMall < 1205    to the left,  improve= 29.11563, (0 missing)
##   Surrogate splits:
##       deck       splits as  RRLRRRR-R,   agree=0.705, adj=0.032, (0 split)
##       HomePlanet splits as  RLRR,        agree=0.700, adj=0.015, (0 split)
##       FoodCourt  < 6922.5  to the right, agree=0.698, adj=0.012, (0 split)
##       VIP        splits as  RLR,         agree=0.698, adj=0.009, (0 split)
##       Age        < 64.5    to the right, agree=0.697, adj=0.006, (0 split)
## 
## Node number 22: 777 observations,    complexity param=0.01013905
##   predicted class=FALSE  expected loss=0.2213642  P(node) =0.1117343
##     class counts:   605   172
##    probabilities: 0.779 0.221 
##   left son=44 (692 obs) right son=45 (85 obs)
##   Primary splits:
##       FoodCourt  < 2866    to the left,  improve=44.810930, (0 missing)
##       deck       splits as  LRRRLLL-L,   improve=22.050440, (0 missing)
##       HomePlanet splits as  LRLR,        improve=15.305190, (0 missing)
##       ailesıra   splits as  LRLRRRLL,    improve= 8.888805, (0 missing)
##       VRDeck     < 1678.5  to the right, improve= 8.710888, (0 missing)
##   Surrogate splits:
##       ailesıra    splits as  LLLLRLLL,    agree=0.893, adj=0.024, (0 split)
##       RoomService < 324     to the left,  agree=0.892, adj=0.012, (0 split)
## 
## Node number 23: 1769 observations,    complexity param=0.01289108
##   predicted class=TRUE   expected loss=0.415489  P(node) =0.254386
##     class counts:   735  1034
##    probabilities: 0.415 0.585 
##   left son=46 (1528 obs) right son=47 (241 obs)
##   Primary splits:
##       HomePlanet   splits as  LRLL,        improve=47.25637, (0 missing)
##       FoodCourt    < 1738.5  to the left,  improve=44.18785, (0 missing)
##       deck         splits as  RRRLLLL-L,   improve=43.43835, (0 missing)
##       Spa          < 205.5   to the right, improve=22.04282, (0 missing)
##       ShoppingMall < 1481.5  to the left,  improve=18.38204, (0 missing)
##   Surrogate splits:
##       deck         splits as  RRRLLLL-L,   agree=0.979, adj=0.842, (0 split)
##       FoodCourt    < 1809.5  to the left,  agree=0.921, adj=0.419, (0 split)
##       ShoppingMall < 8081    to the left,  agree=0.865, adj=0.012, (0 split)
##       Spa          < 518     to the left,  agree=0.865, adj=0.008, (0 split)
## 
## Node number 44: 692 observations
##   predicted class=FALSE  expected loss=0.1618497  P(node) =0.09951107
##     class counts:   580   112
##    probabilities: 0.838 0.162 
## 
## Node number 45: 85 observations
##   predicted class=TRUE   expected loss=0.2941176  P(node) =0.01222318
##     class counts:    25    60
##    probabilities: 0.294 0.706 
## 
## Node number 46: 1528 observations,    complexity param=0.01289108
##   predicted class=TRUE   expected loss=0.4613874  P(node) =0.2197297
##     class counts:   705   823
##    probabilities: 0.461 0.539 
##   left son=92 (255 obs) right son=93 (1273 obs)
##   Primary splits:
##       Spa          < 111     to the right, improve=27.805020, (0 missing)
##       ShoppingMall < 1205    to the left,  improve=20.602200, (0 missing)
##       VRDeck       < 11.5    to the right, improve=15.930710, (0 missing)
##       Age          < 4.5     to the right, improve=13.375610, (0 missing)
##       RoomService  < 103.5   to the right, improve= 9.157256, (0 missing)
##   Surrogate splits:
##       Age < 74.5    to the right, agree=0.834, adj=0.008, (0 split)
## 
## Node number 47: 241 observations
##   predicted class=TRUE   expected loss=0.1244813  P(node) =0.03465631
##     class counts:    30   211
##    probabilities: 0.124 0.876 
## 
## Node number 92: 255 observations
##   predicted class=FALSE  expected loss=0.3254902  P(node) =0.03666954
##     class counts:   172    83
##    probabilities: 0.675 0.325 
## 
## Node number 93: 1273 observations,    complexity param=0.0112978
##   predicted class=TRUE   expected loss=0.418696  P(node) =0.1830601
##     class counts:   533   740
##    probabilities: 0.419 0.581 
##   left son=186 (125 obs) right son=187 (1148 obs)
##   Primary splits:
##       VRDeck       < 129     to the right, improve=15.611210, (0 missing)
##       ShoppingMall < 1205    to the left,  improve=15.326260, (0 missing)
##       RoomService  < 104     to the right, improve=10.677220, (0 missing)
##       Age          < 3.5     to the right, improve= 7.897389, (0 missing)
##       FoodCourt    < 1738.5  to the left,  improve= 6.552242, (0 missing)
## 
## Node number 186: 125 observations
##   predicted class=FALSE  expected loss=0.344  P(node) =0.01797527
##     class counts:    82    43
##    probabilities: 0.656 0.344 
## 
## Node number 187: 1148 observations
##   predicted class=TRUE   expected loss=0.3928571  P(node) =0.1650848
##     class counts:   451   697
##    probabilities: 0.393 0.607
rpart.plot(fit_tree)

preds = predict(fit_tree, newdata = testing_set[-11], type = "class")
y_pred = ifelse(preds == TRUE, 1, 0)
cm = table(y_true, y_pred)
cm
##       y_pred
## y_true   0   1
##      0 623 240
##      1 123 753
(623+753)/(623+753+240+123)
## [1] 0.7912593
fit_tree <- rpart::rpart(Transported ~ ., data = train_set)
summary(fit_tree)
## Call:
## rpart::rpart(formula = Transported ~ ., data = train_set)
##   n= 8693 
## 
##           CP nsplit rel error    xerror        xstd
## 1 0.43244496      0 1.0000000 1.0000000 0.010803454
## 2 0.03429896      1 0.5675550 0.5675550 0.009719864
## 3 0.01000000      4 0.4646582 0.4730012 0.009158661
## 
## Variable importance
##    CryoSleep          Spa       VRDeck  RoomService    FoodCourt ShoppingMall 
##           44           17           15           11            7            4 
##   HomePlanet         deck 
##            2            1 
## 
## Node number 1: 8693 observations,    complexity param=0.432445
##   predicted class=TRUE   expected loss=0.4963764  P(node) =1
##     class counts:  4315  4378
##    probabilities: 0.496 0.504 
##   left son=2 (5656 obs) right son=3 (3037 obs)
##   Primary splits:
##       CryoSleep    splits as  LRL,        improve=920.2004, (0 missing)
##       RoomService  < 0.5    to the right, improve=523.3930, (0 missing)
##       Spa          < 0.5    to the right, improve=504.8747, (0 missing)
##       VRDeck       < 0.5    to the right, improve=462.3201, (0 missing)
##       ShoppingMall < 0.5    to the right, improve=282.7649, (0 missing)
##   Surrogate splits:
##       Spa          < 0.5    to the right, agree=0.722, adj=0.204, (0 split)
##       FoodCourt    < 0.5    to the right, agree=0.706, adj=0.157, (0 split)
##       VRDeck       < 0.5    to the right, agree=0.703, adj=0.150, (0 split)
##       RoomService  < 0.5    to the right, agree=0.692, adj=0.119, (0 split)
##       ShoppingMall < 0.5    to the right, agree=0.685, adj=0.097, (0 split)
## 
## Node number 2: 5656 observations,    complexity param=0.03429896
##   predicted class=FALSE  expected loss=0.3350424  P(node) =0.6506384
##     class counts:  3761  1895
##    probabilities: 0.665 0.335 
##   left son=4 (1432 obs) right son=5 (4224 obs)
##   Primary splits:
##       RoomService < 346.5  to the right, improve=121.3964, (0 missing)
##       Spa         < 266.5  to the right, improve=112.9389, (0 missing)
##       Age         < 12.5   to the right, improve=109.4055, (0 missing)
##       FoodCourt   < 1331   to the left,  improve= 98.1198, (0 missing)
##       VRDeck      < 587.5  to the right, improve= 75.0317, (0 missing)
##   Surrogate splits:
##       HomePlanet splits as  RRLR,       agree=0.785, adj=0.151, (0 split)
##       deck       splits as  RRRLRRRRR,  agree=0.749, adj=0.007, (0 split)
##       Age        < 78.5   to the right, agree=0.747, adj=0.001, (0 split)
## 
## Node number 3: 3037 observations
##   predicted class=TRUE   expected loss=0.1824169  P(node) =0.3493616
##     class counts:   554  2483
##    probabilities: 0.182 0.818 
## 
## Node number 4: 1432 observations
##   predicted class=FALSE  expected loss=0.1571229  P(node) =0.1647302
##     class counts:  1207   225
##    probabilities: 0.843 0.157 
## 
## Node number 5: 4224 observations,    complexity param=0.03429896
##   predicted class=FALSE  expected loss=0.3953598  P(node) =0.4859082
##     class counts:  2554  1670
##    probabilities: 0.605 0.395 
##   left son=10 (1459 obs) right son=11 (2765 obs)
##   Primary splits:
##       Spa          < 205    to the right, improve=166.33280, (0 missing)
##       VRDeck       < 417.5  to the right, improve=127.52200, (0 missing)
##       Age          < 12.5   to the right, improve= 76.46367, (0 missing)
##       FoodCourt    < 2507.5 to the left,  improve= 63.32833, (0 missing)
##       ShoppingMall < 627    to the left,  improve= 59.33765, (0 missing)
##   Surrogate splits:
##       HomePlanet splits as  RLRR,       agree=0.686, adj=0.092, (0 split)
##       deck       splits as  LLLRRRRLR,  agree=0.674, adj=0.055, (0 split)
##       FoodCourt  < 3197.5 to the right, agree=0.662, adj=0.022, (0 split)
##       VRDeck     < 2052   to the right, agree=0.659, adj=0.014, (0 split)
##       Age        < 65.5   to the right, agree=0.656, adj=0.003, (0 split)
## 
## Node number 10: 1459 observations
##   predicted class=FALSE  expected loss=0.2021933  P(node) =0.1678362
##     class counts:  1164   295
##    probabilities: 0.798 0.202 
## 
## Node number 11: 2765 observations,    complexity param=0.03429896
##   predicted class=FALSE  expected loss=0.4972875  P(node) =0.318072
##     class counts:  1390  1375
##    probabilities: 0.503 0.497 
##   left son=22 (807 obs) right son=23 (1958 obs)
##   Primary splits:
##       VRDeck       < 355    to the right, improve=180.83390, (0 missing)
##       FoodCourt    < 2069.5 to the left,  improve= 60.83808, (0 missing)
##       HomePlanet   splits as  LRRL,       improve= 46.66727, (0 missing)
##       ShoppingMall < 1540.5 to the left,  improve= 37.98636, (0 missing)
##       Age          < 7.5    to the right, improve= 34.98638, (0 missing)
##   Surrogate splits:
##       deck        splits as  RRLRRRRRR,  agree=0.712, adj=0.014, (0 split)
##       FoodCourt   < 6922.5 to the right, agree=0.710, adj=0.007, (0 split)
##       Age         < 68.5   to the right, agree=0.709, adj=0.002, (0 split)
##       RoomService < 343    to the right, agree=0.708, adj=0.001, (0 split)
## 
## Node number 22: 807 observations
##   predicted class=FALSE  expected loss=0.2156134  P(node) =0.09283331
##     class counts:   633   174
##    probabilities: 0.784 0.216 
## 
## Node number 23: 1958 observations
##   predicted class=TRUE   expected loss=0.386619  P(node) =0.2252387
##     class counts:   757  1201
##    probabilities: 0.387 0.613
rpart.plot(fit_tree)

preds = predict(fit_tree, newdata = test_set, type = "class")
y_pred = ifelse(preds == TRUE, TRUE, FALSE)
Transported <- as.character(y_pred)
PassengerId <- test$PassengerId
Transported <- as.vector(Transported)
DT_sonuc <- cbind(PassengerId, Transported)
DT_sonuc <- as.data.frame(DT_sonuc)
DT_sonuc$Transported <- str_to_title(DT_sonuc$Transported)
write.csv(DT_sonuc, "DT_sonuc.csv", row.names = FALSE, quote = FALSE)

DECİSİON TREES SONUCU