Spaceship Titanic

Kozmik bir gizemi çözmek için veri bilimi becerilerinize ihtiyaç duyulan 2912 yılına hoş geldiniz. Dört ışık yılı öteden bir sinyal aldık ve işler pek iyi görünmüyor.

Uzay Gemisi Titanik, bir ay önce fırlatılan yıldızlararası bir yolcu gemisiydi. Gemide neredeyse 13.000 yolcu bulunan gemi, güneş sistemimizden göçmenleri yakın yıldızların yörüngesinde bulunan üç yeni yaşanabilir dış gezegene taşımak üzere ilk yolculuğuna çıktı.

Dikkatsiz Uzay Gemisi Titanic, ilk varış noktası olan kavurucu 55 Cancri E’ye giderken Alpha Centauri’yi dönerken, bir toz bulutunun içine gizlenmiş bir uzay-zaman anormalliğiyle çarpıştı. Ne yazık ki 1000 yıl öncesindeki adaşı ile benzer bir kaderle karşılaştı. Gemi sağlam kalmasına rağmen yolcuların neredeyse yarısı alternatif bir boyuta taşındı!

İlk önce elimizdeki train datasını library kodunu kullanarak R yüklememiz gerekiyor.

library(readr)
train <- read_csv("train.csv")
## Rows: 8693 Columns: 14
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (5): PassengerId, HomePlanet, Cabin, Destination, Name
## dbl (6): Age, RoomService, FoodCourt, ShoppingMall, Spa, VRDeck
## lgl (3): CryoSleep, VIP, Transported
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.

Daha sonra str kodunu kullanarak veri setimizdeki kaç tane gözlem olduğunu ve ne anlama geldiğini öğrenelim.

str(train)
## spc_tbl_ [8,693 × 14] (S3: spec_tbl_df/tbl_df/tbl/data.frame)
##  $ PassengerId : chr [1:8693] "0001_01" "0002_01" "0003_01" "0003_02" ...
##  $ HomePlanet  : chr [1:8693] "Europa" "Earth" "Europa" "Europa" ...
##  $ CryoSleep   : logi [1:8693] FALSE FALSE FALSE FALSE FALSE FALSE ...
##  $ Cabin       : chr [1:8693] "B/0/P" "F/0/S" "A/0/S" "A/0/S" ...
##  $ Destination : chr [1:8693] "TRAPPIST-1e" "TRAPPIST-1e" "TRAPPIST-1e" "TRAPPIST-1e" ...
##  $ Age         : num [1:8693] 39 24 58 33 16 44 26 28 35 14 ...
##  $ VIP         : logi [1:8693] FALSE FALSE TRUE FALSE FALSE FALSE ...
##  $ RoomService : num [1:8693] 0 109 43 0 303 0 42 0 0 0 ...
##  $ FoodCourt   : num [1:8693] 0 9 3576 1283 70 ...
##  $ ShoppingMall: num [1:8693] 0 25 0 371 151 0 3 0 17 0 ...
##  $ Spa         : num [1:8693] 0 549 6715 3329 565 ...
##  $ VRDeck      : num [1:8693] 0 44 49 193 2 0 0 NA 0 0 ...
##  $ Name        : chr [1:8693] "Maham Ofracculy" "Juanna Vines" "Altark Susent" "Solam Susent" ...
##  $ Transported : logi [1:8693] FALSE TRUE FALSE FALSE TRUE TRUE ...
##  - attr(*, "spec")=
##   .. cols(
##   ..   PassengerId = col_character(),
##   ..   HomePlanet = col_character(),
##   ..   CryoSleep = col_logical(),
##   ..   Cabin = col_character(),
##   ..   Destination = col_character(),
##   ..   Age = col_double(),
##   ..   VIP = col_logical(),
##   ..   RoomService = col_double(),
##   ..   FoodCourt = col_double(),
##   ..   ShoppingMall = col_double(),
##   ..   Spa = col_double(),
##   ..   VRDeck = col_double(),
##   ..   Name = col_character(),
##   ..   Transported = col_logical()
##   .. )
##  - attr(*, "problems")=<externalptr>

Oluşturulan verilerin raporunu görmek için ilk önce paket bölümünden DataExplorer paketi yüklenir ve library kodu kullanılarak R a yükleriz. Daha sonra create_report kodu ile train data verisinin tablosu çıkarılır.

library(DataExplorer)
create_report(train)
## 
## 
## processing file: report.rmd
## 
  |                                           
  |                                     |   0%
  |                                           
  |.                                    |   2%                                 
  |                                           
  |..                                   |   5% [global_options]                
  |                                           
  |...                                  |   7%                                 
  |                                           
  |....                                 |  10% [introduce]                     
  |                                           
  |....                                 |  12%                                 
  |                                           
  |.....                                |  14% [plot_intro]                    
  |                                           
  |......                               |  17%                                 
  |                                           
  |.......                              |  19% [data_structure]                
  |                                           
  |........                             |  21%                                 
  |                                           
  |.........                            |  24% [missing_profile]               
  |                                           
  |..........                           |  26%                                 
  |                                           
  |...........                          |  29% [univariate_distribution_header]
  |                                           
  |...........                          |  31%                                 
  |                                           
  |............                         |  33% [plot_histogram]                
  |                                           
  |.............                        |  36%                                 
  |                                           
  |..............                       |  38% [plot_density]                  
  |                                           
  |...............                      |  40%                                 
  |                                           
  |................                     |  43% [plot_frequency_bar]            
  |                                           
  |.................                    |  45%                                 
  |                                           
  |..................                   |  48% [plot_response_bar]             
  |                                           
  |..................                   |  50%                                 
  |                                           
  |...................                  |  52% [plot_with_bar]                 
  |                                           
  |....................                 |  55%                                 
  |                                           
  |.....................                |  57% [plot_normal_qq]                
  |                                           
  |......................               |  60%                                 
  |                                           
  |.......................              |  62% [plot_response_qq]              
  |                                           
  |........................             |  64%                                 
  |                                           
  |.........................            |  67% [plot_by_qq]                    
  |                                           
  |..........................           |  69%                                 
  |                                           
  |..........................           |  71% [correlation_analysis]          
  |                                           
  |...........................          |  74%                                 
  |                                           
  |............................         |  76% [principal_component_analysis]  
  |                                           
  |.............................        |  79%                                 
  |                                           
  |..............................       |  81% [bivariate_distribution_header] 
  |                                           
  |...............................      |  83%                                 
  |                                           
  |................................     |  86% [plot_response_boxplot]         
  |                                           
  |.................................    |  88%                                 
  |                                           
  |.................................    |  90% [plot_by_boxplot]               
  |                                           
  |..................................   |  93%                                 
  |                                           
  |...................................  |  95% [plot_response_scatterplot]     
  |                                           
  |.................................... |  98%                                 
  |                                           
  |.....................................| 100% [plot_by_scatterplot]           
## output file: /cloud/project/report.knit.md
## /usr/lib/rstudio-server/bin/quarto/bin/tools/pandoc +RTS -K512m -RTS /cloud/project/report.knit.md --to html4 --from markdown+autolink_bare_uris+tex_math_single_backslash --output /cloud/project/report.html --lua-filter /cloud/lib/x86_64-pc-linux-gnu-library/4.3/rmarkdown/rmarkdown/lua/pagebreak.lua --lua-filter /cloud/lib/x86_64-pc-linux-gnu-library/4.3/rmarkdown/rmarkdown/lua/latex-div.lua --embed-resources --standalone --variable bs3=TRUE --section-divs --table-of-contents --toc-depth 6 --template /cloud/lib/x86_64-pc-linux-gnu-library/4.3/rmarkdown/rmd/h/default.html --no-highlight --variable highlightjs=1 --variable theme=yeti --mathjax --variable 'mathjax-url=https://mathjax.rstudio.com/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML' --include-in-header /tmp/Rtmp1oVMqE/rmarkdown-strf511c347f17.html
## 
## Output created: report.html

veride ilk 6 sütunu görmek için yukarıdaki kodu kullanırız.

ilkaltısütun<- train[,1:6]

ilk 6 sütuna ek olarak farklı bir sütun eklemek için yukarıdaki kodu kullanırız.

ilkaltısütun1<- train[,c(1:6,14)]

Veririyi düzenlemek ve temizlemek için tidyverse paketi kullanırız.

library(tidyverse)
## ── Attaching core tidyverse packages ──────────────────────── tidyverse 2.0.0 ──
## ✔ dplyr     1.1.4     ✔ purrr     1.0.2
## ✔ forcats   1.0.0     ✔ stringr   1.5.1
## ✔ ggplot2   3.4.4     ✔ tibble    3.2.1
## ✔ lubridate 1.9.3     ✔ tidyr     1.3.0
## ── Conflicts ────────────────────────────────────────── tidyverse_conflicts() ──
## ✖ dplyr::between()     masks data.table::between()
## ✖ dplyr::filter()      masks stats::filter()
## ✖ dplyr::first()       masks data.table::first()
## ✖ lubridate::hour()    masks data.table::hour()
## ✖ lubridate::isoweek() masks data.table::isoweek()
## ✖ dplyr::lag()         masks stats::lag()
## ✖ dplyr::last()        masks data.table::last()
## ✖ lubridate::mday()    masks data.table::mday()
## ✖ lubridate::minute()  masks data.table::minute()
## ✖ lubridate::month()   masks data.table::month()
## ✖ lubridate::quarter() masks data.table::quarter()
## ✖ lubridate::second()  masks data.table::second()
## ✖ purrr::transpose()   masks data.table::transpose()
## ✖ lubridate::wday()    masks data.table::wday()
## ✖ lubridate::week()    masks data.table::week()
## ✖ lubridate::yday()    masks data.table::yday()
## ✖ lubridate::year()    masks data.table::year()
## ℹ Use the conflicted package (<http://conflicted.r-lib.org/>) to force all conflicts to become errors

Data verimizde kaç tane NA olduğu ve ortalamayı göstermek için aşağıdaki kodu kullanırız.

library(explore)
train %>% describe_all()
## # A tibble: 14 × 8
##    variable     type     na na_pct unique   min   mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl>  <dbl> <dbl>
##  1 PassengerId  chr       0    0     8693    NA  NA       NA
##  2 HomePlanet   chr     201    2.3      4    NA  NA       NA
##  3 CryoSleep    lgl     217    2.5      3     0   0.36     1
##  4 Cabin        chr     199    2.3   6561    NA  NA       NA
##  5 Destination  chr     182    2.1      4    NA  NA       NA
##  6 Age          dbl     179    2.1     81     0  28.8     79
##  7 VIP          lgl     203    2.3      3     0   0.02     1
##  8 RoomService  dbl     181    2.1   1274     0 225.   14327
##  9 FoodCourt    dbl     183    2.1   1508     0 458.   29813
## 10 ShoppingMall dbl     208    2.4   1116     0 174.   23492
## 11 Spa          dbl     183    2.1   1328     0 311.   22408
## 12 VRDeck       dbl     188    2.2   1307     0 305.   24133
## 13 Name         chr     200    2.3   8474    NA  NA       NA
## 14 Transported  lgl       0    0        2     0   0.5      1

Data verisindeki passengerId sütununu ailenumarası ve ailesıra numarası olarak ayırmak için aşağıdaki kodu kullanıyoruz.

train$ailenum<- str_split_fixed(train$PassengerId, "_",2)

Yukarıdaki işlemin ikinci yolu olarak aşağıdaki kodu kullanabilriz.

train[c('ailenum','ailesıra')]<- str_split_fixed(train$PassengerId, "_",2)

Ailenum ve Ailesıra sütununu veride başa almak için aşağıdaki kodu kullanırız

train<- train[,c(15:16,1:14)]

İlk altı gözlemi görmek için head kodu kullanılır.

head(train)
## # A tibble: 6 × 16
##   ailenum ailesıra PassengerId HomePlanet CryoSleep Cabin Destination     Age
##   <chr>   <chr>    <chr>       <chr>      <lgl>     <chr> <chr>         <dbl>
## 1 0001    01       0001_01     Europa     FALSE     B/0/P TRAPPIST-1e      39
## 2 0002    01       0002_01     Earth      FALSE     F/0/S TRAPPIST-1e      24
## 3 0003    01       0003_01     Europa     FALSE     A/0/S TRAPPIST-1e      58
## 4 0003    02       0003_02     Europa     FALSE     A/0/S TRAPPIST-1e      33
## 5 0004    01       0004_01     Earth      FALSE     F/1/S TRAPPIST-1e      16
## 6 0005    01       0005_01     Earth      FALSE     F/0/P PSO J318.5-22    44
## # ℹ 8 more variables: VIP <lgl>, RoomService <dbl>, FoodCourt <dbl>,
## #   ShoppingMall <dbl>, Spa <dbl>, VRDeck <dbl>, Name <chr>, Transported <lgl>

Veri tablosunda sütunların boş olan veya bilinmeyen boş bilgilere NA değeri eklenemek için aşağıdaki kodları kullanırız.

train $PassengerId<-addNA(train$PassengerId)
train $HomePlanet<-addNA(train$HomePlanet)
train$CryoSleep<-addNA(train$CryoSleep)
train $Cabin<-addNA(train$Cabin)
train $Destination<-addNA(train$Destination)
train $Age<-addNA(train$Age)
train $VIP<-addNA(train$VIP)
train $RoomService<-addNA(train$RoomService)
train $FoodCourt<-addNA(train$FoodCourt)
train $ShoppingMall<-addNA(train$ShoppingMall)
train $Spa<-addNA(train$Spa)
train $VRDeck<-addNA(train$VRDeck)
train $Name<-addNA(train$Name)
train $Transported<-addNA(train$Transported)

Bir sütundaki her bir farklı değeri öğrenmek için unique kodunu kullanırız. Örnek olarak HomePlanet sütununu kullanalım.

unique(train$HomePlanet)
## [1] Europa Earth  Mars   <NA>  
## Levels: Earth Europa Mars <NA>
library(png)

train datasını bitirdikten sonra elimizdeki test datasını library kodunu kullanarak R yüklememiz gerekiyor.

library(readr)
test <- read_csv("test.csv")
## Rows: 4277 Columns: 13
## ── Column specification ────────────────────────────────────────────────────────
## Delimiter: ","
## chr (5): PassengerId, HomePlanet, Cabin, Destination, Name
## dbl (6): Age, RoomService, FoodCourt, ShoppingMall, Spa, VRDeck
## lgl (2): CryoSleep, VIP
## 
## ℹ Use `spec()` to retrieve the full column specification for this data.
## ℹ Specify the column types or set `show_col_types = FALSE` to quiet this message.

Daha sonra str kodunu kullanarak veri setimizdeki kaç tane gözlem olduğunu ve ne anlama geldiğini öğrenelim.

str(test)
## spc_tbl_ [4,277 × 13] (S3: spec_tbl_df/tbl_df/tbl/data.frame)
##  $ PassengerId : chr [1:4277] "0013_01" "0018_01" "0019_01" "0021_01" ...
##  $ HomePlanet  : chr [1:4277] "Earth" "Earth" "Europa" "Europa" ...
##  $ CryoSleep   : logi [1:4277] TRUE FALSE TRUE FALSE FALSE FALSE ...
##  $ Cabin       : chr [1:4277] "G/3/S" "F/4/S" "C/0/S" "C/1/S" ...
##  $ Destination : chr [1:4277] "TRAPPIST-1e" "TRAPPIST-1e" "55 Cancri e" "TRAPPIST-1e" ...
##  $ Age         : num [1:4277] 27 19 31 38 20 31 21 20 23 24 ...
##  $ VIP         : logi [1:4277] FALSE FALSE FALSE FALSE FALSE FALSE ...
##  $ RoomService : num [1:4277] 0 0 0 0 10 0 0 0 0 0 ...
##  $ FoodCourt   : num [1:4277] 0 9 0 6652 0 ...
##  $ ShoppingMall: num [1:4277] 0 0 0 0 635 263 0 0 0 0 ...
##  $ Spa         : num [1:4277] 0 2823 0 181 0 ...
##  $ VRDeck      : num [1:4277] 0 0 0 585 0 60 0 0 0 0 ...
##  $ Name        : chr [1:4277] "Nelly Carsoning" "Lerome Peckers" "Sabih Unhearfus" "Meratz Caltilter" ...
##  - attr(*, "spec")=
##   .. cols(
##   ..   PassengerId = col_character(),
##   ..   HomePlanet = col_character(),
##   ..   CryoSleep = col_logical(),
##   ..   Cabin = col_character(),
##   ..   Destination = col_character(),
##   ..   Age = col_double(),
##   ..   VIP = col_logical(),
##   ..   RoomService = col_double(),
##   ..   FoodCourt = col_double(),
##   ..   ShoppingMall = col_double(),
##   ..   Spa = col_double(),
##   ..   VRDeck = col_double(),
##   ..   Name = col_character()
##   .. )
##  - attr(*, "problems")=<externalptr>

Oluşturulan verilerin raporunu görmek için ilk önce paket bölümünden DataExplorer paketi yüklenir ve library kodu kullanılarak R a yükleriz. Daha sonra create_report kodu ile train data verisinin tablosu çıkarılır.

library(DataExplorer)
create_report(test)
## 
## 
## processing file: report.rmd
## 
  |                                           
  |                                     |   0%
  |                                           
  |.                                    |   2%                                 
  |                                           
  |..                                   |   5% [global_options]                
  |                                           
  |...                                  |   7%                                 
  |                                           
  |....                                 |  10% [introduce]                     
  |                                           
  |....                                 |  12%                                 
  |                                           
  |.....                                |  14% [plot_intro]                    
  |                                           
  |......                               |  17%                                 
  |                                           
  |.......                              |  19% [data_structure]                
  |                                           
  |........                             |  21%                                 
  |                                           
  |.........                            |  24% [missing_profile]               
  |                                           
  |..........                           |  26%                                 
  |                                           
  |...........                          |  29% [univariate_distribution_header]
  |                                           
  |...........                          |  31%                                 
  |                                           
  |............                         |  33% [plot_histogram]                
  |                                           
  |.............                        |  36%                                 
  |                                           
  |..............                       |  38% [plot_density]                  
  |                                           
  |...............                      |  40%                                 
  |                                           
  |................                     |  43% [plot_frequency_bar]            
  |                                           
  |.................                    |  45%                                 
  |                                           
  |..................                   |  48% [plot_response_bar]             
  |                                           
  |..................                   |  50%                                 
  |                                           
  |...................                  |  52% [plot_with_bar]                 
  |                                           
  |....................                 |  55%                                 
  |                                           
  |.....................                |  57% [plot_normal_qq]                
  |                                           
  |......................               |  60%                                 
  |                                           
  |.......................              |  62% [plot_response_qq]              
  |                                           
  |........................             |  64%                                 
  |                                           
  |.........................            |  67% [plot_by_qq]                    
  |                                           
  |..........................           |  69%                                 
  |                                           
  |..........................           |  71% [correlation_analysis]          
  |                                           
  |...........................          |  74%                                 
  |                                           
  |............................         |  76% [principal_component_analysis]  
  |                                           
  |.............................        |  79%                                 
  |                                           
  |..............................       |  81% [bivariate_distribution_header] 
  |                                           
  |...............................      |  83%                                 
  |                                           
  |................................     |  86% [plot_response_boxplot]         
  |                                           
  |.................................    |  88%                                 
  |                                           
  |.................................    |  90% [plot_by_boxplot]               
  |                                           
  |..................................   |  93%                                 
  |                                           
  |...................................  |  95% [plot_response_scatterplot]     
  |                                           
  |.................................... |  98%                                 
  |                                           
  |.....................................| 100% [plot_by_scatterplot]           
## output file: /cloud/project/report.knit.md
## /usr/lib/rstudio-server/bin/quarto/bin/tools/pandoc +RTS -K512m -RTS /cloud/project/report.knit.md --to html4 --from markdown+autolink_bare_uris+tex_math_single_backslash --output /cloud/project/report.html --lua-filter /cloud/lib/x86_64-pc-linux-gnu-library/4.3/rmarkdown/rmarkdown/lua/pagebreak.lua --lua-filter /cloud/lib/x86_64-pc-linux-gnu-library/4.3/rmarkdown/rmarkdown/lua/latex-div.lua --embed-resources --standalone --variable bs3=TRUE --section-divs --table-of-contents --toc-depth 6 --template /cloud/lib/x86_64-pc-linux-gnu-library/4.3/rmarkdown/rmd/h/default.html --no-highlight --variable highlightjs=1 --variable theme=yeti --mathjax --variable 'mathjax-url=https://mathjax.rstudio.com/latest/MathJax.js?config=TeX-AMS-MML_HTMLorMML' --include-in-header /tmp/Rtmp1oVMqE/rmarkdown-strf5130dcfca9.html
## 
## Output created: report.html

Veride ilk 6 sütunu görmek için yukarıdaki kodu kullanırız.

ilkaltısütun2<- test[,1:8]

ilk 6 sütuna ek olarak farklı bir sütun eklemek için yukarıdaki kodu kullanırız.

ilkaltısütun3<- test[,c(1:8,12)]

Veririyi düzenlemek ve temizlemek için tidyverse paketi kullanırız.

library(tidyverse)

Data verimizde kaç tane NA olduğu ve ortalamayı göstermek için aşağıdaki kodu kullanırız.

test%>%describe_all()
## # A tibble: 13 × 8
##    variable     type     na na_pct unique   min   mean   max
##    <chr>        <chr> <int>  <dbl>  <int> <dbl>  <dbl> <dbl>
##  1 PassengerId  chr       0    0     4277    NA  NA       NA
##  2 HomePlanet   chr      87    2        4    NA  NA       NA
##  3 CryoSleep    lgl      93    2.2      3     0   0.37     1
##  4 Cabin        chr     100    2.3   3266    NA  NA       NA
##  5 Destination  chr      92    2.2      4    NA  NA       NA
##  6 Age          dbl      91    2.1     80     0  28.7     79
##  7 VIP          lgl      93    2.2      3     0   0.02     1
##  8 RoomService  dbl      82    1.9    843     0 219.   11567
##  9 FoodCourt    dbl     106    2.5    903     0 439.   25273
## 10 ShoppingMall dbl      98    2.3    716     0 177.    8292
## 11 Spa          dbl     101    2.4    834     0 303.   19844
## 12 VRDeck       dbl      80    1.9    797     0 311.   22272
## 13 Name         chr      94    2.2   4177    NA  NA       NA

Data verisindeki passengerId sütununu ailenumarası ve ailesıra numarası olarak ayırmak için aşağıdaki kodu kullanıyoruz.

test$ailenum<- str_split_fixed(test$PassengerId, "_",2)

Yukarıdaki işlemin ikinci yolu olarak aşağıdaki kodu kullanabilriz

test[c('ailenum','ailesıra')]<- str_split_fixed(test$PassengerId, "_",2)

Ailenum ve Ailesıra sütununu veride başa almak için aşağıdaki kodu kullanırız.

test<- test[,c(14:15,1:13)]

İlk altı gözlemi görmek için head kodu kullanılır.

head(test)
## # A tibble: 6 × 15
##   ailenum ailesıra PassengerId HomePlanet CryoSleep Cabin Destination   Age
##   <chr>   <chr>    <chr>       <chr>      <lgl>     <chr> <chr>       <dbl>
## 1 0013    01       0013_01     Earth      TRUE      G/3/S TRAPPIST-1e    27
## 2 0018    01       0018_01     Earth      FALSE     F/4/S TRAPPIST-1e    19
## 3 0019    01       0019_01     Europa     TRUE      C/0/S 55 Cancri e    31
## 4 0021    01       0021_01     Europa     FALSE     C/1/S TRAPPIST-1e    38
## 5 0023    01       0023_01     Earth      FALSE     F/5/S TRAPPIST-1e    20
## 6 0027    01       0027_01     Earth      FALSE     F/7/P TRAPPIST-1e    31
## # ℹ 7 more variables: VIP <lgl>, RoomService <dbl>, FoodCourt <dbl>,
## #   ShoppingMall <dbl>, Spa <dbl>, VRDeck <dbl>, Name <chr>

Veri tablosunda sütunların boş olan veya bilinmeyen boş bilgilere NA değeri eklenemek için aşağıdaki kodları kullanırız.

test $PassengerId<-addNA(test$PassengerId)
test $HomePlanet<-addNA(test$HomePlanet)
test$CryoSleep<-addNA(test$CryoSleep)
test $Cabin<-addNA(test$Cabin)
test $Destination<-addNA(test$Destination)
test $Age<-addNA(test$Age)
test $VIP<-addNA(test$VIP)
test $RoomService<-addNA(test$RoomService)
test $FoodCourt<-addNA(test$FoodCourt)
test $ShoppingMall<-addNA(test$ShoppingMall)
test $Spa<-addNA(test$Spa)
test $VRDeck<-addNA(test$VRDeck)
test $Name<-addNA(test$Name)

Bir sütundaki her bir farklı değeri öğrenmek için unique kodunu kullanırız. Örnek olarak HomePlanet sütununu kullanalım.

unique(test$HomePlanet)
## [1] Earth  Europa Mars   <NA>  
## Levels: Earth Europa Mars <NA>