R Markdown

Descarga de paquetes

library(lsm)      # Para descargar una base de datos
library(dplyr)
## 
## Adjuntando el paquete: 'dplyr'
## The following objects are masked from 'package:stats':
## 
##     filter, lag
## The following objects are masked from 'package:base':
## 
##     intersect, setdiff, setequal, union
library(moments)  # Para hallar las medidas de forma
library(e1071)
## 
## Adjuntando el paquete: 'e1071'
## The following objects are masked from 'package:moments':
## 
##     kurtosis, moment, skewness
library(ggplot2)
## 
## Adjuntando el paquete: 'ggplot2'
## The following object is masked from 'package:e1071':
## 
##     element

Interpretación: empleamos aquí el chunck para obtener los paquetes que se utilizaran durante todo el trabajo

Base de datos

datosCompleto <- lsm::survey

Interpretación: aqui seleccionamos un conjunto de datos (survey) del paquete (lsm)

Revisando el data Frame

Para visualizar solo una parte de los datos, se pueden utilizar las funciones head y/o tail.

tail(datosCompleto)        #C) Por defecto, solo las últimas 6 observaciones
## # A tibble: 6 × 66
##   Observation ID       Gender Like    Age Smoke Height Weight   BMI School SES  
##         <dbl> <chr>    <chr>  <chr> <dbl> <chr>  <dbl>  <dbl> <dbl> <chr>  <chr>
## 1         795 AC31201… Female <NA>   NA   <NA>    1.64     NA  NA   <NA>   Low  
## 2         796 AC31201… Female TV     13.5 <NA>    1.71     78  26.7 Public <NA> 
## 3         797 AC31201… <NA>   Netw…  15.8 No      1.68     53  18.8 Priva… Medi…
## 4         798 AC31201… Male   <NA>   15.7 No     NA        83  NA   <NA>   Low  
## 5         799 AC31201… Female TV     NA   No      1.76     73  23.6 Priva… Low  
## 6         800 AC31201… Male   TV     16.6 No      1.62     70  26.7 Priva… <NA> 
## # ℹ 55 more variables: Enrollment <chr>, Score <dbl>, MotherHeight <chr>,
## #   MotherAge <dbl>, MotherCHD <dbl>, FatherHeight <chr>, FatherAge <dbl>,
## #   FatherCHD <dbl>, Status <chr>, SemAcum <dbl>, Exam1 <dbl>, Exam2 <dbl>,
## #   Exam3 <dbl>, Exam4 <dbl>, ExamAcum <dbl>, Definitive <dbl>, Expense <dbl>,
## #   Income <dbl>, Gas <dbl>, Course <chr>, Law <chr>, Economic <chr>,
## #   Race <chr>, Region <chr>, EMO1 <dbl>, EMO2 <dbl>, EMO3 <dbl>, EMO4 <dbl>,
## #   EMO5 <dbl>, GOAL1 <chr>, GOAL2 <chr>, GOAL3 <chr>, Pre_STAT1 <dbl>, …

Interpretación: la tabla muestra que no todas las observaciones incluyen datos en todas las variables consideradas para el estudio

Revisar base de datos

Esta función permite ver la estructura. Esta brinda información sobre el tipo de objeto, la cantidad de filas y columnas, así como datos complementarios, tales como los nombres de las variables y su tipo, además de algunas observaciones iniciales correspondientes a cada una.

str(datosCompleto)   #A) Estructura de los datos
## tibble [800 × 66] (S3: tbl_df/tbl/data.frame)
##  $ Observation : num [1:800] 1 2 3 4 5 6 7 8 9 10 ...
##  $ ID          : chr [1:800] "SB11201910010435" "SB11201910004475" "SB11201910011427" "SB11201910041975" ...
##  $ Gender      : chr [1:800] "Female" "Male" "Male" "Male" ...
##  $ Like        : chr [1:800] "TV" "Network" "Network" "TV" ...
##  $ Age         : num [1:800] 21.4 21.1 20.9 18.4 16.6 ...
##  $ Smoke       : chr [1:800] "No" "Yes" "Yes" "Yes" ...
##  $ Height      : num [1:800] 1.58 1.6 1.5 1.53 1.78 1.65 1.73 1.53 1.64 1.52 ...
##  $ Weight      : num [1:800] 75 80 64 49 82 80 90 55 50 78 ...
##  $ BMI         : num [1:800] 30 31.2 28.4 20.9 25.9 ...
##  $ School      : chr [1:800] "Private" "Public" "Private" "Public" ...
##  $ SES         : chr [1:800] "Medium" "High" "High" "Low" ...
##  $ Enrollment  : chr [1:800] "Credit" "Scholarship" "Scholarship" "Credit" ...
##  $ Score       : num [1:800] 81 78 77 70 68 65 54 50 36 35 ...
##  $ MotherHeight: chr [1:800] "Short_M" "Normal_M" "Normal_M" "Tall_M" ...
##  $ MotherAge   : num [1:800] 41 45 45 45 46 46 47 48 48 48 ...
##  $ MotherCHD   : num [1:800] 0 0 0 0 1 0 0 0 0 1 ...
##  $ FatherHeight: chr [1:800] "Normal_F" "Short_F" "Tall_F" "Short_F" ...
##  $ FatherAge   : num [1:800] 40 43 44 45 45 46 46 48 48 49 ...
##  $ FatherCHD   : num [1:800] 1 1 1 2 1 1 1 1 1 1 ...
##  $ Status      : chr [1:800] "Distinguished" "Distinguished" "Distinguished" "Regular" ...
##  $ SemAcum     : num [1:800] 4.25 2.8 4.15 3.2 3.45 2.75 2.7 4.35 4.3 2.8 ...
##  $ Exam1       : num [1:800] 1.5 2.3 3.4 2.5 3.1 3.8 5 4 2.5 2.4 ...
##  $ Exam2       : num [1:800] 5 4.9 3.6 4.2 3.5 4.4 3 2.3 3.3 2.6 ...
##  $ Exam3       : num [1:800] 5 3.7 2 5 5 4.2 3.5 4.6 3.8 4.3 ...
##  $ Exam4       : num [1:800] 4.5 3.3 1.9 2.5 3 5 3.6 4.3 1.9 5 ...
##  $ ExamAcum    : num [1:800] 16 14.2 10.9 14.2 14.6 17.4 15.1 15.2 11.5 14.3 ...
##  $ Definitive  : num [1:800] 4 3.55 2.73 3.55 3.65 ...
##  $ Expense     : num [1:800] 48.9 72.1 85.2 56.6 64.6 63 40.8 65.4 37.3 63 ...
##  $ Income      : num [1:800] 1.61 2.07 2.84 1.55 2.32 2.1 1.69 2.18 1.71 2.1 ...
##  $ Gas         : num [1:800] 27.4 24.2 22.3 23.1 27.3 ...
##  $ Course      : chr [1:800] "Face-to-Face" "Virtual" "Face-to-Face" "Virtual" ...
##  $ Law         : chr [1:800] "Agree" "Agree" "Agree" "Agree" ...
##  $ Economic    : chr [1:800] "Regular" "Good" "Regular" "Bad" ...
##  $ Race        : chr [1:800] "Ethnic" "Ethnic" "Ethnic" "Ethnic" ...
##  $ Region      : chr [1:800] "North" "Center" "North" "Center" ...
##  $ EMO1        : num [1:800] 1 4 3 4 2 3 2 3 4 2 ...
##  $ EMO2        : num [1:800] 2 4 1 2 1 1 4 1 2 2 ...
##  $ EMO3        : num [1:800] 2 1 3 3 2 4 2 4 3 3 ...
##  $ EMO4        : num [1:800] 1 2 3 1 4 2 3 2 1 1 ...
##  $ EMO5        : num [1:800] 4 1 2 2 2 2 1 1 2 2 ...
##  $ GOAL1       : chr [1:800] "Strongly agree" "Undecided" "Agree" "Agree" ...
##  $ GOAL2       : chr [1:800] "Agree" "Disagree" "Disagree" "Undecided" ...
##  $ GOAL3       : chr [1:800] "Strongly agree" "Disagree" "Agree" "Strongly agree" ...
##  $ Pre_STAT1   : num [1:800] 2 1 5 4 1 4 4 2 2 2 ...
##  $ Pre_STAT2   : num [1:800] 4 1 1 3 4 1 2 3 3 5 ...
##  $ Pre_STAT3   : num [1:800] 2 1 3 1 1 5 4 3 3 2 ...
##  $ Pre_STAT4   : num [1:800] 5 1 1 2 2 3 2 3 2 4 ...
##  $ Post_STAT1  : num [1:800] 4 5 5 3 5 2 3 3 2 5 ...
##  $ Post_STAT2  : num [1:800] 5 1 2 2 3 3 2 3 2 3 ...
##  $ Post_STAT3  : num [1:800] 2 3 3 4 3 5 5 4 5 4 ...
##  $ Post_STAT4  : num [1:800] 2 3 3 5 4 4 3 5 5 1 ...
##  $ Pre_IDARE1  : chr [1:800] "Quite a bit" "Quite a bit" "Quite a bit" "Little" ...
##  $ Pre_IDARE2  : chr [1:800] "Little" "Little" "Little" "Nothing" ...
##  $ Pre_IDARE3  : chr [1:800] "Quite a bit" "A lot" "Quite a bit" "Quite a bit" ...
##  $ Pre_IDARE4  : chr [1:800] "Quite a bit" "Nothing" "Quite a bit" "Quite a bit" ...
##  $ Pre_IDARE5  : chr [1:800] "Little" "Quite a bit" "Little" "Nothing" ...
##  $ Post_IDARE1 : chr [1:800] "A lot" "A little" "Nothing" "Quite a bit" ...
##  $ Post_IDARE2 : chr [1:800] "A lot" "Nothing" "Quite a bit" "A little" ...
##  $ Post_IDARE3 : chr [1:800] "A little" "Quite a bit" "Nothing" "A lot" ...
##  $ Post_IDARE4 : chr [1:800] "Quite a bit" "A lot" "Nothing" "Quite a bit" ...
##  $ Post_IDARE5 : chr [1:800] "A lot" "Quite a bit" "Nothing" "A lot" ...
##  $ PSICO1      : chr [1:800] "Frequently" "Frequently" "Sometimes" "Almost always" ...
##  $ PSICO2      : chr [1:800] "Almost always" "Sometimes" "Sometimes" "Frequently" ...
##  $ PSICO3      : chr [1:800] "Frequently" "Sometimes" "Sometimes" "Frequently" ...
##  $ PSICO4      : chr [1:800] "Almost always" "Frequently" "Frequently" "Almost never" ...
##  $ PSICO5      : chr [1:800] "Almost always" "Frequently" "Sometimes" "Sometimes" ...

Interpretación: La información adquirida nos permite clasificar las variables en numéricas (num) o de carácter (chr).

Explorar los nombres de las variables

Esta función nos señala las variables que se consideraron para llevar a cabo el análisis.

names(datosCompleto)    #A) Muestra los nombres de las columnas (variables).
##  [1] "Observation"  "ID"           "Gender"       "Like"         "Age"         
##  [6] "Smoke"        "Height"       "Weight"       "BMI"          "School"      
## [11] "SES"          "Enrollment"   "Score"        "MotherHeight" "MotherAge"   
## [16] "MotherCHD"    "FatherHeight" "FatherAge"    "FatherCHD"    "Status"      
## [21] "SemAcum"      "Exam1"        "Exam2"        "Exam3"        "Exam4"       
## [26] "ExamAcum"     "Definitive"   "Expense"      "Income"       "Gas"         
## [31] "Course"       "Law"          "Economic"     "Race"         "Region"      
## [36] "EMO1"         "EMO2"         "EMO3"         "EMO4"         "EMO5"        
## [41] "GOAL1"        "GOAL2"        "GOAL3"        "Pre_STAT1"    "Pre_STAT2"   
## [46] "Pre_STAT3"    "Pre_STAT4"    "Post_STAT1"   "Post_STAT2"   "Post_STAT3"  
## [51] "Post_STAT4"   "Pre_IDARE1"   "Pre_IDARE2"   "Pre_IDARE3"   "Pre_IDARE4"  
## [56] "Pre_IDARE5"   "Post_IDARE1"  "Post_IDARE2"  "Post_IDARE3"  "Post_IDARE4" 
## [61] "Post_IDARE5"  "PSICO1"       "PSICO2"       "PSICO3"       "PSICO4"      
## [66] "PSICO5"

Interpretación: La ayuda de este chunck nos ayuda a tener una visión más concisa de las variables que se han considerado para realizar la investigación.

Explorar tamaños

Aquí examinamos algunas propiedades de las variables y los objetos.

length(datosCompleto)   #A) Revisando número de variables del objeto
## [1] 66
dim(datosCompleto)      #B Muestra las dimensiones del objeto.
## [1] 800  66
ncol(datosCompleto)     #C) Muestra el número de columnas del objeto.
## [1] 66
nrow(datosCompleto)     #D) Muestra el número de filas del objeto.
## [1] 800

Interpretación: En este lugar se encuentra que existen 66 variables en total, y que cada objeto también tiene 66 variables.

Trabajo con muestras: función corchete

Llevamos a cabo datosCompleto[i,j], en el que i y j representan las filas y columnas que se eliminarán o emplearán, respectivamente.

Muestra1 <- datosCompleto[1:10,2:7]       # A) Un nuevo data frame 
Muestra1
## # A tibble: 10 × 6
##    ID               Gender Like      Age Smoke Height
##    <chr>            <chr>  <chr>   <dbl> <chr>  <dbl>
##  1 SB11201910010435 Female TV       21.4 No      1.58
##  2 SB11201910004475 Male   Network  21.1 Yes     1.6 
##  3 SB11201910011427 Male   Network  20.9 Yes     1.5 
##  4 SB11201910041975 Male   TV       18.4 Yes     1.53
##  5 SB11201910013623 Female TV       16.6 Yes     1.78
##  6 SB11201910038122 Female Network  16.0 No      1.65
##  7 SB11201910037905 Female TV       19.3 Yes     1.73
##  8 SB11201910038140 Female TV       18.6 Yes     1.53
##  9 SB11201910038005 Female TV       17.0 Yes     1.64
## 10 SB11201910037919 Male   TV       19.7 Yes     1.52

Interpretación: Se eligió una muestra de 10 objetos a partir de la población y solo se seleccionaron 6 variables entre las 66 iniciales.

Tipos de variables

Nominales o caracter

R lee muchas de las variables como si fueran de tipo carácter (debido al símbolo chr); sin embargo, algunas presentan definiciones incorrectas y, por lo tanto, tenemos que redefinirlas con la función que sigue:

Codigo <- datosCompleto$ID     #B) Si es tipo caracter, es correcto)
Edad <- datosCompleto$Age      #C) Si es tipo caracter, es incorrecto 
Sexo <- datosCompleto$Gender   #D) Si es tipo caracter, es incorrecto

Interpretación: Podemos observar que se lleva a cabo una verificación en la categorización de las variables, ya que pueden existir equivocaciones.

Numéricas

P1   <- datosCompleto$Exam1 #E) Numérica
P2   <- datosCompleto$Exam2 #F) Numérica
Edad <- datosCompleto$Age   #G) Numérica

Interpretación: Aquí se puede observar que existe una categorización de las variables cuantitativas (numéricas).

Categóricas o factor

Sexo <- as.factor(Sexo)  #H) Convirtiendo a factor
class(Sexo)              #I) Sale: "factor"
## [1] "factor"
str(Sexo)                #J) Sale: Factor w/ 2 levels "Female","Masculino": 1 2 2 2 1 1 1 1 1 2 ...
##  Factor w/ 2 levels "Female","Male": 1 2 2 2 1 1 1 1 1 2 ...
levels(Sexo)             #K) Sale: "Female"  "Masculino"
## [1] "Female" "Male"

Interpretación: Aquí se observa que a la variable “género” le corresponde un número: el 1 para “FEMALE” y el 2 para “MALE”.