library(lsm) # Para descargar una base de datos
library(dplyr)
##
## Adjuntando el paquete: 'dplyr'
## The following objects are masked from 'package:stats':
##
## filter, lag
## The following objects are masked from 'package:base':
##
## intersect, setdiff, setequal, union
library(moments) # Para hallar las medidas de forma
library(e1071)
##
## Adjuntando el paquete: 'e1071'
## The following objects are masked from 'package:moments':
##
## kurtosis, moment, skewness
library(ggplot2)
##
## Adjuntando el paquete: 'ggplot2'
## The following object is masked from 'package:e1071':
##
## element
Interpretación: Aquí usamos el chunk para descargar los paquetes que utilizaremos a lo largo del trabajo
datosCompleto <- lsm::survey
Interpretación: Aquí seleccionamos un conjunto de datos (survey) del paquete (lsm).
Para revisar el data frame se pueden usar las funciones HEAD o TAIL. En mi caso funcionó usar TAIL.
tail(datosCompleto) #C) Por defecto, solo las últimas 6 observaciones
## # A tibble: 6 × 66
## Observation ID Gender Like Age Smoke Height Weight BMI School SES
## <dbl> <chr> <chr> <chr> <dbl> <chr> <dbl> <dbl> <dbl> <chr> <chr>
## 1 795 AC31201… Female <NA> NA <NA> 1.64 NA NA <NA> Low
## 2 796 AC31201… Female TV 13.5 <NA> 1.71 78 26.7 Public <NA>
## 3 797 AC31201… <NA> Netw… 15.8 No 1.68 53 18.8 Priva… Medi…
## 4 798 AC31201… Male <NA> 15.7 No NA 83 NA <NA> Low
## 5 799 AC31201… Female TV NA No 1.76 73 23.6 Priva… Low
## 6 800 AC31201… Male TV 16.6 No 1.62 70 26.7 Priva… <NA>
## # ℹ 55 more variables: Enrollment <chr>, Score <dbl>, MotherHeight <chr>,
## # MotherAge <dbl>, MotherCHD <dbl>, FatherHeight <chr>, FatherAge <dbl>,
## # FatherCHD <dbl>, Status <chr>, SemAcum <dbl>, Exam1 <dbl>, Exam2 <dbl>,
## # Exam3 <dbl>, Exam4 <dbl>, ExamAcum <dbl>, Definitive <dbl>, Expense <dbl>,
## # Income <dbl>, Gas <dbl>, Course <chr>, Law <chr>, Economic <chr>,
## # Race <chr>, Region <chr>, EMO1 <dbl>, EMO2 <dbl>, EMO3 <dbl>, EMO4 <dbl>,
## # EMO5 <dbl>, GOAL1 <chr>, GOAL2 <chr>, GOAL3 <chr>, Pre_STAT1 <dbl>, …
tail(datosCompleto, 2) #D) Solo las últimas 2 observaciones
## # A tibble: 2 × 66
## Observation ID Gender Like Age Smoke Height Weight BMI School SES
## <dbl> <chr> <chr> <chr> <dbl> <chr> <dbl> <dbl> <dbl> <chr> <chr>
## 1 799 AC31201… Female TV NA No 1.76 73 23.6 Priva… Low
## 2 800 AC31201… Male TV 16.6 No 1.62 70 26.7 Priva… <NA>
## # ℹ 55 more variables: Enrollment <chr>, Score <dbl>, MotherHeight <chr>,
## # MotherAge <dbl>, MotherCHD <dbl>, FatherHeight <chr>, FatherAge <dbl>,
## # FatherCHD <dbl>, Status <chr>, SemAcum <dbl>, Exam1 <dbl>, Exam2 <dbl>,
## # Exam3 <dbl>, Exam4 <dbl>, ExamAcum <dbl>, Definitive <dbl>, Expense <dbl>,
## # Income <dbl>, Gas <dbl>, Course <chr>, Law <chr>, Economic <chr>,
## # Race <chr>, Region <chr>, EMO1 <dbl>, EMO2 <dbl>, EMO3 <dbl>, EMO4 <dbl>,
## # EMO5 <dbl>, GOAL1 <chr>, GOAL2 <chr>, GOAL3 <chr>, Pre_STAT1 <dbl>, …
Interpretación: A partir de la siguiente tabla podemos ver que no todas las observaciones registran datos en todas las variables tomadas en cuenta para el estudio.
Con esta función se puede observar la estructura. Esta proporciona información sobre el tipo de objeto, el número de filas y columnas, al igual que información adicional como los nombres de las variables y su tipo seguido de algunas de las observaciones iniciales de cada una de ellas.
str(datosCompleto) #A) Estructura de los datos
## tibble [800 × 66] (S3: tbl_df/tbl/data.frame)
## $ Observation : num [1:800] 1 2 3 4 5 6 7 8 9 10 ...
## $ ID : chr [1:800] "SB11201910010435" "SB11201910004475" "SB11201910011427" "SB11201910041975" ...
## $ Gender : chr [1:800] "Female" "Male" "Male" "Male" ...
## $ Like : chr [1:800] "TV" "Network" "Network" "TV" ...
## $ Age : num [1:800] 21.4 21.1 20.9 18.4 16.6 ...
## $ Smoke : chr [1:800] "No" "Yes" "Yes" "Yes" ...
## $ Height : num [1:800] 1.58 1.6 1.5 1.53 1.78 1.65 1.73 1.53 1.64 1.52 ...
## $ Weight : num [1:800] 75 80 64 49 82 80 90 55 50 78 ...
## $ BMI : num [1:800] 30 31.2 28.4 20.9 25.9 ...
## $ School : chr [1:800] "Private" "Public" "Private" "Public" ...
## $ SES : chr [1:800] "Medium" "High" "High" "Low" ...
## $ Enrollment : chr [1:800] "Credit" "Scholarship" "Scholarship" "Credit" ...
## $ Score : num [1:800] 81 78 77 70 68 65 54 50 36 35 ...
## $ MotherHeight: chr [1:800] "Short_M" "Normal_M" "Normal_M" "Tall_M" ...
## $ MotherAge : num [1:800] 41 45 45 45 46 46 47 48 48 48 ...
## $ MotherCHD : num [1:800] 0 0 0 0 1 0 0 0 0 1 ...
## $ FatherHeight: chr [1:800] "Normal_F" "Short_F" "Tall_F" "Short_F" ...
## $ FatherAge : num [1:800] 40 43 44 45 45 46 46 48 48 49 ...
## $ FatherCHD : num [1:800] 1 1 1 2 1 1 1 1 1 1 ...
## $ Status : chr [1:800] "Distinguished" "Distinguished" "Distinguished" "Regular" ...
## $ SemAcum : num [1:800] 4.25 2.8 4.15 3.2 3.45 2.75 2.7 4.35 4.3 2.8 ...
## $ Exam1 : num [1:800] 1.5 2.3 3.4 2.5 3.1 3.8 5 4 2.5 2.4 ...
## $ Exam2 : num [1:800] 5 4.9 3.6 4.2 3.5 4.4 3 2.3 3.3 2.6 ...
## $ Exam3 : num [1:800] 5 3.7 2 5 5 4.2 3.5 4.6 3.8 4.3 ...
## $ Exam4 : num [1:800] 4.5 3.3 1.9 2.5 3 5 3.6 4.3 1.9 5 ...
## $ ExamAcum : num [1:800] 16 14.2 10.9 14.2 14.6 17.4 15.1 15.2 11.5 14.3 ...
## $ Definitive : num [1:800] 4 3.55 2.73 3.55 3.65 ...
## $ Expense : num [1:800] 48.9 72.1 85.2 56.6 64.6 63 40.8 65.4 37.3 63 ...
## $ Income : num [1:800] 1.61 2.07 2.84 1.55 2.32 2.1 1.69 2.18 1.71 2.1 ...
## $ Gas : num [1:800] 27.4 24.2 22.3 23.1 27.3 ...
## $ Course : chr [1:800] "Face-to-Face" "Virtual" "Face-to-Face" "Virtual" ...
## $ Law : chr [1:800] "Agree" "Agree" "Agree" "Agree" ...
## $ Economic : chr [1:800] "Regular" "Good" "Regular" "Bad" ...
## $ Race : chr [1:800] "Ethnic" "Ethnic" "Ethnic" "Ethnic" ...
## $ Region : chr [1:800] "North" "Center" "North" "Center" ...
## $ EMO1 : num [1:800] 1 4 3 4 2 3 2 3 4 2 ...
## $ EMO2 : num [1:800] 2 4 1 2 1 1 4 1 2 2 ...
## $ EMO3 : num [1:800] 2 1 3 3 2 4 2 4 3 3 ...
## $ EMO4 : num [1:800] 1 2 3 1 4 2 3 2 1 1 ...
## $ EMO5 : num [1:800] 4 1 2 2 2 2 1 1 2 2 ...
## $ GOAL1 : chr [1:800] "Strongly agree" "Undecided" "Agree" "Agree" ...
## $ GOAL2 : chr [1:800] "Agree" "Disagree" "Disagree" "Undecided" ...
## $ GOAL3 : chr [1:800] "Strongly agree" "Disagree" "Agree" "Strongly agree" ...
## $ Pre_STAT1 : num [1:800] 2 1 5 4 1 4 4 2 2 2 ...
## $ Pre_STAT2 : num [1:800] 4 1 1 3 4 1 2 3 3 5 ...
## $ Pre_STAT3 : num [1:800] 2 1 3 1 1 5 4 3 3 2 ...
## $ Pre_STAT4 : num [1:800] 5 1 1 2 2 3 2 3 2 4 ...
## $ Post_STAT1 : num [1:800] 4 5 5 3 5 2 3 3 2 5 ...
## $ Post_STAT2 : num [1:800] 5 1 2 2 3 3 2 3 2 3 ...
## $ Post_STAT3 : num [1:800] 2 3 3 4 3 5 5 4 5 4 ...
## $ Post_STAT4 : num [1:800] 2 3 3 5 4 4 3 5 5 1 ...
## $ Pre_IDARE1 : chr [1:800] "Quite a bit" "Quite a bit" "Quite a bit" "Little" ...
## $ Pre_IDARE2 : chr [1:800] "Little" "Little" "Little" "Nothing" ...
## $ Pre_IDARE3 : chr [1:800] "Quite a bit" "A lot" "Quite a bit" "Quite a bit" ...
## $ Pre_IDARE4 : chr [1:800] "Quite a bit" "Nothing" "Quite a bit" "Quite a bit" ...
## $ Pre_IDARE5 : chr [1:800] "Little" "Quite a bit" "Little" "Nothing" ...
## $ Post_IDARE1 : chr [1:800] "A lot" "A little" "Nothing" "Quite a bit" ...
## $ Post_IDARE2 : chr [1:800] "A lot" "Nothing" "Quite a bit" "A little" ...
## $ Post_IDARE3 : chr [1:800] "A little" "Quite a bit" "Nothing" "A lot" ...
## $ Post_IDARE4 : chr [1:800] "Quite a bit" "A lot" "Nothing" "Quite a bit" ...
## $ Post_IDARE5 : chr [1:800] "A lot" "Quite a bit" "Nothing" "A lot" ...
## $ PSICO1 : chr [1:800] "Frequently" "Frequently" "Sometimes" "Almost always" ...
## $ PSICO2 : chr [1:800] "Almost always" "Sometimes" "Sometimes" "Frequently" ...
## $ PSICO3 : chr [1:800] "Frequently" "Sometimes" "Sometimes" "Frequently" ...
## $ PSICO4 : chr [1:800] "Almost always" "Frequently" "Frequently" "Almost never" ...
## $ PSICO5 : chr [1:800] "Almost always" "Frequently" "Sometimes" "Sometimes" ...
Interpretación: Entre la información obtenida podemos ver que clasifica las variables si son de caracter (chr) o si son numéricas (num).
Esta función nos indica el nombre de las variables que se tomaron en cuenta para realizar el estudio.
names(datosCompleto) #A) Muestra los nombres de las columnas (variables).
## [1] "Observation" "ID" "Gender" "Like" "Age"
## [6] "Smoke" "Height" "Weight" "BMI" "School"
## [11] "SES" "Enrollment" "Score" "MotherHeight" "MotherAge"
## [16] "MotherCHD" "FatherHeight" "FatherAge" "FatherCHD" "Status"
## [21] "SemAcum" "Exam1" "Exam2" "Exam3" "Exam4"
## [26] "ExamAcum" "Definitive" "Expense" "Income" "Gas"
## [31] "Course" "Law" "Economic" "Race" "Region"
## [36] "EMO1" "EMO2" "EMO3" "EMO4" "EMO5"
## [41] "GOAL1" "GOAL2" "GOAL3" "Pre_STAT1" "Pre_STAT2"
## [46] "Pre_STAT3" "Pre_STAT4" "Post_STAT1" "Post_STAT2" "Post_STAT3"
## [51] "Post_STAT4" "Pre_IDARE1" "Pre_IDARE2" "Pre_IDARE3" "Pre_IDARE4"
## [56] "Pre_IDARE5" "Post_IDARE1" "Post_IDARE2" "Post_IDARE3" "Post_IDARE4"
## [61] "Post_IDARE5" "PSICO1" "PSICO2" "PSICO3" "PSICO4"
## [66] "PSICO5"
Interpretación: Con ayuda de este chunk podemos tener una vista más compacta de las variables que han sido tomadas en cuenta para realizar el estudio.
Aquí exploramos ciertas características de los objetos y las variables.
length(datosCompleto) #A) Revisando número de variables del objeto
## [1] 66
dim(datosCompleto) #B Muestra las dimensiones del objeto.
## [1] 800 66
ncol(datosCompleto) #C) Muestra el número de columnas del objeto.
## [1] 66
nrow(datosCompleto) #D) Muestra el número de filas del objeto.
## [1] 800
Interpretación: Aquí obtenemos que hay un total de 66 variables, además de que por cada objeto hay un total de 66 variables.
Ejecutamos datosCompleto[i,j], donde i y j son las filas y columnas que se va a utilizar o quitar, respectivamente.
Muestra1 <- datosCompleto[1:10,2:7] # A) Un nuevo data frame
Muestra1
## # A tibble: 10 × 6
## ID Gender Like Age Smoke Height
## <chr> <chr> <chr> <dbl> <chr> <dbl>
## 1 SB11201910010435 Female TV 21.4 No 1.58
## 2 SB11201910004475 Male Network 21.1 Yes 1.6
## 3 SB11201910011427 Male Network 20.9 Yes 1.5
## 4 SB11201910041975 Male TV 18.4 Yes 1.53
## 5 SB11201910013623 Female TV 16.6 Yes 1.78
## 6 SB11201910038122 Female Network 16.0 No 1.65
## 7 SB11201910037905 Female TV 19.3 Yes 1.73
## 8 SB11201910038140 Female TV 18.6 Yes 1.53
## 9 SB11201910038005 Female TV 17.0 Yes 1.64
## 10 SB11201910037919 Male TV 19.7 Yes 1.52
Interpretación: A partir de la población se escogió una muestra de 10 objetos y se escogieron únicamente 6 variables de las 66 iniciales.
R lee muchas de las variables como de tipo caracter (por el símbolo chr), pero algunas están mal definidas y, por esta razón, debemos redefinirlas con la siguiente función:
Codigo <- datosCompleto$ID #B) Si es tipo caracter, es correcto)
Edad <- datosCompleto$Age #C) Si es tipo caracter, es incorrecto
Sexo <- datosCompleto$Gender #D) Si es tipo caracter, es incorrecto
Interpretación: Aquí podemos ver que hay una verificación en la clasificación de las variables puesto que puede haber errores.
P1 <- datosCompleto$Exam1 #E) Numérica
P2 <- datosCompleto$Exam2 #F) Numérica
Edad <- datosCompleto$Age #G) Numérica
Interpretación: Aquí podemos ver que hay una clasificación de las variables de tipo cuantitativas (numérica).
Sexo <- as.factor(Sexo) #H) Convirtiendo a factor
class(Sexo) #I) Sale: "factor"
## [1] "factor"
str(Sexo) #J) Sale: Factor w/ 2 levels "Female","Masculino": 1 2 2 2 1 1 1 1 1 2 ...
## Factor w/ 2 levels "Female","Male": 1 2 2 2 1 1 1 1 1 2 ...
levels(Sexo) #K) Sale: "Female" "Masculino"
## [1] "Female" "Male"
Interpretación: Aquí podemos ver que a la variable “género” se le asigna un número: 1 para “FEMALE” y 2 para “MALE”.
Se quiere construir una tabla de frecuencias para analizar la distribución de la variable categórica Sexo. La tabla de frecuencias de interés construída con table es la siguiente
Muestra <- datosCompleto[1:100,]
#A) Definiendo y convirtiendo en factor
Sexo <- as.factor(Muestra$Gender)
#B) Calcular tabla de frecuencias
Tabla1 <- table(Sexo)
Tabla1
## Sexo
## Female Male
## 49 51
Interpretación: en esta muestra de 100 personas hay 49 mujeres y 51 hombres.
# Porcentaje de mujeres
(49/100)*100
## [1] 49
Interpretación: Para hacer un comentario debo abrir un chunk, poner numeral dentro del chunk (el programa identifica esto como comentario) y procedemos a hacer el cálculo correspondiente. Se corre el chunk y obtenemos el output
#Variable: ¿FUMA? (Hay que arreglarla)
#B) Tabla de frecuencias
Fuma <- Muestra$Smoke
Tabla3 <- table(Sexo,Fuma)
Tabla3
## Fuma
## Sexo No Yes
## Female 21 28
## Male 24 27
Interp: Aquí están los hombres y mujeres que fuman y no fuman
(28/55)*100
## [1] 50.90909
Interretación: Dentro del grupo de fumadores el porcentaje de mujeres es del 50.9%
Muestra <- datosCompleto[1:100,]
Sexo <- as.factor(Muestra$Gender)
Fuma <- as.factor(Muestra$Smoke)
Se hace con la función ggplot de la Muestra (100)
ggplot(Muestra, aes(x = Sexo)) + #1
#geom_bar() + #2
geom_bar(width=0.5, colour="red", fill="skyblue") + #2
labs(x="Sexo",y= "Frecuencia") + #3
ylim(c(0,60)) + #4
#xlim(c(0,300)) + #4
ggtitle("Diagrama de barras") + #5
# theme_bw() + #6
theme_bw(base_size = 12) + #6
#coord_flip() + #7
geom_text(aes(label=..count..), stat='count', #8
position=position_dodge(0.9),
vjust=-0.5,
size=5.0
) +
facet_wrap(~"Variable Sexo") #9
Interpretación: Aquí podemos observar a través de un gráfico de barras
la frecuencia en las variables de sexo: Female y Male.
ggplot(Muestra, aes(x = Sexo)) + #1
#geom_bar() + #2
geom_bar(width=0.5, colour="purple", fill="green") + #2
labs(x="Sexo",y= "Frecuencia") + #3
ylim(c(0,60)) + #4
#xlim(c(0,300)) + #4
ggtitle("Diagrama de barras") + #5
# theme_bw() + #6
theme_bw(base_size = 12) + #6
#coord_flip() + #7
geom_text(aes(label=..count..), stat='count', #8
position=position_dodge(0.9),
vjust=-0.5,
size=5.0
) +
facet_wrap(~"Variable Sexo") #9
Interpretación: Aquí es lo mismo que en el punto anterior pero jugando un poco más con los colores.
Si se se quiere analizar la distribución de la variable “SEXO” dentro de cada nivel de la variable “FUMA”. En este caso, el diagrama de barras se obtiene así:
ggplot(Muestra, aes(Fuma, fill=Sexo)) + #1
geom_bar()+ #2
labs(x= "Fuma", y="Frecuencias", fill="Sexo") + #3
ylim(c(0,60)) + #4
#xlim(c(0,300)) + #4
ggtitle("Diagrama de barras") + #5
#coord_flip() + #6
#theme_bw() + #7
theme_bw(base_size = 12) #7
Interpretación: Aquí vemos el acumulado de los fumadores y no fumadores según la variable de sexo.
ggplot(Muestra, aes(Fuma, fill=Sexo)) +
geom_bar(position="dodge",colour="yellow") +
labs(x= "Fuma", y="Frecuencias", fill="Sexo") +
ylim(c(0,30)) +
#xlim(c(0,300)) +
ggtitle("Diagrama de barras") +
#theme_bw() +
theme_bw(base_size = 12) +
#coord_flip() +
#guides(fill=FALSE)+ #8
scale_fill_manual(values = c("blue","red")) + #9
geom_text(aes(label=..count..), stat='count', #10
position=position_dodge(0.9),
vjust=-0.5,
size=5.0
)+
facet_wrap(~"Sexo por fumadores y no fumadores") #11
Interpretación: Aquí vemos la cantidad de fumadores y no fumadores según el sexo.
Hacemos el mismo gráfico pero cambiando el orden de las variables.
ggplot(Muestra, aes(Sexo, fill=Fuma)) +
geom_bar(position="dodge",colour="pink") +
labs(x= "Fuma", y="Frecuencias", fill="Sexo") +
ylim(c(0,30)) +
#xlim(c(0,300)) +
ggtitle("Diagrama de barras") +
#theme_bw() +
theme_bw(base_size = 12) +
#coord_flip() +
#guides(fill=FALSE)+ #8
scale_fill_manual(values = c("purple","skyblue")) + #9
geom_text(aes(label=..count..), stat='count', #10
position=position_dodge(0.9),
vjust=-0.5,
size=5.0
)+
facet_wrap(~"Sexo por fumadores y no fumadores") #11
nterpretación: Aquí vemos la frecuencia de hombres y mujeres de acuerdo a la variable “Fuma”.
En este ejemplo vamos a aprender a calcular medidas estadísticas con la ayuda de R. Aquí usamos variables NÚMERICAS.
Muestra2 <- datosCompleto[1:100,]
x <- as.numeric(Muestra2$Exam3) # A) Convirtiendo la variable a numérica
x
## [1] 5.0 3.7 2.0 5.0 5.0 4.2 3.5 4.6 3.8 4.3 3.0 3.8 3.4 3.3 3.5 4.5 3.6 4.0
## [19] 3.4 4.0 4.2 3.5 3.7 4.0 4.0 3.2 2.9 2.9 3.0 3.3 2.8 2.4 3.8 3.3 3.2 2.2
## [37] 2.6 3.2 3.3 1.2 4.2 2.4 5.0 2.8 3.0 3.8 3.2 1.5 2.6 3.8 3.2 3.3 1.4 3.8
## [55] 1.4 3.6 3.6 2.4 2.8 3.1 2.4 1.8 1.6 3.3 4.4 1.0 4.5 2.0 4.2 4.2 3.1 2.3
## [73] 2.6 2.7 2.4 2.2 2.8 2.4 1.9 2.4 1.7 2.9 2.4 2.2 2.8 3.2 3.1 2.7 2.5 3.5
## [91] 3.3 2.1 3.3 2.1 3.7 5.0 3.7 2.0 5.0 5.0
Aquí verificamos la naturaleza de la variable “Exam3”
str (Muestra)
## tibble [100 × 66] (S3: tbl_df/tbl/data.frame)
## $ Observation : num [1:100] 1 2 3 4 5 6 7 8 9 10 ...
## $ ID : chr [1:100] "SB11201910010435" "SB11201910004475" "SB11201910011427" "SB11201910041975" ...
## $ Gender : chr [1:100] "Female" "Male" "Male" "Male" ...
## $ Like : chr [1:100] "TV" "Network" "Network" "TV" ...
## $ Age : num [1:100] 21.4 21.1 20.9 18.4 16.6 ...
## $ Smoke : chr [1:100] "No" "Yes" "Yes" "Yes" ...
## $ Height : num [1:100] 1.58 1.6 1.5 1.53 1.78 1.65 1.73 1.53 1.64 1.52 ...
## $ Weight : num [1:100] 75 80 64 49 82 80 90 55 50 78 ...
## $ BMI : num [1:100] 30 31.2 28.4 20.9 25.9 ...
## $ School : chr [1:100] "Private" "Public" "Private" "Public" ...
## $ SES : chr [1:100] "Medium" "High" "High" "Low" ...
## $ Enrollment : chr [1:100] "Credit" "Scholarship" "Scholarship" "Credit" ...
## $ Score : num [1:100] 81 78 77 70 68 65 54 50 36 35 ...
## $ MotherHeight: chr [1:100] "Short_M" "Normal_M" "Normal_M" "Tall_M" ...
## $ MotherAge : num [1:100] 41 45 45 45 46 46 47 48 48 48 ...
## $ MotherCHD : num [1:100] 0 0 0 0 1 0 0 0 0 1 ...
## $ FatherHeight: chr [1:100] "Normal_F" "Short_F" "Tall_F" "Short_F" ...
## $ FatherAge : num [1:100] 40 43 44 45 45 46 46 48 48 49 ...
## $ FatherCHD : num [1:100] 1 1 1 2 1 1 1 1 1 1 ...
## $ Status : chr [1:100] "Distinguished" "Distinguished" "Distinguished" "Regular" ...
## $ SemAcum : num [1:100] 4.25 2.8 4.15 3.2 3.45 2.75 2.7 4.35 4.3 2.8 ...
## $ Exam1 : num [1:100] 1.5 2.3 3.4 2.5 3.1 3.8 5 4 2.5 2.4 ...
## $ Exam2 : num [1:100] 5 4.9 3.6 4.2 3.5 4.4 3 2.3 3.3 2.6 ...
## $ Exam3 : num [1:100] 5 3.7 2 5 5 4.2 3.5 4.6 3.8 4.3 ...
## $ Exam4 : num [1:100] 4.5 3.3 1.9 2.5 3 5 3.6 4.3 1.9 5 ...
## $ ExamAcum : num [1:100] 16 14.2 10.9 14.2 14.6 17.4 15.1 15.2 11.5 14.3 ...
## $ Definitive : num [1:100] 4 3.55 2.73 3.55 3.65 ...
## $ Expense : num [1:100] 48.9 72.1 85.2 56.6 64.6 63 40.8 65.4 37.3 63 ...
## $ Income : num [1:100] 1.61 2.07 2.84 1.55 2.32 2.1 1.69 2.18 1.71 2.1 ...
## $ Gas : num [1:100] 27.4 24.2 22.3 23.1 27.3 ...
## $ Course : chr [1:100] "Face-to-Face" "Virtual" "Face-to-Face" "Virtual" ...
## $ Law : chr [1:100] "Agree" "Agree" "Agree" "Agree" ...
## $ Economic : chr [1:100] "Regular" "Good" "Regular" "Bad" ...
## $ Race : chr [1:100] "Ethnic" "Ethnic" "Ethnic" "Ethnic" ...
## $ Region : chr [1:100] "North" "Center" "North" "Center" ...
## $ EMO1 : num [1:100] 1 4 3 4 2 3 2 3 4 2 ...
## $ EMO2 : num [1:100] 2 4 1 2 1 1 4 1 2 2 ...
## $ EMO3 : num [1:100] 2 1 3 3 2 4 2 4 3 3 ...
## $ EMO4 : num [1:100] 1 2 3 1 4 2 3 2 1 1 ...
## $ EMO5 : num [1:100] 4 1 2 2 2 2 1 1 2 2 ...
## $ GOAL1 : chr [1:100] "Strongly agree" "Undecided" "Agree" "Agree" ...
## $ GOAL2 : chr [1:100] "Agree" "Disagree" "Disagree" "Undecided" ...
## $ GOAL3 : chr [1:100] "Strongly agree" "Disagree" "Agree" "Strongly agree" ...
## $ Pre_STAT1 : num [1:100] 2 1 5 4 1 4 4 2 2 2 ...
## $ Pre_STAT2 : num [1:100] 4 1 1 3 4 1 2 3 3 5 ...
## $ Pre_STAT3 : num [1:100] 2 1 3 1 1 5 4 3 3 2 ...
## $ Pre_STAT4 : num [1:100] 5 1 1 2 2 3 2 3 2 4 ...
## $ Post_STAT1 : num [1:100] 4 5 5 3 5 2 3 3 2 5 ...
## $ Post_STAT2 : num [1:100] 5 1 2 2 3 3 2 3 2 3 ...
## $ Post_STAT3 : num [1:100] 2 3 3 4 3 5 5 4 5 4 ...
## $ Post_STAT4 : num [1:100] 2 3 3 5 4 4 3 5 5 1 ...
## $ Pre_IDARE1 : chr [1:100] "Quite a bit" "Quite a bit" "Quite a bit" "Little" ...
## $ Pre_IDARE2 : chr [1:100] "Little" "Little" "Little" "Nothing" ...
## $ Pre_IDARE3 : chr [1:100] "Quite a bit" "A lot" "Quite a bit" "Quite a bit" ...
## $ Pre_IDARE4 : chr [1:100] "Quite a bit" "Nothing" "Quite a bit" "Quite a bit" ...
## $ Pre_IDARE5 : chr [1:100] "Little" "Quite a bit" "Little" "Nothing" ...
## $ Post_IDARE1 : chr [1:100] "A lot" "A little" "Nothing" "Quite a bit" ...
## $ Post_IDARE2 : chr [1:100] "A lot" "Nothing" "Quite a bit" "A little" ...
## $ Post_IDARE3 : chr [1:100] "A little" "Quite a bit" "Nothing" "A lot" ...
## $ Post_IDARE4 : chr [1:100] "Quite a bit" "A lot" "Nothing" "Quite a bit" ...
## $ Post_IDARE5 : chr [1:100] "A lot" "Quite a bit" "Nothing" "A lot" ...
## $ PSICO1 : chr [1:100] "Frequently" "Frequently" "Sometimes" "Almost always" ...
## $ PSICO2 : chr [1:100] "Almost always" "Sometimes" "Sometimes" "Frequently" ...
## $ PSICO3 : chr [1:100] "Frequently" "Sometimes" "Sometimes" "Frequently" ...
## $ PSICO4 : chr [1:100] "Almost always" "Frequently" "Frequently" "Almost never" ...
## $ PSICO5 : chr [1:100] "Almost always" "Frequently" "Sometimes" "Sometimes" ...
min(x) #B) Mínimo
## [1] 1
max(x) #C) Máximo
## [1] 5
range(x) #D) Obtenemos (min, max)
## [1] 1 5
length(x) #E) Tamaño
## [1] 100
sum(x) #F) Suma los valores de los datos
## [1] 317.6
mean(x) #G) Media aritmética
## [1] 3.176
median(x) #H) Mediana
## [1] 3.2
var(x) #I) Varianza muestral
## [1] 0.8885091
sqrt(var(x)) #J) Desviación estándar muestral (una forma)
## [1] 0.9426076
sd(x) #K) Desviación estándar muestral (otra forma)
## [1] 0.9426076
skewness(x) #L) Sesgo
## [1] 0.01846742
quantile(x, probs=0.80) #M) 80-ésimo percentil o percentil 85
## 80%
## 4
quantile(x, probs=0.25) #N) Primer cuartil o 25-ésimo percentil
## 25%
## 2.4
quantile(x, probs=0.50) #O) Segundo cuartil o 50-ésimo percentil o mediana
## 50%
## 3.2
quantile(x, probs=0.75) #P) Tercer cuartil o 75-ésimo percentil
## 75%
## 3.8
es la distancia entre dato menor y dato mayor. En este caso el rango es de 4 (5-1).
Es el dato central, de la mitad). En este caso la mediana es de 3.2, es decir que el 50% de las notas del examen 3 es menor/mayor o igual que 3.2
NO SE INTERPETAN
Uno siempre mira si el sesgo es mayor que 0 (+), si es menor que 0 (-) o si es 0 (sesgo=0). Siempre se observa el signo del sesgo. - Sesgo a la derecha –> POSITIVO - Sesgo a la izquierda –> NEGATIVO
Consiste en dividir los datos en porcentajes de 1. Se ordenan los datos. Se interpreta como la mediana PERO va a ser solamente “menor o igual”.
Dividen los datos en porcentajes de 25. Igual que los percentiles.
quantile(x, probs=0.85) #M) 80-ésimo percentil o percentil 85
## 85%
## 4.2
Interpretación: El 85% de las notas es menor o igual que 4.2