A continuación un análisis de diferentes bases de datos:
#Base de dates
##Estadisticas descriptivas
summary(mtcars)
###Graficos de relaciones entre variables:
plot(mtcars$mpg,mtcars$wt)
#Exploracion Analítica de datos EDA ## carga de datos y revisión inicial
## tibble [374 × 13] (S3: tbl_df/tbl/data.frame)
## $ Person ID : num [1:374] 1 2 3 4 5 6 7 8 9 10 ...
## $ Gender : chr [1:374] "Male" "Male" "Male" "Male" ...
## $ Age : num [1:374] 27 28 28 28 28 28 29 29 29 29 ...
## $ Occupation : chr [1:374] "Software Engineer" "Doctor" "Doctor" "Sales Representative" ...
## $ Sleep Duration : chr [1:374] "6.1" "6.2" "6.2" "5.9" ...
## $ Quality of Sleep : num [1:374] 6 6 6 4 4 4 6 7 7 7 ...
## $ Physical Activity Level: num [1:374] 42 60 60 30 30 30 40 75 75 75 ...
## $ Stress Level : num [1:374] 6 8 8 8 8 8 7 6 6 6 ...
## $ BMI Category : chr [1:374] "Overweight" "Normal" "Normal" "Obese" ...
## $ Blood Pressure : chr [1:374] "126/83" "125/80" "125/80" "140/90" ...
## $ Heart Rate : num [1:374] 77 75 75 85 85 85 82 70 70 70 ...
## $ Daily Steps : num [1:374] 4200 10000 10000 3000 3000 3000 3500 8000 8000 8000 ...
## $ Sleep Disorder : num [1:374] 0 0 0 1 1 1 1 0 0 0 ...
## [1] "Person ID" "Gender"
## [3] "Age" "Occupation"
## [5] "Sleep Duration" "Quality of Sleep"
## [7] "Physical Activity Level" "Stress Level"
## [9] "BMI Category" "Blood Pressure"
## [11] "Heart Rate" "Daily Steps"
## [13] "Sleep Disorder"
# Convertir variables categóricas a factor
sleep_data$Gender <- as.factor(sleep_data$Gender)
sleep_data$Occupation <- as.factor(sleep_data$Occupation)
sleep_data$`BMI Category` <- as.factor(sleep_data$`BMI Category`)
sleep_data$`Sleep Disorder` <- as.factor(sleep_data$`Sleep Disorder`)
# Verificar estructura nuevamente
str(sleep_data)## tibble [374 × 13] (S3: tbl_df/tbl/data.frame)
## $ Person ID : num [1:374] 1 2 3 4 5 6 7 8 9 10 ...
## $ Gender : Factor w/ 2 levels "Female","Male": 2 2 2 2 2 2 2 2 2 2 ...
## $ Age : num [1:374] 27 28 28 28 28 28 29 29 29 29 ...
## $ Occupation : Factor w/ 11 levels "Accountant","Doctor",..: 10 2 2 7 7 10 11 2 2 2 ...
## $ Sleep Duration : chr [1:374] "6.1" "6.2" "6.2" "5.9" ...
## $ Quality of Sleep : num [1:374] 6 6 6 4 4 4 6 7 7 7 ...
## $ Physical Activity Level: num [1:374] 42 60 60 30 30 30 40 75 75 75 ...
## $ Stress Level : num [1:374] 6 8 8 8 8 8 7 6 6 6 ...
## $ BMI Category : Factor w/ 4 levels "Normal","Normal Weight",..: 4 1 1 3 3 3 3 1 1 1 ...
## $ Blood Pressure : chr [1:374] "126/83" "125/80" "125/80" "140/90" ...
## $ Heart Rate : num [1:374] 77 75 75 85 85 85 82 70 70 70 ...
## $ Daily Steps : num [1:374] 4200 10000 10000 3000 3000 3000 3500 8000 8000 8000 ...
## $ Sleep Disorder : Factor w/ 2 levels "0","1": 1 1 1 2 2 2 2 1 1 1 ...
Las variables tipo factor representan categorías (por ejemplo: sexo, ocupación).
Las variables numéricas representan cantidades medibles (edad, pasos diarios, presión arterial).
Esta clasificación es importante porque determina el tipo de gráfico adecuado.
grafico1 <- plot_ly(
data = sleep_data,
x = ~Gender,
y = ~`Sleep Duration`,
type = "box",
color = ~Gender
)
grafico1Permite observar cómo se distribuye la duración del sueño según género.
El gráfico tipo boxplot (diagrama de caja) permite visualizar dispersión y posibles valores atípicos (valores extremos).
grafico2 <- plot_ly(
data = sleep_data,
x = ~`Physical Activity Level`,
y = ~`Quality of Sleep`,
type = "scatter",
mode = "markers",
color = ~Gender
)
grafico2Es un gráfico de dispersión (muestra relación entre dos variables numéricas).
Permite explorar si mayor actividad física parece asociarse con mejor calidad del sueño.
No se establece causalidad (no significa que una variable cause la otra).
grafico3 <- plot_ly(
data = sleep_data,
x = ~`BMI Category`,
y = ~`Daily Steps`,
type = "box",
color = ~`BMI Category`
)
grafico3Permite observar diferencias en actividad diaria según categoría de IMC.
Desde la gestión sanitaria, esto puede sugerir patrones conductuales asociados al riesgo cardiovascular.