A continuación un análisis de diferentes bases de datos:

#Base de dates

mtcars

##Estadisticas descriptivas

summary(mtcars)

###Graficos de relaciones entre variables:

plot(mtcars$mpg,mtcars$wt)

#Exploracion Analítica de datos EDA ## carga de datos y revisión inicial

library(readxl)
library(dplyr)
library(plotly)
sleep_data <- read_excel("Sleep_health_and_lifestyle_dataset.xlsx")
head(sleep_data)
str(sleep_data)
## tibble [374 × 13] (S3: tbl_df/tbl/data.frame)
##  $ Person ID              : num [1:374] 1 2 3 4 5 6 7 8 9 10 ...
##  $ Gender                 : chr [1:374] "Male" "Male" "Male" "Male" ...
##  $ Age                    : num [1:374] 27 28 28 28 28 28 29 29 29 29 ...
##  $ Occupation             : chr [1:374] "Software Engineer" "Doctor" "Doctor" "Sales Representative" ...
##  $ Sleep Duration         : chr [1:374] "6.1" "6.2" "6.2" "5.9" ...
##  $ Quality of Sleep       : num [1:374] 6 6 6 4 4 4 6 7 7 7 ...
##  $ Physical Activity Level: num [1:374] 42 60 60 30 30 30 40 75 75 75 ...
##  $ Stress Level           : num [1:374] 6 8 8 8 8 8 7 6 6 6 ...
##  $ BMI Category           : chr [1:374] "Overweight" "Normal" "Normal" "Obese" ...
##  $ Blood Pressure         : chr [1:374] "126/83" "125/80" "125/80" "140/90" ...
##  $ Heart Rate             : num [1:374] 77 75 75 85 85 85 82 70 70 70 ...
##  $ Daily Steps            : num [1:374] 4200 10000 10000 3000 3000 3000 3500 8000 8000 8000 ...
##  $ Sleep Disorder         : num [1:374] 0 0 0 1 1 1 1 0 0 0 ...
colnames(sleep_data)
##  [1] "Person ID"               "Gender"                 
##  [3] "Age"                     "Occupation"             
##  [5] "Sleep Duration"          "Quality of Sleep"       
##  [7] "Physical Activity Level" "Stress Level"           
##  [9] "BMI Category"            "Blood Pressure"         
## [11] "Heart Rate"              "Daily Steps"            
## [13] "Sleep Disorder"

Exploración Estructural de Variables

# Convertir variables categóricas a factor
sleep_data$Gender <- as.factor(sleep_data$Gender)
sleep_data$Occupation <- as.factor(sleep_data$Occupation)
sleep_data$`BMI Category` <- as.factor(sleep_data$`BMI Category`)
sleep_data$`Sleep Disorder` <- as.factor(sleep_data$`Sleep Disorder`)

# Verificar estructura nuevamente
str(sleep_data)
## tibble [374 × 13] (S3: tbl_df/tbl/data.frame)
##  $ Person ID              : num [1:374] 1 2 3 4 5 6 7 8 9 10 ...
##  $ Gender                 : Factor w/ 2 levels "Female","Male": 2 2 2 2 2 2 2 2 2 2 ...
##  $ Age                    : num [1:374] 27 28 28 28 28 28 29 29 29 29 ...
##  $ Occupation             : Factor w/ 11 levels "Accountant","Doctor",..: 10 2 2 7 7 10 11 2 2 2 ...
##  $ Sleep Duration         : chr [1:374] "6.1" "6.2" "6.2" "5.9" ...
##  $ Quality of Sleep       : num [1:374] 6 6 6 4 4 4 6 7 7 7 ...
##  $ Physical Activity Level: num [1:374] 42 60 60 30 30 30 40 75 75 75 ...
##  $ Stress Level           : num [1:374] 6 8 8 8 8 8 7 6 6 6 ...
##  $ BMI Category           : Factor w/ 4 levels "Normal","Normal Weight",..: 4 1 1 3 3 3 3 1 1 1 ...
##  $ Blood Pressure         : chr [1:374] "126/83" "125/80" "125/80" "140/90" ...
##  $ Heart Rate             : num [1:374] 77 75 75 85 85 85 82 70 70 70 ...
##  $ Daily Steps            : num [1:374] 4200 10000 10000 3000 3000 3000 3500 8000 8000 8000 ...
##  $ Sleep Disorder         : Factor w/ 2 levels "0","1": 1 1 1 2 2 2 2 1 1 1 ...

Las variables tipo factor representan categorías (por ejemplo: sexo, ocupación).

Las variables numéricas representan cantidades medibles (edad, pasos diarios, presión arterial).

Esta clasificación es importante porque determina el tipo de gráfico adecuado.

Visualización básica interactiva

Gráfico 1: Duración del sueño por género

grafico1 <- plot_ly(
  data = sleep_data,
  x = ~Gender,
  y = ~`Sleep Duration`,
  type = "box",
  color = ~Gender
)

grafico1

Permite observar cómo se distribuye la duración del sueño según género.

El gráfico tipo boxplot (diagrama de caja) permite visualizar dispersión y posibles valores atípicos (valores extremos).

Gráfico 2: Nivel de actividad física vs calidad del sueño

grafico2 <- plot_ly(
  data = sleep_data,
  x = ~`Physical Activity Level`,
  y = ~`Quality of Sleep`,
  type = "scatter",
  mode = "markers",
  color = ~Gender
)

grafico2

Es un gráfico de dispersión (muestra relación entre dos variables numéricas).

Permite explorar si mayor actividad física parece asociarse con mejor calidad del sueño.

No se establece causalidad (no significa que una variable cause la otra).

Gráfico 3: Distribución de pasos diarios por categoría de IMC

grafico3 <- plot_ly(
  data = sleep_data,
  x = ~`BMI Category`,
  y = ~`Daily Steps`,
  type = "box",
  color = ~`BMI Category`
)

grafico3

Permite observar diferencias en actividad diaria según categoría de IMC.

Desde la gestión sanitaria, esto puede sugerir patrones conductuales asociados al riesgo cardiovascular.

Gráficos de relaciones entre variables

plot(mtcars$mpg, mtcars$wt)