Executive summary

This study explores the complex relationships between lifestyle factors—age, occupations, and BMI—and their impact on sleep duration, quality, and the prevalence of sleep disorders. By integrating comprehensive sleep metrics with these lifestyle variables, the research aims to uncover actionable insights for improving sleep health and overall well-being.

Sleep is a fundamental aspect of both physical and mental health, impacting productivity, cognitive function, and the risk of long-term diseases. Chronic sleep deprivation and poor sleep quality have been linked to a range of serious health issues, such as cardiovascular disease, obesity, diabetes, and mental health disorders. With the increasing prevalence of stress and sedentary lifestyles, understanding how daily behaviors influence sleep has become more crucial than ever.

This research seeks to explore three main objectives: first, to understand the impact of lifestyle behaviors on sleep, identifying those that either promote or hinder restorative sleep; second, to examine the long-term health risks associated with sleep deficiencies; and third, to provide data-driven recommendations for improving sleep health through behavioral interventions. Specifically, the study analyzes lifestyle factors such as physical activity, stress levels, and BMI, all of which play a direct role in sleep health. By assessing physical activity levels, we aim to understand how different types of activity affect sleep duration and quality, potentially leading to better exercise recommendations for improved rest. The research also investigates the impact of stress on sleep, as high stress levels are a leading cause of sleep disturbances, with the goal of identifying stress-reduction strategies. Additionally, the relationship between BMI and sleep quality is explored, as both obesity and being underweight are linked to sleep disorders, offering a foundation for targeted weight-management interventions. Lastly, the research examines the connection between sleep and cardiovascular health. Poor sleep can exacerbate conditions like high blood pressure and abnormal heart rates, while cardiovascular issues may also disrupt sleep. By analyzing blood pressure and heart rate in relation to sleep patterns, the study aims to provide early warnings of health conditions tied to poor sleep or unhealthy lifestyles. Through these insights, this research aims to inform strategies for improving sleep health and overall well-being.

This research addresses the growing sleep health crisis in modern society, where stress and sedentary habits undermine restorative sleep. By identifying key drivers of sleep deficiency, the study paves the way for targeted health interventions, fostering better habits that not only enhance sleep but also reduce the risk of chronic diseases.

Data background

This study leverages one dataset: “Sleep Health and Lifestyle Dataset” collected by Lacsika Tharmalingam. These datasets are publicly available on Kaggle, a platform widely used for health-related research and data analysis. The Sleep Health and Lifestyle Dataset provides comprehensive information on sleep and lifestyle habits, including variables such as sleep duration, sleep quality, physical activity levels, stress levels, BMI, blood pressure, heart rate, and the presence of sleep disorders. The Student Sleep Pattern dataset focuses on students’ sleep behaviors, adding a valuable dimension to understanding sleep trends among young adults.

The key features analyzed include:
- Sleep-related factors: Duration, quality, and disorders.
- Lifestyle factors: Physical activity levels, stress, and BMI.
- Health indicators: Blood pressure and heart rate.

The goal of this research is to explore how lifestyle factors—such as physical activity, stress levels, and BMI—affect sleep duration, quality, and the prevalence of sleep disorders. Poor sleep has well-documented links to adverse health outcomes, including cardiovascular disease, obesity, and mental health issues. By understanding these relationships, the study aims to contribute to the development of interventions that promote better sleep health.

Through this analysis, we aim to:
1. Uncover patterns between lifestyle factors and sleep quality.
2. Identify correlations that can predict sleep disorders based on lifestyle habits.
3. Provide actionable insights into preventative measures and recommendations for improving sleep health.

This analysis provides a crucial step toward understanding the interplay between lifestyle habits and sleep, offering potential benefits for both individual health and public health initiatives.

Data cleaning

Data transformation and cleaning are essential steps in preparing the dataset for meaningful analysis. For variables such as age, particularly for a specific population like university students and workers, it may have been beneficial to group or transform age into categories (e.g., 20-29 years, 30-39, 40-49, 50+ years) to better capture trends related to age-specific behaviors. Similarly, physical activity data might have been categorized into specific activity levels (e.g.,good, poor) to assess how different activity levels impact sleep quality and duration.

Data cleaning was also a critical process, especially for handling missing or inconsistent data. Incomplete data points could have been addressed through either removal or imputation, depending on the context and the extent of missing information. For example, if only a few data points were missing for a variable, imputation techniques (such as mean imputation or predictive modeling) could have been used to maintain dataset integrity. On the other hand, rows with excessive missing values might have been removed to ensure the reliability of the analysis.

Additionally, continuous variables, such as BMI, were categorized into groups (e.g., underweight, normal weight, overweight, and obese) to simplify analysis and enhance interpretability. This categorization makes it easier to identify patterns and relationships between BMI and sleep quality while allowing for more straightforward comparisons across groups. Such transformations not only streamline the analysis process but also help in presenting the findings in a more accessible and actionable manner.

#Loading data from csv file
sleep_data <- read.csv("Sleep_health_and_lifestyle_dataset.csv")
str(sleep_data)
## 'data.frame':    374 obs. of  13 variables:
##  $ Person.ID              : int  1 2 3 4 5 6 7 8 9 10 ...
##  $ Gender                 : chr  "Male" "Male" "Male" "Male" ...
##  $ Age                    : int  27 28 28 28 28 28 29 29 29 29 ...
##  $ Occupation             : chr  "Software Engineer" "Doctor" "Doctor" "Sales Representative" ...
##  $ Sleep.Duration         : num  6.1 6.2 6.2 5.9 5.9 5.9 6.3 7.8 7.8 7.8 ...
##  $ Quality.of.Sleep       : int  6 6 6 4 4 4 6 7 7 7 ...
##  $ Physical.Activity.Level: int  42 60 60 30 30 30 40 75 75 75 ...
##  $ Stress.Level           : int  6 8 8 8 8 8 7 6 6 6 ...
##  $ BMI.Category           : chr  "Overweight" "Normal" "Normal" "Obese" ...
##  $ Blood.Pressure         : chr  "126/83" "125/80" "125/80" "140/90" ...
##  $ Heart.Rate             : int  77 75 75 85 85 85 82 70 70 70 ...
##  $ Daily.Steps            : int  4200 10000 10000 3000 3000 3000 3500 8000 8000 8000 ...
##  $ Sleep.Disorder         : chr  "None" "None" "None" "Sleep Apnea" ...

This data contains life factors as variables: Person Number, Gender, Age, Occupation, Sleep Duration, Sleep Quality, Physical Activity Level, Stress Level, BMI Category, Blood Pressure, Heart Rate, Daily Steps and Sleep Disorder. For our project, we needed to create a new column using existing variables. The participation of the data research consisted of people in the minimum age of 27 and the maximum age of 59. So we decided to group the age category 10 years each.

sleep_data <- sleep_data %>%
  mutate(AgeCategory = case_when(
    Age < 20 ~ "-20",
    Age >= 20 & Age < 30 ~ "20-29",
    Age >= 30 & Age < 40 ~ "30-39",
    Age >= 40 & Age < 50 ~ "40-49",
    Age >= 50 ~ "50+"))
Age Count
20-29 30-39 40-49 50+
19 142 117 96

We also made another data called ‘sleep_summary’ which contains the mean and standard deviation of sleep duration to make a pointrange graph.

sleep_summary <- sleep_data %>%
  group_by(Occupation) %>%
  summarize(
    Sleep.Duration_mean = mean(Sleep.Duration, na.rm = TRUE),
    Sleep.Duration_sd = sd(Sleep.Duration, na.rm = TRUE),
    sleep_degree = ifelse(Sleep.Duration_mean >= 7, "Good", "Poor")) #according to proper sleep duration: 7 hours

To use blood pressure data, we needed to separate the variables since the data was in the complete form: ‘120/80. By using ’separate’, we converted the blood pressure data into systolic and diastolic. This will help in visualizing the blood pressure in point. This was not in the course, so we searched this in Google. [https://www.rdocumentation.org/packages/tidyr/versions/1.3.1/topics/separate]

sleep_data <- sleep_data %>%
  separate(Blood.Pressure, into = c("Systolic", "Diastolic"), sep = "/", convert = TRUE)

Finally, there is a research from the American College of Cardiology, saying sleeping less than 7 hours can increase the risk of causing hypertension. To check this, we made another data which filters people sleep less than 7 hours.

filtered_data <- sleep_data %>%
  filter(Sleep.Duration < 7)

Individual figures

Figure 1

The first figure is about how differences in the sleep duration are shown across age categories, separated by gender. The reason we chose a boxplot is because a boxplot effectively displays the distribution and variability of sleep duration, highlighting medians, quartiles, and outliers. Also, Faceting by age and gender emphasizes group-specific trends.

figure1 <- ggplot(sleep_data, aes(x = Gender, y = Sleep.Duration, fill = Gender)) +
  geom_boxplot() +
  facet_wrap(~AgeCategory) +
  theme_minimal() +
  theme(legend.position="none") +
  labs(title = "Sleep Duration by Gender and Age Category",
       x = "Gender",
       y = "Sleep Duration (hours)")

figure1

ggsave(figure1, filename = "images/figure1.png", 
       width = 10, height = 6, units = "in", bg = "white")

Unfortunately, there is no data collected from male whose age is over 50. Since there are only two variables in the legend, we decided to remove the legend from the figure. Red declares female group while blue for male. We can observe the sleep duration becomes longer as people aged. Seeing the median of sleep duration, female’s one increases from 6.5 hours to over 8 hours. Male’s case also increases from 6 hours to 7.3 hours approximately. We can say that physical issues cause people to sleep more than when they were younger. The another takeaway here is that there are too many outliers in the 30–39 and 40–49 age groups. It indicates a greater variability in sleep duration within these groups compared to others and could be individual problems. They are mostly at the peak of their working lives which results in both short sleep duration due to working and long sleep duration to relieve stress.

Figure 2

The second figure is a pointrange chart showing sleep duration by occupation. Pointrange charts highlight the means and variability for each group, while the use of color separates those with good and poor sleep. Could be the same reason we used the boxplot for the first one, but this case has more variables, so we used pointrange. To be honest, we didn’t think about the color designation. Since the red color means warning, next time we need to select blue for “Good”, and red for “Poor”

figure2 <- ggplot(data = sleep_summary, 
       mapping = aes(x = Occupation, 
                     y = Sleep.Duration_mean, 
                     color = sleep_degree)) +
  geom_pointrange(mapping = aes(
      ymin = Sleep.Duration_mean - Sleep.Duration_sd, 
      ymax = Sleep.Duration_mean + Sleep.Duration_sd)) +
  coord_flip() + #To see the name of occupation clearly
  theme_minimal() +
  labs(title = "Sleep Duration by Occupation",
       x = "Occupation",
       y = "Sleep Duration (hours)",
      color = "Sleep Degree")

figure2

ggsave(figure2, filename = "images/figure2.png", 
       width = 10, height = 6, units = "in", bg = "white")

Since we chose 7 hours as a proper sleep duration, those who sleep more than 7 hours are set to have good sleep degree. In the result, only 4 occupations(nurse, lawyer, engineer, accountant) are above the proper sleep duration. However, we need to see the medical occupations(nurse and doctor) have a long range and near of 7 hours of sleep duration. This means that those occupation has individual differences in sleep duration. Most occupations has the poor sleep degree. This figure will help to understand the occupational stress related to sleep duration, particularly in natural science professions

Figure 3

The last figure is a scatterplot. It is ideal for visualizing relationships between two continuous variables, with color and size encoding additional dimensions like BMI category and sleep quality. Furthermore, the filter data highlights how insufficient sleep amplifies health risks.

Figure 3-1: Blood pressure by BMI Category and its Sleep Quality

figure3_1 <- ggplot(sleep_data) +
  geom_jitter(aes(x = Systolic,
                  y = Diastolic,
                  color = BMI.Category,
                  size = Quality.of.Sleep),
              width = 1.5, height = 1, alpha = 0.3) + #alpha = 0.3 to see overlapped points clearly
  scale_color_brewer(palette = "Set2") +
  labs(title = "Blood Pressure by BMI Category",
       x = "Systolic Blood Pressure",
       y = "Diastolic Blood Pressure",
       color = "BMI Category",
       size = "Quality of Sleep") +
  theme_minimal()

figure3_1

ggsave(figure3_1, filename = "images/figure3_1.png", 
       width = 10, height = 6, units = "in", bg = "white")

First to know is that we used color palette function in scale_color_brewer. Set palletes is also called Qualitative palettes which are best suited to represent nominal or categorical data. The Set2 color was suitable to show BMI category into the colorized form. Green for Normal BMI, Orange for Normal weight(higher than Normal BMI), Blue for obese and Purple for overweight. As we know, blood pressure is consistent. If systolic blood pressure increases, diastolic blood pressure also increases to a similar degree. That is why the graph is shown to be the top right. Normal blood pressure indicates under 120/80. The degree of hypertension is over 140/90. Those in the middle of them are a risk group of hypertension. People who is both normal and normal weight usually placed in a normal blood pressure category, but most of the overweight is at the risk of hypertension. We also checked how Quality of Sleep related to Blood pressure. We predicted that the point size of overweight people would be small meaning low sleep quality, but there was no difference between underweight and overweight.

Figure 3-2: Blood Pressure whose sleep less than 7 hours

figure3_2 <- ggplot(filtered_data) +
  geom_jitter(aes(x = Systolic,
                  y = Diastolic,
                  color = BMI.Category),
              width = 1.5, height = 1, alpha = 0.5, size = 3) +
    scale_color_brewer(palette = "Set2") +
    labs(title = "Blood Pressure (Sleep Duration less than 7 hours)",
         x = "Systolic Blood Pressure",
         y = "Diastolic Blood Pressure",
         color = "BMI Category") +
    theme_minimal()

figure3_2

ggsave(figure3_2, filename = "images/figure3_2.png", 
       width = 10, height = 6, units = "in", bg = "white")

There is one research from ACC(American Colleges of Cardiology) declares that sleeping less than 7 hours can increase the risk of hypertension. [https://www.segye.com/newsView/20230330515246] As we mentioned above, over 120/80 blood pressure means they are at the risk group of hypertension. This figure indicates that people who sleep less than 7 hours have over 120/80 blood pressure are mostly overweight in BMI category. The degree of obesity is also related to hypertension, therefore, this figure can help related researches.

Conclusion

With the figures above, we can say that there is a definite relationship between sleep and life factors. When people are young, they used to sleep shorter than the proper sleep duration. However this continuously changes as they get older. This is checked with the fact that the median sleep duration of female over 50-year-old is about 8 hours. To compare males and females, females tend to has dynamically individual differences in sleep duration. Occupational stress can affect people to sleep less. There are 11 different occupations and only 4 occupations turn out having good sleep degree. The result is telling that those who work for sales parts and natural science parts and teachers sleep less than 7 hours. We also focus that there is a big deviation in software engineer, nurse and doctor. We cannot ignore the individual differences here. Further research should be processed to see what are the other factors. Blood pressure and BMI result in the predictable way. People with high blood pressure mostly have a position in obese or overweight. Also those people tend to sleep less than 7 hours, which increase the risk of hypertension. On the other hand, BMI category do nothing with the quality of sleep.

To be healthy, we need to consider a lot of factors in our life. But with the result we distract, we would like to recommend to sleep more than 7 hours which decreases hypertension as earlier as possible, when we are young. Do exercise to lose some weight also help having healthier sleeping life.