Introduction

Hello! This is team Grandmothering. Precisely, this is Taisiia Zhikharevich, Svetlana Martianova, Ekaterina Semyonova and Ekaterina Prokhorova. We would like to present a result of out half-year long work: this is a portfolio of all the projects we did for the Data analysis course in 2024.

For these projects we chose to explore Czechia via the European Social Survey dataset (https://www.europeansocialsurvey.org/). We chose the round 10 of this survey, conducted in 2020. The main theme of this round was: “Democracy, digital social contacts”. After exploring the set of questions and variables provided, we decided to choose “COVID-19 and Politics” as the focus of our analysis. We find this topic interesting because the COVID-19 period created new challenges for governments of all countries of the world, forced them to work in crisis (Lipscy, 2020), and created a new paradigm for research. In the project, we looked at this topic from different angles - how satisfied the Czech citizens were with the government’s decisions and what influenced their opinion, which parties are popular with different groups of the population, how labor policy was conducted and so on.

Structure overview

To navigate these projects more easily, we would like to provide you with some overview of the structure of these projects. Each project has some essential parts, such as:

  • Theoretical framework, a review of existing literature on the topic, justifying the hypotheses and the choice of variables for analysis,

  • Research question,

  • Preparing and cleaning the dataset, where variables are ordered and recoded to suit the analysis,

  • Variable overview, where the meaning and type of variables is provided, as well as visualisations and descriptive statistics,

  • List of contributions, where the contributions of each individual team member is mentioned.

library(dplyr)
library(ggplot2)
library(rstatix)
library(stats)
library(sjPlot)
library(car)
library(foreign)
library(tidyr)
library(stringr)
library(RColorBrewer)
library(knitr)
library(kableExtra)
library(rcompanion)
library(sjstats)
library(psych)
library(DescTools)
library(pwr)
library(corrplot)
library(gridExtra)
library(table1)

cz <-read.spss("C:/Users/Taisia/Desktop/арчик/practicals/ESS10-subset.sav", use.value.labels = T, to.data.frame = T)

Project 1: Descriptive statistics and graphs

Theoretical framework

The coronavirus pandemic has had a strong impact on many areas of people’s lives in different countries, including politics. The Czech case seemed interesting to us for several reasons. As noted by Klimovský et al. (2021), despite the fact that the Czech government had no experience in dealing with such situations, during the first wave they tried to take quick action and imposed restrictions, which made it possible to successfully cope with the spread of the virus. However, as it became clear later, the government did not have an accurate strategy and was criticized for chaotic decisions and poor communication. So, in the summer and autumn, restrictions were weakened and when signs of the second wave appeared, the government did not react immediately and tried to pretend that everything was under control, which led to a great deterioration of the situation and loss of trust of citizens.

Diving more into the political component, from the article by Havlík & Kluknavská (2022) we see that in recent years the populist ANO party has dominated in the Czech Republic, but it has turned people away from itself due to illiberal rhetoric, accusations of corruption by the head and weak policies to combat the pandemic. As a result, the party failed in the elections in 2021, and the parties opposed to populism, which created two coalitions, on the contrary, won the majority of seats in parliament. In the project, it will be interesting for us to look at the situation and problems described above using the available data.

Reseacrh question

How the COVID-19 pandemic affected the attitude of Czech citizens towards politics?

Preparing and cleaning the dataset

We deleted the missing variables (NA) from our dataset and replaced some values for a more convenient analysis. For the trstplt and gvhanc19 variables we changed some of the values so the whole scale would be a numeric one that goes from 0 to 10. For the prtvtecz/prtclecz and respc19 variables we only left two levels: “ANO 2011”/“Other” and “Yes”/“No” respectfully. These replacements are justified because: [prtvtecz/prtclecz] We are only interested in the dynamic of support for the ANO 2011 party; [respc19] It is possible to analyze the attitude of a person towards COVID-19 policies whether they really had COVID-19 or think they have COVID-19, so it’s relevant to unite them. We also converted all the variables to appropriate types.

cz1 <- cz%>%
  select(prtvtecz, prtclecz, trstplt, gvhanc19, respc19)

cz1 <- drop_na(cz1)
cz1$trstplt <- cz1$trstplt %>% str_replace_all("No trust at all", "0")
cz1$trstplt <- cz1$trstplt %>% str_replace_all("Complete trust", "10")
cz1$gvhanc19 <- cz1$gvhanc19 %>% str_replace_all("Extremely dissatisfied", "0")
cz1$gvhanc19 <- cz1$gvhanc19 %>% str_replace_all("Extremely satisfied", "10")


cz1$respc19 <- as.character(cz1$respc19)

cz1$respc19[cz1$respc19 == "Yes, I tested positive for COVID-19"] <- "Yes"
cz1$respc19[cz1$respc19 == "Yes, I think I had COVID-19 but was not tested/did not test positive "] <- "Yes"
cz1$respc19[cz1$respc19 == "No, I have not had COVID-19"] <- "No"

cz1$prtvtecz1 <- cz1$prtvtecz %>% str_replace_all("ODS", "Other")
cz1$prtvtecz <- cz1$prtvtecz %>% str_replace_all("Svoboda a přímá demokracie", "Other")
cz1$prtvtecz <- cz1$prtvtecz %>% str_replace_all("Česká pirátská strana", "Other")
cz1$prtvtecz <- cz1$prtvtecz %>% str_replace_all("ČSSD", "Other")
cz1$prtvtecz <- cz1$prtvtecz %>% str_replace_all("KSČM", "Other")
cz1$prtvtecz <- cz1$prtvtecz %>% str_replace_all("KDU-ČSL", "Other")
cz1$prtvtecz <- cz1$prtvtecz %>% str_replace_all("Starostové a nezávislí", "Other")
cz1$prtvtecz <- cz1$prtvtecz %>% str_replace_all("TOP 09", "Other")

cz1$prtclecz <- cz1$prtclecz %>% str_replace_all("ODS", "Other")
cz1$prtclecz <- cz1$prtclecz %>% str_replace_all("Svoboda a přímá demokracie", "Other")
cz1$prtclecz <- cz1$prtclecz %>% str_replace_all("Česká pirátská strana", "Other")
cz1$prtclecz <- cz1$prtclecz %>% str_replace_all("ČSSD", "Other")
cz1$prtclecz <- cz1$prtclecz %>% str_replace_all("KSČM", "Other")
cz1$prtclecz <- cz1$prtclecz %>% str_replace_all("KDU-ČSL", "Other")
cz1$prtclecz <- cz1$prtclecz %>% str_replace_all("Starostové a nezávislí", "Other")
cz1$prtclecz <- cz1$prtclecz %>% str_replace_all("TOP 09", "Other")

cz1$trstplt <- as.numeric(cz1$trstplt)
cz1$gvhanc19 <- as.numeric(cz1$gvhanc19)
cz1$respc19 <- as.factor(cz1$respc19)
cz1$prtvtecz <- as.factor(cz1$prtvtecz)
cz1$prtclecz <- as.factor(cz1$prtclecz)

Variable overview

Label = c("prtvtecz", "prtclecz", "trstplt", "gvhanc19", "respc19") 
Meaning = c("Which political party did respondent vote for", "The views of which political party the respondent adheres to.", "Trust in politicians", "How well is the government doing to limit the spread of COVID-19 according to resdondents opinion", "Whether the respondent was ill with COVID-19")
Level_Of_Measurement <- c("Categorical Nominal", "Categorical Nominal", "Numeric Interval", "Numeric Interval", "Categorical Nominal")
df <- data.frame(Label, Meaning, Level_Of_Measurement, stringsAsFactors = FALSE)

kable(df) %>% 
  kable_styling(bootstrap_options=c("bordered", "responsive","striped"), full_width = FALSE)
Label Meaning Level_Of_Measurement
prtvtecz Which political party did respondent vote for Categorical Nominal
prtclecz The views of which political party the respondent adheres to. Categorical Nominal
trstplt Trust in politicians Numeric Interval
gvhanc19 How well is the government doing to limit the spread of COVID-19 according to resdondents opinion Numeric Interval
respc19 Whether the respondent was ill with COVID-19 Categorical Nominal

Overall, we have 5 variables for analysis: 3 categorical and 2 numeric ones. The table above shows the name of the variable, its interpretation and type.

mode <- function (x) {
  u <- unique(x)
  tab <- tabulate(match(x, u))
  u[tab == max(tab)]
}
t.trstplt <- c(mean(cz1$trstplt), mode(cz1$trstplt), median(cz1$trstplt), sd(cz1$trstplt), var(cz1$trstplt), max(cz1$trstplt), min(cz1$trstplt))

names(t.trstplt) <- c("mean", "mode", "median", "sd", "variance", "max", "min")
t.gvhanc19 <- c(mean(cz1$gvhanc19), mode(cz1$gvhanc19), median(cz1$gvhanc19), sd(cz1$gvhanc19), var(cz1$gvhanc19), max(cz1$gvhanc19), min(cz1$gvhanc19))

names(t.gvhanc19) <- c("mean", "mode", "median", "sd", "variance", "max", "min")


t.prtvtecz<-c("none", "Other", "none", "none", "none", "none", "none")
names(t.prtvtecz) <- c("mean", "mode", "median", "sd", "variance", "max", "min")


t.prtclecz<-c("none", "Other", "none", "none", "none", "none", "none")
names(t.prtclecz) <- c("mean", "mode", "median", "sd", "variance", "max", "min")


t.respc19<-c("none", "No", "none", "none", "none", "none", "none")
names(t.respc19) <- c("mean", "mode", "median", "sd", "variance", "max", "min")

desc_stat_table<-data.frame(t.trstplt, t.gvhanc19,t.prtvtecz, t.prtclecz, t.respc19)
kable(desc_stat_table)%>% 
  kable_styling(bootstrap_options=c("bordered", "responsive","striped"))
t.trstplt t.gvhanc19 t.prtvtecz t.prtclecz t.respc19
mean 4.251248 5.600666 none none none
mode 5.000000 5.000000 Other Other No
median 4.000000 6.000000 none none none
sd 2.537407 2.768923 none none none
variance 6.438436 7.666933 none none none
max 10.000000 10.000000 none none none
min 0.000000 0.000000 none none none

Categorical variables

As we see in the table, for the variables prtvtecz and prtclecz the most popular value is Other. It means that respondents give their preferences to other political parties, not ANO 2011. For the variable respc19, mode is ‘No’. It shows that most of the respondents did not have COVID-19.

Numeric variables

For the variable trstplt, median is 4. It means that half of the respondents have a level of trust in politicians of less than 4, the other half - more than 4. Mode is 5, so it is the most popular answer. Mean is 4,25, the level of trust in politicians is not so high in Czech Republic. Sd 2,53 and variance 6,4, so there is a noticeable difference between the respondents’ trust levels.

For the variable gvhanc19, median is 6, half of the respondents think that government doing to limit the spread of COVID-19 of less than 6, the other half - more than 4. Most popular opinion here (mode) - 5. Mean is 5,6, so it is the average level of assessment of government actions. Sd 2,8 and variance 7,6. So, there is a significant difference between the respondents’ views.

Distribution of the variables

  1. prtvtecz: Which political party did respondent vote for
ggplot(cz1) +
  geom_bar(aes(x=prtvtecz, fill=prtvtecz)) +
  xlab("Political party (vote)") + 
  ylab("Number of people")+
  theme_bw()+
scale_fill_brewer(palette="Oranges")

Conclusion for Political party (vote) barplot: There are more people who vote not for ANO 2011, but ANO 2011 voters is the largest group within all voters (in category Other we compare voters for all parties except ANO 2011)

  1. prtclecz: The views of which political party the respondent adheres to
ggplot(cz1)+
  geom_bar(aes(x=prtclecz, fill=prtclecz))+
  xlab("Political party (views)") +
  ylab("Number of people")+
  theme_bw()+
scale_fill_brewer(palette="Oranges")

Conclusion for Political party (views) barplot: There are more people who have same views not as ANO 2011, but ANO 2011 viewers is the largest group within all viewers (in category Other we compare for all parties except ANO 2011). Also in comparison with voters there are less people who choose ANO 2011 views

3.respc19: Whether the respondent was ill with COVID-19

ggplot(cz1)+
  geom_bar(aes(x=respc19, fill=respc19))+
  xlab("COVID-19")+
  ylab("Number of people")+
    theme_bw()+
scale_fill_brewer(palette="Oranges")

Conclusion for COVID-19 barplot: there are more people who haven’t COVID-19 disease

  1. trstplt: Trust in politicians
ggplot(cz1)+ 
  geom_histogram(aes(x=trstplt), fill="#FEB24C", color="#FEB24C", binwidth = 1, alpha=0.5)+
  theme_bw()+
  xlab("Trust in politicians")+
  ylab("Number of people")+
  geom_vline(aes(xintercept = mean(cz1$trstplt), color = 'mean'), linetype="solid", size=1) +
  geom_vline(aes(xintercept = median(cz1$trstplt), color = 'median'), linetype="solid", size=1)+
  geom_vline(aes(xintercept = mode(cz1$trstplt), color = 'mode'), linetype="solid",size=1)

Conclusion for trust in politicians barplot: Level of trust in politician is skewed to the right, so most of people tend to trust in politician in the middle of a scale, but there are some who have high level of trust

  1. gvhanc19: How well is the government doing to limit the spread of COVID-19 according to resdondents opinion
ggplot(cz1)+
  geom_histogram(aes(x=gvhanc19), fill="#FEB24C", color="#FEB24C", binwidth = 1, alpha=0.5)+
  xlab("Satisfaction with governmental anti-coronavirus policies")+
  ylab("Number of people")+
  geom_vline(aes(xintercept = mean(cz1$gvhanc19), color = 'mean'), linetype="solid", size=1) +
  geom_vline(aes(xintercept = median(cz1$gvhanc19), color = 'median'), linetype="solid", size=1)+
  geom_vline(aes(xintercept = mode(cz1$gvhanc19), color = 'mode'), linetype="solid",size=1)+
  theme_bw()

Conclusion for satisfaction with governmental anti-coronavirus policies: satisfaction with governmental anti-coronavirus is skewed to the left, so most of people tend to have satisfaction in the middle of a scale, but there are some who have low level of satisfaction.

Graphs

Scatterplot

Is there a connection between the assessment of government actions in the field of COVID-19 and trust in politicians? We assume that the higher the respondent evaluates the actions of politicians to prevent the spread of COVID-19, the more they will trust politicians.

ggplot(cz1) + 
  geom_point(aes(x = gvhanc19, y = trstplt), color = "#FEB24C", alpha = 0.07, size = 9) +
  geom_smooth(aes(x = gvhanc19, y = trstplt), method = "lm", se = FALSE, color = "red") +
  theme_bw() +
  xlab("Assessment of government actions") +
  ylab("Trust in politicians") +
  ggtitle("The level of trust in politicians, depending on the assessment \nof government actions in the field of COVID-19") +
  scale_fill_brewer(palette = "Oranges")

Scatterplot conclusion: Correlation between variables is positive and weak. this means that there is a certain relationship between the level of trust in politicians and the assessment of their actions. The higher the score, the greater the trust. Regression line shows that.

Boxplot

Are there differences in satisfaction with the government’s actions in connection with COVID-19 among those who were sick and those who were not sick with coronavirus? In order to understand whether the experience of Coronavirus disease has influenced satisfaction with the actions of the government (at the time of the ESS study, it was dominated by the ANO party). We assume that those who were ill and those who were not ill may have different opinions about the anti-coronavirus policy, because those who were ill, unlike those who were not ill, faced not only the experience of avoiding diagnosis, but also its treatment

ggplot(cz1)+
  geom_boxplot(aes(y=gvhanc19, x=respc19, fill=respc19),alpha=0.5)+
   theme_bw()+
  scale_fill_brewer(palette = "Oranges")+
  labs(y="satisfaction with the Government's actions \n in connection with COVID-19",y=NULL,
       x="the respondent has COVID-19", 
       title = "Comparing the level of satisfaction with government \n 
       actions due to whether the respondent had COVID-19")+
  theme(legend.position="none",  plot.title = element_text(face = "bold",hjust = 0.5))

Boxplot conclusion: The graph shows that the response range of those who were ill and those who were not ill with COVID-19 is approximately the same. At the same time, those who were not sick with coronavirus are more satisfied with the government’s anti-coronavirus policy

Stacked barplot

Does the fact of the respondent having contracted COVID-19 make them less likely to support ANO 2011?

Czech party ANO 2011 is a populist party that was widely supported from 2013 to 2020. However, during the COVID-19 years, this party saw a decline in popularity, losing its seats. This is explained partly by the fact that the party’s COVID-19 policies were very ineffective (Havlík & Kluknavská, 2022).

We decided to use variables prtclecz (which party the respodent feels closer to) and respc19 (has the respondent contracted COVID-19). We decided to only leave two levels for each of the variables to show the relationship more clearly. For the prtclecz variable we divided respondents in two groups: ones who supported ANO 2011 (“ANO 2011”), and ones who supported a different party (“Other”). For respc19 variable we simplified it to people who had COVID-19 or think they had COVID-19 (“yes”) and people who didn’t have COVID-19 (“no”)

ggplot(cz1, aes(x = respc19, fill = prtclecz)) +
  geom_bar(position = "stack") +
  geom_text(stat='count', aes(label=..count..), position=position_stack(vjust=0.5)) +
  labs(title = "Does the fact of the respondent having contracted COVID-19 \nmake them less likely to support ANO 2011?",
       x = "has the respondent ever contracted COVID-19",
       y = "number of respondents",
       fill = "party which the respondent feels closer to")+
  theme_bw()+
  scale_fill_brewer(palette = "Oranges")

Stacked barplot conclusion: Here in the graph it is evident that people who have not contracted COVID-19 feel closer for ANO 2011 more than people who have not contracted it. There is a significant difference between respondents who were ill and who not in the context of felling closer to ANO 2011.

Conclusion

Based on the analysis, we can draw the following conclusions: 1) The most popular party in the Czech Republic is the ANO, but at the same time there are about twice as many people who vote for any other party, 2) Most Czech residents have not had COVID-19, 3) Trust in politicians and satisfaction with government actions during COVID-19 are at an average level, while the higher the satisfaction with government actions, the higher the trust, 4) Those who have not had COVID-19 are more satisfied with the actions of the government and vote for ANO more often than those who have had it.

List of contributions

  • Martianova Svetlana - theoretical framework, knitting html,
  • Semyonova Ekaterina - boxplot, table with descriptive statistics,
  • Prokhorova Ekaterina - scatterplot, describing variables (distribution, types),
  • Zhikharevich Taisiia - cleaning dataset, stacked barplot.

Project 2: Chi-square, t-test, ANOVA

Theoretical framework

The COVID-19 pandemic affected many countries and governments and Czechia is not an exception. As evident in Andoh (2020) and Havlík (2022) the pandemic affected Czech labor market and political sphere in a number of ways we would like to explore in our project. For example, an interesting contradiction in data is that Czechia had a very low rate of people losing their jobs because of COVID-19, but then the COVID-19 policies in particular were the ones that led to ANO 2011 populist party’s downfall in 2021 elections. We would like to explore this topic more thouroughly.

Research question

How the COVID-19 pandemic affected politiсal and labour market situation in Czech Republic?

Preparing and cleaning the dataset

We cleaned the dataset and made separate datasets for different tests (czc for chi-square, czt for t-test and cza for ANOVA).

We left only two levels (“Yes” and “No”) for the variable respc19 (whether respondent had COVID-19), as we decided that, for the purposes of our research, the fact that the respondent thinks they had COVID-19 is more important than whether they actually had it.

We also recoded the variable prtclecz (which party the respondent feels closer to) to include the coalitions formed in 2021, not each individual party.

We chose people over 17, to only take into account respondents who had the right to vote.

czc <- cz %>%
  select(respc19, hapljc19)

czc <- drop_na(czc)
czc$respc19 <- as.character(czc$respc19)
czc$respc19[czc$respc19 == "Yes, I tested positive for COVID-19"] <- "Yes"
czc$respc19[czc$respc19 == "Yes, I think I had COVID-19 but was not tested/did not test positive "] <- "Yes"
czc$respc19[czc$respc19 == "No, I have not had COVID-19"] <- "No"

czc$respc19 <- as.factor(czc$respc19)


czc <- drop_na(czc)
czc <- droplevels(czc)

cza <- cz%>%
  select(prtclecz, agea)

cza <- drop_na(cza)

cza$prtclecz <- cza$prtclecz %>% str_replace_all("ODS", "SPOLU")
cza$prtclecz <- cza$prtclecz %>% str_replace_all("TOP 09", "SPOLU")
cza$prtclecz <- cza$prtclecz %>% str_replace_all("KDU-ČSL", "SPOLU")
cza$prtclecz <- cza$prtclecz %>% str_replace_all("Svoboda a přímá demokracie", "SPD")
cza$prtclecz <- cza$prtclecz %>% str_replace_all("ČSSD", "Left")
cza$prtclecz <- cza$prtclecz %>% str_replace_all("KSČM", "Left")
cza$prtclecz <- cza$prtclecz %>% str_replace_all("Starostové a nezávislí", "PirSTAN")
cza$prtclecz <- cza$prtclecz %>% str_replace_all("Česká pirátská strana", "PirSTAN")

cza$prtclecz <- as.factor(cza$prtclecz)
cza$agea <- as.numeric(cza$agea)

cza <- filter(cza, agea > 17)
cza[cza=="Other"]<-NA
cza <- drop_na(cza)
cza <- droplevels(cza)

czt <- cz%>%
  select(gvhanc19, respc19)
czt <- drop_na(czt)

czt$gvhanc19 <- czt$gvhanc19 %>% str_replace_all("Extremely dissatisfied", "0")
czt$gvhanc19 <- czt$gvhanc19 %>% str_replace_all("Extremely satisfied", "10")
czt$gvhanc19<-as.factor(czt$gvhanc19)
czt$gvhanc19<-as.numeric(czt$gvhanc19)


czt$respc19 <- as.character(czt$respc19)
czt$respc19[czt$respc19 == "Yes, I tested positive for COVID-19"] <- "Yes"
czt$respc19[czt$respc19 == "Yes, I think I had COVID-19 but was not tested/did not test positive "] <- "Yes"
czt$respc19[czt$respc19 == "No, I have not had COVID-19"] <- "No"
czt$respc19<-as.factor(czt$respc19)

Variable overview

Variable = c("respc19", "hapljc19", "prtclecz", "gvhanc19", "agea") 
Meaning = c("Whether the respondent was ill with COVID-19",  "Things happened since start of COVID-19: was made redundant/lost job", "Which political party respondents feel closer to", "How well is the government doing to limit the spread of COVID-19 according to resdondents opinion","Age of respondent")
Type <- c("Categorical Nominal, Binary", "Categorical Nominal, Binary", "Categorical Nominal", "Quasi-interval", "Ratio")
Class <- c(class(czc$respc19), class(czc$hapljc19), class(cza$prtclecz), class(czt$gvhanc19), class(cza$agea))
df1 <- data.frame(Variable, Meaning, Type, Class, stringsAsFactors = FALSE)

kable(df1) %>% 
  kable_styling(bootstrap_options=c("bordered", "responsive","striped"), full_width = FALSE)
Variable Meaning Type Class
respc19 Whether the respondent was ill with COVID-19 Categorical Nominal, Binary factor
hapljc19 Things happened since start of COVID-19: was made redundant/lost job Categorical Nominal, Binary factor
prtclecz Which political party respondents feel closer to Categorical Nominal factor
gvhanc19 How well is the government doing to limit the spread of COVID-19 according to resdondents opinion Quasi-interval numeric
agea Age of respondent Ratio numeric

Chi-squared test

Theoretical framework

For our chi-square test we decided to choose two variables: respc19 (Whether the respondent was ill with COVID-19, categorical nominal, two levels) and hapljc19 (Things happened since start of COVID-19: was made redundant/lost job, categorical nominal, two levels). We are interested in these variables, because COVID-19 significantly changed the labor market in Chezh Republic. After the COVID-19 total employment decreased by 87,5 thousand year-on-tear (Hedvicakova & Kozubikova, 2021).

Variable overview

The categories of variables are mutually exclusive, the respondents can be assigned to only one of the groups, and the studied groups are independent (trusting the data collectors). We also created a plot that shows distribution of the variables.

t1 <- as.table(table(czc$respc19, czc$hapljc19))
kable(t1, caption="Contingency table")
Contingency table
Not marked Marked
No 1629 39
Yes 717 37
plot_xtab(czc$respc19, czc$hapljc19,
          margin = "row", 
          bar.pos = "stack", 
          title = "Distribution of those who lost and did not lose their jobs and whether they were ill COVID-19",
          axis.titles = "Whether the respondent was ill with COVID-19",
          legend.title = "Was made redundant/lost job",
          legend.labels = c("No", "Yes"),
          geom.colors = "Oranges",
          show.summary = T)+
  theme_bw()

We can see a great difference between them, so we can conduct a chi-square test. We can see that majority of respondents didn’t lose their job, however, the percentage of those who lost the job is bigger for people that had COVID-19

Test

H0 - There is no relation between the loosing job and COVID-19 illness H1 - There is a relation between the losing job and COVID-19 illness

chisq.test(czc$respc19, czc$hapljc19)
## 
##  Pearson's Chi-squared test with Yates' continuity correction
## 
## data:  czc$respc19 and czc$hapljc19
## X-squared = 10.446, df = 1, p-value = 0.001229

After conducting a chi-square test, we see that p-value is less than 0,05. This means that we can reject the H0 and say that there is relation between variables. The presence of COVID-19 in a respondent affects his work situation. So, we have to check residuals.

Residuals

chi3<-as.table(chisq.test(czc$respc19, czc$hapljc19)$expected)
chi4<-as.table(chisq.test(czc$respc19, czc$hapljc19)$stdres)
kable(chi3, caption = "Expected values")
Expected values
Not marked Marked
No 1615.6598 52.34021
Yes 730.3402 23.65979
kable(chi4,  caption = "Standardized residuals")
Standardized residuals
Not marked Marked
No 3.357914 -3.357914
Yes -3.357914 3.357914

Looking at the expected values, we can confirm the assumption of 5 observations minimum being in each cell is correct. This means that the test conducted was relevant.

Then, we should inspect the standardized residuals. We can see that every category made a significant contribution to the results, however:

  • we expected to see less people who didn’t have COVID-19 and didn’t lose their job;
  • we expected to see less people who had COVID-19 and lost their job;
  • we expected to see more people who didn’t have COVID-19 and lost their job;
  • we expected to see more people who had COVID-19 and didn’t lose their job.

All-in-all, this confirms the results that the test conducted gave us.

However, as we can see from the image in Variable overview, the effect size is only 0.07, which is considered to be small. Though our findings are significant, they are not of a big effect on a population.

Conclusion

The test results showed us that there is a connection between job loss and alleged COVID-19 disease. We can assume that this is due to the fact that people fell ill in the workplace and lost their jobs due to long-term treatment or transfer to a remote format, in such conditions or human condition that labor productivity decreases.

T-test

Theoretical framework

We are interested in studying the difference in satisfaction with government actions during COVID-19, depending on whether the respondent was ill with COVID-19 or not, since the health of the population and the understanding that others are healthy is one of the most important factors for satisfaction with government actions (Chen et. al., 2021). We assume that there will be a difference between the groups of those who were ill and those who were not ill, since those who were ill will be included in the context of treatment for coronavirus, acquaintance with other patients - living sick role (Parsons, 1951).

Variable overview

For the analysis, we chose respc19 as the grouping variable (Whether the respondent was ill with COVID-19, categorical nominal, two levels), and the dependent variable is gvhanc19 (Satisfaction with government actions in connection with COVID-19, quasi-interval, 10 levels). Let’s look at the variables chosen for analysis through building a boxplot.

ggplot(czt, aes(x = respc19, y = gvhanc19, fill=respc19)) + 
  geom_boxplot() +
  stat_summary(fun.y = mean, geom = "point", size = 2, col = "red") +
  theme_bw()+
  labs(title = "The respondent's COVID-19 and satisfaction \nwith how well the state is coping with the pandemic", x = "Whether the respondent have COVID-19", y = "Satisfaction with government actions \nin connection with COVID-19")+
  scale_fill_brewer(palette="Oranges")+
  theme(legend.position="none")

Boxplot showed that the variance looks the same and we don’t have outliers, but there is a difference in means (red dot).

Assumptions

Let us check the assumptions to perform the t-test correctly.

Normality of data

Since our data is quasi-interval, it will not be very convenient for us to use a histogram to check the distribution, so we will check the normality of the data via q-q plot.

qqnorm(czt$gvhanc19, ylim = c(0, 100)); qqline(czt$gvhanc19, ylim = c(0, 100), col= 2)

The q-q plot shows that the data practically does not deviate from the line of normality, but for verification it is worth looking at skew and kurtosis.

describeBy(czt, group = czt$respc19)
## 
##  Descriptive statistics by group 
## group: No
##          vars    n mean   sd median trimmed  mad min max range  skew kurtosis
## gvhanc19    1 1646 6.83 2.59      7    6.99 2.97   1  11    10 -0.48    -0.51
## respc19     2 1646 1.00 0.00      1    1.00 0.00   1   1     0   NaN      NaN
##            se
## gvhanc19 0.06
## respc19  0.00
## ------------------------------------------------------------ 
## group: Yes
##          vars   n mean   sd median trimmed  mad min max range  skew kurtosis
## gvhanc19    1 743 6.67 2.52      7    6.88 2.97   1  11    10 -0.64     -0.2
## respc19     2 743 2.00 0.00      2    2.00 0.00   2   2     0   NaN      NaN
##            se
## gvhanc19 0.09
## respc19  0.00

Skew and kurtosis fall into the range from -2 to 2 (even from -1 to 1) from which we can still conclude that the data is normal and conduct a t-test.

Equality of variances

Let’s check equality of variances with Levene’s test.

H0: variances of satisfaction with the government’s actions during COVID-19 among those who were sick with COVID-19 and those who were not sick are equal. H1: variances of satisfaction with the government’s actions during COVID-19 among those who were sick with COVID-19 and those who were not sick are not equal.

leveneTest(czt$gvhanc19 ~ czt$respc19)
## Levene's Test for Homogeneity of Variance (center = median)
##         Df F value Pr(>F)
## group    1  2.3891 0.1223
##       2387

Test show that the data have high probability to occur if null hypothesis is true. Thus, the variances are equal. Let’s doublecheck it with Bartlett’s test.

bartlett.test(czt$gvhanc19 ~ czt$respc19)
## 
##  Bartlett test of homogeneity of variances
## 
## data:  czt$gvhanc19 by czt$respc19
## Bartlett's K-squared = 0.78052, df = 1, p-value = 0.377

Test show that the data have high probability to occur if null hypothesis is true. Thus, the variances are equal.

Test

H0: the mean satisfaction with the government’s actions during COVID-19 among those who were sick with COVID-19 and those who were not sick is equal

H1: the mean satisfaction with the government’s actions during COVID-19 among those who were sick with COVID-19 and those who were not sick is not equal

t.test(czt$gvhanc19 ~ czt$respc19, var.equal = T)
## 
##  Two Sample t-test
## 
## data:  czt$gvhanc19 by czt$respc19
## t = 1.4277, df = 2387, p-value = 0.1535
## alternative hypothesis: true difference in means between group No and group Yes is not equal to 0
## 95 percent confidence interval:
##  -0.06063575  0.38528918
## sample estimates:
##  mean in group No mean in group Yes 
##          6.829891          6.667564

P-value > 0,05 thus we cannot reject null hypothesis and means are equal. On average, those who were ill and those who were not ill have the same satisfaction with the actions of the government during COVID-19. We have to check it with non-parametric test.

Non-parametric test

Our data is not continuous, the samples are independent, so we use the Mann-Whitney test

wilcox.test(czt$gvhanc19 ~ czt$respc19, paired = F)
## 
##  Wilcoxon rank sum test with continuity correction
## 
## data:  czt$gvhanc19 by czt$respc19
## W = 633591, p-value = 0.1534
## alternative hypothesis: true location shift is not equal to 0

P-value > 0,05 thus we cannot reject null hypothesis and means are equal. On average, those who were ill and those who were not ill have the same satisfaction with the actions of the government during COVID-19.

Conclusion

Based on the analysis, we can conclude that in the Czech case, the difference in satisfaction with the actions of politicians during COVID-19 among those who were ill and those who were not ill with it is not statistically significant. Our assumptions turned out to be incorrect, so we believe that this topic should be studied in more detail in the future.

ANOVA

Theoretical framework

Continuing the research on people’s political preferences that was started in the last project, we found an interesting work linking age and political preferences/voting behaviour. Holland (2013) talks about three main theories of the relationship between age and political preferences. Firstly, the idea of the preference of candidates is close to their age, as well as the hypothesis that young people are more liberal than people of the older generation, which will influence their preferences. And finally, the idea that older people prefer moderate/traditional candidates, and people of the younger generation, on the contrary, those who oppose the traditional party system. The study was conducted in an American context and it will be interesting for us to look at the situation in the Czech Republic, focusing on what kind of party people feel closer to and their age.

Research hypothesis: The age mean will be higher in leftist and populist parties and lower in centrist and liberal.

Variable overview

For our analysis we chose agea variable (age of respondent, continuous) and prtclecz variable that defines to which party respondents feel closer to (categorical). Above you can see that we have united some parties based on the existing coalitions of the Czech Republic that were mentioned in one of the articles (Havlík & Kluknavská, 2022).

describeBy(cza$agea, cza$prtclecz, mat = TRUE) %>% 
  select(prtclecz = group1, N=n, Mean=mean, SD=sd, Median=median, Min=min, Max=max, 
                Skew=skew, Kurtosis=kurtosis, st.error = se) %>% 
  kable(align=c("lrrrrrrrr"), digits=2, row.names = FALSE,
        caption="Party/Coalition feel closer to by age") %>% 
  kable_styling(bootstrap_options=c("bordered", "responsive","striped"), full_width = FALSE)
Party/Coalition feel closer to by age
prtclecz N Mean SD Median Min Max Skew Kurtosis st.error
ANO 2011 216 46.33 12.46 47 18 72 -0.36 -0.60 0.85
Left 117 49.09 12.75 52 18 71 -0.57 -0.64 1.18
PirSTAN 79 36.11 13.88 32 18 65 0.61 -0.87 1.56
SPD 79 38.78 12.05 39 18 62 0.06 -1.17 1.36
SPOLU 119 40.82 13.37 40 18 72 0.26 -0.88 1.23

From the table we can see that groups are of comparable size, skew close to 0 in almost all cases, however, standard deviation is quite big.

ggplot(cza)+
  geom_boxplot(aes(x = prtclecz, y = agea, fill=prtclecz, alpha = 0.5))+
  theme_bw()+
scale_fill_brewer(palette="Oranges")+
  labs(title = "The age of people based on which party/coalition they feel closer to", x = "Which party/coalition feel closer to", y = "Age") +
  theme(legend.position="none", plot.title = element_text(face = "bold",hjust = 0.5))

From the plot, we see that the agea variable is distributed quite normally across the groups, the medians are close to the centre of boxes, also we do not see outliers. Another thing that we see is that response range for different groups looks quite similar (however, slightly less for PirSTAN and SPD). Also we can say that those who feel closer to ANO and left parties are older than people in other groups.

Test

H0: the means of all groups are equal; Ha: at least one of the groups have different mean.

res_aov <- aov(data = cza, agea~prtclecz)
summary(res_aov)
##              Df Sum Sq Mean Sq F value   Pr(>F)    
## prtclecz      4  12319  3079.8    18.7 1.68e-14 ***
## Residuals   605  99665   164.7                     
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Results show that p-value < 0.05, therefore, we can reject the null hypothesis in favor of the hypothesis that there is a difference in means between the groups. The difference in the age across groups of political parties that people feel closer to is statistically significant.

Assumptions

Now we will check the assumptions for ANOVA.

The variances are equal

leveneTest(agea ~ prtclecz, data = cza)
## Levene's Test for Homogeneity of Variance (center = median)
##        Df F value Pr(>F)
## group   4   0.773  0.543
##       605

The Levene test shows us a p-value > 0.05, so we can’t reject the null hypothesis of variances being equal.

Independence

The dataset can guarantee us the randomness and the absence of “before-after” design. There is also no relationship between the groups, because one person could only choose one party that they felt closer to.

Normality of residuals

First, we check the skew and kurtosis.

describe(res_aov$residuals)[11:12]
##     skew kurtosis
## X1 -0.06    -0.73

They are both less than 2, which is normal.

Next, we perform a Shapiro-Wilk test.

shapiro.test(res_aov$residuals)
## 
##  Shapiro-Wilk normality test
## 
## data:  res_aov$residuals
## W = 0.98874, p-value = 0.0001233

However, the p-value < 0.05 and we should reject the null hypothesis of the residuals being normal. However, this test is really sensitive and is used primarily on small samples, so we should visualize our data to make better judgements about its normality.

Now we visualize our data with a qqplot and a histogram.

plotNormalHistogram(res_aov$residuals, prob = FALSE, linecol="orange")

qqPlot(res_aov$residuals,
  id = FALSE
  )

On the histogram we see that the distribution is fairly close to normal. The qqplot also suggests that the data is close to a normal distribution, despite some deviations here and there. Therefore, we can proceed.

Post hoc analysis

Since our first test is parametric and significant and the varicances are equal, we choose the Tukey test to perform post hoc analysis.

TukeyHSD(res_aov)
##   Tukey multiple comparisons of means
##     95% family-wise confidence level
## 
## Fit: aov(formula = agea ~ prtclecz, data = cza)
## 
## $prtclecz
##                        diff         lwr       upr     p adj
## Left-ANO 2011      2.760684  -1.2703466  6.791714 0.3326364
## PirSTAN-ANO 2011 -10.219409 -14.8366694 -5.602149 0.0000000
## SPD-ANO 2011      -7.548523 -12.1657833 -2.931263 0.0000899
## SPOLU-ANO 2011    -5.518207  -9.5272050 -1.509210 0.0017011
## PirSTAN-Left     -12.980093 -18.0937939 -7.866392 0.0000000
## SPD-Left         -10.309207 -15.4229078 -5.195506 0.0000005
## SPOLU-Left        -8.278891 -12.8508608 -3.706921 0.0000093
## SPD-PirSTAN        2.670886  -2.9165840  8.258356 0.6864433
## SPOLU-PirSTAN      4.701202  -0.3951489  9.797553 0.0867080
## SPOLU-SPD          2.030316  -3.0660350  7.126667 0.8117327
Tukey <- TukeyHSD(res_aov)
par(mar = c(5, 13, 3, 1))
plot(Tukey, las = 2)

The post hoc analysis shows us that most of the pairwise comparisons are significant, except for these:

SPD-PirSTAN, SPOLU-PirSTAN, SPOLU-SPD and Left-ANO 2011.

This can be explained by the fact that the first three parties are the parties with similar ideologies (e. g. right-wing) and the last pair with the fact that ANO 2011 still has an overwhelming amount of support in all age groups.

Then, we calculate the effect size for this test by looking at the omega-squared.

anova_stats(res_aov) 
## term      |  df |     sumsq |   meansq | statistic | p.value | etasq | partial.etasq | omegasq | partial.omegasq | epsilonsq | cohens.f | power
## -----------------------------------------------------------------------------------------------------------------------------------------------
## prtclecz  |   4 | 12319.152 | 3079.788 |    18.695 |  < .001 | 0.110 |         0.110 |   0.104 |           0.104 |     0.104 |    0.352 |     1
## Residuals | 605 | 99665.215 |  164.736 |           |         |       |               |         |                 |           |          |

It is 0.104 which is a moderate effect size. This means that the age only moderately corresponds with the choice of the political party that the respondent feels closer to.

Non-parametric test

Then, we perform a non-parametric test (Kruskal-Wallis) which turns out to be significant and calculate its effect size, which is moderate.

kruskal.test(agea~prtclecz, data = cza)
## 
##  Kruskal-Wallis rank sum test
## 
## data:  agea by prtclecz
## Kruskal-Wallis chi-squared = 66.552, df = 4, p-value = 1.212e-13
kruskal_effsize(cza, agea~prtclecz)
## # A tibble: 1 × 5
##   .y.       n effsize method  magnitude
## * <chr> <int>   <dbl> <chr>   <ord>    
## 1 agea    610   0.103 eta2[H] moderate

Then, we perform post hoc analysis for the non-parametric test. It is the same as in the parametric test.

DunnTest(agea ~ prtclecz, data = cza,
         method = "holm")
## 
##  Dunn's test of multiple comparisons using rank sums : holm  
## 
##                  mean.rank.diff    pval    
## Left-ANO 2011          38.62055 0.16856    
## PirSTAN-ANO 2011     -129.03167 2.3e-07 ***
## SPD-ANO 2011          -97.70889 0.00015 ***
## SPOLU-ANO 2011        -71.49753 0.00189 ** 
## PirSTAN-Left         -167.65222 6.4e-10 ***
## SPD-Left             -136.32944 8.6e-07 ***
## SPOLU-Left           -110.11808 1.1e-05 ***
## SPD-PirSTAN            31.32278 0.52770    
## SPOLU-PirSTAN          57.53415 0.09777 .  
## SPOLU-SPD              26.21136 0.52770    
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Conclusion

Thus, we can say that the difference in the age means between groups of political parties that people feel closer to is statistically significant for almost all pairs. For those parties with no statistical significance, we can assume that similar ideologies or a very large audience coverage of different ages play a role. Even though anova is non-directional and we can’t make definitive judgements of data, looking at the presented graphs our hypothesis turned out to be at least somewhat true. Medians of leftist and populist parties turned out to be higher than in other groups. Considering that from pairwise comparisons we see that the adherents of the ANO and the left parties are older (the comparisons are negative, so the first party in each pair has a lower age than the adherents of the second party in the pair, in almost all cases we see that the second party was the ANO or the left) than the rest. Thus, our hypothesis is confirmed.

Conclusion

Based on our analysis, we can draw the following conclusions: 1) There is statistically significant relations between status of the COVID-19 patient and the working status of Czech citizens 2) People of different ages choose different parties to which they feel closer to, the older generation prefers leftist and populist parties 3) The status of a COVID-19 patient does not affect satisfaction with the actions of politicians during the pandemic Thus, we found out what factors influencing the political situation in the Czech Republic during the coronavirus pandemic.

List of contributions

  • Prokhorova Ekaterina - Chi-squared test, theoretical framework,

  • Semyonova Ekaterina - The t-test, theoretical framework, knitting html,

  • Martianova Svetlana - ANOVA (boxplot, assumptions, F-test), theoretical framework,

  • Zhikharevich Taisiia - ANOVA (assumptions, non-parametric test, post hoc tests), theoretical framework, knitting html.

Project 3: Linear regression, additive model

Theoretical framework

Monitoring the public and satisfaction with government’s handling of COVID-19

During the COVID-19, there was a great public discussion about citizens’ privacy. According to Ioannou, Tussyadiah (2021), citizens who show greater trust in the government tend to be more flexible about restrictions on their privacy for public safety purposes. They can see such measures as a necessary evil in the fight against the pandemic and trust the Government to take appropriate measures. Another study shows that citizens are not ready to share all personal information in full and prefer to be able to manage the data provided, regardless of their trust in the government.(Simko, Chang, Jiang, Calo, Roesner, Kohno, 2022). In our opinion, the relationship between these variables exists. We assume that the more a citizen trusts the government, the more they will agree to provide personal data for the purpose of monitoring COVID-19.

Hypothesis: those who preferred monitoring the lives of citizens were more satisfied with the actions of the government

Conformity (governmental rules or people’s own decisions) and satisfaction with government’s handling of COVID-19

Hypothesis: those who turned out to be more conformist and believed that it was always more important to follow the decisions of the government were more satisfied with its activities during COVID-19

News about politics and current affairs and satisfaction with government’s handling of COVID-19

One of the articles we found looked at the relationship between the frequency of news consumption and political trust. As a result, they concluded that in free countries, people who follow the news are more likely to have less political trust, while in non-free societies the relationship is reversed (Jiang & Zhang, 2021). Another article states that the higher people’s satisfaction with the government’s actions during the pandemic, the higher their trust in it (Wu et al., 2021). We are interested in the relationship between news consumption and satisfaction with government actions during pandemic, however, we didn’t find articles with this pure relationship. Therefore, based on the results of studies mentioned above, our hypothesis is that with an increase in the time of news consumption, satisfaction with government’s handling of COVID-19 decreases. We think so because, according to Freedom House data, which was used in the mentioned article to determine the level of freedom, the Czech Republic belongs to the category of free societies. Because of this, the level of trust should be lower and since less satisfaction with government actions leads to less trust, we assume that reading the news will also lead to less satisfaction with actions.

Hypothesis: With an increase in the time of news consumption, satisfaction with government’s handling of COVID-19 decreases

COVID-19 status and satisfaction with government’s handling of COVID-19

We believe that the fact of COVID-19 disease can also become a predictor, because according to Chen et. al. (2021), this is one of the important factors influencing the opinion about government policy during the pandemic in connection with the concept of sick role and worries about the general level of morbidity, worries about loved ones who may to get started. For these reasons, we decided to make the respc19 variable an additional predictor.

Hypothesis: People who had COVID-19 are less likely to be satisfied with government’s handling of COVID-19

Reseacrh question

What are the predictors of satisfaction with government’s handling of COVID-19 in Czechia?

Preparing and cleaning the dataset

We recoded several values for ease of analysis: respc19 was reduced to two levels by combining two different “Yes” answers into one, in the remaining variables we recoded the extreme points from the text format to 0 and 10.

cz3 <-cz

cz3$respc19 <- as.character(cz3$respc19)
cz3$respc19[cz3$respc19 == "Yes, I tested positive for COVID-19"] <- "Yes"
cz3$respc19[cz3$respc19 == "Yes, I think I had COVID-19 but was not tested/did not test positive "] <- "Yes"
cz3$respc19[cz3$respc19 == "No, I have not had COVID-19"] <- "No"
cz3$respc19 <- as.factor(cz3$respc19)

cz3$nwspol <- as.numeric(as.character(cz3$nwspol))

cz3$panfolru <- as.character(cz3$panfolru)
cz3$panfolru[cz3$panfolru == "Much more important to follow government rules"] <- "0"
cz3$panfolru[cz3$panfolru == "Much more important to make your own decisions"] <- "10"
cz3$panfolru <- as.numeric(cz3$panfolru)

cz3$panmonpb <- as.character(cz3$panmonpb)
cz3$panmonpb[cz3$panmonpb == "Much more important to monitor and track the public"] <- "0"
cz3$panmonpb[cz3$panmonpb == "Much more important to maintain public privacy"] <- "10"
cz3$panmonpb <- as.numeric(cz3$panmonpb)

cz3$gvhanc19 <- cz3$gvhanc19 %>% str_replace_all("Extremely dissatisfied", "0")
cz3$gvhanc19 <- cz3$gvhanc19 %>% str_replace_all("Extremely satisfied", "10")
cz3$gvhanc19<-as.numeric(cz3$gvhanc19)

Variable overview

First of all, let’s look closer at the variables that we have selected for analysis.

Variable = c("panmonpb", "gvhanc19", "nwspol", "panfolru", "respc19") 
Meaning = c("More important for governments to monitor and track the public or to maintain public privacy when fighting a pandemic", "How satisfied with government's handling of COVID-19", "News about politics and current affairs, watching, reading or listening, in minutes", "More important to follow government rules or to make own decisions when fighting a pandemic", "Whether the respondent was ill with COVID-19")
Type <- c("Quasi-interval", "Quasi-interval", "Ratio", "Quasi-interval", "Binary")
Class <- c(class(cz3$panmonpb), class(cz3$gvhanc19), class(cz3$nwspol), class(cz3$panfolru), class(cz3$respc19))
df1 <- data.frame(Variable, Meaning, Type, Class, stringsAsFactors = FALSE)

kable(df1) %>% 
  kable_styling(bootstrap_options=c("bordered", "responsive","striped"), full_width = FALSE)
Variable Meaning Type Class
panmonpb More important for governments to monitor and track the public or to maintain public privacy when fighting a pandemic Quasi-interval numeric
gvhanc19 How satisfied with government’s handling of COVID-19 Quasi-interval numeric
nwspol News about politics and current affairs, watching, reading or listening, in minutes Ratio numeric
panfolru More important to follow government rules or to make own decisions when fighting a pandemic Quasi-interval numeric
respc19 Whether the respondent was ill with COVID-19 Binary factor

In total, we selected 5 variables for analysis, since the ESS data almost does not contain continuous variables, we selected quasi-interval variables that have 10 categories and can be used for analysis as interval variables. As a categorical predictor, we will use the variable respc19 and panfolru as a continuous predictor.

News about politics and current affairs

As this variable is the only non-quasi-interval here, we would like to look into it in more detail.

polnews <- c(mean(cz3$nwspol, na.rm = T), DescTools::Mode(cz3$nwspol, na.rm = T), median(cz3$nwspol, na.rm = T), sd(cz3$nwspol, na.rm = T), min(cz3$nwspol, na.rm = T), max(cz3$nwspol, na.rm = T),skew(cz3$nwspol, na.rm = T), Kurt(cz3$nwspol,na.rm = T), parameters::standard_error(cz3$nwspol, na.rm = T))
names(polnews) <- c("mean", "mode", "median",  "sd",  "min",  "max",  "skew",  "kurtosis",  "st.error")

kable(polnews, align=c("lrrrrrrrr"), digits=2) %>% 
  kable_styling(bootstrap_options=c("bordered", "responsive","striped"), full_width = FALSE)
x
mean 65.65
mode 60.00
median 45.00
sd 81.44
min 0.00
max 960.00
skew 4.04
kurtosis 23.90
st.error 1.70

From the table we see that the mean is 65.65, median is 45, but a maximum value is 960. Also our data is skewed positively (4 > 0.5) and the distribution is sharp (kurtosis > 1). So the data is not normal, on average people spend 66 minutes a day on news, however, we have outliers with larger news consumption.

ggplot(subset(cz3, !is.na(cz3$nwspol)), aes(x= nwspol)) +
geom_histogram(fill="#FEB24C", color="#FEB24C", alpha=0.5, binwidth = 15)+
theme_bw() +
labs(y = "Number of people", x='Minutes')+
  ggtitle("Time spent on news about politics and current affairs")+
theme(plot.title = element_text(size=18, face='bold', hjust=0.4))+
scale_x_continuous(breaks = c(0,100,200,300,400,500,600,700,800,900))+
geom_vline(aes(xintercept=mean(nwspol, na.rm=T), color="mean"), linetype="solid", linewidth = 1)+ geom_vline(aes(xintercept=median(nwspol, na.rm=T), color="median"), linetype="dashed", linewidth = 1)+ geom_vline(aes(xintercept=Mode(nwspol), color="mode"), linetype="dotted")+ 
  scale_color_manual(name = "Measurement", values = c(median = "#E31A1C", mean = "#800026", mode = "#4b4b4b"))

Again we see that our data is skewed to the right, however, the number of people that spend more than 400 minutes on news is really small. Most people spend not more than 70 minutes. Also, compared to people who spend several hours on the news, many more people spend almost no time on it (more than 200 respondents at 0 point).

Other continuous variables

For panmonpb 0 is “Much more important to monitor and track the public”, 10 is Much more important to maintain public privacy. For gvhanc19 0 is “Extremely dissatisfied with the government’s handling of the coronavirus pandemic”, 10 is “Extremely satisfied”. For panfolru 0 is “Much more important to follow government rules”, 10 is “Much more important to make your own decisions”.

cz3 %>%
  summarise_at(vars(c(gvhanc19, panmonpb, panfolru)), list(Mean= ~mean(., na.rm = T), Median=~median(., na.rm = T), Mode = ~DescTools::Mode(., na.rm = T), SD=~sd(., na.rm = T), Min=~min(., na.rm = T), Max=~max(., na.rm = T)))%>%
  pivot_longer(cols = gvhanc19_Mean:panfolru_Max, names_to = c("Variable", "Statistics"), names_sep = "_", values_to = "Values")%>%
  pivot_wider(names_from = Variable, values_from = Values)%>% 
  kable(align=c("lrrrrrrrr"), digits=2, row.names = FALSE) %>% 
  kable_styling(bootstrap_options=c("bordered","responsive","striped"), full_width = FALSE)
Statistics gvhanc19 panmonpb panfolru
Mean 5.23 7.01 5.75
Median 5.00 7.00 6.00
Mode 5.00 10.00 5.00
SD 2.51 2.46 2.81
Min 0.00 0.00 0.00
Max 10.00 10.00 10.00

Mode and median for satisfaction with the government’s handling of the pandemic is 5, mean is 5.23, standard deviation is 2.5. We can see that people occupy a middle position, that is, they are not fully satisfied with how the government is coping, but it cannot be said that they are not satisfied with it at all. For importance for governments to monitor and track the public the mean and median is 7, and the mide is 10. So people mostly think that it is much more important to maintain public privacy for governments than monitor the public. For following government rules or making own decisions the mean and median is around 6, mode is 5. On average people think that it is important to make their own decisions when fighting a pandemic, but following government rules is also important because answers lay almost in the middle.

media <- ggplot(subset(cz3, is.na(cz$panmonpb) == F), aes(x = as.factor(panmonpb))) +
geom_bar(fill='#FEB24C',color = "#FC4E2A", alpha=0.7) +
theme_bw() +
labs(x = "Monitoring the public \nor maintaining public privacy", y = "Number of people")

restricting <- ggplot(subset(cz3, is.na(cz3$gvhanc19) == F), aes(x = as.factor(gvhanc19))) +
geom_bar(fill='#FED976',color = "#FD8D3C", alpha=0.7) +
theme_bw() +
labs(x = "How satisfied with government's \nhandling of COVID-19", y = "Number of people")

decisions <- ggplot(subset(cz3, is.na(cz3$panfolru) == F), aes(x = as.factor(panfolru))) +
geom_bar(fill='#BD0026',color = "#800026", alpha=0.7) +
theme_bw() +
labs(x = "Follow government rules or to make own decisions", y = "Number of people")

grid.arrange(media,restricting,decisions, nrow = 2)

From plots we see that more people think that it is important to maintain public privacy for governments, also they are more satisfied with the government’s handling of the pandemic, but extreme values (9 and 10) we see rarely. For following government rules or making own decisions as we earlier discussed people mostly answer in the middle, however, more people think that it is important to make their own decisions when fighting a pandemic.

Categorical variable

respc19. This categorical variable shows whether the respondent was ill with COVID-19 or not. It has 2 calegories: Yes and No. It the graph and table below we can observe frequency of answers

cz1 <- cz3 %>%
  select(respc19, medcrgvc, panresmo, nwspol, panfolru, gvhanc19)

cz1 <- na.omit(cz1)

ggplot(cz1)+
  geom_bar(aes(x=respc19, fill=respc19))+
  scale_fill_brewer(palette = "YlOrRd")+
  xlab("COVID-19")+
  ylab("Number of people")+theme_bw()

respc19_1 <- data.frame(cz1 %>% 
                          group_by(respc19) %>% 
                          summarise(n=n()))

print(respc19_1)
##   respc19    n
## 1      No 1448
## 2     Yes  652

As we can see, more than half of the respondents did not have COVID-19.

Here we want to examine whether the experience of COVID-19 has influenced satisfaction with the actions of the government. We assume that citizens who were ill and who not, may have a different opinions about the anti-coronavirus policy

ggplot(cz1)+
  geom_boxplot(aes(y=gvhanc19, x=respc19, fill=respc19),alpha=0.5)+
  scale_fill_brewer(palette = "YlOrRd")+
  theme_bw()+
  labs(y="satisfaction with the Government's actions \n in connection with COVID-19",y=NULL,
       x="the respondent has COVID-19", 
       title = "Comparing the level of satisfaction with government \n 
       actions due to whether the respondent had COVID-19")+
  theme(legend.position="none",  plot.title = element_text(face = "bold",hjust = 0.5))

The graph shows that those who have not been sick with coronavirus are more satisfied with the government’s policy to combat coronavirus.

Correlation matrix

Let’s look at the correlations between our outcome variable (gvhanc19 - satisfaction with the government’s handling of the pandemic) and other variables.

cor_mat <- cor(cz3[c("gvhanc19", "nwspol", "panmonpb", "panfolru")], use="pairwise.complete.obs", method="spearman")

corrplot(cor_mat, method="color", addCoef.col = "black", tl.col="black", diag=FALSE, type="lower", number.digits = 2, col = COL1('YlOrRd', 10))

Correlation between the satisfaction with the Government’s actions in connection with COVID-19 and the view that it is important to follow government rules or make your own decisions in the fight against the pandemic is medium and negative (-0.31). For for the opinion that it is more important for government to monitor and track the public or to maintain public privacy when fighting a pandemic and satisfaction with the government’s actions during COVID-19 the correlation is small and negative (-0.13), and for news consumption and satisfaction with the government’s actions during COVID-19 is really small and positive (0.05). So we can say that variable about following government rules or making your own decisions (panfolru) have decent association with satisfaction with the government’s handling of the pandemic and the more people believe that decisions should be made on their own, the less satisfied they are. As for correlations between other variables, the only medium correlation is between view that it is important to follow government rules or make your own decisions in the fight against the pandemic and opinion that it is more important for government to monitor and track the public or to maintain public privacy when fighting a pandemic and it is positive (0.3). In general, this is quite logical since people who prefer to make decisions themselves probably also want the government not to monitor them but to preserve public privacy in the fight against the pandemic.

Our hypothesis about the connection between conformity and satisfaction with the actions of the state has been confirmed: there is a medium effect between them. The hypothesis about the connection between the opinion about the importance of monitoring the public and satisfaction with the actions of the state was not confirmed: the effect was too weak to say that they correlate

We created a scatterplot in order to determine a linear relationship between variables

ggplot(cz1)+
  geom_point(aes(x=panfolru, y=gvhanc19), color="#FEB24C", alpha=0.04, size=9)+
  geom_smooth(aes(x=panfolru, y=gvhanc19), color="#FD8D3C")+
  labs(title="Relationship between importance of following the rules and level of trust", x="Importance of following the rules", y="Satisfaction with government’s handling of COVID-19")+
  theme_bw()+
  theme(plot.title=element_text(size=15, face='bold', hjust=0.5))
## `geom_smooth()` using method = 'gam' and formula = 'y ~ s(x, bs = "cs")'

In the graph we see that relationship is not quite linear.

Linear regression models

Building of the models

So, with our variables of choice (gvhanc19, panfolru and respc19) we build two regression models. First, we build a model (m1) with one continuous predictor (panfolru). Then, we build a model (m2) with two predictors: continuous (panfolru) and categorical (respc19).

ss = cz3 %>% select(gvhanc19, panfolru, respc19)
ss <- drop_na(ss)

m1<-lm(gvhanc19 ~ panfolru, data = ss)
m2<-lm(gvhanc19 ~ panfolru+respc19, data = ss)

Comparison

In order to decide, which model has the best fit, we must 1) look at their determination coefficient (R^2), 2) compare them using the anova() function.

print(summary(m1)$r.squared)
## [1] 0.1138955
summary(m2)$r.squared
## [1] 0.1178138
  1. The determination coefficient of m1 is 0.114, while the determination coefficient of m2 is a little higher at 0.118.
anova(m1,m2)
## Analysis of Variance Table
## 
## Model 1: gvhanc19 ~ panfolru
## Model 2: gvhanc19 ~ panfolru + respc19
##   Res.Df   RSS Df Sum of Sq      F   Pr(>F)   
## 1   2362 12839                                
## 2   2361 12783  1    56.776 10.487 0.001219 **
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1
  1. The p-value is less than 0.5, so the second model is significantly better than the first one. Also we see that second model has less RSS than first model, so this is another confirmation that second model is better

Interpretation

To interpret the model, we shall look to the determination coefficient once more. It is 0.118 for this model, which means that our predictors explain about 12% of the variance of gvhanc19.

Well, this isn’t a really good model. However, it checks out with some theoretical points. While this model isn’t perfect, it’s great that it can explain at least some of the variance that is going on with gvhanc19.

We assume that the conformity of citizens explains satisfaction with the actions of the government only by such a small fraction, because there are still many factors that may be more influential. Such factors may include: retirement age, because during the pandemic the government was generous to them and allocated bonuses, which could improve their opinion of the government; entrepreneurial activity, because the government also conducted an active business assistance program during the pandemic; working status, because the Czech government reacted rather late and poorly to layoffs and “endless” vacations, which could make its actions less attractive to employees; political orientation, because stereotypes about the weakness and chaotic nature of populist leaders who mainly decided the fate of the Czech Republic are quite common; family status, because there are large and single-parent families in the Czech Republic they are most at risk of poverty due to weak government support; access to information, because In the Czech Republic, there were problems informing citizens about the course of COVID-19 and its consequences, which also led to distrust (Greer et al, 2021; Klimovsky, 2021)

Let’s look closer at some of the coefficients of this model’s equation. To do this, it’s nice to look at a pretty table. So let’s build it!

sjPlot::tab_model(m2)
  gvhanc19
Predictors Estimates CI p
(Intercept) 7.04 6.82 – 7.26 <0.001
panfolru -0.29 -0.33 – -0.26 <0.001
respc19 [Yes] -0.34 -0.54 – -0.13 0.001
Observations 2364
R2 / R2 adjusted 0.118 / 0.117

Our p-value is less than 0.05. Cool!

The coefficients we are interested in are estimates of the intercept, panfolru and respc19[yes].

The intercept estimate is 7.04. This means that when the panfolru is equal to 0 (when the person thinks it’s the most important to follow government rules) and when the respc19 is “No” (when the respondent didn’t have COVID-19), the gvhanc19 is equal to 7.04 (so the person is leaning towards being satisfied with the governmental policies during COVID-19).

The panfolru estimate is -0.29. This means that when the panfolru variable is increased by 1, gvhanc19 is decreased by 0.29 (so when the person thinks it’s more important to think for themselves, they are a little less likely to be satisfied with the governmental policies for COVID-19).

The respc19[yes] estimate is equal to -0.34. This means that when the person has had COVID-19 (or thought they had COVID-19), they are less likely to be satisfied with the governmental policies for COVID-19.

Now, let’s write out the equation!

\[ gvhanc19 = 7.04 - 0.29*panfolru - 0.34*respc19[yes] \]

Conclusion

To summarize, let’s look at how things stand with our hypotheses.

  1. Our hypothesis about the consumption of political news and satisfaction with how the government is handling the pandemic has not been confirmed. From the correlation matrix, we can see that the relationship is very weak and positive (0.05).

  2. The hypothesis about the connection between the opinion about the importance of monitoring the public and satisfaction with the actions of the state was not confirmed: the effect was too weak (-0,13).

  3. Our hypothesis about the connection between conformity and satisfaction with the actions of the state has been confirmed: there is a medium effect between them. But regression analysis showed that this relationship explains the variation of outcome by only 12%, even with the added predictor in the form of a variable saying whether the respondent was sick with COVID-19 or not, which made the model more significant. We can relate this to the influence of many other factors.

Thus, we can conclude that of the predictors we have considered, only the conformity of the population can explain satisfaction with the actions of the Government of the Czech Republic - and explain only partially.

List of contributions

  • Zhikharevich Taisia - regression, knitting html,

  • Martianova Svetlana - literature review, plots and descriptive statistics for continuous variables, correlation matrix,

  • Prokhorova Ekaterina - literature review, scatter plot, boxplot, plot and descriptive statistics for categorical variable,

  • Semyonova Ekaterina - literature review, regression, conclusion.

Project 4: Linear regression, interactive model

Theoretical framework

For our model, we use variables from the previous project, but also add a variable related to trust in political parties (trstprt):

Trust in political parties and satisfaction with government’s handling of COVID-19

There have been many studies on political trust during the coronavirus pandemic. One of the articles where the ESS data was also used considers political trust and people’s satisfaction with the Government’s actions to combat the pandemic in Belgium (Gugushvili et. al., 2023). They concluded that political trust is one of the strongest predictors and people with the highest levels of trust are more satisfied with government actions. However, in the article political trust consists of 3 components (questions), in our project we will use only trust in political parties as one of the components of political trust.

Hypothesis: The higher the level of trust in political parties, the greater the satisfaction with government’s handling of COVID-19.

Interaction effect of COVID-19 status and Trust in political parties

Speaking about the relationship between COVID-19 status (whether a person was ill or not) and trust in political parties, one of the studies examines these variables in the context of Germany and the UK (Delhey et. al., 2021). They claim that institutional trust (including political trust) is strengthened due to health-related insecurity. “Health-related insecurity” refers to a positive test result for COVID-19, the presence of symptoms, or if loved ones had symptoms of COVID-19. No literature has been found on the presence of COVID-19 as a moderator of the relationship between trust in political parties and satisfaction. However, since it is assumed that the status of the COVID-19 affects trust (greater trust for those who were ill), as well as satisfaction, but in the opposite direction (less satisfaction for those who were ill) and there is a relationship between trust and satisfaction (presumably direct), it will be interesting to consider the influence of the moderator.

Hypothesis: People with higher trust in political parties have higher level of satisfaction with government’s handling of COVID-19 but satisfaction is less for those who had COVID-19.

The remaining variables and the literature for them are taken from the previous project:

COVID-19 status and satisfaction with government’s handling of COVID-19

We believe that the fact of COVID-19 disease can also become a predictor, because according to Chen et. al. (2021), this is one of the important factors influencing the opinion about government policy during the pandemic in connection with the concept of sick role and worries about the general level of morbidity, worries about loved ones who may to get started. For these reasons, we decided to make the respc19 variable an additional predictor.

Hypothesis: People who had COVID-19 are less likely to be satisfied with government’s handling of COVID-19.

Conformity (governmental rules or people’s own decisions) and satisfaction with government’s handling of COVID-19

During the COVID-19, there was a great public discussion about citizens’ privacy. According to Ioannou, Tussyadiah (2021), citizens who show greater trust in the government tend to be more flexible about restrictions on their privacy for public safety purposes. They can see such measures as a necessary evil in the fight against the pandemic and trust the Government to take appropriate measures.

Hypothesis: those who turned out to be more conformist and believed that it was always more important to follow the decisions of the government were more satisfied with its activities during COVID-19

Reseacrh question

Research question 1: What are the predictors of satisfaction with government’s handling of COVID-19 in Czechia? Research question 2: What factors moderate the relationship between trust in political parties and satisfaction with government’s handling of COVID-19?

Preparing and cleaning the dataset

We recoded several values for ease of analysis: respc19 was reduced to two levels by combining two different “Yes” answers into one, in the remaining variables we recoded the extreme points from the text format to 0 and 10.

cz$respc19 <- as.character(cz$respc19)
cz$respc19[cz$respc19 == "Yes, I tested positive for COVID-19"] <- "Yes"
cz$respc19[cz$respc19 == "Yes, I think I had COVID-19 but was not tested/did not test positive "] <- "Yes"
cz$respc19[cz$respc19 == "No, I have not had COVID-19"] <- "No"
cz$respc19 <- as.factor(cz$respc19)

cz$panfolru <- as.character(cz$panfolru)
cz$panfolru[cz$panfolru == "Much more important to follow government rules"] <- "0"
cz$panfolru[cz$panfolru == "Much more important to make your own decisions"] <- "10"
cz$panfolru <- as.numeric(cz$panfolru)

cz$gvhanc19 <- cz$gvhanc19 %>% str_replace_all("Extremely dissatisfied", "0")
cz$gvhanc19 <- cz$gvhanc19 %>% str_replace_all("Extremely satisfied", "10")
cz$gvhanc19<-as.numeric(cz$gvhanc19)


cz$trstprt<- as.character(cz$trstprt)
cz$trstprt[cz$trstprt=="No trust at all"] <- "0"
cz$trstprt[cz$trstprt=="Complete trust"] <- "10"
cz$trstprt<- as.numeric(cz$trstprt)

cz4 <- cz %>% select(respc19, panfolru, gvhanc19, trstprt)

Variable overview

Let’s first look at the variables used in this project and their meaning, as well as other properties. For gvhanc19 0 is “Extremely dissatisfied with the government’s handling of the coronavirus pandemic”, 10 is “Extremely satisfied”. For panfolru 0 is “Much more important to follow government rules”, 10 is “Much more important to make your own decisions”. For trstprt 0 is “No trust at all”, 10 is “Complete trust”.

Variable = c("trstprt", "gvhanc19", "panfolru", "respc19") 
Meaning = c("Trust in political parties", "How satisfied with government's handling of COVID-19", "More important to follow government rules or to make own decisions when fighting a pandemic", "Whether the respondent was ill with COVID-19")
Type <- c("Quasi-interval", "Quasi-interval", "Quasi-interval", "Binary")
Class <- c(class(cz4$trstprt), class(cz4$gvhanc19), class(cz4$panfolru), class(cz4$respc19))
df1 <- data.frame(Variable, Meaning, Type, Class, stringsAsFactors = FALSE)

kable(df1) %>% 
  kable_styling(bootstrap_options=c("bordered", "responsive","striped"), full_width = FALSE)
Variable Meaning Type Class
trstprt Trust in political parties Quasi-interval numeric
gvhanc19 How satisfied with government’s handling of COVID-19 Quasi-interval numeric
panfolru More important to follow government rules or to make own decisions when fighting a pandemic Quasi-interval numeric
respc19 Whether the respondent was ill with COVID-19 Binary factor

Next, we will look at some descriptive statistics regarding our numeric variables, to give us a better understanding of their distribution.

Mode and median for satisfaction with the government’s handling of the pandemic is 5, mean is 5.24, standard deviation is 2.5. We can see that people occupy a middle position, that is, they are not fully satisfied with how the government is coping, but it cannot be said that they are not satisfied with it at all. For trust in political party, the mean and median are close to 3, and the mode is 2.6. So people mostly don’t trust political parties as much. For following government rules or making own decisions the mean and median is around 6, mode is 5. On average people think that it is important to make their own decisions when fighting a pandemic, but following government rules is also important because answers lay almost in the middle.

cz4 %>%
  summarise_at(vars(c(gvhanc19, trstprt, panfolru)), list(Mean= ~mean(., na.rm = T), Median=~median(., na.rm = T), Mode = ~DescTools::Mode(., na.rm = T), SD=~sd(., na.rm = T), Min=~min(., na.rm = T), Max=~max(., na.rm = T)))%>%
  pivot_longer(cols = gvhanc19_Mean:panfolru_Max, names_to = c("Variable", "Statistics"), names_sep = "_", values_to = "Values")%>%
  pivot_wider(names_from = Variable, values_from = Values)%>% 
  kable(align=c("lrrrrrrrr"), digits=2, row.names = FALSE) %>% 
  kable_styling(bootstrap_options=c("bordered","responsive","striped"), full_width = FALSE)
Statistics gvhanc19 trstprt panfolru
Mean 5.23 3.83 5.75
Median 5.00 4.00 6.00
Mode 5.00 5.00 5.00
SD 2.51 2.60 2.81
Min 0.00 0.00 0.00
Max 10.00 10.00 10.00

From these histograms we can see, for example, that not a lot of respondents chose high values for trusting political parties, and that a lot of respondents chose to answer that it is important to make your own decisions, rather than follow the government rules during the pandemic.

trust <- ggplot(subset(cz4, is.na(cz4$trstprt) == F), aes(x = as.factor(trstprt))) +
geom_bar(fill='#FDD0A2',color = "#000000", alpha=0.7) +
theme_bw() +
labs(x = "Trust in political parties", y = "Number of people")

restricting <- ggplot(subset(cz4, is.na(cz4$gvhanc19) == F), aes(x = as.factor(gvhanc19))) +
geom_bar(fill='#FDAE6B',color = "#000000", alpha=0.7) +
theme_bw() +
labs(x = "How satisfied with government's handling of COVID-19", y = "Number of people")

decisions <- ggplot(subset(cz4, is.na(cz4$panfolru) == F), aes(x = as.factor(panfolru))) +
geom_bar(fill='#FD8D3C',color = "#000000", alpha=0.7) +
theme_bw() +
labs(x = "Follow government rules or to make own decisions", y = "Number of people")

grid.arrange(trust,restricting,decisions, nrow = 2)

Linear regression models

Building models

Let’s build two models: one additive, one interactive.

cz4 <- drop_na(cz4)

m41<-lm(gvhanc19 ~ panfolru+respc19+trstprt, data = cz4)
m42<-lm(gvhanc19 ~ trstprt*respc19+panfolru, data = cz4)

tab_model(m41, m42)
  gvhanc19 gvhanc19
Predictors Estimates CI p Estimates CI p
(Intercept) 6.12 5.86 – 6.38 <0.001 6.27 6.00 – 6.54 <0.001
panfolru -0.28 -0.31 – -0.24 <0.001 -0.27 -0.31 – -0.24 <0.001
respc19 [Yes] -0.42 -0.62 – -0.22 <0.001 -0.98 -1.35 – -0.61 <0.001
trstprt 0.21 0.18 – 0.25 <0.001 0.17 0.13 – 0.21 <0.001
trstprt × respc19 [Yes] 0.14 0.06 – 0.22 <0.001
Observations 2339 2339
R2 / R2 adjusted 0.165 / 0.164 0.169 / 0.168

Looking at this beautiful table, we can say that the models have some differences. For example, additive model’s determination coefficient is 0.164, which means that it explains 16.4% of existing variance, while interactive model’s determination coefficient is 0.168, which means that it explains 16.8% of the variance. Let’s compare and see, which one is better.

Comparison

Comparing the models, we can say that the interactive model fits best (as we can see, the difference between models is significant, as indicated by p-value being drastically less that 0.05, and the interactive model’s RSS parameter is lower, which means it has a better fit).

anova(m41, m42)
## Analysis of Variance Table
## 
## Model 1: gvhanc19 ~ panfolru + respc19 + trstprt
## Model 2: gvhanc19 ~ trstprt * respc19 + panfolru
##   Res.Df   RSS Df Sum of Sq      F    Pr(>F)    
## 1   2335 11965                                  
## 2   2334 11901  1    63.477 12.449 0.0004264 ***
## ---
## Signif. codes:  0 '***' 0.001 '**' 0.01 '*' 0.05 '.' 0.1 ' ' 1

Let’s interpret the interactive model then.

Let’s look closer at its formula:

\[ gvhanc19 = 6.27 - 0.27*panfolru - 0.98*respc19[Yes] + 0.17*trstprt + 0.14*trstprt*respc19[Yes] \]

Looking at the coefficients, we can say that:

  • gvhanc19 is 6.27 when panfolru = 0 (respondent thinks they have to make their own decisions) and respc19 is “No” (respondent doesn’t think they had COVID-19),
  • when panfolru increases by 1 (respondent is closer to thinking they have to make their own decisions), gvhanc19 decreases by 0.27 (their satisfaction with government’s handling of the pandemic decreases),
  • when respc19 is “Yes” (respondent thinks they had COVID-19), gvhanc19 decreases by 0.98 (their satisfaction with government’s handling of the pandemic decreases),
  • when trstprt increases by 1 (the respondent is closer to trusting political parties more), gvhanc19 increases by 0.17 (their satisfaction with government’s handling of the pandemic increases),
  • when trstprt increases by 1 (the respondent is closer to trusting political parties more) while respc19 is “Yes” (respondent thinks they had COVID-19), gvhanc19 increases by 0.14 (their satisfaction with government’s handling of the pandemic increases).

Interaction plots

plot_model(m42, type="int", colors = c("#FDAE6B", "#D94800"))+
  theme_bw()+
  aes(color=group, linetype = group)+
  labs(x="Trust in political parties", y="Satisfaction with government’s \n handling of COVID-19", color="COVID-19 status", linetype="COVID-19 status", title = "Interaction effect of COVID-19 status \n and trust in political parties")+
  theme(plot.title=element_text(face='bold', size=18, hjust = 0,5))

We see that the greater the trust, the higher the level of satisfaction. However, when the level of trust is low and average, people who have been ill with covid rate satisfaction lower than those who have not been ill. When the level of trust in the parties becomes higher than 7, people who have had covid, on the contrary, rate satisfaction with how the government handled the pandemic higher. The difference is not very big, but nevertheless it can be said that people who strongly trust the parties are more satisfied with how the government is coping with the pandemic and the impact of whether they were sick or not is not so great.

Conclusion

Let’s see if our hypotheses have been confirmed:

  1. Our hypothesis about trust in political parties and satisfaction with the actions of the state has been confirmed: when trust increases, satisfaction also increases.

  2. The hypothesis about moderation relationships was partially confirmed. When the trust is low or medium people who had Covid indeed rate lower satisfaction (but it’s still increases with increase in trust. However, when the trust is high, people who had Covid rate their satisfaction with government’s handling of pandemic higher.

  3. The hypothesis about the connection of COVID-19 status and satisfaction has also been confirmed: people who had COVID-19 are less satisfied with government’s handling of pandemic.

  4. The hypothesis about the connection between conformity and satisfaction with the actions of the state has been confirmed. People with lower levels of conformity (closer to think that they have to make their own decisions) have lower satisfaction. Also all our predictors are significant, but even the best model with interaction explains only 17% of the variance in satisfaction.

To put it simply and briefly, people who were ill and had low trust are less satisfied, but more satisfied when they have high trust than those who were not ill. Also, those who tend to make their own decisions rather than rely on the government are less satisfied with their actions.

List of contributions

  • Prokhorova Ekaterina - descriptive statistics,
  • Semyonova Ekaterina - comparing the model fit,
  • Martianova Svetlana - theoretical framework, interaction plots,
  • Zhikharevich Taisiia - interactive model, knitting html.

Conclusion & acknowledgements

Most Czech citizens have an average level of satisfaction with the government’s actions during the pandemic, while people with a high level of satisfaction trusted politicians more. This satisfaction was influenced by the fact that the respondent had the disease (but this requires further research and verification, since not all of our tests showed this result), as well as the conformity of the respondent (those who believed that during a pandemic it is worth following their own decisions, rather than following the decisions of the government). We also noticed that there is a statistically significant relationship between COVID-19 disease and layoffs, which is also a result of the decisions of the government in the field of labor, so we believe that these facts require more detailed consideration.

Our other focus was the political preferences of Czech residents. We found out that the choice of the party was often influenced by age - leftists and populists in general were older than centrists and liberals. We also consider it important to consider in more detail the situation with the change in political views after many people had COVID-19 - our data showed that the ANO party was less popular with those who had COVID-19, but this requires further confirmation.

We would like to thank our teachers and assistants that were helping us on this tough journey through the world of data analysis: Tatiana Tkacheva, Elina Tsigeman, Ekaterina Titova, Taisiia Sludzskaya, Ustinya Goryachkina, Valeriia Rusak, as well as the creators of this course, Olesya Volchenko and Anna Shirokanova.

References

Chen, C. W. S., Lee, S., Dong, M. C., & Taniguchi, M. (2021). What factors drive the satisfaction of citizens with governments’ responses to COVID-19? International Journal of Infectious Diseases, 102, 327–331. doi:10.1016/j.ijid.2020.10.050

Delhey, J., Steckermeier, L. C., Boehnke, K., Deutsch, F., Eichhorn, J., Kühnen, U., & Welzel, C. (2021). A virus of distrust. Existential insecurity and trust during the Coronavirus pandemic, 80.

Freedom House (2020) Freedom in Czechia https://freedomhouse.org/country/czechia/freedom-world/2020

Greer, S.L. et al. (2021) Coronavirus politics: The comparative politics and policy of covid-19. University of Michigan Press.

Gugushvili, D., Spruit, D., Van de Walle, S., Baudewyns, P., & Meuleman, B. (2023). How satisfied are Belgians with the government’s handling of the COVID-19 pandemic? Evidence from the European Social Survey. How satisfied are Belgians with the government’s handling of the COVID-19 pandemic? Evidence from the European Social Survey.

Havlík, V., & Kluknavská, A. (2022). The populist vs anti‐populist divide in the time of pandemic: The 2021 Czech national election and its consequences for European politics. JCMS: Journal of Common Market Studies, 60, 76-87.

Hedvičáková, M., & Kozubíková, Z. (2021). Impacts of COVID-19 on the labour market - evidence from the Czech Republic. Hradec Economic Days. https://doi.org/10.36689/uhk/hed/2021-01-023

Holland, J. L. (2013). Age gap? The influence of age on voting behavior and political preferences in the American electorate. Washington State University.

Idnes.cz (2020) Vláda omezila volný pohyb do 1. dubna, brzdí s EET, odpouští i odvody. https://www.idnes.cz/zpravy/domaci/jednani-vlady-koronavirus-pomoc-podnikatelum-zakaz-pohybu-pendleri.A200323_051446_domaci_kop

Ioannou, A., & Tussyadiah, I. (2021). Privacy and surveillance attitudes during health crises: Acceptance of surveillance and privacy protection behaviours. Technology in Society, 67, 101774. https://doi.org/10.1016/j.techsoc.2021.101774

Jiang, A., & Zhang, T. H. (2021). Political freedom, news consumption, and patterns of political trust: evidence from East and Southeast Asia, 2001-2016. Political Science, 73(3), 250-269.

Klimovský, D., Nemec, J., & Bouckaert, G. (2021). The COVID-19 pandemic in the Czech Republic and Slovakia. Scientific Papers of the University of Pardubice-Series D, Faculty of Economics and Administration, volume 29, issue: 1.

Lipscy, P. Y. (2020). COVID-19 and the Politics of Crisis. International Organization, 74(S1), E98–E127. doi:10.1017/S0020818320000375

Simko, L., Chang, J., Jiang, M., Calo, R., Roesner, F., & Kohno, T. (2022). COVID-19 Contact Tracing and Privacy: A Longitudinal Study of Public opinion. Digital Threats, 3(3), 1–36. https://doi.org/10.1145/3480464

Wu, C., Shi, Z., Wilkes, R., Wu, J., Gong, Z., He, N., … & Nicola Giordano, G. (2021). Chinese citizen satisfaction with government performance during COVID-19. Journal of Contemporary China, 30(132), 930-944.