It is completely normal to feel overwhelmed when starting something new, especially coding. The process is a different way of thinking, but it is also very logical and something you can absolutely learn. We are going to break down the foundational concepts of R into small, clear, manageable pieces.
Think of the R code boxes in this assignment as a direct line to a powerful data processor. You give R a precise command, click Run, and R gives you a precise result back.
This assignment combines two important goals:
You do not need to install R or RStudio on your own computer for this assignment. You will use RStudio through the course Posit Cloud workspace.
This is an R Markdown file. It contains regular written instructions and gray R code boxes called code chunks.
Action Item: To run a code chunk, click the small
green play button on the far right of the code chunk,
or put your cursor inside the code chunk and press
Cmd/Ctrl + Enter.
The data files for this assignment are already included in the Posit Cloud project. You do not need to download the Arbuthnot or present-day birth records from another website.
In this course, R is not just something we use to get answers from a computer.
We will use R the way applied researchers, data analysts, and statisticians use it: to organize data, calculate summaries, create graphs, check data, run statistical methods, and write conclusions.
Before we can do that, you need to understand the basic building blocks of R:
$ operatorThis first assignment is about becoming comfortable with those building blocks.
Let’s start with the most basic command.
R is a powerful calculator. If you give it a math problem, it will solve it.
Action Item: Run the code chunk below.
# R can perform basic arithmetic.
8 + 3
## [1] 11
log(2)
## [1] 0.6931472
((131 * 3) / (6^3)) + (1 / pi)
## [1] 2.137754
This is the core interaction: you provide a command, and R executes it.
The # symbol marks off text as a
comment. R does not run comments as code. Comments are
useful because they let you leave notes to yourself inside your
code.
This is one of the most important concepts in coding.
Very often, you will want to save a value or a result so you can use it again.
To do this, you create an object, which is also commonly called a variable.
Think of an object as a labeled storage box.
The arrow, <-, is the assignment
operator. It is a command that means:
“Put the value on the right into a box with the label on the left.”
Let’s examine this in detail.
# Create x and y.
x <- 8 + 3
y <- log(2)
# View the values stored in x and y.
x
## [1] 11
y
## [1] 0.6931472
Here is the precise instruction you are giving to R when you run
x <- 8 + 3:
8 + 3.x.11 inside the object named
x.Now, the R environment remembers that the object named x
contains the value 11.
The power of this is that you can now reference the object in your commands.
# We can use objects in calculations.
x + y
## [1] 11.69315
R processes this by thinking:
“First, I need to resolve the objects named x and
y. Then I can add their stored values.”
You can also create a new object based on existing objects.
# Define and return z.
z <- x * y
z
## [1] 7.624619
Important Note: R is case-sensitive. This means
x and X are not the same. If you try to run a
name that has not been defined with the exact capitalization, R will
return an error.
If you use the same object name again, R will overwrite the previous value.
Run the code below.
# A semicolon can be used to separate commands on the same line.
x <- 8 + 3; x
## [1] 11
x <- 21; x
## [1] 21
In this example, the object x first stores the value 11.
Then x is redefined to store the value 21.
For now, the main lesson is simple: object names matter. If you reuse an object name, R replaces the old value with the new value.
R comes with many built-in commands called functions.
A function is a specialized tool that performs a specific job.
Think of a function as a verb in the R language. It performs an action.
Some common functions include:
log(): Calculates a natural logarithm.mean(): Calculates the average of a set of
numbers.sd(): Calculates the standard deviation of a set of
numbers.plot(): Creates a plot.To use a function, you must follow a specific and consistent syntax:
function_name(argument)
There are three main parts:
function_name): This is
the exact name of the tool you want to use.(): These are required.
The parentheses signal to R that you are executing a function.argument): This is the input
the function needs to do its work. You place the input inside the
parentheses.Let’s see this with the log() function.
# We use the log() function with the argument 2.
log(2)
## [1] 0.6931472
R recognizes log as a function, sees the value
2 inside the parentheses, and returns the natural logarithm
of 2.
So far, we have stored one piece of information inside an object.
But what if we have several values?
A vector is a special type of object that stores multiple values in a single object.
Think of a vector like an ice cube tray:
To create a vector in R, we use the c() function.
The c stands for combine or
concatenate.
The c() function takes individual items and combines
them into one vector.
Run the code below.
# Define and return vectors a and b.
a <- c(4.1, 6.7, 8.2, 1.8)
b <- 2 * a
a
## [1] 4.1 6.7 8.2 1.8
b
## [1] 8.2 13.4 16.4 3.6
In this example:
a stores four numbers.b stores values that are twice the values in
a.R can perform calculations on whole vectors. That is one reason R is so useful for data analysis.
We can use functions to summarize vectors.
Run the code below.
# Calculate the mean and standard deviation of a.
mean(a)
## [1] 5.2
sd(a)
## [1] 2.829605
We can also create a simple plot.
# Plot a against b.
plot(a, b)
The plot should appear in the Plots pane in the lower-right portion of RStudio.
R also has help pages. To look up what a function does, type a question mark followed by the function name.
# Help for the mean() function.
?mean
This code chunk has eval=FALSE, so it will show the
command but will not run automatically when you knit the document.
$
OperatorIn scientific analysis, we rarely work with only one number at a time.
We usually work with datasets.
In R, datasets are often stored in a data frame.
A data frame is a structure that represents a table or spreadsheet.
It has:
For example, in a patient dataset:
Now we will use a real historical dataset.
The arbuthnot data frame contains Dr. John Arbuthnot’s
baptism records.
Dr. Arbuthnot was an 18th century physician, writer, and mathematician. He was interested in the ratio of newborn boys to newborn girls, so he gathered baptism records for children born in London for every year from 1629 to 1710.
The data file is already included in this Posit Cloud project. The hidden setup chunk loaded it for you.
Run the code below to view the data.
# View the Arbuthnot data frame.
arbuthnot
## year boys girls
## 1 1629 5218 4683
## 2 1630 4858 4457
## 3 1631 4422 4102
## 4 1632 4994 4590
## 5 1633 5158 4839
## 6 1634 5035 4820
## 7 1635 5106 4928
## 8 1636 4917 4605
## 9 1637 4703 4457
## 10 1638 5359 4952
## 11 1639 5366 4784
## 12 1640 5518 5332
## 13 1641 5470 5200
## 14 1642 5460 4910
## 15 1643 4793 4617
## 16 1644 4107 3997
## 17 1645 4047 3919
## 18 1646 3768 3395
## 19 1647 3796 3536
## 20 1648 3363 3181
## 21 1649 3079 2746
## 22 1650 2890 2722
## 23 1651 3231 2840
## 24 1652 3220 2908
## 25 1653 3196 2959
## 26 1654 3441 3179
## 27 1655 3655 3349
## 28 1656 3668 3382
## 29 1657 3396 3289
## 30 1658 3157 3013
## 31 1659 3209 2781
## 32 1660 3724 3247
## 33 1661 4748 4107
## 34 1662 5216 4803
## 35 1663 5411 4881
## 36 1664 6041 5681
## 37 1665 5114 4858
## 38 1666 4678 4319
## 39 1667 5616 5322
## 40 1668 6073 5560
## 41 1669 6506 5829
## 42 1670 6278 5719
## 43 1671 6449 6061
## 44 1672 6443 6120
## 45 1673 6073 5822
## 46 1674 6113 5738
## 47 1675 6058 5717
## 48 1676 6552 5847
## 49 1677 6423 6203
## 50 1678 6568 6033
## 51 1679 6247 6041
## 52 1680 6548 6299
## 53 1681 6822 6533
## 54 1682 6909 6744
## 55 1683 7577 7158
## 56 1684 7575 7127
## 57 1685 7484 7246
## 58 1686 7575 7119
## 59 1687 7737 7214
## 60 1688 7487 7101
## 61 1689 7604 7167
## 62 1690 7909 7302
## 63 1691 7662 7392
## 64 1692 7602 7316
## 65 1693 7676 7483
## 66 1694 6985 6647
## 67 1695 7263 6713
## 68 1696 7632 7229
## 69 1697 8062 7767
## 70 1698 8426 7626
## 71 1699 7911 7452
## 72 1700 7578 7061
## 73 1701 8102 7514
## 74 1702 8031 7656
## 75 1703 7765 7683
## 76 1704 6113 5738
## 77 1705 8366 7779
## 78 1706 7952 7417
## 79 1707 8379 7687
## 80 1708 8239 7623
## 81 1709 7840 7380
## 82 1710 7640 7288
You should see three columns:
yearboysgirlsEach row represents a different year.
The row numbers on the far left are not part of Arbuthnot’s data. R adds those as an index to help you keep track of rows.
You can see the dimensions of a data frame by using
dim().
# Check the dimensions of the Arbuthnot data frame.
dim(arbuthnot)
## [1] 82 3
This output tells you the number of rows and columns.
You can see the names of the columns by using
names().
# Check the column names of the Arbuthnot data frame.
names(arbuthnot)
## [1] "year" "boys" "girls"
The dim() and names() commands are both
functions. Each one takes the name of the data frame as its
argument.
$Now suppose we want to look at only one column from the data frame.
To do this, we use the dollar sign $.
Think of $ as an extractor tool.
It selects one column from a data frame.
The syntax is:
DataFrameName$ColumnName
Here is an example.
# This command extracts just the boys column from the Arbuthnot data frame.
arbuthnot$boys
## [1] 5218 4858 4422 4994 5158 5035 5106 4917 4703 5359 5366 5518 5470 5460 4793
## [16] 4107 4047 3768 3796 3363 3079 2890 3231 3220 3196 3441 3655 3668 3396 3157
## [31] 3209 3724 4748 5216 5411 6041 5114 4678 5616 6073 6506 6278 6449 6443 6073
## [46] 6113 6058 6552 6423 6568 6247 6548 6822 6909 7577 7575 7484 7575 7737 7487
## [61] 7604 7909 7662 7602 7676 6985 7263 7632 8062 8426 7911 7578 8102 8031 7765
## [76] 6113 8366 7952 8379 8239 7840 7640
The command arbuthnot shows the whole data frame.
The command arbuthnot$boys shows only the
boys column.
What command would you use to extract just the counts of girls baptized?
Write and run your code in the chunk below.
#`arbuthnot$girls`
arbuthnot$girls
## [1] 4683 4457 4102 4590 4839 4820 4928 4605 4457 4952 4784 5332 5200 4910 4617
## [16] 3997 3919 3395 3536 3181 2746 2722 2840 2908 2959 3179 3349 3382 3289 3013
## [31] 2781 3247 4107 4803 4881 5681 4858 4319 5322 5560 5829 5719 6061 6120 5822
## [46] 5738 5717 5847 6203 6033 6041 6299 6533 6744 7158 7127 7246 7119 7214 7101
## [61] 7167 7302 7392 7316 7483 6647 6713 7229 7767 7626 7452 7061 7514 7656 7683
## [76] 5738 7779 7417 7687 7623 7380 7288
What is the third entry in the vector of counts of girls baptized? 4102 Use your output from Exercise 1 to answer this question in the matching eLC quiz.
We can use functions such as mean() and
sd() on columns from a data frame.
For example, we can find the mean and standard deviation of the number of boys baptized yearly between 1629 and 1710.
# Mean and standard deviation of the boys column.
mean(arbuthnot$boys)
## [1] 5907.098
sd(arbuthnot$boys)
## [1] 1652.754
What is the mean number of girls baptized yearly between 1629 and 1710?
Write and run your code in the chunk below.
# Mean and standard deviation of the girls column.
mean(arbuthnot$girls)
## [1] 5534.646
sd(arbuthnot$girls)
## [1] 1592.137
What is the standard deviation of the number of girls baptized yearly between 1629 and 1710? 1592.137 Write and run your code in the chunk below.
# Standard deviation of the girls column.
sd(arbuthnot$girls)
## [1] 1592.137
Use your output from Exercises 3 and 4 to answer the matching eLC quiz questions.
R has powerful functions for making graphics.
We can create a simple plot of the number of girls baptized each year.
# Plot the number of girls baptized each year.
plot(x = arbuthnot$year, y = arbuthnot$girls)
By default, R creates a scatterplot with each point shown separately.
If we want to connect the data points with lines, we can add a third
argument: type = "l".
# Plot the number of girls baptized each year and connect the points with lines.
plot(x = arbuthnot$year, y = arbuthnot$girls, type = "l")
Is there an apparent trend in the number of girls baptized over the years?
You do not need to write new code for this question. Use the plot you created above to answer this question in the matching eLC quiz.
Now suppose we want to find the total number of baptisms each year.
We could calculate this one year at a time, but there is a faster way.
If we add the vector of boys and the vector of girls, R computes all of the sums at once.
# Add the boys and girls columns.
arbuthnot$boys + arbuthnot$girls
## [1] 9901 9315 8524 9584 9997 9855 10034 9522 9160 10311 10150 10850
## [13] 10670 10370 9410 8104 7966 7163 7332 6544 5825 5612 6071 6128
## [25] 6155 6620 7004 7050 6685 6170 5990 6971 8855 10019 10292 11722
## [37] 9972 8997 10938 11633 12335 11997 12510 12563 11895 11851 11775 12399
## [49] 12626 12601 12288 12847 13355 13653 14735 14702 14730 14694 14951 14588
## [61] 14771 15211 15054 14918 15159 13632 13976 14861 15829 16052 15363 14639
## [73] 15616 15687 15448 11851 16145 15369 16066 15862 15220 14928
We can also plot the total number of baptisms over time.
# Plot total baptisms over time.
plot(arbuthnot$year, arbuthnot$boys + arbuthnot$girls, type = "l")
Next, we can compute the ratio of boys to girls in 1629.
# Ratio of boys to girls in 1629.
5218 / 4683
## [1] 1.114243
We can also compute the ratio for every year.
# Ratio of boys to girls for every year.
arbuthnot$boys / arbuthnot$girls
## [1] 1.114243 1.089971 1.078011 1.088017 1.065923 1.044606 1.036120 1.067752
## [9] 1.055194 1.082189 1.121656 1.034884 1.051923 1.112016 1.038120 1.027521
## [17] 1.032661 1.109867 1.073529 1.057215 1.121267 1.061719 1.137676 1.107290
## [25] 1.080095 1.082416 1.091371 1.084565 1.032533 1.047793 1.153901 1.146905
## [33] 1.156075 1.085988 1.108584 1.063369 1.052697 1.083121 1.055242 1.092266
## [41] 1.116143 1.097744 1.064016 1.052778 1.043112 1.065354 1.059647 1.120575
## [49] 1.035467 1.088679 1.034100 1.039530 1.044237 1.024466 1.058536 1.062860
## [57] 1.032846 1.064054 1.072498 1.054359 1.060974 1.083128 1.036526 1.039092
## [65] 1.025792 1.050850 1.081931 1.055748 1.037981 1.104904 1.061594 1.073219
## [73] 1.078254 1.048981 1.010673 1.065354 1.075460 1.072132 1.090022 1.080808
## [81] 1.062331 1.048299
The proportion of baptisms that were boys in 1629 is:
# Proportion of baptisms that were boys in 1629.
5218 / (5218 + 4683)
## [1] 0.5270175
The parentheses matter. We want to divide the number of boys by the total number of baptisms.
We can also compute the proportion of boys for all years at once.
# Proportion of baptisms that were boys for every year.
arbuthnot$boys / (arbuthnot$boys + arbuthnot$girls)
## [1] 0.5270175 0.5215244 0.5187705 0.5210768 0.5159548 0.5109082 0.5088698
## [8] 0.5163831 0.5134279 0.5197362 0.5286700 0.5085714 0.5126523 0.5265188
## [15] 0.5093518 0.5067868 0.5080341 0.5260366 0.5177305 0.5139059 0.5285837
## [22] 0.5149679 0.5322023 0.5254569 0.5192526 0.5197885 0.5218447 0.5202837
## [29] 0.5080030 0.5116694 0.5357262 0.5342132 0.5361942 0.5206108 0.5257482
## [36] 0.5153557 0.5128359 0.5199511 0.5134394 0.5220493 0.5274422 0.5232975
## [43] 0.5155076 0.5128552 0.5105507 0.5158214 0.5144798 0.5284297 0.5087122
## [50] 0.5212285 0.5083822 0.5096910 0.5108199 0.5060426 0.5142178 0.5152360
## [57] 0.5080788 0.5155165 0.5174905 0.5132301 0.5147925 0.5199527 0.5089677
## [64] 0.5095857 0.5063659 0.5123973 0.5196766 0.5135590 0.5093183 0.5249190
## [71] 0.5149385 0.5176583 0.5188268 0.5119526 0.5026541 0.5158214 0.5181790
## [78] 0.5174052 0.5215362 0.5194175 0.5151117 0.5117899
Make a plot of the proportion of boys over time.
Write and run your code in the chunk below.
# Proportion of baptisms that were boys over time.
arbuthnot$boys / (arbuthnot$boys + arbuthnot$girls)
## [1] 0.5270175 0.5215244 0.5187705 0.5210768 0.5159548 0.5109082 0.5088698
## [8] 0.5163831 0.5134279 0.5197362 0.5286700 0.5085714 0.5126523 0.5265188
## [15] 0.5093518 0.5067868 0.5080341 0.5260366 0.5177305 0.5139059 0.5285837
## [22] 0.5149679 0.5322023 0.5254569 0.5192526 0.5197885 0.5218447 0.5202837
## [29] 0.5080030 0.5116694 0.5357262 0.5342132 0.5361942 0.5206108 0.5257482
## [36] 0.5153557 0.5128359 0.5199511 0.5134394 0.5220493 0.5274422 0.5232975
## [43] 0.5155076 0.5128552 0.5105507 0.5158214 0.5144798 0.5284297 0.5087122
## [50] 0.5212285 0.5083822 0.5096910 0.5108199 0.5060426 0.5142178 0.5152360
## [57] 0.5080788 0.5155165 0.5174905 0.5132301 0.5147925 0.5199527 0.5089677
## [64] 0.5095857 0.5063659 0.5123973 0.5196766 0.5135590 0.5093183 0.5249190
## [71] 0.5149385 0.5176583 0.5188268 0.5119526 0.5026541 0.5158214 0.5181790
## [78] 0.5174052 0.5215362 0.5194175 0.5151117 0.5117899
What do you see? Use your plot to answer the matching eLC quiz question.
We can add a new variable to a data frame.
The code below creates a new variable called total in
the arbuthnot data frame.
# Add a total variable to the Arbuthnot data frame.
arbuthnot$total <- arbuthnot$boys + arbuthnot$girls
# View the updated data frame.
arbuthnot
## year boys girls total
## 1 1629 5218 4683 9901
## 2 1630 4858 4457 9315
## 3 1631 4422 4102 8524
## 4 1632 4994 4590 9584
## 5 1633 5158 4839 9997
## 6 1634 5035 4820 9855
## 7 1635 5106 4928 10034
## 8 1636 4917 4605 9522
## 9 1637 4703 4457 9160
## 10 1638 5359 4952 10311
## 11 1639 5366 4784 10150
## 12 1640 5518 5332 10850
## 13 1641 5470 5200 10670
## 14 1642 5460 4910 10370
## 15 1643 4793 4617 9410
## 16 1644 4107 3997 8104
## 17 1645 4047 3919 7966
## 18 1646 3768 3395 7163
## 19 1647 3796 3536 7332
## 20 1648 3363 3181 6544
## 21 1649 3079 2746 5825
## 22 1650 2890 2722 5612
## 23 1651 3231 2840 6071
## 24 1652 3220 2908 6128
## 25 1653 3196 2959 6155
## 26 1654 3441 3179 6620
## 27 1655 3655 3349 7004
## 28 1656 3668 3382 7050
## 29 1657 3396 3289 6685
## 30 1658 3157 3013 6170
## 31 1659 3209 2781 5990
## 32 1660 3724 3247 6971
## 33 1661 4748 4107 8855
## 34 1662 5216 4803 10019
## 35 1663 5411 4881 10292
## 36 1664 6041 5681 11722
## 37 1665 5114 4858 9972
## 38 1666 4678 4319 8997
## 39 1667 5616 5322 10938
## 40 1668 6073 5560 11633
## 41 1669 6506 5829 12335
## 42 1670 6278 5719 11997
## 43 1671 6449 6061 12510
## 44 1672 6443 6120 12563
## 45 1673 6073 5822 11895
## 46 1674 6113 5738 11851
## 47 1675 6058 5717 11775
## 48 1676 6552 5847 12399
## 49 1677 6423 6203 12626
## 50 1678 6568 6033 12601
## 51 1679 6247 6041 12288
## 52 1680 6548 6299 12847
## 53 1681 6822 6533 13355
## 54 1682 6909 6744 13653
## 55 1683 7577 7158 14735
## 56 1684 7575 7127 14702
## 57 1685 7484 7246 14730
## 58 1686 7575 7119 14694
## 59 1687 7737 7214 14951
## 60 1688 7487 7101 14588
## 61 1689 7604 7167 14771
## 62 1690 7909 7302 15211
## 63 1691 7662 7392 15054
## 64 1692 7602 7316 14918
## 65 1693 7676 7483 15159
## 66 1694 6985 6647 13632
## 67 1695 7263 6713 13976
## 68 1696 7632 7229 14861
## 69 1697 8062 7767 15829
## 70 1698 8426 7626 16052
## 71 1699 7911 7452 15363
## 72 1700 7578 7061 14639
## 73 1701 8102 7514 15616
## 74 1702 8031 7656 15687
## 75 1703 7765 7683 15448
## 76 1704 6113 5738 11851
## 77 1705 8366 7779 16145
## 78 1706 7952 7417 15369
## 79 1707 8379 7687 16066
## 80 1708 8239 7623 15862
## 81 1709 7840 7380 15220
## 82 1710 7640 7288 14928
The data frame should now include a new column called
total.
In addition to simple mathematical operators such as addition, subtraction, multiplication, and division, we can ask R to make comparisons.
For example, we can ask whether the number of boys baptized was greater than the number of girls baptized in each year.
# Ask whether boys outnumbered girls in each year.
arbuthnot$boys > arbuthnot$girls
## [1] TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE
## [16] TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE
## [31] TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE
## [46] TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE
## [61] TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE TRUE
## [76] TRUE TRUE TRUE TRUE TRUE TRUE TRUE
This command returns logical values:
TRUE if the statement is true for that year,FALSE if the statement is not true for that year.We can also sort a data frame based on the values of one variable.
# Sort the Arbuthnot data frame by the number of boys.
arbuthnot[order(arbuthnot$boys), ]
## year boys girls total
## 22 1650 2890 2722 5612
## 21 1649 3079 2746 5825
## 30 1658 3157 3013 6170
## 25 1653 3196 2959 6155
## 31 1659 3209 2781 5990
## 24 1652 3220 2908 6128
## 23 1651 3231 2840 6071
## 20 1648 3363 3181 6544
## 29 1657 3396 3289 6685
## 26 1654 3441 3179 6620
## 27 1655 3655 3349 7004
## 28 1656 3668 3382 7050
## 32 1660 3724 3247 6971
## 18 1646 3768 3395 7163
## 19 1647 3796 3536 7332
## 17 1645 4047 3919 7966
## 16 1644 4107 3997 8104
## 3 1631 4422 4102 8524
## 38 1666 4678 4319 8997
## 9 1637 4703 4457 9160
## 33 1661 4748 4107 8855
## 15 1643 4793 4617 9410
## 2 1630 4858 4457 9315
## 8 1636 4917 4605 9522
## 4 1632 4994 4590 9584
## 6 1634 5035 4820 9855
## 7 1635 5106 4928 10034
## 37 1665 5114 4858 9972
## 5 1633 5158 4839 9997
## 34 1662 5216 4803 10019
## 1 1629 5218 4683 9901
## 10 1638 5359 4952 10311
## 11 1639 5366 4784 10150
## 35 1663 5411 4881 10292
## 14 1642 5460 4910 10370
## 13 1641 5470 5200 10670
## 12 1640 5518 5332 10850
## 39 1667 5616 5322 10938
## 36 1664 6041 5681 11722
## 47 1675 6058 5717 11775
## 40 1668 6073 5560 11633
## 45 1673 6073 5822 11895
## 46 1674 6113 5738 11851
## 76 1704 6113 5738 11851
## 51 1679 6247 6041 12288
## 42 1670 6278 5719 11997
## 49 1677 6423 6203 12626
## 44 1672 6443 6120 12563
## 43 1671 6449 6061 12510
## 41 1669 6506 5829 12335
## 52 1680 6548 6299 12847
## 48 1676 6552 5847 12399
## 50 1678 6568 6033 12601
## 53 1681 6822 6533 13355
## 54 1682 6909 6744 13653
## 66 1694 6985 6647 13632
## 67 1695 7263 6713 13976
## 57 1685 7484 7246 14730
## 60 1688 7487 7101 14588
## 56 1684 7575 7127 14702
## 58 1686 7575 7119 14694
## 55 1683 7577 7158 14735
## 72 1700 7578 7061 14639
## 64 1692 7602 7316 14918
## 61 1689 7604 7167 14771
## 68 1696 7632 7229 14861
## 82 1710 7640 7288 14928
## 63 1691 7662 7392 15054
## 65 1693 7676 7483 15159
## 59 1687 7737 7214 14951
## 75 1703 7765 7683 15448
## 81 1709 7840 7380 15220
## 62 1690 7909 7302 15211
## 71 1699 7911 7452 15363
## 78 1706 7952 7417 15369
## 74 1702 8031 7656 15687
## 69 1697 8062 7767 15829
## 73 1701 8102 7514 15616
## 80 1708 8239 7623 15862
## 77 1705 8366 7779 16145
## 79 1707 8379 7687 16066
## 70 1698 8426 7626 16052
In the previous sections, you recreated some displays and preliminary analyses of Arbuthnot’s baptism data.
Now you will repeat some of those steps using present-day birth records from the United States.
The data are stored in a data frame called present.
The present.csv file is already included in the Posit
Cloud project. The hidden setup chunk loaded it for you.
Run the code below to view the data.
# View the present-day birth records.
present
## year boys girls
## 1 1940 1211684 1148715
## 2 1941 1289734 1223693
## 3 1942 1444365 1364631
## 4 1943 1508959 1427901
## 5 1944 1435301 1359499
## 6 1945 1404587 1330869
## 7 1946 1691220 1597452
## 8 1947 1899876 1800064
## 9 1948 1813852 1721216
## 10 1949 1826352 1733177
## 11 1950 1823555 1730594
## 12 1951 1923020 1827830
## 13 1952 1971262 1875724
## 14 1953 2001798 1900322
## 15 1954 2059068 1958294
## 16 1955 2073719 1973576
## 17 1956 2133588 2029502
## 18 1957 2179960 2074824
## 19 1958 2152546 2051266
## 20 1959 2173638 2071158
## 21 1960 2179708 2078142
## 22 1961 2186274 2082052
## 23 1962 2132466 2034896
## 24 1963 2101632 1996388
## 25 1964 2060162 1967328
## 26 1965 1927054 1833304
## 27 1966 1845862 1760412
## 28 1967 1803388 1717571
## 29 1968 1796326 1705238
## 30 1969 1846572 1753634
## 31 1970 1915378 1816008
## 32 1971 1822910 1733060
## 33 1972 1669927 1588484
## 34 1973 1608326 1528639
## 35 1974 1622114 1537844
## 36 1975 1613135 1531063
## 37 1976 1624436 1543352
## 38 1977 1705916 1620716
## 39 1978 1709394 1623885
## 40 1979 1791267 1703131
## 41 1980 1852616 1759642
## 42 1981 1860272 1768966
## 43 1982 1885676 1794861
## 44 1983 1865553 1773380
## 45 1984 1879490 1789651
## 46 1985 1927983 1832578
## 47 1986 1924868 1831679
## 48 1987 1951153 1858241
## 49 1988 2002424 1907086
## 50 1989 2069490 1971468
## 51 1990 2129495 2028717
## 52 1991 2101518 2009389
## 53 1992 2082097 1982917
## 54 1993 2048861 1951379
## 55 1994 2022589 1930178
## 56 1995 1996355 1903234
## 57 1996 1990480 1901014
## 58 1997 1985596 1895298
## 59 1998 2016205 1925348
## 60 1999 2026854 1932563
## 61 2000 2076969 1981845
## 62 2001 2057922 1968011
## 63 2002 2057979 1963747
What years are included in the present data set?
1940-1949 Write and run code that helps you answer this question.
# View the years included in the present-day birth records.
present
## year boys girls
## 1 1940 1211684 1148715
## 2 1941 1289734 1223693
## 3 1942 1444365 1364631
## 4 1943 1508959 1427901
## 5 1944 1435301 1359499
## 6 1945 1404587 1330869
## 7 1946 1691220 1597452
## 8 1947 1899876 1800064
## 9 1948 1813852 1721216
## 10 1949 1826352 1733177
## 11 1950 1823555 1730594
## 12 1951 1923020 1827830
## 13 1952 1971262 1875724
## 14 1953 2001798 1900322
## 15 1954 2059068 1958294
## 16 1955 2073719 1973576
## 17 1956 2133588 2029502
## 18 1957 2179960 2074824
## 19 1958 2152546 2051266
## 20 1959 2173638 2071158
## 21 1960 2179708 2078142
## 22 1961 2186274 2082052
## 23 1962 2132466 2034896
## 24 1963 2101632 1996388
## 25 1964 2060162 1967328
## 26 1965 1927054 1833304
## 27 1966 1845862 1760412
## 28 1967 1803388 1717571
## 29 1968 1796326 1705238
## 30 1969 1846572 1753634
## 31 1970 1915378 1816008
## 32 1971 1822910 1733060
## 33 1972 1669927 1588484
## 34 1973 1608326 1528639
## 35 1974 1622114 1537844
## 36 1975 1613135 1531063
## 37 1976 1624436 1543352
## 38 1977 1705916 1620716
## 39 1978 1709394 1623885
## 40 1979 1791267 1703131
## 41 1980 1852616 1759642
## 42 1981 1860272 1768966
## 43 1982 1885676 1794861
## 44 1983 1865553 1773380
## 45 1984 1879490 1789651
## 46 1985 1927983 1832578
## 47 1986 1924868 1831679
## 48 1987 1951153 1858241
## 49 1988 2002424 1907086
## 50 1989 2069490 1971468
## 51 1990 2129495 2028717
## 52 1991 2101518 2009389
## 53 1992 2082097 1982917
## 54 1993 2048861 1951379
## 55 1994 2022589 1930178
## 56 1995 1996355 1903234
## 57 1996 1990480 1901014
## 58 1997 1985596 1895298
## 59 1998 2016205 1925348
## 60 1999 2026854 1932563
## 61 2000 2076969 1981845
## 62 2001 2057922 1968011
## 63 2002 2057979 1963747
What are the dimensions of the present data frame?
Write and run your code in the chunk below.
# View the dimensions of the present data frame.
present
## year boys girls
## 1 1940 1211684 1148715
## 2 1941 1289734 1223693
## 3 1942 1444365 1364631
## 4 1943 1508959 1427901
## 5 1944 1435301 1359499
## 6 1945 1404587 1330869
## 7 1946 1691220 1597452
## 8 1947 1899876 1800064
## 9 1948 1813852 1721216
## 10 1949 1826352 1733177
## 11 1950 1823555 1730594
## 12 1951 1923020 1827830
## 13 1952 1971262 1875724
## 14 1953 2001798 1900322
## 15 1954 2059068 1958294
## 16 1955 2073719 1973576
## 17 1956 2133588 2029502
## 18 1957 2179960 2074824
## 19 1958 2152546 2051266
## 20 1959 2173638 2071158
## 21 1960 2179708 2078142
## 22 1961 2186274 2082052
## 23 1962 2132466 2034896
## 24 1963 2101632 1996388
## 25 1964 2060162 1967328
## 26 1965 1927054 1833304
## 27 1966 1845862 1760412
## 28 1967 1803388 1717571
## 29 1968 1796326 1705238
## 30 1969 1846572 1753634
## 31 1970 1915378 1816008
## 32 1971 1822910 1733060
## 33 1972 1669927 1588484
## 34 1973 1608326 1528639
## 35 1974 1622114 1537844
## 36 1975 1613135 1531063
## 37 1976 1624436 1543352
## 38 1977 1705916 1620716
## 39 1978 1709394 1623885
## 40 1979 1791267 1703131
## 41 1980 1852616 1759642
## 42 1981 1860272 1768966
## 43 1982 1885676 1794861
## 44 1983 1865553 1773380
## 45 1984 1879490 1789651
## 46 1985 1927983 1832578
## 47 1986 1924868 1831679
## 48 1987 1951153 1858241
## 49 1988 2002424 1907086
## 50 1989 2069490 1971468
## 51 1990 2129495 2028717
## 52 1991 2101518 2009389
## 53 1992 2082097 1982917
## 54 1993 2048861 1951379
## 55 1994 2022589 1930178
## 56 1995 1996355 1903234
## 57 1996 1990480 1901014
## 58 1997 1985596 1895298
## 59 1998 2016205 1925348
## 60 1999 2026854 1932563
## 61 2000 2076969 1981845
## 62 2001 2057922 1968011
## 63 2002 2057979 1963747
What are the variable or column names in the present
data frame? Year, boys, girls Write and run your code in the chunk
below.
# View the column names in the present1 data frame.
present
## year boys girls
## 1 1940 1211684 1148715
## 2 1941 1289734 1223693
## 3 1942 1444365 1364631
## 4 1943 1508959 1427901
## 5 1944 1435301 1359499
## 6 1945 1404587 1330869
## 7 1946 1691220 1597452
## 8 1947 1899876 1800064
## 9 1948 1813852 1721216
## 10 1949 1826352 1733177
## 11 1950 1823555 1730594
## 12 1951 1923020 1827830
## 13 1952 1971262 1875724
## 14 1953 2001798 1900322
## 15 1954 2059068 1958294
## 16 1955 2073719 1973576
## 17 1956 2133588 2029502
## 18 1957 2179960 2074824
## 19 1958 2152546 2051266
## 20 1959 2173638 2071158
## 21 1960 2179708 2078142
## 22 1961 2186274 2082052
## 23 1962 2132466 2034896
## 24 1963 2101632 1996388
## 25 1964 2060162 1967328
## 26 1965 1927054 1833304
## 27 1966 1845862 1760412
## 28 1967 1803388 1717571
## 29 1968 1796326 1705238
## 30 1969 1846572 1753634
## 31 1970 1915378 1816008
## 32 1971 1822910 1733060
## 33 1972 1669927 1588484
## 34 1973 1608326 1528639
## 35 1974 1622114 1537844
## 36 1975 1613135 1531063
## 37 1976 1624436 1543352
## 38 1977 1705916 1620716
## 39 1978 1709394 1623885
## 40 1979 1791267 1703131
## 41 1980 1852616 1759642
## 42 1981 1860272 1768966
## 43 1982 1885676 1794861
## 44 1983 1865553 1773380
## 45 1984 1879490 1789651
## 46 1985 1927983 1832578
## 47 1986 1924868 1831679
## 48 1987 1951153 1858241
## 49 1988 2002424 1907086
## 50 1989 2069490 1971468
## 51 1990 2129495 2028717
## 52 1991 2101518 2009389
## 53 1992 2082097 1982917
## 54 1993 2048861 1951379
## 55 1994 2022589 1930178
## 56 1995 1996355 1903234
## 57 1996 1990480 1901014
## 58 1997 1985596 1895298
## 59 1998 2016205 1925348
## 60 1999 2026854 1932563
## 61 2000 2076969 1981845
## 62 2001 2057922 1968011
## 63 2002 2057979 1963747
In what year did we see the largest total number of births in the United States?
To answer this, create a new total variable in the
present data frame. Then sort the data frame by the
total variable.
Write and run your code in the chunk below.
#```{r create_total_present} # Add a total variable to the present2 data frame. present\(total <- present\)boys + present$girls
#View the updated data frame.
#A. 3
#B. 8
#C. 11
#D. 83 #Answer: C.
#In R, what does the assignment operator <- do?
#A. It checks whether two values are equal.
#B. It stores the value on the right in the object named on the
left.
#C. It extracts a column from a data frame.
#D. It creates a plot.
Answer: B
#After running the creating_objects chunk, what value is
stored in x?
#A. 8
#B. 3
#C. 11
#D. log(2)
Answer: C
#After running the creating_objects chunk, what value is
stored in y?
#A. Approximately 0.6931472
#B. 2
#C. 11
#D. 21
Answer: D
#Which statement best explains why x and X
are different object names in R?
#A. R ignores capitalization.
#B. R is case-sensitive.
#C. R only allows lowercase object names.
#D. R only allows uppercase object names.
Answer: B
#What does the c() function do in this assignment?
#A. It creates a plot.
#B. It combines multiple values into a vector.
#C. It calculates a standard deviation.
#D. It extracts a column from a data frame.
Answer:B
#After running the create_vectors chunk, what values are
stored in a?
#A. 8.2, 13.4, 16.4, 3.6
#B. 4.1, 6.7, 8.2, 1.8
#C. 2, 2, 2, 2
#D. x, y, z
Answer:B
#After running the create_vectors chunk, what values are
stored in b?
#A. 4.1, 6.7, 8.2, 1.8
#B. 8.2, 13.4, 16.4, 3.6
#C. 2, 4, 6, 8
#D. 5.2
Answer:B
#After running the summarize_vector chunk, what value
does R return for mean(a)?
#A. 2.8
#B. 4.1
#C. 5.2
#D. 8.2
Answer:C
#Which statement best describes a data frame?
#A. A single number stored in R.
#B. A table-like structure made of rows and columns.
#C. A symbol used to extract a column.
#D. A plot created by R.
Answer:B
#After running dim(arbuthnot), how many rows and columns
are in the arbuthnot data frame?
#A. 3 rows and 82 columns
#B. 82 rows and 3 columns
#C. 82 rows and 4 columns
#D. 1710 rows and 3 columns
Answer:B
#Which variables are included in the original arbuthnot
data frame?
#A. year, boys, and
girls
#B. year, male, and female
#C. boys, girls, and total
#D. age, sex, and births
Answer:A
#What does this command do?
#r arbuthnot$boys
A. It shows the entire arbuthnot data frame.
B. It extracts the boys column from the
arbuthnot data frame.
C. It creates a new column called boys.
D. It calculates the mean number of boys.
Answer:B
What command extracts just the counts of girls baptized from the
arbuthnot data frame?
A. girls$arbuthnot
B. arbuthnot$girls
C. mean(arbuthnot$girls)
D. names(arbuthnot)
Answer:B
What is the third entry in the vector of counts of girls baptized?
A. 3919
B. 4102
C. 4457
D. 4683
Answer:B
Which command calculates the mean number of girls baptized yearly between 1629 and 1710?
A. mean(arbuthnot)
B. mean(girls)
C. mean(arbuthnot$girls)
D. arbuthnot$mean(girls)
Answer:C
Based on the plot of girls baptized over time, which statement is most reasonable?
A. The number of girls baptized appears to generally increase over
time.
B. The number of girls baptized is exactly the same every year.
C. The number of girls baptized decreases to zero.
D. The plot cannot be created from these data.
Answer:A
What does the command below create?
arbuthnot$total <- arbuthnot$boys + arbuthnot$girls
A. A new data frame called total.
B. A new variable called total inside the
arbuthnot data frame.
C. A plot of total births.
D. A help file.
Answer:B
What years are included in the present data set?
A. 1629 to 1710
B. 1900 to 2000
C. 1940 to 2002
D. 2000 to 2020
Answer:C
In what year did the present data show the largest total
number of births in the United States?
A. 1940
B. 1955
C. 1961
D. 2002
Answer:D
In later assignments, you will use the tools that we have learned here to answer real statistical questions. For example, you will load datasets, describe variables, create graphs, check assumptions, run statistical tests, and write conclusions in context.
For now, the goal is to become comfortable enough with R that those later tasks feel possible.