R Markdown

This is an R Markdown document. Markdown is a simple formatting syntax for authoring HTML, PDF, and MS Word documents. For more details on using R Markdown see http://rmarkdown.rstudio.com.

When you click the Knit button a document will be generated that includes both content as well as the output of any embedded R code chunks within the document. You can embed an R code chunk like this:

Note that the echo = FALSE parameter was added to the code chunk to prevent printing of the R code that generated the plot.

1.(a,b)

url <- "https://raw.githubusercontent.com/JeffSackmann/tennis_atp/master/atp_matches_2016.csv"

destination <- "atp_matches_2016.csv"

download.file(url, destfile = destination, mode = "wb")

atp_data <- read.csv(destination)

head(atp_data)
##   tourney_id tourney_name surface draw_size tourney_level tourney_date
## 1  2016-M020     Brisbane    Hard        32             A     20160104
## 2  2016-M020     Brisbane    Hard        32             A     20160104
## 3  2016-M020     Brisbane    Hard        32             A     20160104
## 4  2016-M020     Brisbane    Hard        32             A     20160104
## 5  2016-M020     Brisbane    Hard        32             A     20160104
## 6  2016-M020     Brisbane    Hard        32             A     20160104
##   match_num winner_id winner_seed winner_entry       winner_name winner_hand
## 1       271    105062          NA              Mikhail Kukushkin           R
## 2       272    103285          NA           PR    Radek Stepanek           R
## 3       273    106071           7                  Bernard Tomic           R
## 4       275    104471          NA            Q        Ivan Dodig           R
## 5       276    106298          NA                  Lucas Pouille           R
## 6       277    105676           6                   David Goffin           R
##   winner_ht winner_ioc winner_age loser_id loser_seed loser_entry
## 1       183        KAZ       28.0   104797         NA            
## 2       185        CZE       37.1   105583         NA            
## 3       193        AUS       23.2   103917         NA            
## 4       183        CRO       31.0   117352         NA           Q
## 5       185        FRA       21.8   106415         NA           Q
## 6       180        BEL       25.0   105064         NA            
##           loser_name loser_hand loser_ht loser_ioc loser_age       score
## 1      Denis Istomin          R      188       UZB      29.3     6-2 7-5
## 2      Dusan Lajovic          R      180       SRB      25.5     6-0 6-3
## 3      Nicolas Mahut          R      190       FRA      33.9     6-4 6-3
## 4    Oliver Anderson          U       NA       AUS      17.6     6-3 6-2
## 5 Yoshihito Nishioka          L      170       JPN      20.2 4-6 6-3 7-5
## 6    Thomaz Bellucci          L      188       BRA      28.0     6-4 6-4
##   best_of round minutes w_ace w_df w_svpt w_1stIn w_1stWon w_2ndWon w_SvGms
## 1       3   R32      84     1    3     67      36       27       20      10
## 2       3   R32      67     3    2     48      25       18       16       8
## 3       3   R32      69     8    0     59      34       28       14      10
## 4       3   R32      67    11    2     49      30       24       13       9
## 5       3   R32     143    17    2     95      64       53       15      16
## 6       3   R32      82     3    4     57      30       23       16      10
##   w_bpSaved w_bpFaced l_ace l_df l_svpt l_1stIn l_1stWon l_2ndWon l_SvGms
## 1         3         3     6    0     53      32       22       12      10
## 2         2         2     0    2     46      25       15        8       7
## 3         4         5     4    1     50      29       21       10       9
## 4         0         0     3    1     52      30       22        9       8
## 5         2         4     4    3    120      64       42       30      15
## 6         1         2     7    2     66      40       25       15      10
##   l_bpSaved l_bpFaced winner_rank winner_rank_points loser_rank
## 1         4         7          65                762         61
## 2         4         8         197                252         76
## 3         3         6          18               1675         71
## 4         3         6          87                636        813
## 5        12        15          78                672        117
## 6         7        10          16               1880         37
##   loser_rank_points
## 1               781
## 2               678
## 3               710
## 4                25
## 5               495
## 6              1105

1.(c)

I chose this data because it relates to tennis, a sport my family and I love. Although the tennis_atp is a collection of ATP (Association of Tennis Professionals) data from 1968 to today, my eyes fell on the data for 2016, because that’s when one of my favorite players Andy Murray, who had the best season of his career. He reached two grand slam finals (Australian and French Open), won Wimbledon (2nd in his career), took gold at the Olympics in Brazil, won the ATP Finals (the final tournament of the season among the top 8 best tennis players in the world) and most importantly he finished #1 in the ATP rankings or simply the best player in the world at the time. I also want to add that I respect this athlete, he has an incredible work ethic, and he is a shining example that patience and work on yourself will help you overcome difficulties on the way to achieving your goal and will lead you to the top.

2.(a)

This data contains historical records of ATP matches, sourced by Jeff Sackman who is the founder of TennisAbstract, a comprehensive database of professional tennis results and statistics. Also, he has written about the sport for The Wall Street Journal, ESPN, and Tennis Magazine, in addition to his own blog, Heavy Topspin.

2.(b,c)

This dataset contain 2,941 entries and 49 total columns.

ncol(atp_data)
## [1] 49
nrow(atp_data)
## [1] 2941

Speaking of types of variables, I would like to separate them into three categories such as: . Match Details Examples: tourney_id (id for tournament), tourney_name, surface (hard, grass or clay), draw_size (nuber of players in the tournament), tourney_level (Grand Slam or Masters), etc.

. Player info Examples: winner_name,loser_name, winner_hand/loser_hand (left/right), winner_id, loser_id, winner_seed,loser_seed, etc.

. Match stats Examples: winner_age, loser_age, w_ace/l_ace (number of aces hit by w or l), w_df/l_df (number of double faults), w_svpt/l_svpt (total serve points), w_bpSaved/ l_bpSaved (break points saved), etc.

Data was collected from the official ATP records. The Association of Tennis Professionals maintains records of all the tournaments its supervise. Records inclued tournament details, match results and players stats.

Sampling straegies.

Simple random sampling. In the OpenIntro Statistics, example of simple random sampling was given on MLB player to be more clear on their salaries. In the case of tennis data, one of the best options is to randomly select matches from each tournament from the entire year, in this case, every game have a chance to be elected without subjective opinion, what can be very important in knowing the fact of my preferance to a spacific player.

Stratified sampling. This stategy will be a little more complicated than random sampling, but more proficient. For stratified sampling I will devide tournaments into strata by level (Masters, Grand Slam, Davis Cup) and $ tourney_level will help me with that. After devision into strata follow simple random sampling will provide proportional amount of matches by differetn level. Important to say that instead of level of tournaments I can use surface, but in this case number of matches will be unproportional due to different amount of tournaments on each surface. Ussualy, Hard surface take over the half of the tournametns, about 55%-60% of the seson. Clay about a third and Grass take the rest, around 10% of the whole season.

Cluster sampling. In this case, tournaments will be clusters; they won’t be divided by surface or level, but by name (tourney_name) or ID number (tourney_id). Then randomly select clusters and include all matches from selected tournaments.

3.(a)

In my opinion, a sampling strategy is not necessary in case of this data, due its reliability and the fact that it is just census of all the matches from 2016

  1. How player’s seeding influence on winning ? Outcome: winner_name. Explanatory: winner_seed, loser_seed. How player’s age impact on the winning ? Outcome: winner_name. Explanatory: winner_age, loser_age.

5.(a)

library(stargazer)
## 
## Please cite as:
##  Hlavac, Marek (2022). stargazer: Well-Formatted Regression and Summary Statistics Tables.
##  R package version 5.2.3. https://CRAN.R-project.org/package=stargazer
stargazer(atp_data[c("winner_seed", "loser_seed", "winner_age", "loser_age")], type = "text", summary = TRUE, digits = 2)
## 
## ============================================
## Statistic     N   Mean  St. Dev.  Min   Max 
## --------------------------------------------
## winner_seed 1,359 7.48    6.65     1    33  
## loser_seed   714  8.94    7.42     1    33  
## winner_age  2,941 27.89   4.23   16.70 37.80
## loser_age   2,941 27.69   4.42   17.10 37.80
## --------------------------------------------

5.(b)

library(ggplot2)


ggplot(atp_data, aes(x = winner_age)) + geom_density(fill = "pink", alpha = 0.4) + geom_density(aes(x = loser_age), fill = "yellow", alpha = 0.4) + labs(title = "Density Distribution of Winner and Loser Ages", x = "Age", y = "Density") + theme_minimal()

ggplot(atp_data, aes(x = winner_seed)) +  geom_density(fill = "purple", alpha = 0.4) +  geom_density(aes(x = loser_seed), fill = "gray", alpha = 0.4) + labs(title = "Density Distribution of Winner and Loser Seeds", x = "Seed Number", y = "Density") + theme_minimal()
## Warning: Removed 1582 rows containing non-finite outside the scale range
## (`stat_density()`).
## Warning: Removed 2227 rows containing non-finite outside the scale range
## (`stat_density()`).

5.(c)

ggplot(atp_data, aes(x = winner_seed, y = winner_age)) + geom_point(alpha = 0.6, color = "green") + labs(title = "Scatter Plot of Winner Seed vs Winner Age", x = "Winner Seed", y = "Winner Age") + theme_minimal()
## Warning: Removed 1582 rows containing missing values or values outside the scale range
## (`geom_point()`).

# Scatter plot: Loser Seed vs. Loser Age
ggplot(atp_data, aes(x = loser_seed, y = loser_age)) + geom_point(alpha = 0.6, color = "red") + labs(title = "Scatter Plot of Loser Seed vs Loser Age", x = "Loser Seed", y = "Loser Age") +theme_minimal()
## Warning: Removed 2227 rows containing missing values or values outside the scale range
## (`geom_point()`).

6.(a)

Speaking of target relations between seeding and match outcome. As we can see winner average seed is 7.48 and loser’s is 8.94, base on this information we can make a conclusion that despite the fact that difference isn’t that big, on average player with lower seed is more likely to win. Difference in standard deviation also pretty similar 6.65 for winner to 7.42 for loser, which also correlates with the statement above, that mostly likely player with lower seed will win, but again difference is very small.

Speaking of age we will see that AVERAGE winner age is 27.89 and loser’s age is 27.69, this numbers can tell us that age difference is the minimal, and I think that age doesn’t influence outcome of the match base on the data that I got. Standard deviation says exact same thing, numbers for winner are 4.23 and for loser are 4.42, what also proof that age doesn’t influence influence on outcome that much.

If I will need to predict an outcome for the match, I would rather base my opinion on seed difference than age.

About prior work and research on the data that I work with, there is a research age and performance trends in men’s prof. Tennis from 1991 to 2012 from Journal of Quantitative Analysis in Sport. They found that younger players concur with the tour, and being in your 3Os as a player in the tour means to be old. Compare this research to mine, first of all, my work based only on one season in 2016, not two decades, and very important is the fact that I was looking primary on the seeded player, best 32 tennis athletes in the world, not a whole tour in general.