2024-06-04

Background of Data

The data set I wanted to for project 1 I got from kaggle.com that was listed in the project 1 pdf. I’m always interested in music and though this would be a nice one to explore, manipulate, and visualized. The csv provided details and metrics for each track.

Includes:

-charting and playlist metrics

-track and artist name

-number of streams

-other metric like bpm, danceability, and energy

Idenity Problems and Objectives

A problem that comes with data sets online is sometimes you have to clean them to make sure its ready for exploration and visualization. An object of this project is to use data manipulation to find out what songs are the most popular and what makes a song popular to then use the insights to help produce more popular songs.

Loading Libraries

library(dplyr)
## 
## Attaching package: 'dplyr'
## The following objects are masked from 'package:stats':
## 
##     filter, lag
## The following objects are masked from 'package:base':
## 
##     intersect, setdiff, setequal, union
library(ggplot2)
library(plotly)
## 
## Attaching package: 'plotly'
## The following object is masked from 'package:ggplot2':
## 
##     last_plot
## The following object is masked from 'package:stats':
## 
##     filter
## The following object is masked from 'package:graphics':
## 
##     layout

Loading Data

music = read.csv("Popular_Spotify_Songs.csv")

Data Cleaning

Data cleaning is extremely important especially when you find a csv online. Checked if there are any NA values and there were not. I also wanted to change streams from char to a numerical value.

sum_na = sum(is.na(music))
print(sum_na)
## [1] 0
data_clean_music = music%>%
  mutate(streams = as.numeric(streams))%>%
  na.omit()

Data Wrangling

Data wranlging is used to get insights and to produce new and analyzes exsiting metric from the data frame. An important library for this to have is dplyr to be able to select, mutate, and arange, There are many function and way to weangle data.

Using my objective for this lab I made questions and used data manipulation to answer them to get results to help answer the question but to help solve the objective.

Getting full date format

When I read the csv file I saw there were seperate columns for released date, month, and year. I wanted a column with full date year/month/day format. Data wrangling helped me achive that and here is the R code that I used.

data_clean_music = data_clean_music %>%
  mutate(released_date = as.Date(paste(released_year, released_month, released_day, sep = "-"), format = "%Y-%m-%d"))
print(data_clean_music$released_date[1:5])
## [1] "2023-07-14" "2023-03-23" "2023-06-30" "2019-08-23" "2023-05-18"

Data Exploration: Getting insights

My objective to find what makes a song popular. I was curious to see if all the popular songs in the data are from today’s hit made in the year 2024 and 2023 but appon further data exploration I found a that the older song in the data frame was relased in 1930-01-01 and the newest song was released in “2023-07-14”. Here is the code in R that helped me find the oldest and newest song.

oldest_song = min(data_clean_music$released_date)
newest_song = max(data_clean_music$released_date)

Data Visulization

Code Used for 3D Plot

# music_3d = data_clean_music %>%
#    select(streams, bpm, energy_.)
# xax = list(
# title = "Streams",
# titlefont = list(family = "Modern Computer Roman")
#  )
# yax = list(
#    title = "BPM",
#    titlefont = list(family = "Modern Computer Roman")
# )
#  zax = list(
#    title = "Energy",
#    titlefont = list(family = "Modern Computer Roman")
#  )

Code Continued

# plot =  plot_ly(data = music_3d, 
#          x= ~streams, y= ~bpm, z= ~energy_.,
#          type = "scatter3d", mode = "markers",
#          marker = list(size = 3,
#                        color = 'red'))%>%
#    layout(
#      title = "Streams VS. (BPM, Energy)",
#      scene = list(xaxis = xax, yaxis = yax, zaxis = zax))
# plot

Analyzing Findging

  • Plotly is a important data visualization tool that can make interactive graphs for better analysis

  • Goal is to analyze the metric of populars songs that have large amounts of streams

  • I concluded that songs with large amount of streams have higher BPM and average energy

Conclusion

To concluded, The objective was to clean the data found on kaggle and manipulate it and explore to find what makes a song popular and have large quantities of streams. The data was insightful, in my finding I found metrics of songs with high amounts of stream what months songs are released correlated to higher streams.There is much more exploration that could be done; I explored and manipulated what would work for me and my objective for project 1.