Final projects for STAT 107: Data Science Discovery and IS 204: Research Design for Information Sciences at UIUC


Languages and Tools Used

python pandas seaborn scikit_learn vega-lite




Exploring the Spotify Top 100 Hit Songs 2010 to 2022 Dataset

The second project uses another Kaggle dataset for Spotify’s top 100 hit songs per year from 2010 to 2022. This data was extracted directly from the Spotify API and specifically their ‘Top Hits’ playlist for each year (with 100 songs per year). The dataset has 23 columns and 2400 rows with 13 track audio features consisting of danceability, energy, key, loudness, mode, speechiness, acousticness, instrumentalness, liveness, valence, tempo, duration, and time signature.

Research Questions:
Analysis Results:

image tooltip here Table 1 – Summary Statistics

Figure 1.1 – Exploratory Quantitative Analysis on Popularity Factors

Figure 1.2 – Exploratory Quantitative Analysis on Average Duration of Songs

Figure 2 – Linear Regression Analysis of Artist Popularity and Track Popularity

Figure 3 – Correlation between Popular Genres