Scrape the 500 Greatest Albums of All Time using Puppeteer Sharp
Home
About
Articles
As a keen Spotify user that is always in search of new albums and artists to listen to, I was particularly interested when I heard that Rolling Stone magazine announced a new, updated list of their 500 Greatest Albums of All Time for 2020.
I thought it would be a good idea, whilst at home during the pandemic, to listen to all 500 albums in descending order. One way to accomplish this is to scroll through the list of albums, listening to them one by one, but the programmer in me thought to go one better! Why not scrape the HTML so I can create a private document in Google Sheets , marking each of the albums as I listen to them and writing a comment on each one?
I decided to scrape the HTML using Puppeteer Sharp , the .NET port of the popular Puppeteer Node library by Google, mostly because I’m a .NET developer and also because I’m already familiar with it. It’s a really useful library that will enable a developer to “puppeteer” a web browser and automate a series of tasks.
Once the Puppeteer Sharp package is installed in a new project (in this case a console application) , we can begin to scrape the HTML of all of the pages that contain the 500 greatest albums of all time.
Let’s begin by downloading the default version of Chromium onto the local machine and opening a new browser:
We’ll need a list of all of the pages that contain the 500 albums to iterate through:
Now that we have the list of pages in a dictionary object, we can iterate through each page, scraping the albums one by one (50 per page) using the elements in the HTML to indicate where each album and it’s metadata are placed. We’ll assign the album metadata to variables and output these to the console window for the purpose of this example:
Please note: In this example we’re using an ethical user-agent header. This is to let the website know this is a web scraping script and to provide contact information. To find out more on ethical web scraping, please visit The Ultimate Guide To Ethical Web Scraping .
Please be aware that the HTML elements (class names, etc.) used in this code snippet may change over time, and these will need to be updated accordingly if a class name or element no longer exists.
Once we’re done, let’s close the browser and write to the console that the scrape is complete:
While the official list of the 500 Greatest Albums of All Time should be the first place you visit for this information, I hope this article is a useful resource for how to use Puppeteer Sharp to scrape the list of albums with the aim of listening through all of them, one by one in descending order.
Have you ever wondered if you could use the .NET CLI and Visual Studio Code to build a new .NET application, instead of relying on Visual Studio to do all of the heavy lifting? Read more
Back in 2019 I wrote an article on Best Practices for Writing Unit Tests in C# for Bulletproof Code. This has become one of my more popular articles, and despite it approaching 2 years old, the best practices mentioned are still relevant today. I touched upon the popular mocking framework Moq, bu... Read more
MediatR, by it’s definition, is a simple, unambitious mediator implementation in .NET. It was released in 2014 by Jimmy Bogard and is a useful package that can be used to implement the popular mediator pattern in .NET projects. It’s available on NuGet and is also open-source on GitHub. Read more
RSS