I first wrote this article in 2021, with an example based on Stack Overflow job listings. That part of the site has since been shut down, and Selenium has changed too, so the old code no longer worked. That is why the example was rewritten and tested in October 2026.
When is scraping allowed?
Scraping saves hours of copying by hand. It is most often used for:
- monitoring prices at suppliers or competitors;
- moving your own data from an old website to a new one when the old system has no export;
- collecting public data for analysis and research;
- checking your own website, for example the titles and links on every page.
Being publicly visible does not mean something can be taken and used however you like. Before you start, check:
- robots.txt. At the site's address followed by
/robots.txt, the owner lists which parts of the site robots should not open. - The site's terms of use. Many sites explicitly prohibit automated downloading.
- Copyright. Do not copy other people's texts, descriptions and images to your own site. It is not allowed, and it does not help in search either, because Google filters out duplicates. Read more in the article about duplicate content.
- Personal data. Names, email addresses and phone numbers are personal data and are protected by data protection law.
- Server load. Pause between requests, as in the examples below. A program that opens hundreds of pages per second can bring down a small website.
- Whether there is an API. If the site offers an API for the data you need, use it. It is faster, more reliable and allowed.
Setup
The example was made on Windows, and the same code works on Linux and macOS.
- Install Python from python.org. On the first installer screen, tick Add python.exe to PATH, so you do not have to add the path by hand.
- I write code in Visual Studio Code (visualstudio.com) with the Python extension (ms-python.python), which VS Code will offer to install.
- Create a project folder, open it in VS Code, and install the libraries in the terminal:
pip install requests beautifulsoup4 selenium
The pip command downloads libraries that do not come with Python, so you do not have to hunt for them online. For the second approach you also need Firefox installed. In the past you had to download geckodriver.exe by hand and copy it into the project folder; since version 4.6, Selenium does this itself.
For practice we use quotes.toscrape.com, a site made specifically for learning web scraping. It has 100 famous quotes across 10 pages, plus a version whose content is built by JavaScript. The goal is to put all the quotes and their authors into one table.
Which approach should you choose?
Open the page in your browser and press Ctrl+U to see its source code. If the data you need is visible in the source, the first approach is enough. If it is not, JavaScript fills in the page after loading, and you need the second approach.
| requests and BeautifulSoup | Selenium and Firefox | |
|---|---|---|
| How it works | Downloads the page HTML and reads it | Controls a real browser, as if you were clicking |
| Speed | Fast | Slower, because the page is actually rendered |
| JavaScript pages | Does not see content built by JavaScript | Sees everything the user sees |
| Logins, clicks, forms | Harder | Easy |
| When to use it | Whenever it works | When the first approach cannot see the data |
First approach: requests and BeautifulSoup
In the project folder, create a file quotes.py with the following code:
import csv
import random
import time
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
wait_from = 2
wait_to = 4
url = "https://quotes.toscrape.com/"
headers = {"User-Agent": "Mozilla/5.0 (example from programiranje.co.rs)"}
with open("quotes.csv", "w", newline="", encoding="utf-8-sig") as csv_file:
writer = csv.writer(csv_file)
writer.writerow(["quote", "author"])
while url:
response = requests.get(url, headers=headers, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for quote in soup.select("div.quote"):
text = quote.select_one("span.text").get_text(strip=True)
author = quote.select_one("small.author").get_text(strip=True)
writer.writerow([text, author])
next_link = soup.select_one("li.next > a")
url = urljoin(url, next_link["href"]) if next_link else None
if url:
time.sleep(random.randint(wait_from, wait_to))
print("Done!")
Run it with python quotes.py. In about thirty seconds a quotes.csv file with 100 quotes will appear in the folder. Here is what each part does:
wait_fromandwait_toset how many seconds the program waits between two pages, randomly each time. We do not want to open pages too quickly, because that loads the server, and the site may block us.requests.getdownloads the page HTML.timeoutstops the program from waiting forever for a site that does not respond, andraise_for_status()stops it if the site returns an error, instead of quietly saving an empty table.soup.select("div.quote")finds all elements matching a CSS selector. You find the selector by right-clicking the page, choosing Inspect and looking at which element and class hold the data. Here each quote is in adivwith the classquote, the text inspan.textand the author insmall.author.li.next > ais the “Next” link at the bottom of the page. While it exists, the program moves on to the next page. When it is gone, we have reached the last page and the loop ends. Theurljoinfunction turns the link's relative address (/page/2/) into a full one.
Second approach: Selenium and Firefox
The version of the site at /js/ looks the same, but the quotes are inserted by JavaScript. Its source code contains no quotes, so the first approach returns an empty table there. Selenium opens a real Firefox, waits for the page to render and clicks “Next”, just like a person would:
import csv
import random
import time
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC
from selenium.webdriver.support.ui import WebDriverWait
wait_from = 2
wait_to = 4
browser = webdriver.Firefox()
browser.get("https://quotes.toscrape.com/js/")
wait = WebDriverWait(browser, 10)
with open("quotes-js.csv", "w", newline="", encoding="utf-8-sig") as csv_file:
writer = csv.writer(csv_file)
writer.writerow(["quote", "author"])
while True:
wait.until(EC.presence_of_element_located((By.CSS_SELECTOR, "div.quote")))
for quote in browser.find_elements(By.CSS_SELECTOR, "div.quote"):
text = quote.find_element(By.CSS_SELECTOR, "span.text").text
author = quote.find_element(By.CSS_SELECTOR, "small.author").text
writer.writerow([text, author])
next_links = browser.find_elements(By.CSS_SELECTOR, "li.next > a")
if not next_links:
break
time.sleep(random.randint(wait_from, wait_to))
first_quote = browser.find_element(By.CSS_SELECTOR, "div.quote")
next_links[0].click()
wait.until(EC.staleness_of(first_quote))
browser.quit()
print("Done!")
webdriver.Firefox()opens a new Firefox window controlled by the program. At the end,browser.quit()closes it.find_elements(By.CSS_SELECTOR, ...)is today's way of searching. Selenium 3 usedfind_elements_by_css_selector(...), and that was the code in the first version of this article. Selenium 4 removed those functions, which is why old examples from the internet now throw an error.WebDriverWaitwaits up to 10 seconds for the quotes to appear on the page. That is more reliable than a fixed delay, because it does not depend on the site's speed.staleness_of(first_quote)waits after the click until the old page is gone. Without it, the program could read the same page twice before the new one loads.
If you do not want to watch the browser window, Firefox can also run invisibly. Instead of browser = webdriver.Firefox(), write:
options = webdriver.FirefoxOptions()
options.add_argument("-headless")
browser = webdriver.Firefox(options=options)
How do you open the result in Excel?
Both examples save a CSV file with utf-8-sig encoding. That marker at the start of the file tells Excel the text is UTF-8, so accented and non-Latin characters display correctly. The csv module puts quotation marks around text that contains a comma, so each quote stays in a single cell.
If Excel puts a whole row into one column, it is because of regional settings: in many European countries Excel expects a semicolon as the separator. In that case open the file via Data → From Text/CSV, or write csv.writer(csv_file, delimiter=";") in the code.
Common problems
AttributeError: 'WebDriver' object has no attribute 'find_elements_by_css_selector': the code was written for Selenium 3. Replace it withfind_elements(By.CSS_SELECTOR, ...)as in the example above.- The table is empty: the page builds its content with JavaScript, so use Selenium, or the CSS selector has changed. Check it again with Inspect.
- Error 403 or 429: the site has blocked you or is asking you to slow down. Increase the pauses and check whether the site allows automated downloading at all.
UnicodeEncodeErroron Windows: the file was opened withoutencoding. Always specifyencoding="utf-8"or"utf-8-sig".- The program used to work and now it does not: the site changed its layout and the selectors no longer match. Scraping programs need maintenance, because websites do not ask before they change.
Is it easy to extract data from a web page?
With a little code it is, as the examples show. This example uses Python, but the same can be done in PHP, C# and other programming languages. The harder part starts with real websites: logins, bot protection, frequent layout changes. I have built such projects for clients, and you can see examples in my portfolio under browser automation and web scraping.
I hope this guide helped. If you get stuck somewhere, leave a comment below and I will help.
Need a program that does this for you?
I build programs that collect data from websites, fill in forms and move data into Excel or a database, while respecting the rules of the sites they work with.
Leave a Comment