Requests.get() does not seem to be returning the expected bytes for Wikipedia image URLs, such as https://upload.wikimedia.org/wikipedia/commons/0/05/20100726_Kalamitsi_Beach_Ionian_Sea_Lefkada_island_Greece.jpg:
import wikipedia import requests page = wikipedia.page("beach") first_image_link = page.images[0] req = requests.get(first_image_link) req.content b'<!DOCTYPE html>n<html lang="en">n<meta charset="utf-8">n<title>Wikimedia Error</title>n<style>n*...
Advertisement
Answer
Most websites block requests that come in without a valid browser as a User-Agent. Wikimedia is one such.
import requests headers={'User-Agent': 'Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/102.0.0.0 Safari/537.36'} res = requests.get('https://upload.wikimedia.org/wikipedia/commons/0/05/20100726_Kalamitsi_Beach_Ionian_Sea_Lefkada_island_Greece.jpg', headers=headers) res.content
which will give you expected output