UnicodeDecodeError 🐍 Python

UnicodeDecodeError: 'utf-8' codec can't decode byte 0x92 in position 10: invalid start byte

Python read bytes that aren’t valid in the encoding it assumed (usually UTF-8).

Seen on: Python

Meaning

The file was saved in another encoding (Windows-1252/latin-1, UTF-16), or it’s binary. Byte 0x92 is a Windows “smart quote”; 0xff 0xfe at the start indicates UTF-16.

Common causes

  • CSV exported from Excel in cp1252
  • Reading a binary file in text mode
  • UTF-16 file
  • Default encoding differs on Windows

⚡ Quick fix

  1. Pass the right encoding: open(p, encoding="cp1252")
  2. Try encoding="utf-8-sig" for BOM files
  3. Use errors="replace" only when data loss is acceptable
  4. Open binaries with "rb"

Detailed fix by platform

Python

  1. pd.read_csv("export.csv", encoding="cp1252")

How to diagnose

  1. Byte — Which byte value? (0x92/0x96 → cp1252, 0xff 0xfe → UTF-16)
  2. File — Text or binary?

🧠 Still stuck? Analyze your error

Paste the full message, response headers or stack trace — we'll detect the platform and point to the most likely cause.