Encrypt and Decrypt PDFs with Python pypdf
Learn how to add user and owner passwords, set PDF permissions, decrypt protected files, and remove encryption using the Python pypdf library.
Working with password-protected PDFs in Python often means using pypdf for both encrypting and decrypting documents. This article covers adding a password to a PDF, opening an existing protected file, setting permissions, and removing encryption.
Installing pypdf
pypdf is a pure-Python library, so installation is straightforward with pip:
pip install pypdf
pypdf is the successor to PyPDF2 and shares most of the API. Check the pypdf release notes for the Python version required by the release you install.
Encrypting a PDF with a user password
To encrypt a PDF, create a PdfWriter, add pages from a source PdfReader, and call encrypt() before writing the output file.
from pypdf import PdfReader, PdfWriter reader = PdfReader('original.pdf') writer = PdfWriter() for page in reader.pages: writer.add_page(page) writer.encrypt(user_password='secret') with open('protected.pdf', 'wb') as output_file: writer.write(output_file)
encrypt() accepts two password arguments: user_password and owner_password. The user password is what a reader must enter to open the document. The owner password is the master password used to change or remove encryption and to define what the user is allowed to do.
If you provide only user_password, pypdf uses the same value as the owner password too. That means the same password can open the document with owner-level permissions. If you need separate behavior, pass an explicit owner_password.
Setting owner password and permissions
Passing an owner password explicitly separates the user and owner passwords:
writer.encrypt(user_password='user123', owner_password='owner456')
You can also restrict user actions with the permissions argument. The permission flags are available in pypdf.constants.Permission:
from pypdf import PdfWriter from pypdf.constants import Permission writer = PdfWriter() # ... add pages ... writer.encrypt( user_password='user123', owner_password='owner456', permissions=Permission.PRINT )
Permission provides flags such as PRINT, COPY, MODIFY, ANNOTATE, and FILL_FORMS. Combine them with bitwise OR:
permissions = Permission.PRINT | Permission.COPY
If you omit permissions, all permissions are granted to the user.
The following table summarizes the common permission flags and what they control:
| Flag | Allows the user to |
|---|---|
PRINT | Print the document |
COPY | Copy text and images from the document |
MODIFY | Modify the document content |
ANNOTATE | Add or modify annotations |
FILL_FORMS | Fill in form fields |
Decrypting a password-protected PDF
To read an encrypted PDF, create a PdfReader and call decrypt() with the user or owner password. The return value is an integer: 0 means the password was incorrect, 1 means the user password was accepted, and 2 means the owner password was accepted.
from pypdf import PdfReader reader = PdfReader('protected.pdf') result = reader.decrypt('user123') if result == 0: print('Incorrect password') else: for page in reader.pages: print(page.extract_text())
After a successful decrypt(), you can access pages, extract text, or copy pages to a new writer.
Handling errors and edge cases
If you try to access pages from an encrypted PDF without calling decrypt(), pypdf raises FileNotDecryptedError. This error comes from pypdf.errors and can be confusing if you expected an automatic password prompt.
Some PDFs are encrypted with an empty user password. pypdf treats an empty string as a valid password, so decrypt('') may succeed for such files.
To remove encryption, decrypt the PDF and write it to a new PdfWriter without calling encrypt():
from pypdf import PdfReader, PdfWriter reader = PdfReader('protected.pdf') reader.decrypt('user123') writer = PdfWriter() for page in reader.pages: writer.add_page(page) with open('unprotected.pdf', 'wb') as f: writer.write(f)
This creates a new PDF with no password protection.
Security considerations and limitations
PDF password protection is not a strong security boundary. It is designed to prevent casual access, not to withstand determined attackers. The practical strength depends on the encryption algorithm, the password, and the PDF reader.
pypdf's encryption support is password-based. It may not handle certificate-based encryption or proprietary encryption schemes; for those files, you may need another tool. Check the pypdf documentation for the supported algorithms and current defaults before picking an algorithm for encrypt().
Performance and memory considerations
pypdf's encryption and write operations are not incremental: they rewrite the PDF, so large files can require significant memory and I/O. For very large documents, consider whether you can split the work or otherwise reduce the file size before encryption.
Preserving metadata and structure
The page-by-page PdfWriter approach creates a new PDF; document-level metadata is not automatically copied by add_page(). If metadata such as the title or author matters, copy it to the writer explicitly or use pypdf's metadata-related methods and verify the output. Some interactive elements, such as JavaScript actions, may also be lost when a PDF is rewritten, so test encrypted output with your target PDF reader.
Calling encrypt() sets encryption state on the writer. If you need a different password or permission set, create a fresh PdfWriter, add the pages again, and call encrypt() with the new settings.