Skip to main content

Command Palette

Search for a command to run...

Export PDF Metadata to JSON Using Python

Updated
•3 min read•View as Markdown
Export PDF Metadata to JSON Using Python
P
PDF Python Hub is dedicated to teaching PDF programming, automation, and engineering with Python through practical, project-based courses. Whether you're a complete beginner or an experienced Python developer, the goal is to help you build the skills needed to work with PDF documents confidently—from reading and extracting data to building enterprise document processing and AI-powered workflows. What You'll Learn The curriculum covers every stage of PDF processing, including: -PDF fundamentals -Document extraction -PDF manipulation -Forms, tables, and images -OCR and scanned documents -Security and redaction -Enterprise automation -AI-powered document processing How You'll Learn Every course is designed to be: -Beginner-friendly with clear, step-by-step explanations -Focused on practical, real-world applications -Built around hands-on projects and assessments -Structured as a complete learning path from beginner to expert Our Mission Our mission is simple: To make PDF processing with Python accessible, practical, and enjoyable by teaching real skills that can be applied to everyday automation tasks and professional software development. Whether you're learning for your career, your business, or your own projects, you'll gain the knowledge and confidence to build powerful PDF automation solutions with Python. Learn more: https://payhip.com/PDFPythonHub/blog/news

The Python script below uses PyMuPDF to read basic information from a PDF, saves that information as a JSON file, and downloads the file to your computer.

Install pymupdf

If pymupdf isn't already installed in your Colab environment, run:

%pip install -q -U pymupdf

Import the dependencies

import pymupdf
import json
from google.colab import files

Upload the PDF

uploaded = files.upload()

Open the PDF

doc = pymupdf.open("sample.pdf")

This opens the PDF so Python can access its information.

Replace "sample.pdf" with the name of your PDF.

Collect the information

The script creates a dictionary containing:

  • page_number — the total number of pages

  • metadata — information such as the PDF title, author, subject, creator, and creation date

data = {
    "page_number": doc.page_count,
    "metadata": doc.metadata
}

Save the information as JSON

with open("pdf_information.json", "w", encoding="utf-8") as f:
    json.dump(data, f, indent=4, ensure_ascii=False)
  • "w" — opens the file in write mode, meaning new content is written to the file.

  • encoding="utf-8" — specifies UTF-8 character encoding, allowing the file to correctly handle characters from different languages and special characters.

  • as f — gives the opened file a short name (f).

  • json.dump() —writes Python data to JSON.

  • indent=4 — formats the JSON with 4 spaces of indentation, making it easier for humans to read.

  • ensure_ascii=False — keeps non-ASCII characters as they are instead of converting them into Unicode escape sequences. For example, é stays as é rather than becoming \u00e9.

  • with — automatically closes the JSON file when the block finishes, even if something goes wrong while writing.

Download the JSON file

Because this is running in Google Colab, you can download the generated file directly:

files.download("pdf_information.json")

Complete Example

%pip install -q -U pymupdf

import pymupdf
import json
from google.colab import files

uploaded = files.upload()

doc = pymupdf.open("sample.pdf")

data = {
    "page_number": doc.page_count,
    "metadata": doc.metadata
}

with open("pdf_information.json", "w", encoding="utf-8") as f:
    json.dump(data, f, indent=4, ensure_ascii=False)

files.download("pdf_information.json")

Example Output

The resulting JSON might look like:

{
    "page_number": 4,
    "metadata": {
        "format": "PDF 1.5",
        "title": "Artificial Intelligence (AI) ",
        "author": "PDF Python Hub",
        "subject": "An Overview of Artificial Intelligence, Machine Learning, Generative AI, Ethical Frameworks, and Future Outlook",
        "keywords": "Artificial Intelligence, AI, Machine Learning, Deep Learning, Generative AI, Large Language Models, AI Ethics",
        "creator": "Google Docs",
        "producer": "Skia/PDF m153 Google Docs Renderer",
        "creationDate": "D:20260816101230+02'00'",
        "modDate": "D:20260818145823+02'00'",
        "trapped": "",
        "encryption": null
    }
}


Open the notebook

Open the Google Colab notebook for this PDF mini-guide and run the code as you follow along.

[Open in Google Colab]


More PDF Python Resources

Want to learn how to open a PDF file?

Check out this mini-guide:

Open Your First PDF with PyMuPDF