Back to Blog
Python

Python PyGithub: Commits, Releases, and Repository Contents

Use PyGithub to fetch repository contents, commits, and releases from the GitHub API, including pagination, rate-limit handling, and a combined script.

PyGithubGitHub APIPythonGitAutomation
Illustration of PyGithub retrieving commits, releases, and files from a GitHub repository.

PyGithub is a popular Python client for the GitHub REST API. When you need commit history, file contents, or release metadata from a repository, the Repository object exposes all three through the same authenticated session. This article covers the practical patterns for fetching each data set, handling pagination, and combining them in a single script.

Authentication and Repository Access

Every PyGithub operation starts with a Github instance and a Repository handle. GitHub removed password authentication for the REST API, so a personal access token or a GitHub App token is required for anything beyond anonymous access.

from github import Github, Auth auth = Auth.Token("github_pat_...") g = Github(auth=auth) repo = g.get_repo("owner/repository")

The Auth.Token class is the current recommended way to pass credentials. If you work against GitHub Enterprise, pass the server URL as well:

g = Github(auth=auth, base_url="https://github.example.com/api/v3")

Anonymous access is possible with Github() and no auth, but the unauthenticated rate limit is a small fraction of the authenticated one, so any script that iterates over commits or releases should use a token. Keep the token out of source control and read it from an environment variable.

Reading Repository Contents

repo.get_contents(path) is the entry point for file and directory access. For a file it returns a single ContentFile; for a directory it returns a list of ContentFile objects.

root = repo.get_contents("") for item in root: print(item.type, item.path)

Each item has a type of "file" or "dir". To read a file's raw bytes, use decoded_content:

readme = repo.get_contents("README.md") text = readme.decoded_content.decode("utf-8") print(text[:500])

decoded_content returns bytes, so decode with the file's actual encoding. For a file at a specific tag or branch, pass the ref parameter:

content = repo.get_contents("version.txt", ref="v1.2.0")

Two limitations matter in practice. First, get_contents makes one API request per path, so walking a large tree recursively generates many requests. For a full tree listing, the Git trees API (repo.get_git_tree(tree_sha, recursive=True), where tree_sha is the SHA of a commit's root tree) is more efficient. Second, the contents API has a size limit for individual files; very large files should be fetched through the raw URL or the Git blobs API instead.

Working with Commits

repo.get_commits() returns a PaginatedList of commit objects, ordered newest first. You can filter by path, author, or starting ref.

commits = repo.get_commits(path="src/") for commit in commits: print(commit.sha, commit.commit.message.splitlines()[0])

The path argument limits results to commits that touched that path. The sha argument accepts a branch name or commit SHA and starts the listing from that ref; a tag name is often accepted as a starting ref as well.

Useful fields on the commit objects returned by get_commits():

for commit in commits: author = commit.commit.author print(commit.sha) print(author.name, author.date) print(commit.commit.message)

The default commit listing does not include per-commit stats or files in the GitHub API response. Call repo.get_commit(commit.sha) to fetch full detail, including changed files and line counts:

detail = repo.get_commit(commit.sha) print(detail.stats.additions, detail.stats.deletions) for f in detail.files: print(f.filename, f.status, f.additions, f.deletions)

Working with Releases

repo.get_releases() returns published releases, newest first. Each Release exposes the tag, title, body, and publication date.

releases = repo.get_releases() for release in releases: print(release.tag_name, release.name, release.published_at) print(release.body)

Release assets are not loaded automatically. Call release.get_assets() to enumerate them:

for release in releases: for asset in release.get_assets(): print(asset.name, asset.size, asset.browser_download_url)

For the most recent release, repo.get_latest_release() is more direct. It raises GithubException when the repository has no releases, so guard the call:

try: latest = repo.get_latest_release() except GithubException: latest = None

To fetch a release by tag, use repo.get_release("v1.2.0"), which also raises if the tag does not exist.

Pagination and Rate Limits

Every list-returning method in PyGithub returns a PaginatedList, which loads results lazily in pages of 30 by default. Iterating it issues additional requests as needed. This matters for two reasons: memory and rate limits.

commits = repo.get_commits() print(commits.totalCount) for page_index in range(commits.totalCount // 30 + 1): page = commits.get_page(page_index) for commit in page: print(commit.sha)

get_page(n) fetches page n explicitly, which is useful when you only need a range of results. You can also reduce request count by raising per_page:

commits = repo.get_commits(per_page=100)

Check the remaining quota before a long loop:

core = g.get_rate_limit().core print(core.remaining, core.limit, core.reset)

When the limit is exhausted, PyGithub raises RateLimitExceededException. For batch jobs, wait until the reset time and retry, or structure the loop to stop once core.remaining is low.

A Combined Example: Latest Release, Version File, and Commits

The following script combines the three capabilities: it reads a version file from the latest release tag, fetches the latest release metadata, and lists the commits reachable from that tag.

import os from github import Github, Auth, GithubException auth = Auth.Token(os.environ["GITHUB_TOKEN"]) g = Github(auth=auth) repo = g.get_repo("owner/repository") try: latest = repo.get_latest_release() except GithubException: print("No releases found") raise SystemExit(1) print(f"Latest release: {latest.tag_name}") version_file = repo.get_contents("version.txt", ref=latest.tag_name) print("Version at release:", version_file.decoded_content.decode().strip()) # This starts at the tag's commit and walks backward through ancestors. commits = repo.get_commits(sha=latest.tag_name, per_page=100) for commit in commits: first_line = commit.commit.message.splitlines()[0] print(f"{commit.sha[:8]} {first_line}")

This pattern is useful for changelogs that start from a release tag, CI pipelines that need release metadata and version-file contents, and tools that inspect commits reachable from a release. Note that get_commits(sha=latest.tag_name) starts from the tag ref, so the list contains the tag's ancestors rather than commits that landed after the tag. To find commits introduced after a tag, compare the tag with another ref using the compare commits API instead.

Handling Missing Data and Large Repositories

Several edge cases produce surprising failures. A repository with no commits raises GithubException on get_commits(). A path that does not exist raises UnknownObjectException, a subclass of GithubException. Empty directories are not returned by the contents API because Git does not track them.

For large repositories, prefer the trees API for full tree snapshots and the blobs API for individual large files, since both are designed for bulk reads. The contents API remains the right tool for targeted file access, especially when you need decoded_content or a specific ref.

When you combine commits, releases, and contents in one script, keep the pagination and rate-limit behavior in mind: each list iteration consumes quota, and a long loop over a large repository can exhaust the hourly allowance. Reading g.get_rate_limit() before and during the loop gives you a way to stop cleanly instead of failing mid-run.

PyGithub Guide: Fetching Commits, Releases, and Repository Contents | RYUSLOG DEV