Back to Blog
Python

Python SciPy Sparse Matrices and Linear Algebra

Learn how SciPy sparse matrices store only nonzero values, choose among CSR, CSC, COO, LIL, and DIA formats, and solve sparse linear systems and eigenvalue problems.

scipysparse matriceslinear algebraspsolvenumerical computing
A visual metaphor for sparse matrices showing a grid with only a few highlighted cells, representing nonzero entries, and a linear algebra solver icon.

Sparse matrices are essential when your data is large and mostly zeros. This guide explains how to create them with SciPy, choose the right sparse format, and solve linear systems and eigenvalue problems efficiently using scipy.sparse and scipy.sparse.linalg.

Why Sparse Matrices Matter

Many real-world datasets—graph adjacency matrices, finite-element meshes, recommendation systems—are sparse: most entries are zero. Storing them as dense NumPy arrays wastes memory and compute. For example, a 100,000 x 100,000 matrix with 0.1% nonzero entries (10 million values) would consume about 80 GB as a dense float64 array, while a sparse format would need only on the order of 100–200 MB depending on the format and index precision. Sparse structures store only nonzero values and their positions, enabling operations that scale with the number of nonzeros rather than the full dimensions.

SciPy provides a family of sparse matrix classes in scipy.sparse, each with different strengths. The choice of format directly affects the performance of arithmetic, slicing, and conversions.

Creating Sparse Matrices

The simplest way to create a sparse matrix is from a dense NumPy array using csr_matrix or coo_matrix:

import numpy as np from scipy.sparse import csr_matrix, coo_matrix dense = np.array([[0, 0, 1], [2, 0, 0], [0, 3, 0]]) sparse_csr = csr_matrix(dense) print(sparse_csr)

Output:

  (0, 2)\t1
  (1, 0)\t2
  (2, 1)\t3

For large matrices, constructing from coordinate lists avoids building a dense array. Use coo_matrix with row, column, and data arrays:

rows = [0, 1, 2] cols = [2, 0, 1] data = [1, 2, 3] sparse_coo = coo_matrix((data, (rows, cols)), shape=(3, 3))

COO is efficient for assembly but not for arithmetic. Convert to CSR or CSC before operations:

sparse_csr = sparse_coo.tocsr()

Choosing the Right Sparse Format

SciPy offers several formats, each optimized for different tasks. The most common are:

FormatBest forTypical use case
CSR (Compressed Sparse Row)Row slicing, matrix-vector products, general arithmeticMost linear algebra operations
CSC (Compressed Sparse Column)Column slicing, solving systems with column-oriented solversWhen column access is frequent
COO (Coordinate)Fast assembly, incremental constructionBuilding matrices from data
LIL (List of Lists)In-place element assignmentModifying individual entries
DIA (Diagonal)Storing banded matricesFinite-difference discretizations

For most linear algebra, csr_matrix is a good default: it supports efficient matrix-vector products and row slicing, and it is accepted by the solvers in scipy.sparse.linalg. If you need to modify entries frequently, start with lil_matrix and convert to csr when done.

Basic Linear Algebra Operations

Sparse matrices support the same arithmetic operators as dense arrays, but with different performance characteristics.

Matrix-vector multiplication is fast and uses the sparsity structure:

v = np.array([1, 2, 3]) result = sparse_csr.dot(v) # or sparse_csr @ v

Matrix-matrix multiplication works as expected, but the result may have more nonzeros than either input:

product = sparse_csr @ sparse_csr.T

Element-wise operations like addition and subtraction are also supported, but the result's sparsity pattern is the union of the operands. If both matrices are large and sparse, the result may have many more nonzeros.

Solving Linear Systems

The scipy.sparse.linalg module provides spsolve for direct solving of sparse linear systems using sparse LU factorization. It is a good fit for moderate-sized systems; for very large systems, the memory required by the factorization may become a bottleneck, and iterative solvers are often a better choice.

import numpy as np from scipy.sparse import csr_matrix from scipy.sparse.linalg import spsolve A = csr_matrix([[4, 1, 0], [1, 3, 1], [0, 1, 2]]) b = np.array([1, 2, 3]) x = spsolve(A, b)

Pass A as a CSR or CSC matrix when possible. spsolve returns a dense array.

For very large systems, iterative solvers like cg (conjugate gradient) or gmres are more memory-efficient because they do not factor the matrix. These are available in scipy.sparse.linalg and accept a linear operator or sparse matrix. cg is designed for symmetric positive-definite systems; gmres works for a broader class of nonsymmetric problems.

from scipy.sparse.linalg import cg x, info = cg(A, b, tol=1e-6)

Eigenvalue Problems

Finding eigenvalues of sparse matrices is common in physics and graph analysis. scipy.sparse.linalg.eigs computes a few eigenvalues for general matrices, while eigsh is optimized for symmetric matrices.

from scipy.sparse.linalg import eigsh # A is a sparse symmetric matrix eigenvalues, eigenvectors = eigsh(A, k=3, which='SM')

eigsh uses ARPACK and requires a symmetric matrix. For non-symmetric matrices, use eigs. The k parameter specifies how many eigenvalues to compute; which selects the smallest ('SM') or largest ('LM') magnitude. ARPACK is generally most reliable for largest-magnitude eigenvalues, so which='SM' can require more iterations.

These functions operate on the sparse representation, so they avoid forming the dense matrix that numpy.linalg.eig would require.

Performance and Memory Considerations

The main benefit of sparse matrices is reduced memory usage and faster operations when the sparsity is high. However, certain operations can be unexpectedly expensive:

  • Converting between formats (for example, csr to csc) requires reordering the nonzeros and can be costly for large matrices.
  • Slicing a CSR matrix by rows is efficient, but slicing by columns is slow. Use CSC for column slicing.
  • Element-wise operations that create many new nonzeros, such as adding two matrices with disjoint sparsity patterns, can consume significant memory.
  • Matrix-vector multiplication is O(nnz) and very fast, but matrix-matrix multiplication can produce a much denser result if the sparsity patterns are unfavorable.

Always check the number of nonzeros after operations using .nnz to avoid accidental densification.

Common Pitfalls and Best Practices

A frequent mistake is converting a sparse matrix to a dense array with .toarray() just to inspect it. For large matrices this defeats the purpose. Use A.nnz or inspect a small slice instead.

Another issue is reaching for numpy.linalg.solve, which expects dense ndarray inputs and does not exploit sparsity. Use spsolve or an iterative solver from scipy.sparse.linalg.

When building a sparse matrix incrementally, avoid repeatedly inserting into a csr_matrix. Use lil_matrix or coo_matrix for assembly and convert once.

Finally, be mindful of the data type. Sparse matrices default to float64. If your data is integer or boolean, you may need to cast explicitly to avoid unexpected type promotion in operations.

By understanding the formats and the available linear algebra routines, you can handle large-scale problems that would be impossible with dense arrays.

Working With SciPy Sparse Matrices: Formats, Solvers, and Eigenvalues | RYUSLOG DEV