Python NumPy Mean, Median, Sum, Variance, and Standard Deviation
Compute mean, median, sum, variance, and standard deviation with Python NumPy. Covers axis handling, sample vs population variance, NaN-aware functions, dtype overflow, and empty arrays.
NumPy provides vectorized functions for computing common descriptive statistics: np.mean, np.median, np.sum, np.var, and np.std. They operate directly on arrays without explicit Python loops, which makes them the standard choice for basic statistical work in scientific Python.
Basic Usage of np.mean, np.median, and np.sum
The simplest case is a one-dimensional array. np.mean computes the arithmetic average, np.median finds the middle value, and np.sum totals all elements.
import numpy as np data = np.array([4, 8, 6, 5, 3, 7]) print(np.mean(data)) # 5.5 print(np.median(data)) # 5.5 print(np.sum(data)) # 33
np.mean and np.sum accept an optional dtype parameter to control the output type. For integer arrays, np.mean defaults to float64, while np.sum preserves the array's integer type unless you specify another dtype. This matters when you need a specific precision or when values are large.
Understanding the Axis Parameter for Multi-dimensional Arrays
For 2D or higher-dimensional arrays, the axis parameter determines along which dimension the operation is performed. Without axis, the function operates on the flattened array. With axis=0, the operation is column-wise; with axis=1, it is row-wise.
matrix = np.array([[1, 2, 3], [4, 5, 6]]) print(np.mean(matrix)) # 3.5 (all elements) print(np.mean(matrix, axis=0)) # [2.5 3.5 4.5] (per column) print(np.mean(matrix, axis=1)) # [2. 5.] (per row)
The same axis logic applies to np.sum, np.median, np.var, and np.std. For higher-dimensional arrays, you can pass a tuple to axis to operate on multiple axes at once. This is useful for reducing a 3D array to a 2D result without reshaping.
Variance and Standard Deviation: Population vs Sample
np.var and np.std compute the population variance and standard deviation by default, dividing by N. To obtain the sample statistics, which divide by N-1, pass ddof=1. This is a common source of confusion when comparing results with Python's statistics module or with spreadsheet functions.
data = np.array([2, 4, 4, 4, 5, 5, 7, 9]) print(np.var(data)) # 4.0 (population) print(np.var(data, ddof=1)) # 4.571 (sample) print(np.std(data)) # 2.0 print(np.std(data, ddof=1)) # 2.138
The ddof parameter is the delta degrees of freedom. Use the population version when your array represents the entire dataset; use ddof=1 when it is a sample that estimates a larger population.
Handling NaN Values with NaN-aware Functions
NumPy provides np.nanmean, np.nanmedian, np.nansum, np.nanvar, and np.nanstd to ignore NaN values. These are useful for real-world datasets with missing values.
data_with_nan = np.array([1.0, np.nan, 3.0, 4.0]) print(np.nanmean(data_with_nan)) # 2.6667 print(np.nanmedian(data_with_nan)) # 3.0 print(np.nansum(data_with_nan)) # 8.0
The NaN-aware functions skip NaN entries and compute the statistic over the remaining values. They also support the axis parameter. If all values in a slice are NaN, np.nansum returns 0.0; the other NaN-aware functions return NaN and NumPy may emit a RuntimeWarning.
Performance Considerations: Vectorization vs Python Loops
NumPy's statistical functions are vectorized and implemented in C, avoiding Python-level loop overhead. For large arrays, this can be much faster than explicit Python loops. Exact timings depend on hardware and array size, so use the following example as a demonstration, not a benchmark.
import time large_array = np.random.rand(10_000_000) # Vectorized start = time.time() mean_vec = np.mean(large_array) end = time.time() print(f'Vectorized: {end - start:.4f} seconds') # Python loop start = time.time() total = 0 for value in large_array: total += value mean_loop = total / len(large_array) end = time.time() print(f'Loop: {end - start:.4f} seconds')
For small arrays, the overhead of calling NumPy can be comparable to a Python loop, but for production workloads with large arrays, vectorization is the preferred approach.
Memory Usage and dtype Considerations
The memory footprint of the result depends on the input array's dtype. Integer sums can overflow silently when the result exceeds the range of the array's integer type. Choose a wider dtype when you know the result may not fit.
small = np.array([200, 200], dtype=np.uint8) print(np.sum(small)) # 144 due to 8-bit overflow print(np.sum(small, dtype=np.uint16)) # 400
Similarly, for integer and float inputs, np.mean returns a floating-point result. For integer arrays it defaults to float64; for high-precision needs you can control the computation with dtype=np.float64 or use an array with a wider floating-point type when available.
Edge Cases: Empty Arrays and Degenerate Inputs
When an array is empty, np.sum returns 0. np.mean, np.median, np.var, and np.std return NaN, usually accompanied by a RuntimeWarning. Check for empty inputs before applying statistical functions if your data pipeline can produce them.
empty = np.array([]) print(np.mean(empty)) # nan with RuntimeWarning print(np.sum(empty)) # 0
Choosing Between NumPy and Python's statistics Module
Python's standard library includes a statistics module with functions such as statistics.mean, statistics.median, statistics.variance, and statistics.stdev. These are useful for small lists and avoid a NumPy dependency, but they lack vectorization and axis support. For arrays, matrices, or data that benefits from vectorized operations, NumPy is the better choice.