Python memoryview cast: Reinterpreting Buffer Data
Learn how Python's memoryview.cast reinterprets buffer data without copying, including syntax, shape rules, zero-copy tradeoffs, and common pitfalls.
When working with binary data in Python, you often need to reinterpret the same underlying bytes as a different type without copying. The memoryview object provides this capability through its cast method, which changes the interpretation of the buffer's contents while sharing the same memory. This article explains how memoryview.cast works, when it is useful, and where it can lead to subtle bugs.
What memoryview.cast Does
A memoryview is a Python object that accesses the underlying memory of another object (such as bytes, bytearray, or an array.array) without copying it. The cast method returns a new memoryview that reinterprets the same memory region using a different format and shape. This is a zero-copy operation: no data is copied, only the metadata describing the buffer changes.
The method signature is:
memoryview.cast(format, shape=None)
The format parameter is a single-element native struct format string that defines the new element type (for example, 'B' for unsigned char, 'h' for short, 'i' for int, and 'd' for double). The optional shape parameter is a tuple that defines the new dimensions of the view. If omitted, the view is treated as one-dimensional.
Syntax and Parameters of cast
The cast method is called on an existing memoryview instance. The resulting view shares the same buffer but interprets it according to the new format and shape. The destination format must be a valid single-element native struct format, and the total byte size of the new view (the product of all shape dimensions times the new element size) must match the original buffer's byte length.
Here is a minimal example:
import array arr = array.array('h', [0x0102, 0x0304]) # two 2-byte shorts view = memoryview(arr) byte_view = view.cast('B') # reinterpret as 4 unsigned bytes print(byte_view.tolist()) # e.g., [2, 1, 4, 3] depending on endianness
The shape parameter allows you to create multi-dimensional views. For instance, you can reshape a flat buffer into a 2D matrix without copying:
buf = bytearray(12) mv = memoryview(buf) matrix = mv.cast('I', shape=(3, 1)) # 3 rows, 1 column of 4-byte ints
But note that the total number of bytes must match: shape[0] * shape[1] * itemsize must equal the original buffer's length.
How Format and Shape Change
The format string determines the element type and size. Common native formats include 'B' (unsigned char), 'b' (signed char), 'h' (short), 'i' (int), 'l' (long), 'q' (long long), 'f' (float), and 'd' (double). The shape tuple defines the dimensions; a one-element tuple yields a 1D view, while a longer tuple yields a multidimensional view.
When you cast, the original buffer's length in bytes is fixed, so the new view's total byte size must preserve that length. For example, a 16-byte buffer can be cast as:
- 16 bytes with format
'B'and shape(16,) - 8 shorts with format
'h'and shape(8,) - 4 ints with format
'i'and shape(4,) - 2 doubles with format
'd'and shape(2,)
If you provide a shape whose total byte size does not match the buffer size, Python raises a ValueError. This prevents accidental size mismatches.
Reinterpreting Bytes as Numeric Types
A common use case is reading a binary file or network packet and interpreting a sequence of bytes as integers or floats. Instead of using struct.unpack, which constructs new Python objects, you can use memoryview.cast to get a zero-copy view that you can index and slice.
For example, suppose you have a bytearray containing 8 bytes that represent two 32-bit integers:
data = bytearray([0x01, 0x00, 0x00, 0x00, 0x02, 0x00, 0x00, 0x00]) mv = memoryview(data) int_view = mv.cast('i') # two ints print(int_view[0], int_view[1]) # 1, 2 on little-endian
You can also use slicing on the cast view, but be aware that slicing a memoryview returns a new memoryview, not a copy. This can be useful for processing chunks without copying.
Zero-Copy Behavior and Performance
The main benefit of memoryview.cast is that it avoids copying the underlying buffer. For example, list(mv.cast('i')) creates new Python int objects, but it does not duplicate the bytearray's memory. If you keep a memoryview and iterate over it instead of building a list, you can process large binary data without allocating separate Python objects for the whole dataset.
Performance gains are not automatic. Creating a memoryview and casting it has overhead. For small data, struct.unpack may be faster because it avoids the extra view layer. For large data, keeping the data in a shared buffer can be worthwhile, especially if you need to inspect or pass the buffer to another buffer-aware function multiple times.
Another important note: memoryview.cast does not change byte order. The interpretation follows the native endianness of the platform. If you need a specific byte order, handle it separately, for example with struct or by converting the bytes manually.
Limitations and Pitfalls
One common pitfall is assuming that cast changes the underlying data. It does not; it only changes how the bytes are interpreted. If you modify the cast view, you modify the original buffer. This is expected, but it can be surprising if you forget that the view is not a copy.
Another limitation is that you can only call cast on a memoryview, not directly on the original object. The original must expose the buffer protocol so it can be wrapped by memoryview(). Some exporters, such as bytes, produce read-only memoryviews; attempting to write through a read-only view raises TypeError.
The destination format must be a single-element native struct format. You cannot use compound formats like '2i'; instead, use a shape of (2,) with format 'i'. The cast method does not support nested structures or arrays within a single element.
cast also requires a C-contiguous source view. If the source is not contiguous, or if the destination shape does not match the byte length of the buffer, Python raises an exception. For ordinary bytearray, bytes, and array.array objects, the memory layout is straightforward, but check the documentation for less common buffer exporters.
Choosing Between cast and Other Conversion Methods
For simple conversions, struct.unpack is often more readable and lets you control byte order explicitly. For example, if the two integers in data are little-endian:
import struct values = struct.unpack('<2i', data)
This returns a tuple of Python ints, but it constructs new objects rather than keeping a view of the buffer. memoryview.cast is better when you need to keep the data in a buffer-like form, perform repeated indexing, or pass the view to functions that expect a buffer (like numpy.frombuffer).
If you are working with large binary data and need to process it in a memory-efficient way, memoryview.cast can be a good choice. But if you need to handle endianness or convert to a list of Python objects, struct might be simpler.
Another alternative is array.array with frombytes, which also creates a copy. The choice depends on whether you need zero-copy semantics and whether you can tolerate platform-dependent endianness.
Advanced Use: Casting with Multidimensional Shape
When working with image data or multi-dimensional arrays, memoryview.cast can reshape a flat buffer into a matrix-like view. This is useful when you want to pass a buffer to a C extension or a library that expects a specific shape without copying.
For example, a 24-byte buffer can be interpreted as a 2x3 array of 4-byte floats:
buf = bytearray(24) mv = memoryview(buf) float_matrix = mv.cast('f', shape=(2, 3)) float_matrix[0, 1] = 3.14
This modifies the original buffer. The shape is stored in the memoryview, and indexing returns scalar values. You cannot change the shape after creation; you would need to create a new cast view.
One limitation is that the destination shape must cover the whole buffer: the product of the dimensions times the element size must equal the buffer length. If you need a view that does not cover the entire buffer with a single fixed item size, cast is not the right tool; a library such as numpy may be a better fit.
Remember that memoryview.cast relies on the buffer protocol. Once you have a memoryview from a buffer-aware exporter such as bytearray, bytes, array.array, or many other built-in types, you can call cast when the view is C-contiguous and the destination format and shape fit the existing buffer. Always check the documentation for the specific object you are wrapping.