Python Int Conversion: Syntax, Edge Cases, and Pitfalls
Learn how Python's int() converts strings, floats, and other values to integers, handles invalid input, and works with custom bases.
Python's int() constructor is the primary tool for converting other types to integers. Whether you're parsing user input, reading configuration values, or normalizing data from an API, understanding how int() behaves with different inputs—and where it fails—is essential for writing robust code. This article covers the mechanics of converting values with int(), focusing on practical usage, error handling, and edge cases that trip up developers.
Using int() with Strings and Numbers
The most common use of int() is converting a string that represents an integer to an int object. With no base argument, the constructor accepts a string, a bytes or bytearray object, a float, another integer, or an object whose __int__() method returns an integer. For strings, the default base is 10: the input must look like an integer literal, with an optional sign and surrounding whitespace allowed. For example:
value = int("42") print(value) # 42 negative = int("-7") print(negative) # -7 with_spaces = int(" 100 ") print(with_spaces) # 100
The string may optionally have a leading + sign. However, it cannot contain commas, decimal points, or underscores unless the underscores form a valid Python grouping separator. Attempting to convert "3.14" raises a ValueError, as does passing an empty string.
When the argument is a float, int() truncates toward zero, discarding the fractional part. This is not rounding; int(3.99) yields 3, and int(-3.99) yields -3. If you need rounding instead of truncation, use round() before conversion, or choose math.floor() / math.ceil() explicitly.
Handling Invalid Input and ValueError
int() raises a ValueError when the input cannot be parsed as an integer. This is the most common failure point in production code, especially when dealing with external data. A robust conversion routine should catch this exception and decide how to proceed—whether to skip the value, use a default, or log the error.
def safe_int(value, default=0): try: return int(value) except (ValueError, TypeError): return default
Note that TypeError is raised when the argument is None or an unsupported type like a list. Catching both exceptions covers most real-world inputs. However, overusing a generic safe_int can mask data quality issues. In data pipelines, it's often better to let the exception propagate so the problem is visible early.
Converting with Custom Base (int(x, base))
The two-argument form int(x, base) interprets the string x in the given base. base must be 0 or an integer between 2 and 36 inclusive. This is useful for parsing hexadecimal, binary, or octal strings from text. Base-2, -8, and -16 strings may include the matching prefix (0b, 0o, or 0x); base 0 tells Python to infer the base from the prefix, as with an integer literal in code.
hex_value = int("ff", 16) print(hex_value) # 255 binary_value = int("1010", 2) print(binary_value) # 10 with_prefix = int("0x1A", 16) print(with_prefix) # 26 # Using base 0 to auto-detect auto = int("0x1A", 0) print(auto) # 26
With base=0, a string such as "010" is not treated as decimal 10; it is interpreted as a code literal and is invalid. If you need octal, pass an explicit base such as int("010", 8). The base argument only applies to string, bytes, or bytearray inputs; passing a float or integer with a base raises TypeError.
Converting Floats and Truncation Behavior
As mentioned, int() truncates a float toward zero. This behavior is consistent across positive and negative numbers, but it can surprise developers who expect rounding. Consider a scenario where you're converting a price from a floating-point calculation to cents:
price = 19.99 cents = int(price * 100) print(cents) # 1998, not 1999
The result is 1998 because floating-point representation of 19.99 * 100 is slightly less than 1999.0. This is a classic precision issue. For financial calculations, use Decimal or round explicitly before converting. If you must use int() on a float, be aware that it truncates, not rounds.
Performance and Runtime Considerations
int() is a built-in implemented at the C level in CPython and is generally fast. The main conversion cost is string parsing, so longer strings take longer than short ones. In ordinary applications this overhead is negligible. Calling int() on a value that is already an integer does not parse a string, but it still costs a function call and type check. If a hot loop converts millions of values, profile before trying to optimize.
Common Pitfalls and Edge Cases
Several edge cases cause unexpected ValueError or TypeError exceptions. One is converting a string with a decimal point—int("3.0") fails even though the value is mathematically an integer. Another is converting a string with leading zeros, which works fine with the default base 10 (int("007") returns 7). However, in Python 3, a string like "0b101" is not automatically interpreted as binary unless you pass base=2 or base=0.
Another pitfall is using int() on None, float('nan'), or float('inf'). int(None) raises TypeError, int(float('nan')) raises ValueError, and int(float('inf')) raises OverflowError. These exceptions are distinct and often need separate handling.
Choosing the Right Conversion Approach
For most cases, int() is the correct choice. But there are alternatives: float() for floating-point conversion, eval() for arbitrary expressions (which you should avoid for security reasons), and ast.literal_eval() for safe evaluation of literals. If you need to parse a string that may contain a decimal, you can try int() first and fall back to float():
def parse_number(s): try: return int(s) except ValueError: return float(s)
This returns an int for decimal-free strings and a float for values like "3.14" or "1e3", but it still raises ValueError for invalid input. The choice between int() and float() depends on the expected data format. For configuration values that should always be integers, int() with error handling is sufficient. For user input that might contain decimals, a more flexible parser is needed.
Underscores and Input Validation
Python 3.6 introduced underscores as visual separators in numeric literals, and int() also accepts them in strings when they form a valid integer literal. For example, int("1_000_000") returns 1000000. Underscores must be placed where Python integer-literal syntax permits them; they cannot appear at the start or end of the literal, and repeated underscores are invalid.
When parsing external data, remember that a single misplaced underscore makes the conversion fail. Also, int() parses integer literals, not arbitrary Unicode text: it does not accept every character that str.isdigit() returns true for. For strict validation, use a regex that matches the exact input format you expect.