Convert XML to Dictionary and Back with Python xmltodict
Learn how to use Python's xmltodict to convert XML to dictionaries and back, including attribute handling, namespaces, round-trip limitations, and when to prefer ElementTree.
Python’s xmltodict library converts XML into Python dictionaries and can serialize those dictionaries back to XML. This guide covers parsing, serialization, attributes, namespaces, and the limitations to keep in mind before using it in production.
Why Use xmltodict for XML and Dictionary Conversion
xmltodict is designed for developers who want to work with XML as if it were JSON. Instead of navigating element trees with ElementTree, you can read and update values by key in a dictionary (or OrderedDict). That convenience also extends in the other direction: when you have a dictionary and need XML output, unparse can generate it without writing element-building code by hand.
Installing xmltodict
Install the library with pip:
pip install xmltodict
The package has no required third-party dependencies, so it works in most Python environments.
Converting XML to a Dictionary
The parse function accepts an XML string or a file-like object and returns an OrderedDict by default. Here is a minimal example:
import xmltodict xml_string = ''' <book> <title>Learning Python</title> <author>Mark Lutz</author> <price>39.99</price> </book> ''' book_dict = xmltodict.parse(xml_string) print(book_dict)
The output is:
OrderedDict([('book', OrderedDict([('title', 'Learning Python'), ('author', 'Mark Lutz'), ('price', '39.99')]))])
Each element becomes a key, and its text content becomes the value. If the same child element appears multiple times, xmltodict collects the values into a list. If you need a single occurrence to be a list too, pass a force_list mapping to parse:
book_dict = xmltodict.parse(xml_string, force_list={'author': True})
If you prefer a regular dict to an OrderedDict, pass dict_constructor=dict.
Converting a Dictionary Back to XML
To serialize a dictionary to XML, use unparse. The dictionary must have exactly one top-level key, which becomes the root element.
import xmltodict book_dict = {'book': {'title': 'Learning Python', 'author': 'Mark Lutz', 'price': '39.99'}} xml_output = xmltodict.unparse(book_dict, pretty=True) print(xml_output)
The pretty=True argument adds indentation. You can also pass indent to control the indentation string. By default, unparse writes an XML declaration; pass full_document=False if you need to omit it.
Handling Attributes and Text Nodes
XML attributes are represented with a leading @ in dictionary keys. For example:
xml_string = ''' <book id='123'> <title>Learning Python</title> <author>Mark Lutz</author> </book> ''' book_dict = xmltodict.parse(xml_string) print(book_dict)
Output:
OrderedDict([('book', OrderedDict([('@id', '123'), ('title', 'Learning Python'), ('author', 'Mark Lutz')]))])
When an element has both attributes and text content, the text is stored under the #text key:
<message id='42'>Hello, world!</message>
This parses to:
OrderedDict([('message', OrderedDict([('@id', '42'), ('#text', 'Hello, world!')]))])
You can customize these keys with the attr_prefix and cdata_key parameters in both parse and unparse.
Dealing with Namespaces
By default, xmltodict does not process namespaces, so keys retain the namespace prefix exactly as it appears in the XML:
xml_string = ''' <ns:book xmlns:ns='http://example.com/ns'> <ns:title>Learning Python</ns:title> </ns:book> ''' book_dict = xmltodict.parse(xml_string) print(book_dict)
Output:
OrderedDict([('ns:book', OrderedDict([('ns:title', 'Learning Python')]))])
If you set process_namespaces=True, the keys use the full namespace URI with the local name:
book_dict = xmltodict.parse(xml_string, process_namespaces=True) print(book_dict)
Output:
OrderedDict([('http://example.com/ns:book', OrderedDict([('http://example.com/ns:title', 'Learning Python')]))])
The namespace_separator parameter controls the separator between the URI and local name. When unparsing namespace-qualified XML, pass the same namespace configuration so the output is consistent.
Round-Trip Limitations and Common Pitfalls
XML-to-dictionary-to-XML round trips are not always lossless. xmltodict makes pragmatic choices that can change the original document. Keep these limitations in mind:
- XML attribute order is not meaningful in most XML applications, and a round trip may not preserve the original attribute order exactly. Do not rely on it.
- Comments and processing instructions are dropped.
- By default, CDATA sections are treated as plain text, so the
<![CDATA[...]]>wrapper is not preserved. unparsewrites an XML declaration by default; passfull_document=Falseto omit it.- Formatting, indentation, and insignificant whitespace are not preserved.
Consider this XML:
<root> <item>value</item> </root>
After a round trip, the output might be:
<root><item>value</item></root>
If you need to keep the exact byte representation or comments from the original file, xmltodict is not the right tool.
Performance and Memory Considerations
xmltodict builds a complete dictionary in memory. For large XML files, this can consume significant memory and increase latency. If you are processing very large documents, consider a streaming parser such as xml.etree.ElementTree.iterparse or lxml with iterative parsing. For configuration files, API responses, and other small-to-medium XML documents, xmltodict is often sufficient.
When to Choose xmltodict Over ElementTree
ElementTree is part of the standard library and gives you more direct control over namespace handling, element iteration, and streaming. It is often the better choice when you are working with very large documents or need to avoid loading the entire XML tree into memory. Keep in mind that the standard-library ElementTree also has limitations when preserving comments and original formatting; if those details matter, consider lxml or a lower-level parser.
Use xmltodict when:
- You are integrating with an API that returns XML and you want to treat it like JSON.
- You need to convert a small XML document to a dictionary for testing or data transformation.
- You want to generate XML from a Python dict without writing verbose ElementTree code.
Choose ElementTree when:
- You need a streaming parser that avoids building a full document tree in memory.
- You need iterative control over elements and namespace events.
- You want to stay within the standard library and do not need a dict-like view of the document.
The decision depends on whether you value convenience and readability over exact fidelity and scalability.