From fbffea5557e6eb972bb713d4b2de76d93d6b5221 Mon Sep 17 00:00:00 2001 From: Serhiy Storchaka Date: Sun, 13 Sep 2026 15:37:18 +0300 Subject: [PATCH] gh-93618: Document the memory usage of the incremental parsers (GH-156763) It was said that iterparse() can be useful for reading a large document without holding it wholly in memory, but the tree is only built incrementally, it is not freed incrementally. Document how to remove the processed elements, and that a custom target does not build a tree at all. (cherry picked from commit e7a393797a0df47a001673f4510c51aa272bc002) Co-authored-by: Serhiy Storchaka --- Doc/library/xml.etree.elementtree.rst | 37 +++++++++++++++++++++++++-- 1 file changed, 35 insertions(+), 2 deletions(-) diff --git a/Doc/library/xml.etree.elementtree.rst b/Doc/library/xml.etree.elementtree.rst index c8c42f66dcfed62..fff5382d5b4bcf1 100644 --- a/Doc/library/xml.etree.elementtree.rst +++ b/Doc/library/xml.etree.elementtree.rst @@ -160,8 +160,37 @@ some storage device. In such cases, blocking reads are unacceptable. Because it's so flexible, :class:`XMLPullParser` can be inconvenient to use for simpler use-cases. If you don't mind your application blocking on reading XML data but would still like to have incremental parsing capabilities, take a look -at :func:`iterparse`. It can be useful when you're reading a large XML document -and don't want to hold it wholly in memory. +at :func:`iterparse`. + +Note that both parsers build the tree incrementally: it is not freed +incrementally, so every parsed element is kept until the whole document is +read. To keep the memory usage low, get rid of the data which is not needed +any more. + +If the processed elements are large, it is enough to clear them. +This works wherever they are in the tree, +but the emptied elements are left in it:: + + for event, elem in ET.iterparse(source): + if elem.tag == 'record': + process(elem) + elem.clear() + +If an element has a large number of children, +remove the processed children from it:: + + for event, elem in ET.iterparse(source, events=('start', 'end')): + if event == 'start' and elem.tag == 'parent': + parent = elem + elif event == 'end' and elem.tag == 'child': + process(elem) + parent.remove(elem) + +These examples are not universal, +they only give an idea for two common cases. +If you do not need a tree at all, +parse with :class:`XMLParser` and a custom target instead; +it is not built then, and nothing has to be removed. Where *immediate* feedback through events is wanted, calling method :meth:`XMLPullParser.flush` can help reduce delay; @@ -635,6 +664,10 @@ Functions for applications where blocking reads can't be made. For fully non-blocking parsing, see :class:`XMLPullParser`. + The tree is only built incrementally, it is not freed incrementally: + every parsed element is kept until the whole document is read. + See :ref:`elementtree-pull-parsing` for how to keep the memory usage low. + .. note:: :func:`iterparse` only guarantees that it has seen the ">" character of a