VAITP Dataset

← Back to the dataset

CVE-2026-80047

Remote Python file written to cache before trust check in Transformers load_custom_generate().

  • CVSS 7.8
  • 273
  • Input Validation and Sanitization
  • Remote

A vulnerability in Hugging Face Transformers (versions >= 4.49.0 and <= 5.8.1) allows remote Python files to be written to local disk without user consent when using GenerativePreTrainedModel.load_custom_generate(). The function fetches and caches a remote module file before performing the required trust_remote_code consent check, inverting the security model enforced by other code-loading paths (such as AutoConfig, AutoModel, and AutoTokenizer). As a result, attacker‑controlled Python code from custom_generate/generate.py is copied into the user’s ~/.cache/huggingface/modules directory even if the user declines the trust prompt. Although execution is correctly gated, the file write is not reversible and can persist across sessions. This can lead to persistent, unauthorized files on disk and stale cache collisions where cached attacker code may later be executed during trusted model loads. The issue stems from an unconditional file write in dynamic_module_utils.py prior to any trust verification.

CWE
273
CVSS base score
7.8
Published
2026-09-01
OWASP
A08 Software and Data Integrity Failures
Orthogonal defect classification
Checking
Code defect classification
Missing Check
Category
Input Validation and Sanitization
Subcategory
Insecure Parsing or Deserialization
Accessibility scope
Remote
Impact
Unauthorized Access
Affected component
Hugging Face Transformers
Fixed by upgrading
Yes

Solution

Upgrade Hugging Face Transformers to v5.8.2 (or any later release where the fix is included).

Vulnerable code sample

import os
import hashlib
import urllib.request
from pathlib import Path

CACHE_DIR = Path.home() / ".cache" / "huggingface" / "modules"

def load_custom_generate(model_id: str, trust_remote_code: bool = False):
    """
    Load a custom generation script for a model, fetching it from a remote URL.
    """
    # Build URL to remote module (e.g., https://example.com/{model_id}/custom_generate.py)
    remote_url = f"https://example.com/{model_id}/custom_generate.py"

    # Compute cache filename based on URL hash
    filename_hash = hashlib.sha256(remote_url.encode()).hexdigest()
    cache_path = CACHE_DIR / f"{filename_hash}.py"

    # VULNERABLE: writes remote file before consent check
    os.makedirs(CACHE_DIR, exist_ok=True)
    with urllib.request.urlopen(remote_url) as resp, open(cache_path, "wb") as f:
        f.write(resp.read())

    # Prompt user for trust if not already trusted
    if not trust_remote_code:
        raise PermissionError("Remote code execution not permitted. Set trust_remote_code=True to proceed.")

    # At this point the file is already on disk; import it
    spec = __import__(cache_path.stem)
    return spec.generate

Patched code sample

import os
import hashlib
import urllib.request
from pathlib import Path

CACHE_DIR = Path.home() / ".cache" / "huggingface" / "modules"

def load_custom_generate(model_id: str, trust_remote_code: bool = False):
    """
    Load a custom generation script for a model, fetching it from a remote URL.
    """
    remote_url = f"https://example.com/{model_id}/custom_generate.py"
    filename_hash = hashlib.sha256(remote_url.encode()).hexdigest()
    cache_path = CACHE_DIR / f"{filename_hash}.py"

    # FIX: defer writing remote file until after trust verification
    if not trust_remote_code:
        raise PermissionError("Remote code execution not permitted. Set trust_remote_code=True to proceed.")

    os.makedirs(CACHE_DIR, exist_ok=True)
    with urllib.request.urlopen(remote_url) as resp, open(cache_path, "wb") as f:
        f.write(resp.read())

    spec = __import__(cache_path.stem)
    return spec.generate

Payload

__VAITP_MODEL_REFUSED__

Cite this entry

@misc{vaitp:cve202680047,
  title        = {{Remote Python file written to cache before trust check in Transformers load_custom_generate().}},
  author       = {Bogaerts, Fr\'ed\'eric and Ivaki, Naghmeh and Fonseca, Jos\'e},
  year         = {2026},
  note         = {VAITP Python Vulnerability Dataset, entry CVE-2026-80047},
  howpublished = {\url{https://netpack.pt/vaitp/vulnerability/CVE-2026-80047/}}
}
Introducing the "VAITP dataset": a specialized repository of Python vulnerabilities and patches, meticulously compiled for the use of the security research community. As Python's prominence grows, understanding and addressing potential security vulnerabilities become crucial. Crafted by and for the cybersecurity community, this dataset offers a valuable resource for researchers, analysts, and developers to analyze and mitigate the security risks associated with Python. Through the comprehensive exploration of vulnerabilities and corresponding patches, the VAITP dataset fosters a safer and more resilient Python ecosystem, encouraging collaborative advancements in programming security.

The supreme art of war is to subdue the enemy without fighting.

Sun Tzu – “The Art of War”

:: Shaping the future through research and ingenuity ::