VAITP Dataset

← Back to the dataset

CVE-2026-91943

SSRF in Crawl4AI PDFContentScrapingStrategy allows internal network access via malicious PDF URLs.

  • CVSS 8.3
  • 918
  • Input Validation and Sanitization
  • Remote

Crawl4AI before 0.9.3 contains a server-side request forgery vulnerability in PDFContentScrapingStrategy where _get_pdf_path() re-downloads targets with Python requests without egress validation. Authenticated attackers can supply URLs that redirect to internal addresses or use DNS rebinding to access internal services, exfiltrating responses through PDF text extraction in crawl results.

CWE
918
CVSS base score
8.3
Published
2026-09-15
OWASP
A10 Server-Side Request Forgery
Orthogonal defect classification
Checking
Code defect classification
Missing Check
Category
Input Validation and Sanitization
Subcategory
Insecure Direct Object References (IDOR)
Accessibility scope
Remote
Impact
Information Disclosure
Affected component
Python
Fixed by upgrading
Yes

Solution

Upgrade Crawl4AI to version 0.9.3 or later.

Vulnerable code sample

import requests
from urllib.parse import urljoin

class PDFContentScrapingStrategy:
    def __init__(self, base_url):
        self.base_url = base_url

    def _get_pdf_path(self, pdf_url):
        # VULNERABLE: no validation of pdf_url, allows SSRF
        response = requests.get(pdf_url, timeout=10)
        response.raise_for_status()
        return response.content

    def scrape(self, pdf_url):
        pdf_bytes = self._get_pdf_path(pdf_url)
        # pretend to extract text from PDF bytes
        return pdf_bytes[:100]

Patched code sample

import requests
from urllib.parse import urlparse, urljoin

class PDFContentScrapingStrategy:
    def __init__(self, base_url):
        self.base_url = base_url

    def _get_pdf_path(self, pdf_url):
        # FIX: validate scheme and netloc to prevent SSRF
        parsed = urlparse(pdf_url)
        if parsed.scheme not in ("http", "https"):
            raise ValueError("Unsupported URL scheme")
        if not parsed.netloc:
            raise ValueError("URL must have a network location")
        response = requests.get(pdf_url, timeout=10)
        response.raise_for_status()
        return response.content

    def scrape(self, pdf_url):
        pdf_bytes = self._get_pdf_path(pdf_url)
        # pretend to extract text from PDF bytes
        return pdf_bytes[:100]

Payload

__VAITP_MODEL_REFUSED__

Cite this entry

@misc{vaitp:cve202691943,
  title        = {{SSRF in Crawl4AI PDFContentScrapingStrategy allows internal network access via malicious PDF URLs.}},
  author       = {Bogaerts, Fr\'ed\'eric and Ivaki, Naghmeh and Fonseca, Jos\'e},
  year         = {2026},
  note         = {VAITP Python Vulnerability Dataset, entry CVE-2026-91943},
  howpublished = {\url{https://netpack.pt/vaitp/vulnerability/CVE-2026-91943/}}
}
Introducing the "VAITP dataset": a specialized repository of Python vulnerabilities and patches, meticulously compiled for the use of the security research community. As Python's prominence grows, understanding and addressing potential security vulnerabilities become crucial. Crafted by and for the cybersecurity community, this dataset offers a valuable resource for researchers, analysts, and developers to analyze and mitigate the security risks associated with Python. Through the comprehensive exploration of vulnerabilities and corresponding patches, the VAITP dataset fosters a safer and more resilient Python ecosystem, encouraging collaborative advancements in programming security.

The supreme art of war is to subdue the enemy without fighting.

Sun Tzu – “The Art of War”

:: Shaping the future through research and ingenuity ::