CVE-2026-91943
SSRF in Crawl4AI PDFContentScrapingStrategy allows internal network access via malicious PDF URLs.
- CVSS 8.3
- 918
- Input Validation and Sanitization
- Remote
Crawl4AI before 0.9.3 contains a server-side request forgery vulnerability in PDFContentScrapingStrategy where _get_pdf_path() re-downloads targets with Python requests without egress validation. Authenticated attackers can supply URLs that redirect to internal addresses or use DNS rebinding to access internal services, exfiltrating responses through PDF text extraction in crawl results.
- CWE
- 918
- CVSS base score
- 8.3
- Published
- 2026-09-15
- OWASP
- A10 Server-Side Request Forgery
- Orthogonal defect classification
- Checking
- Code defect classification
- Missing Check
- Category
- Input Validation and Sanitization
- Subcategory
- Insecure Direct Object References (IDOR)
- Accessibility scope
- Remote
- Impact
- Information Disclosure
- Affected component
- Python
- Fixed by upgrading
- Yes
Solution
Upgrade Crawl4AI to version 0.9.3 or later.
Vulnerable code sample
import requests
from urllib.parse import urljoin
class PDFContentScrapingStrategy:
def __init__(self, base_url):
self.base_url = base_url
def _get_pdf_path(self, pdf_url):
# VULNERABLE: no validation of pdf_url, allows SSRF
response = requests.get(pdf_url, timeout=10)
response.raise_for_status()
return response.content
def scrape(self, pdf_url):
pdf_bytes = self._get_pdf_path(pdf_url)
# pretend to extract text from PDF bytes
return pdf_bytes[:100]Patched code sample
import requests
from urllib.parse import urlparse, urljoin
class PDFContentScrapingStrategy:
def __init__(self, base_url):
self.base_url = base_url
def _get_pdf_path(self, pdf_url):
# FIX: validate scheme and netloc to prevent SSRF
parsed = urlparse(pdf_url)
if parsed.scheme not in ("http", "https"):
raise ValueError("Unsupported URL scheme")
if not parsed.netloc:
raise ValueError("URL must have a network location")
response = requests.get(pdf_url, timeout=10)
response.raise_for_status()
return response.content
def scrape(self, pdf_url):
pdf_bytes = self._get_pdf_path(pdf_url)
# pretend to extract text from PDF bytes
return pdf_bytes[:100]Payload
__VAITP_MODEL_REFUSED__
Cite this entry
@misc{vaitp:cve202691943,
title = {{SSRF in Crawl4AI PDFContentScrapingStrategy allows internal network access via malicious PDF URLs.}},
author = {Bogaerts, Fr\'ed\'eric and Ivaki, Naghmeh and Fonseca, Jos\'e},
year = {2026},
note = {VAITP Python Vulnerability Dataset, entry CVE-2026-91943},
howpublished = {\url{https://netpack.pt/vaitp/vulnerability/CVE-2026-91943/}}
}
Introducing the "VAITP dataset": a specialized repository of Python vulnerabilities and patches, meticulously compiled for the use of the security research community. As Python's prominence grows, understanding and addressing potential security vulnerabilities become crucial. Crafted by and for the cybersecurity community, this dataset offers a valuable resource for researchers, analysts, and developers to analyze and mitigate the security risks associated with Python. Through the comprehensive exploration of vulnerabilities and corresponding patches, the VAITP dataset fosters a safer and more resilient Python ecosystem, encouraging collaborative advancements in programming security.
The supreme art of war is to subdue the enemy without fighting.
Sun Tzu – “The Art of War”
:: Shaping the future through research and ingenuity ::
