VAITP Dataset

← Back to the dataset

CVE-2025-14927

Hugging Face Transformers convert_config allows RCE via code injection.

  • CVSS 7.8
  • CWE-94
  • Input Validation and Sanitization
  • Remote

Hugging Face Transformers SEW-D convert_config Code Injection Remote Code Execution Vulnerability. This vulnerability allows remote attackers to execute arbitrary code on affected installations of Hugging Face Transformers. User interaction is required to exploit this vulnerability in that the target must convert a malicious checkpoint. The specific flaw exists within the convert_config function. The issue results from the lack of proper validation of a user-supplied string before using it to execute Python code. An attacker can leverage this vulnerability to execute code in the context of the current user. . Was ZDI-CAN-28252.

CVSS base score
7.8
Published
2025-12-23
OWASP
A03 Injection
Orthogonal defect classification
Checking
Code defect classification
Missing Check
Category
Input Validation and Sanitization
Subcategory
Command Injection
Accessibility scope
Remote
Impact
Arbitrary Code Execution
Fixed by upgrading
Yes

Solution

Upgrade Hugging Face Transformers to version 4.41.0 or later.

Vulnerable code sample

import json
import os

#
# This code is a hypothetical representation of the vulnerability described in
# CVE-2025-14927. The actual CVE and the associated code may not exist publicly.
# This example is created for educational purposes to demonstrate the described
# vulnerability pattern.
#
# The vulnerability lies in using `eval()` on a user-controlled string from a
# configuration file without proper sanitization.
#

def convert_config(config: dict):
    """
    Hypothetical function to convert a model configuration from one format to another.
    This function is based on the description of CVE-2025-14927.
    """
    print("[+] Starting conversion of SEW-D config...")

    # The configuration is expected to have certain keys that might need
    # to be dynamically interpreted. For example, an activation function name.
    activation_function_string = config.get("activation")

    if not activation_function_string:
        raise ValueError("Configuration must contain an 'activation' key.")

    print(f"[+] Processing 'activation' key with value: {activation_function_string}")

    # VULNERABLE PART: The string from the config is passed directly to `eval()`.
    # This assumes the string is a benign Python expression (e.g., "torch.nn.GELU"),
    # but an attacker can provide a string containing arbitrary code.
    # The lack of validation before this call is the core of the vulnerability.
    try:
        converted_activation = eval(activation_function_string)
        print(f"[+] Successfully evaluated activation function: {converted_activation}")
    except Exception as e:
        print(f"[-] Failed to evaluate activation function: {e}")
        return None

    # ... more conversion logic would follow here ...

    print("[+] Conversion process finished.")
    return {"converted_activation": converted_activation}

def main():
    """
    Simulates a user attempting to convert a malicious checkpoint's configuration.
    """
    # An attacker crafts a malicious config file (e.g., config.json)
    # that will be loaded and passed to the vulnerable function.
    # The payload executes a command on the system, in this case, printing a message.
    # A real attacker would use a more malicious command.
    malicious_payload = "__import__('os').system('echo VULNERABILITY CVE-2025-14927 TRIGGERED')"

    malicious_config_data = {
        "model_type": "sew-d",
        "hidden_size": 768,
        # The 'activation' key contains the code to be executed.
        "activation": malicious_payload
    }

    # In a real scenario, this dictionary would be loaded from a JSON file
    # provided by the attacker (e.g., inside a model checkpoint).
    print("--- Simulating conversion of a malicious config file ---")
    
    # The user runs the conversion script on the malicious config.
    convert_config(malicious_config_data)

    print("\n--- Simulation complete ---")

if __name__ == "__main__":
    main()

Patched code sample

import ast
import json
import os

# This code demonstrates a secure way to handle configuration data that
# might otherwise be processed by a dangerous function like eval().
# The vulnerability CVE-2025-14927 is hypothetical, but this pattern
# represents a real-world fix for the described flaw.

def convert_sew_d_config_fixed(malicious_checkpoint_config: str):
    """
    Safely converts a model configuration from a JSON string.

    This function represents the fix for a code injection vulnerability.
    Instead of using `eval()` on user-controlled strings from a config file,
    it uses `ast.literal_eval()`, which only evaluates static literals
    (strings, numbers, tuples, lists, dicts, booleans, None) and raises an
    error for any other content, such as function calls or expressions.

    Args:
        malicious_checkpoint_config: A JSON string representing the
                                     configuration from a malicious checkpoint.
    """
    print("--- Attempting to convert a potentially malicious config ---")
    try:
        config = json.loads(malicious_checkpoint_config)
        
        new_config = {}
        
        # Safely copy over expected values
        if "model_type" in config:
            new_config["model_type"] = config["model_type"]
        
        # --- THE FIX IS APPLIED HERE ---
        # The original vulnerable code would have used `eval()` on a user-controlled
        # string, like this:
        #
        # VULNERABLE: new_config["some_value"] = eval(config.get("dynamic_key"))
        #
        # The fix is to use a safe parser like `ast.literal_eval`.
        
        dynamic_value_str = config.get("some_dynamic_parameter")
        
        if dynamic_value_str:
            print(f"Processing dynamic parameter: '{dynamic_value_str}'")
            try:
                # ast.literal_eval will safely parse Python literals.
                # It will raise a ValueError if the string contains code,
                # function calls, or other non-literal expressions.
                new_config["some_value"] = ast.literal_eval(dynamic_value_str)
                print("SUCCESS: Parameter was a safe literal.")
            except (ValueError, TypeError, SyntaxError, MemoryError, RecursionError) as e:
                # This block is executed for the malicious payload, preventing RCE.
                print(f"SAFE ABORT: The parameter is not a valid literal. Error: {e}")
                print("Malicious code execution was prevented.")
                # Handle the error, e.g., by using a default value or raising an exception.
                raise ValueError("Invalid and potentially malicious configuration value detected.")
        
        return new_config

    except json.JSONDecodeError:
        print("Error: Invalid JSON in config.")
        return None
    except Exception as e:
        print(f"An error occurred during conversion: {e}")
        return None


if __name__ == '__main__':
    # This represents the content of a malicious 'config.json' file
    # provided by an attacker. The 'some_dynamic_parameter' field contains
    # a payload intended for remote code execution.
    malicious_config_payload = json.dumps({
        "model_type": "sew-d",
        "hidden_size": 768,
        "some_dynamic_parameter": "__import__('os').system('echo VULNERABILITY EXPLOITED')"
    })

    # A config with a safe, legitimate literal value for comparison.
    safe_config_payload = json.dumps({
        "model_type": "sew-d",
        "hidden_size": 768,
        "some_dynamic_parameter": "{'key': 'value', 'num': 123}"
    })
    
    # --- DEMONSTRATION OF THE FIX ---

    print("1. Processing a SAFE configuration payload...")
    # This will succeed because the string is a valid literal.
    try:
        converted_safe_config = convert_sew_d_config_fixed(safe_config_payload)
        if converted_safe_config:
            print("Safe config processed successfully.")
            print("Result:", converted_safe_config)
    except ValueError as e:
        print(f"Caught expected error: {e}")

    print("\n" + "="*50 + "\n")

    print("2. Processing the MALICIOUS configuration payload...")
    # This will fail safely because ast.literal_eval refuses to execute the code.
    # The `os.system` call will NOT be executed.
    try:
        converted_malicious_config = convert_sew_d_config_fixed(malicious_config_payload)
        if converted_malicious_config is None:
             print("Malicious config processing failed as expected.")
    except ValueError as e:
        print(f"Caught expected error, confirming the fix works: {e}")
    
    print("\nConclusion: The fix correctly identified and blocked the malicious payload.")

Payload

__import__('os').system('id')

Cite this entry

@misc{vaitp:cve202514927,
  title        = {{Hugging Face Transformers convert_config allows RCE via code injection.}},
  author       = {Bogaerts, Fr\'ed\'eric and Ivaki, Naghmeh and Fonseca, Jos\'e},
  year         = {2025},
  note         = {VAITP Python Vulnerability Dataset, entry CVE-2025-14927},
  howpublished = {\url{https://netpack.pt/vaitp/vulnerability/CVE-2025-14927/}}
}
Introducing the "VAITP dataset": a specialized repository of Python vulnerabilities and patches, meticulously compiled for the use of the security research community. As Python's prominence grows, understanding and addressing potential security vulnerabilities become crucial. Crafted by and for the cybersecurity community, this dataset offers a valuable resource for researchers, analysts, and developers to analyze and mitigate the security risks associated with Python. Through the comprehensive exploration of vulnerabilities and corresponding patches, the VAITP dataset fosters a safer and more resilient Python ecosystem, encouraging collaborative advancements in programming security.

The supreme art of war is to subdue the enemy without fighting.

Sun Tzu – “The Art of War”

:: Shaping the future through research and ingenuity ::