VAITP Dataset

← Back to the dataset

CVE-2025-46725

Langroid uses pandas eval(), allowing remote code execution via crafted queries.

  • CVSS 8.1
  • CWE-94
  • Input Validation and Sanitization
  • Remote

Langroid is a Python framework to build large language model (LLM)-powered applications. Prior to version 0.53.15, `LanceDocChatAgent` uses pandas eval() through `compute_from_docs()`. As a result, an attacker may be able to make the agent run malicious commands through `QueryPlan.dataframe_calc]`) compromising the host system. Langroid 0.53.15 sanitizes input to the affected function by default to tackle the most common attack vectors, and added several warnings about the risky behavior in the project documentation.

CVSS base score
8.1
Published
2025-05-20
OWASP
A03 Injection
Orthogonal defect classification
Interface
Code defect classification
Incorrect Functionality
Category
Input Validation and Sanitization
Subcategory
Insecure Parsing or Deserialization
Accessibility scope
Remote
Impact
Arbitrary Code Execution
Affected component
Langroid
Fixed by upgrading
Yes

Solution

Upgrade to Langroid version 0.53.15 or later.

Vulnerable code sample

import pandas as pd

class QueryPlan:
    def __init__(self, dataframe_calc):
        """Vulnerable function that demonstrates the security issue."""
        self.dataframe_calc = dataframe_calc

        class LanceDocChatAgent:
            def compute_from_docs(self, query_plan):
                """
                Computes results based on a query plan using pandas eval.
                POTENTIALLY UNSAFE: Uses pandas eval which can execute arbitrary code.
                """
                try:
            # Simulate loading data into a pandas DataFrame
            # In a real application, this would load data from documents.
                    data = {'col1': [1, 2, 3], 'col2': [4, 5, 6]}
                    df = pd.DataFrame(data)

            # Execute the calculation specified in the query plan using pandas eval
                    result = df.eval(query_plan.dataframe_calc)
                    return result
                    except Exception as e:
                        return f"Error during calculation: {e}"

# Example usage (VULNERABLE)
                        if __name__ == '__main__':
                            agent = LanceDocChatAgent()

    # Malicious query plan injecting code execution
                            malicious_calc = "__import__('os').system('echo EXPLOITED > /tmp/pwned.txt')"
                            malicious_plan = QueryPlan(dataframe_calc=malicious_calc)
                            result = agent.compute_from_docs(malicious_plan)
                            print(f"Result: {result}")  # Print the result
    # After execution, a file named 'pwned.txt' containing 'EXPLOITED' will be created in /tmp if the execution was successful.

Patched code sample

import pandas as pd
import re

def safe_eval(expression, local_dict=None, global_dict=None):
    """
    Safely evaluates a mathematical expression.  Restricts access to
    potentially dangerous functions and attributes.  Allows basic
    arithmetic, comparison operators, and explicitly defined variables.

    Args:
        expression (str): The expression to evaluate.
        local_dict (dict, optional): Local variables to use in the evaluation. Defaults to None.
        global_dict (dict, optional): Global variables to use in the evaluation. Defaults to None.

    Returns:
        The result of the evaluation, or None if the expression is unsafe.
    """

    # Whitelist of allowed characters and functions
    allowed_chars = r"[0-9+\-*/().%<>=!&|\s,:a-zA-Z_]"
    allowed_functions = ['abs', 'round', 'min', 'max', 'sum', 'len'] # add more as needed

    # Simple check for potentially malicious code injections
    if not re.match(f"^[{allowed_chars}]*$", expression):
        print("Unsafe characters detected in expression.")
        return None

    # Check for use of potentially dangerous attributes and functions.
    # This is a more comprehensive check than just character whitelisting.
    blacklist = [
        "_",  # Avoid access to hidden attributes and methods
        "import",
        "exec",
        "eval",
        "compile",
        "getattr",
        "setattr",
        "deleteattr",
        "globals",
        "locals",
        "vars",
        "open",
        "read",
        "write",
        "system",
        "os",
        "subprocess",
        "__", # double underscores
    ]

    for item in blacklist:
        if item in expression.lower():  # Lowercase for case-insensitive matching
            print(f"Blacklisted keyword '{item}' detected in expression.")
            return None
    
    # Check if potentially malicious functions are called
    pattern = r'\b(' + '|'.join(allowed_functions) + r')\b'
    function_calls = re.findall(pattern, expression) # list of all function calls
    
    # Check for usage of pandas methods in calculation
    pandas_method_blacklist = [
        "apply",
        "map",
        "pipe",
        "aggregate",
        "transform",
        "groupby",
        "rolling",
        "expanding",
        "shift",
        "diff",
        "pct_change"
    ]

    for item in pandas_method_blacklist:
        if item in expression.lower():
            print(f"Potentially dangerous pandas method '{item}' detected.")
            return None

    try:
        # Create a safe namespace by limiting builtins
        safe_globals = {k: __builtins__.__dict__[k] for k in allowed_functions if k in __builtins__.__dict__}

        # Add specified global variables, if provided
        if global_dict:
            safe_globals.update(global_dict)

        # Evaluate the expression with the safe namespace
        result = eval(expression, safe_globals, local_dict)
        return result
    except (SyntaxError, NameError, TypeError, ZeroDivisionError) as e:
        print(f"Error evaluating expression: {e}")
        return None
    except Exception as e:
        print(f"An unexpected error occurred: {e}")
        return None

def compute_from_docs(df, calc_string):
    """
    Computes a new column in the DataFrame based on a safe calculation string.

    Args:
        df (pd.DataFrame): The DataFrame to operate on.
        calc_string (str): The calculation string to apply.

    Returns:
        pd.DataFrame: The DataFrame with the new column, or the original DataFrame if the calculation fails.
    """
    
    #Sanitize input calc_string using safe_eval function
    result = safe_eval(calc_string)
    if result is None:
        print("The calc string is not safe. Returning original DataFrame")
        return df

    try:
        df['new_column'] = result
        return df
    except Exception as e:
        print(f"Error during calculation: {e}")
        return df


if __name__ == '__main__':
    # Example usage (safe):
    data = {'col1': [1, 2, 3], 'col2': [4, 5, 6]}
    df = pd.DataFrame(data)

    # Simulate user-provided calculation string
    safe_calculation = "df['col1'] + df['col2']"
    df_safe = compute_from_docs(df.copy(), safe_calculation)
    print("Safe calculation result:")
    print(df_safe)

    # Example of a blocked malicious calculation:
    malicious_calculation = "__import__('os').system('ls -l')"
    df_unsafe = compute_from_docs(df.copy(), malicious_calculation)
    print("\nUnsafe calculation attempt:")
    print(df_unsafe)

    malicious_calculation_2 = "df['col1'].apply(lambda x: x**2)"
    df_unsafe_2 = compute_from_docs(df.copy(), malicious_calculation_2)
    print("\nUnsafe calculation attempt 2:")
    print(df_unsafe_2)

    # A calculation that requires usage of pandas functions
    # df['col1'].mean() will cause an issue when applying the fix
    # due to the restriction of pandas methods.
    pandas_calculation = "df['col1'].mean()"
    df_pandas_calculation = compute_from_docs(df.copy(), pandas_calculation)
    print("\nPandas function calculation attempt:")
    print(df_pandas_calculation)

    complex_calculation = "(df['col1'] + df['col2']) * abs(min(df['col1']))"
    df_complex = compute_from_docs(df.copy(), complex_calculation)
    print("\nComplex calculation attempt:")
    print(df_complex)

Payload

__import__('os').system('touch /tmp/pwned')

Cite this entry

@misc{vaitp:cve202546725,
  title        = {{Langroid uses pandas eval(), allowing remote code execution via crafted queries.
}},
  author       = {Bogaerts, Fr\'ed\'eric and Ivaki, Naghmeh and Fonseca, Jos\'e},
  year         = {2025},
  note         = {VAITP Python Vulnerability Dataset, entry CVE-2025-46725},
  howpublished = {\url{https://netpack.pt/vaitp/vulnerability/CVE-2025-46725/}}
}
Introducing the "VAITP dataset": a specialized repository of Python vulnerabilities and patches, meticulously compiled for the use of the security research community. As Python's prominence grows, understanding and addressing potential security vulnerabilities become crucial. Crafted by and for the cybersecurity community, this dataset offers a valuable resource for researchers, analysts, and developers to analyze and mitigate the security risks associated with Python. Through the comprehensive exploration of vulnerabilities and corresponding patches, the VAITP dataset fosters a safer and more resilient Python ecosystem, encouraging collaborative advancements in programming security.

The supreme art of war is to subdue the enemy without fighting.

Sun Tzu – “The Art of War”

:: Shaping the future through research and ingenuity ::