Files have always been a security boundary. Applications have spent decades defending against malicious documents, path traversal, parser flaws, archive bombs, embedded scripts, unsafe deserialization, and executable content disguised as ordinary attachments. AI agents do not replace those risks. They introduce a new one by allowing the contents of a file to influence what the application decides to do next. A conventional document-processing system opens a file, extracts data, and passes the result through predetermined application logic. An AI agent can read the same file, infer instructions from its contents, choose a tool, generate arguments for that tool, and continue acting based on what it found. A malicious document no longer needs to exploit the document parser itself if it can persuade the agent to misuse capabilities the application has already granted.
When a File Becomes an Instruction Channel
Indirect prompt injection is one of the clearest examples of this problem. An attacker can place instructions inside content that an AI system is expected to read, including PDFs, Word documents, spreadsheets, email attachments, source files, and plain text documents. The user may ask an agent to summarize or inspect the file, but the model receives the attacker’s instructions as part of the same context it is expected to interpret. Those instructions can appear in visible text, hidden formatting, encoded content, metadata, images, or material extracted during document processing.
That creates a security problem that traditional file scanning is poorly equipped to solve. Antivirus software may correctly determine that a PDF contains no malicious executable code. A sandbox may confirm that the document never launches a process. A content-disarm system may remove macros and embedded objects. The file can still be dangerous to an AI agent if its contents direct the agent to retrieve confidential information, call another tool, alter a configuration, move a file, or transmit data elsewhere. The malicious behavior happens after the file is parsed, inside the agent workflow rather than inside the document itself.
Tool Access Changes the Impact
Prompt injection against a chatbot may alter an answer, but prompt injection against an agent can alter an action. Modern agent frameworks can expose filesystem operations, browsers, shell commands, code execution, databases, email, source-control systems, cloud services, ticketing platforms, and internal APIs to a model. Once those capabilities are available, the security impact of malicious file content depends heavily on what the agent is allowed to do after reading it.
Microsoft demonstrated this risk in research published in May 2026 involving its Semantic Kernel framework. Researchers identified vulnerabilities in model-controlled file operations that could be chained into arbitrary file access and host-level compromise. One issue, CVE-2026-25592, involved file-transfer functionality tied to a Python execution plugin. A download function exposed to the model allowed model-generated input to influence the destination path of a file written to the host. Without adequate path restrictions, injected instructions could direct the agent to place content outside the intended workspace. Related upload functionality also allowed arbitrary local paths to be selected, creating a route for sensitive files to be transferred into the agent’s execution environment.
The significance of those findings goes beyond one framework. The model did not bypass a filesystem permission system through a traditional exploit. The application had exposed legitimate filesystem capabilities to the model and allowed model-generated parameters to control them. That pattern is becoming increasingly important in agent security, since ordinary application features can become security-sensitive once a model is allowed to select how they are used.
The Filename Is Part of the Attack Surface
The content of a file is not the only untrusted input in these workflows. Filenames, directory names, archive members, metadata fields, MIME types, spreadsheet cells, image text, repository paths, and document properties can all enter model context. An application that asks an agent to inspect an uploaded directory may send filenames directly to the model. If the agent can rename, move, delete, upload, or execute files, a malicious filename can become part of the instruction chain.
Archives create similar problems by combining nested directories, conflicting extensions, symbolic links, oversized decompressed content, and attacker-controlled names with an agent that may interpret README files or other documentation as guidance. The safe assumption is that information derived from an untrusted file remains untrusted after extraction, regardless of whether it appears in visible text, metadata, filenames, or another representation.
Parsing Can Change What the Agent Sees
AI systems frequently rely on several processing stages before a model sees anything. PDFs may go through text extraction or optical character recognition. Office documents may be converted into HTML or plain text. Images may pass through vision models. Spreadsheets may be converted into structured tables. Archives may be unpacked, and software repositories may be divided into chunks before indexing or retrieval.
Those transformations can expose content that a human user never sees. Hidden text may appear in extracted output. Material positioned outside the visible page may still be passed to the model. Metadata ignored by one parser may be preserved by another. Encoded content may be decoded automatically during preprocessing. A security review has to account for both the original file and the representation ultimately placed into model context.
Temporary Files Create Permanent Problems
Agents that work with documents often create temporary copies in object storage, containers, working directories, caches, conversion services, and evaluation systems. A single upload may produce several derivative files before the model ever processes it. Temporary directories with predictable paths can create collisions between sessions, reused storage can expose remnants from earlier tasks, and overly broad permissions can allow one process to inspect another user’s files. Cleanup failures can also leave sensitive artifacts available long after the original conversation ends.
The problem becomes more serious when an agent is allowed to control filesystem paths directly. Path traversal protections, canonicalization, fixed working directories, randomized internal filenames, per-session isolation, storage quotas, and explicit cleanup rules become part of the agent security model. Microsoft’s remediation for the Semantic Kernel issue followed this direction by removing sensitive file-transfer functions from the capabilities directly exposed to the model and adding stronger path restrictions for programmatic use.
File Reads Can Be Just as Dangerous as File Writes
An agent with unrestricted filesystem reads can search user directories, inspect configuration files, collect source code, retrieve API keys, read environment files, and combine information from several locations before sending the result through another tool. A malicious document can exploit that authority through indirect prompt injection by instructing the agent to locate files that the attacker could never access directly. From the operating system’s perspective, the resulting read may appear legitimate if the agent runs with the user’s permissions.
This exposes a weakness in many existing access-control models. A user may legitimately have permission to read payroll records, source code, cloud configuration files, and internal tickets. A PDF downloaded from the internet has none of those permissions. If both are placed into the same model context, the application needs a way to preserve that distinction. Allowing untrusted file content to inherit the authority of the user creates a path from data ingestion to privileged action.
Untrusted Content Should Not Inherit User Authority
Agent security research is increasingly focused on information-flow controls for this reason. One approach is to assign trust or sensitivity labels to content and carry those labels through later tool interactions. External documents can remain marked as untrusted, system policies can remain trusted, and sensitive retrieved data can retain its classification. A runtime policy can then prevent an agent from combining untrusted instructions with confidential data and an external destination, regardless of whether the model believes the action is appropriate.
This moves enforcement away from prompt wording and back into application logic. Rather than trusting the model to distinguish a legitimate instruction from hostile content every time, the application can apply hard restrictions based on where the data came from, what privileges are involved, and which destination the agent is trying to reach.
Sandboxing Only Works if the Boundary Holds
Running an agent inside a container does little if the container has access to the user’s home directory, host credentials, cloud tokens, Docker sockets, unrestricted network connections, or file-transfer functions that can write outside the workspace. The model may never need to escape the sandbox if the application already provides bridges around it.
Security reviews need to examine every interface between the isolated environment and the host, including filesystem mounts, upload and download APIs, environment variables, credential brokers, tool servers, network access, and output directories. The question is not simply whether the agent is sandboxed, but what the sandbox can still reach.
Treat Model-Selected File Operations as Untrusted Input
A useful security principle is to treat any filesystem parameter selected by a model as attacker-controlled input. Paths, filenames, URLs, archive members, storage destinations, file types, command arguments, tool parameters, and references to other files should pass through the same validation that an application would apply to direct user input.
File access should rely on explicit working directories and restricted path sets. Paths should be resolved before authorization checks so alternate representations cannot escape permitted directories. Archive extraction should enforce limits on decompressed size, file count, directory depth, and destination paths. Agent identities should receive access to the smallest practical set of files, with read and write permissions separated wherever possible.
Tool design can reduce the attack surface further. A generic function that allows the model to read any path on the filesystem creates far more risk than a function that retrieves a document by an authorized identifier. A general file-write function creates more exposure than a tool restricted to a temporary workspace. Shell access amplifies the consequences of successful prompt injection compared with narrowly defined application functions with fixed schemas and explicit authorization checks.
The Security Decision Has to Happen Before Execution
Detecting malicious prompts is useful, but detection alone cannot be the control that protects files and downstream systems. Applications need a separate policy layer capable of evaluating what the model is attempting to do before the requested tool receives control.
A document being summarized should not gain authority to initiate arbitrary network requests. Reading an attachment should not grant permission to search the user’s entire filesystem. Extracting an archive should not authorize its contents to invoke a shell. Creating a report should not require unrestricted write access across the host. These restrictions are strongest when they are enforced at the filesystem, operating system, tool, or runtime layer rather than expressed only as instructions telling the model what it should avoid.
File Handling Is Becoming a Core Agent Security Problem
File handling is becoming a central part of AI agent security architecture as agents gain access to local documents, repositories, cloud storage, business applications, code execution, and endpoint resources. Security teams now need to examine more than whether an uploaded file contains malware. They need to determine what the agent will extract from that file, what authority the extracted content receives, which tools become reachable after ingestion, where temporary copies are created, what filesystem paths the model can influence, and whether a malicious document can move the agent from reading information into taking action.
The next generation of malicious files may never exploit a parser, execute a macro, or contain executable code. Some may contain nothing more than instructions written for the AI system opening them. Once that system can read sensitive files, call privileged tools, use stored credentials, and interact with external services, those instructions can be enough to turn an ordinary document into an attack path.
How Can Netizen Help?
Founded in 2013, Netizen is an award-winning technology firm that develops and leverages cutting-edge solutions to create a more secure, integrated, and automated digital environment for government, defense, and commercial clients worldwide. Our innovative solutions transform complex cybersecurity and technology challenges into strategic advantages by delivering mission-critical capabilities that safeguard and optimize clients’ digital infrastructure. One example of this is our popular “CISO-as-a-Service” offering that enables organizations of any size to access executive level cybersecurity expertise at a fraction of the cost of hiring internally.
Netizen also operates a state-of-the-art 24x7x365 Security Operations Center (SOC) that delivers comprehensive cybersecurity monitoring solutions for defense, government, and commercial clients. Our service portfolio includes cybersecurity assessments and advisory, hosted SIEM and EDR/XDR solutions, software assurance, penetration testing, cybersecurity engineering, and compliance audit support. We specialize in serving organizations that operate within some of the world’s most highly sensitive and tightly regulated environments where unwavering security, strict compliance, technical excellence, and operational maturity are non-negotiable requirements. Our proven track record in these domains positions us as the premier trusted partner for organizations where technology reliability and security cannot be compromised.
Netizen holds ISO 27001, ISO 9001, ISO 20000-1, and CMMI Level III SVC registrations demonstrating the maturity of our operations. We are a proud Service-Disabled Veteran-Owned Small Business (SDVOSB) certified by U.S. Small Business Administration (SBA) that has been named multiple times to the Inc. 5000 and Vet 100 lists of the most successful and fastest-growing private companies in the nation. Netizen has also been named a national “Best Workplace” by Inc. Magazine, a multiple awardee of the U.S. Department of Labor HIRE Vets Platinum Medallion for veteran hiring and retention, the Lehigh Valley Business of the Year and Veteran-Owned Business of the Year, and the recipient of dozens of other awards and accolades for innovation, community support, working environment, and growth.
Looking for expert guidance to secure, automate, and streamline your IT infrastructure and operations? Start the conversation today.



