← Latest papers
💻 computer science

Prevalence of Cross-File Dependencies in Terraform and Their Resolution by Security Scanners

This paper analyzes a large-scale Terraform dataset to reveal that cross-file dependencies are prevalent and often involve security-sensitive resources, while demonstrating through controlled experiments that modern security scanners possess varying but significant capabilities to resolve these complex inter-file relationships, challenging the assumption that they cannot handle multi-file reasoning.

Original authors: Moatasem M. Draz

Published 2026-09-21
📖 6 min read🧠 Deep dive

Original authors: Moatasem M. Draz

Original paper licensed under CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/). ✨ This is an AI-generated explanation of the paper below. It is not written or endorsed by the authors. For technical accuracy, refer to the original paper. Read full disclaimer

In the modern digital world, the massive servers and networks that power our apps, banks, and streaming services are no longer built by hand. Instead, engineers write instructions in a computer language to tell the cloud exactly what to build. This practice is called Infrastructure as Code. It allows teams to define complex systems in text files, ensuring that every server, firewall, and storage bucket is created with the same precision every time. One of the most popular tools for this job is called Terraform. It works like a blueprint reader, taking these text instructions and turning them into real, functioning technology. However, just as a house plan might be split across dozens of pages with notes referencing other pages, these code files are rarely isolated. They are woven together, with one file calling upon another, passing values back and forth, and building a system that exists only when all the pieces are assembled.

The security of these systems depends on the accuracy of these instructions. A single mistake in the code, such as accidentally leaving a digital door unlocked, can expose sensitive data to the entire internet. For years, researchers and security experts have worried that the tools designed to scan these blueprints for mistakes might be blind to the connections between files. The prevailing belief was that these scanners could only read one file at a time, missing the dangerous secrets hidden in the spaces between them. If a file said "use the default setting" and another file defined that default as "open to everyone," the scanner might see the first file as safe and the second as safe in isolation, failing to spot the dangerous combination. This paper set out to test that assumption against a massive collection of real-world code and to see if the tools were truly as limited as everyone thought.

The researchers began by gathering a vast library of over 62,000 public projects that use Terraform. They wanted to understand how often these projects actually rely on connections between different files. By mapping out the relationships in these projects, they discovered that cross-file connections are not rare exceptions but a standard part of how these systems are built. In fact, more than one-third of all the projects they studied contained at least one connection where one file depended on another. These connections were not evenly spread out; they were heavily concentrated in a smaller number of complex projects, while many simpler projects had very few or none. The most common way these files were linked was by pointing to shared folders, where a project would reach up or across its directory structure to find a common piece of code it needed. This pattern showed that the way engineers organize their work naturally creates a web of dependencies that spans many files.

The study then asked a critical question: do these connections often lead to the most sensitive parts of a system? The researchers looked for instances where a file depended on another file that controlled security settings, such as who could access a database or which computers were allowed to talk to each other. They found that in over 9,000 of the projects, these cross-file connections did indeed point to these critical security controls. This meant that the safety of these systems often relied on the ability to trace a value from one file, through a connection, to a security rule in another file. If a tool could not follow that path, it would be looking at the wrong picture, potentially missing a vulnerability that existed only because of how the files were stitched together.

With the scale of the problem established, the researchers turned to the tools themselves to see if they could actually follow these paths. They designed a controlled experiment using four of the most popular security scanners available. They created a series of test cases that mimicked the real-world connections they had found, ranging from simple links between files in the same folder to complex chains where a value passed through multiple modules before reaching its final destination. They tested whether the scanners could spot a security flaw when the dangerous setting was hidden in a different file, or if the scanners would only see the danger when it was written directly in the same file.

The results challenged the long-held belief that these tools were fundamentally blind to cross-file connections. The study found that the capabilities of the tools were not a simple yes or no, but a spectrum. Every single tool tested was able to follow the simplest connections, such as a variable defined in one file and used in another within the same folder. They all successfully spotted the danger in these basic scenarios. However, as the connections became more complex, the tools began to diverge. Only one of the four scanners, Checkov, was able to follow the most intricate paths, including those where values were passed through multiple layers of modules or where settings were overridden by special configuration files. The other tools could handle the direct links but often lost the trail when the connection involved intermediate steps or specific file types like variable override files.

This finding suggests that the common assumption that scanners cannot follow cross-file references is too broad to be accurate. It is not that the tools are incapable of the task, but rather that their ability to do so varies significantly depending on the specific tool and the complexity of the connection. For teams relying on the tools that only handle the simplest cases, there is a genuine risk that they are missing security issues hidden in the more complex wiring of their code. The researchers concluded that while the tools have made progress, the industry needs to be more precise about what each tool can and cannot see. Rather than assuming a tool will catch every error, teams should understand the specific limits of their chosen scanner, especially when their code relies on deep layers of connections. The study provides a clear map of these limits, showing that while the simplest cross-file dangers are caught, the deeper, more complex ones require either a more capable tool or a different approach to ensure safety.

Drowning in papers in your field?

Get daily digests of the most novel papers matching your research keywords — with technical summaries, in your language.

Try Digest →