Input / Output
Soufflé reads input facts from tab-separated files that populate the extensional database (EDB). Each relation declared in a Datalog program can be loaded from a corresponding .facts file, where each line represents one tuple with columns separated by tabs. Output relations are written to tab-separated files after evaluation completes.
EDB facts file convention: For a relation named edge, place your input tuples in edge.facts in the same directory as your Datalog source. Each line is one tuple: value1 value2 …
Soufflé offers multiple execution modes to suit different stages of program development and deployment. The interpreter provides a rapid prototyping environment where Datalog programs can be tested and refined without the overhead of compilation. This mode is ideal for exploring the design space of a static analysis, enabling quick iteration on logical rules and relations. The interpreter evaluates queries directly, giving immediate feedback on program behavior and output. For production use, the compiler translates Datalog specifications into optimized parallel C++ programs, leveraging the full performance of modern hardware. This dual approach allows developers to move seamlessly from experimentation to high-performance execution.
Input and output in Soufflé follow a straightforward convention based on tab-separated files. The extensional database, or EDB, is populated by reading facts from files that correspond to declared relations. Each line in a facts file represents a single tuple, with columns separated by tabs, making it easy to prepare data using standard spreadsheet or scripting tools. After evaluation completes, output relations are written to similarly structured files, allowing results to be consumed by downstream analysis pipelines or visualization tools. This simple file-based interface keeps the toolchain lightweight and interoperable, avoiding the need for complex database setup or proprietary formats while maintaining clarity and reproducibility.
The feedback-directed compilation infrastructure in Soufflé represents a sophisticated approach to optimizing Datalog execution. By analyzing runtime behavior and profiling the performance of generated code, the system can identify bottlenecks and apply targeted optimizations. This iterative process refines the compiled C++ program over multiple passes, adapting data structures and evaluation strategies to the specific characteristics of the input program and its data. The result is a specialized, staged compilation pipeline that produces efficient executables tailored to the analysis task at hand. Such techniques draw on the principles of partial evaluation and Futamura projections, translating high-level logical specifications into low-level code that runs with minimal overhead.
Soufflé is designed with scalability as a core requirement, particularly for large-scale static analysis tasks such as Java taint tracking and security checks. The tool’s parallel execution model distributes work across multiple cores and processors, enabling it to handle the substantial relation sizes common in real-world program analysis. Logical relations are stored and processed using specialized data structures that support efficient set operations and incremental evaluation. This foundation allows Soufflé to serve as a practical platform for deep design space exploration in program analysis, where researchers and engineers can rapidly prototype and test new analysis algorithms expressed in Datalog, confident that the underlying synthesis engine will produce performant, parallelized C++ code.
Interpreter
The interpreter evaluates a Datalog program directly without generating C++ code. It is useful for rapid prototyping and debugging, as it provides immediate feedback with lower startup overhead than the compilation path.
souffle program.dl
By default, the interpreter runs in single-threaded mode. Use the -j flag to enable parallel evaluation:
souffle -j4 program.dl
Compiler
The compiler translates a Datalog specification into a standalone parallel C++ executable. This two-stage process first synthesizes C++ source code, then invokes a C++ compiler to produce a native binary optimized for the target machine.
souffle -c program.dl
The resulting executable accepts the same input-fact files and produces the same output as the interpreter, but typically runs significantly faster on large datasets due to compiled optimizations.
Feedback-Directed Compilation
Soufflé supports feedback-directed compilation, where profiling data from a previous run informs the synthesis of a more efficient executable. This iterative approach allows the tool to tune data structures and execution strategies based on observed workload characteristics.
souffle --profile=program.prof program.dl
souffle -c --auto-schedule=program.prof program.dl
Profiling
The built-in profiler records per-relation timing and tuple-count statistics during execution. Profiling output helps identify performance bottlenecks in a Datalog program and guides manual or automatic tuning.
| Option | Description |
|---|---|
--profile=<file> | Write profiling data to the specified file |
--profile-frequency | Enable fine-grained frequency counters in the profile |
Warning Options
Soufflé emits warnings for common pitfalls such as unbound variables, singleton relations, or missing fact files. Warnings can be controlled via command-line flags:
| Option | Effect |
|---|---|
-w | Enable all warnings |
-Werror | Treat warnings as errors |
--no-warn | Suppress all warnings |