File tools
Read, extract, generate, and store files from agents with the filesystem, file-tools, and file-generation integrations
Overview
Three built-in integrations cover file work, and they do different jobs:
- file-tools — read data out of files. Extract rows from CSV, Excel, and JSON, text and tables from PDFs, and content from Word documents, then export results back to Excel.
- file-generation — produce documents. Render PDF, DOCX, XLSX, PPTX, Markdown, JSON, and TXT from structured content blocks or a Jinja2 template.
- filesystem — read and write files on the server itself, restricted to paths an administrator allows.
Enable only the ones a project needs. Each is an MCP server the platform ships and manages, so agents discover the tools and call them during reasoning like any other tool.
File extraction is separate from what happens automatically at a START node. When a user uploads a tabular file as a workflow input, the agent already gets a SQL query tool over that data with no setup. See Database tools. Reach for file-tools when an agent needs to pull data out of a file mid-workflow.
file-tools
Extract structured data from files and export results to Excel. It exposes ten tools.
Getting a file in
upload_file— upload a file from a local path and get back afile_idUUID. Pass thatfile_idto the extraction tools.preview_file— peek at a file's structure without extracting everything. Returns column names and a few sample rows, so an agent can decide what to do before pulling the whole file.
Extracting rows
extract_rows— auto-detects the file type (Excel, CSV, or JSON) and returns rows as an array of objects. Use this when you don't want to care about format.extract_csv_rows— extract rows from a CSV file.extract_excel_rows— extract rows from an Excel file (.xlsx,.xls).list_sheets— list the sheets available in an Excel workbook.
PDFs and Word
extract_pdf_text— extract text page by page. Returns an array of pages, each withpage_numandtext.extract_pdf_tables— scan every page for tables and return them as rows.extract_docx— extract content from a Word.docxdocument. Defaults tomode='tables'for table extraction.
Sending results back out
export_to_excel— merge ForEach results, one row per input plus its output, and export as an Excel file uploaded to storage.
extract_pdf_tables and export_to_excel are built for the ForEach node: extract rows, fan out over them, then merge the results back into a single spreadsheet.
file-generation
Generate documents from structured content blocks. It exposes five tools.
render_document— render a document and return its storage metadata:key,content_type,filename, andsize. Supports PDF, DOCX, XLSX, MD, JSON, and TXT.render_from_template— render a Jinja2 template with supplied data and produce a PDF, HTML, or PPTX file.describe_document— read back a presentation you already generated, slide by slide, in the same shape an edit would take.revise_document— change a presentation you already generated without re-sending it. Pass thekeyreturned byrender_document.upload_branding_image— upload a header or footer image. Returns a key you can pass asheader_imageorfooter_imageon later renders.
describe_document and revise_document are what make iteration cheap. An agent can generate a deck, read back what landed on each slide, and revise just the slides that need changing, instead of regenerating the whole document.
Templates and branding
An administrator can pre-load DOCX, PPTX, and XLSX templates plus header and footer branding images when connecting the integration. Documents then render against your own house style rather than a default.
filesystem
Read and write files on the server, with path-based access control. Every file operation is restricted to directories an administrator has allowed.
Configuration
Add an allowed_paths credential
In Settings → Credentials, add an allowed_paths credential.
List the permitted directories
Set comma-separated absolute paths the workflow may access, for example /data/uploads,/data/exports.
Operations stay inside those paths
Every read and write is restricted to those directories. Anything outside them is refused.
allowed_paths is the only thing standing between an agent and the rest of the server's filesystem. Scope it to the narrowest set of directories the workflow actually needs, and never include a path holding secrets, credentials, or other tenants' data.
Use cases
Invoice data into a spreadsheet
A ForEach node walks a folder of invoices. For each one, extract_pdf_tables pulls the line items, an agent normalises them, and export_to_excel merges every result into a single reconciliation sheet.
Branded report generation
An agent gathers findings, then render_document produces a PDF with your header and footer images applied. describe_document and revise_document let a reviewer request changes without regenerating the file.
Contract review
extract_docx pulls the clause tables out of a Word contract, an agent checks them against policy, and a Human Task node routes anything uncertain to a person.