Firecrawl llms.txt Generator: Setup, Limits, Specs
A technical guide to generating llms.txt and llms-full.txt using Firecrawl, detailing API endpoints, credit enforcement, and markdown specifications.
The Firecrawl llms.txt generator automates the creation of standard llms.txt and expanded llms-full.txt files by combining Firecrawl web scraping with GPT-4-mini processing. The standalone generator endpoint hosted at llmstxt.firecrawl.dev was deprecated after June 30, 2025, shifting production generation to Firecrawl core crawl and scrape endpoints.12

Standard vs. Full: Output Formatting and the llms.txt Specification
According to the llms.txt markdown specification, an H1 header containing the name of the project or site is the only required section in the file. Everything else follows a strict sequential order: an optional byte-order mark (BOM), the mandatory H1 title, a blockquote summary containing essential background context, optional non-heading explanatory sections, and zero or more H2-delimited file lists.4
File lists inside llms.txt follow a precise link format: a markdown hyperlink name, optionally followed by a colon and notes describing the target file. By convention, an H2 section named 'Optional' designates secondary documentation that autonomous agents can skip when operating under tight context limits. Firecrawl produces both this index file (llms.txt) and a concatenated full-text document (llms-full.txt) containing the actual extracted contents of those linked resources.14
How the Generator Works: Firecrawl Engine and GPT-4-mini Processing
The generation workflow relies on a dual-stage architecture. Firecrawl handles crawling and raw page extraction across target domains, while GPT-4-mini (or gpt-4o-mini) structures and summarizes extracted content to conform to markdown standards. Developers deploying the open-source generator locally or via custom infrastructure configure an environment with four specific variables: FIRECRAWL_API_KEY, SUPABASE_URL, SUPABASE_KEY, and OPENAI_API_KEY.1
An example repository demonstrating how to generate LLMs.txt files in Python is publicly accessible at mendableai/create-llmstxt-py on GitHub.1
API Endpoints, Parameter Overrides, and the June 30, 2025 Deprecation
The public standalone generator exposes simple GET endpoints. A request to https://llmstxt.firecrawl.dev/[YOUR_URL_HERE] returns the standard index, while requesting https://llmstxt.firecrawl.dev/{YOUR_URL}/full returns the llms-full.txt version. Passing an API key via query parameter removes usage limits and unlocks complete results. You append the key directly as ?FIRECRAWL_API_KEY=YOUR_API_KEY on the standard path or after /full.15
In the Alpha generator API, query behavior is guided by specific defaults. The 'maxUrls' parameter accepts integer values from 1 to 100, defaulting to 10 pages crawled. Full-text generation of llms-full.txt is disabled by default in the Alpha API, and the processing ceiling across the Alpha tool is capped at 5,000 URLs. The standalone endpoint at llmstxt.firecrawl.dev has not been maintained since June 30, 2025, so production applications belong on the core Firecrawl crawling APIs.16
Credit Mechanics, Crawl Limits, and 402 Error Handling
Billing operates on consumption units: each page crawled costs 1 credit. However, using JSON format on crawl and scrape endpoints imposes an additional charge of 4 credits per page. The default crawl ceiling is 10,000 pages.3
Before executing a crawl, Firecrawl validates your account balance against the crawl limit. If your remaining credit balance cannot cover the ceiling (such as 10,000 credits for the default limit), the API blocks execution immediately and returns an HTTP 402 Payment Required status code. To avoid premature 402 errors on smaller accounts, explicitly set your crawl limit to match actual URL volume rather than leaving the default setting active.3
Completed crawl results remain available via the API for 24 hours. When responses exceed 10MB, the API fragments output and supplies a 'next' URL parameter, which you must call sequentially to fetch each subsequent 10MB chunk.3
Firecrawl Tier Allowances, Concurrency, and Overage Pricing
Running large documentation crawls requires selecting an account tier that supports the required rate limits and credit volume. Free accounts provide 1,000 credits per month with a limit of 2 concurrent requests and a /crawl rate limit of 2 requests per minute.7
| Plan | Annual Monthly Price | Monthly Credits | Concurrent Requests | /crawl Rate Limit | Overage Credit Price |
|---|---|---|---|---|---|
| Free | $0 | 1,000 | 2 | 2/min | N/A |
| Hobby | $16 | 5,000 | 5 | 20/min | $5 per 1,000 |
| Standard | $83 | 100,000 | 25 | 100/min | $5 per 2,000 |
| Growth | $333 | 500,000 | 50 | 1,000/min | $5 per 2,500 |
| Scale | $599 | 1,000,000 | 100 | 2,000/min | $5 per 5,000 |

Teams planning automated documentation updates should configure explicit crawl page limits on core Firecrawl endpoints to prevent unexpected HTTP 402 aborts. Replace any hardcoded calls pointing to llmstxt.firecrawl.dev, migrating pipeline generation logic to standard API crawl requests or custom scripts backed by mendableai/create-llmstxt-py.13
Questions
How to generate an llms.txt file?
You can generate an llms.txt file by submitting a target site URL to the Firecrawl generator endpoint or running the open-source script from the mendableai/create-llmstxt-py repository. The pipeline crawls site URLs, extracts page content, and summarizes resources into standard markdown using GPT-4-mini.
How can I convert a website to LLM text?
A website can be converted into clean LLM context by querying the full-text endpoint at https://llmstxt.firecrawl.dev/{YOUR_URL}/full or initiating a crawl job using the Firecrawl Crawl API.
Why did my Firecrawl crawl fail with a 402 error?
Firecrawl pre-checks account balances against the maximum crawl limit before starting any job, which defaults to 10,000 pages. If your available credits are lower than this limit, the API aborts the request immediately with an HTTP 402 Payment Required error.