Firecrawl includePaths regex vs Sume's literal path prefixes

Firecrawl's includePaths are regex on the URL pathname. Sume's include_paths are literal prefixes starting with /, max 10 items. How to rewrite one.

4 min readSume
All posts

Sume's include_paths and exclude_paths are literal path prefixes, not regular expressions. Each item must start with /, be at most 200 characters, and each list holds at most 10 items. Firecrawl's includePaths are RE2-style regex matched against the pathname, so a pattern like /blog/.* must become the prefix /blog/.

Firecrawl's behavior is from its crawl guide and Sume's from OpenAPI and the crawl_site description, read 2026-09-30.

How do I rewrite a regex as prefixes?

Take the fixed start of each pattern, and split alternations into separate items.

Rewriting includePaths for Sume include_paths, read 2026-09-30.
Firecrawl patternSume include_paths
/blog/.*["/blog/"]
/docs/(api|sdk)/.*["/docs/api/", "/docs/sdk/"]
A pattern ending in a file extensionNo equivalent; filter the returned URLs
regexOnFullURL: trueNo equivalent; only paths are matched

What are the limits on the lists?

Ten items per list and 200 characters per item, and each item must match the schema pattern ^/. A full URL such as https://example.com/blog/ fails validation, so pass only the path.

What does a request look like?

Crawl is a write call, so it needs an Idempotency-Key header.

const res = await fetch("https://api.sume.com/v1/firecrawl/crawl", {
  method: "POST",
  headers: {
    Authorization: "Bearer " + process.env.SUME_API_KEY,
    "Content-Type": "application/json",
    "Idempotency-Key": "crawl-blog-001",
  },
  body: JSON.stringify({
    url: "https://example.com",
    limit: 20,
    include_paths: ["/blog/"],
    exclude_paths: ["/blog/tag/"],
  }),
});
console.log(res.status, await res.json());

What do I do with a pattern I cannot express?

Run crawl_map (limit 1-100) on the origin, which returns links with titles and descriptions, filter that list with your own regex in code, and crawl_scrape the survivors. That keeps the regex logic on your side and avoids crawling pages you will discard. The cap on crawl size is covered in the crawl limit post.

Sources

Related posts

More in Developers

All Developers posts

Written by Sume