pi-firecrawl-lite

A lightweight Pi extension that brings Firecrawl search, scrape, map, and crawl tools to your agent.

Packages

Package details

extension

Install pi-firecrawl-lite from npm and Pi will load the resources declared by the package manifest.

$ pi install npm:pi-firecrawl-lite
Package
pi-firecrawl-lite
Version
0.1.4
Published
Oct 3, 2026
Downloads
483/mo · 483/wk
Author
jameslindfors
License
MIT
Types
extension
Size
58 KB
Dependencies
0 dependencies · 3 peers
Pi manifest JSON
{
  "extensions": [
    "./index.ts"
  ]
}

Security note

Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.

README

pi-firecrawl-lite

Firecrawl for Pi, with a small prompt footprint and the full v2 capability set.

pi-firecrawl-lite adds reliable web search and page scraping to Pi while keeping less-common crawl and site-mapping tools out of the model’s default tool list. When a task needs them, the model can call firecrawl_load to activate the relevant tools for the rest of the session.

Install

Install the package in Pi’s normal extension directory or add it to a project:

pi install npm:pi-firecrawl-lite

For a project-local installation, add the package to package.json and use the extension entry already included in its pi manifest:

{
  "pi": {
    "extensions": ["./node_modules/pi-firecrawl-lite/index.ts"]
  }
}

The extension has no runtime dependencies. Pi, its TUI package, and TypeBox are peer dependencies supplied by the host.

Configure Firecrawl

/firecrawl is the main configuration entry point. Its menu provides Status, Configure global…, Configure project…, and Session overrides…. Global and project URL/timeout changes save immediately to the explicitly named destination; there is no session-then-save detour. Back returns to the main menu; Cancel or dismissing a dialog exits without mutation. Successful actions return to the conversation. Without UI, bare /firecrawl displays status.

The extension resolves each nonsecret setting independently, in this order:

  1. Session override (temporary, cleared on session initialization/reload).
  2. FIRECRAWL_API_URL / FIRECRAWL_TIMEOUT_MS environment variables.
  3. Project saved defaults.
  4. Global saved defaults.
  5. Built-in/inferred defaults.

The project file is exactly <Pi session working directory>/.pi/firecrawl.json. The extension does not search Git roots, parent directories, or ancestor .pi folders: launching from a nested directory uses that nested directory's project file. Global defaults are <Pi agent directory>/firecrawl.json (normally ~/.pi/agent/firecrawl.json); the configurable PI_CODING_AGENT_DIR is respected. Project settings may be committed/shared, but contain only nonsecret URL and timeout fields.

Without an explicit URL, a configured API key selects Firecrawl Cloud at https://api.firecrawl.dev/v2; otherwise the endpoint is http://localhost:3002/v2. The default timeout is 120,000 ms. API keys come only from FIRECRAWL_API_KEY or temporary session overrides and are never saved.

Supported environment variables:

export FIRECRAWL_API_URL=http://localhost:3002
export FIRECRAWL_API_KEY=fc-your-key       # optional for local deployments
export FIRECRAWL_TIMEOUT_MS=120000         # optional

The URL must be HTTP or HTTPS and must point to the host root or end in /v2. v1 URLs are rejected. Keys are never shown in tool output, status, warnings, notifications, saved settings, or errors.

Explicit commands remain supported. Unqualified setters are session-only; unqualified save/settings commands continue to target global settings for compatibility. Add --scope global or --scope project to nonsecret persisted operations:

/firecrawl status
/firecrawl url http://localhost:3002                 # temporary session override
/firecrawl key                                       # temporary; may prompt
/firecrawl timeout 120000                            # temporary
/firecrawl reset                                     # session overrides only
/firecrawl url https://firecrawl.example --scope global
/firecrawl timeout 5000 --scope project              # persist immediately
/firecrawl save --scope project                      # save explicit session URL/timeout only
/firecrawl settings --scope global                  # view that file, not effective values
/firecrawl settings clear --scope project --yes      # delete exact scope without UI

URL/timeout setters accept scope flags before or after the value. Omitted values prompt only when UI is available; non-interactive calls show usage. key is session-only and cannot take a scope; submitting a blank key in its prompt clears authentication for this session, including an environment-provided key. Invalid, missing, or duplicate scope values and unsupported options are rejected before prompting or mutation, without echoing sensitive option values. save copies only explicitly supplied session URL/timeout overrides; it never copies effective environment/inferred values or keys, and it does not clear session overrides.

Status reports effective URL, authentication presence (configured/not configured), timeout, source labels (session, environment, project, global, default), and the project/global paths. Both direct persistent setters and save identify higher-priority masking sources per saved field rather than claiming the saved value is effective. Viewing saved settings shows the selected file independently from effective configuration.

Configuration menus remain accessible when effective configuration is invalid. A successful save is retained even if a higher-priority value, such as an invalid FIRECRAWL_TIMEOUT_MS, still prevents effective configuration from resolving. Feedback distinguishes that successful persistence from the unresolved configuration; use a valid session override or correct the environment value to restore operation.

settings clear requires confirmation naming the scope and exact file path. Clearing removes only that file; session overrides, environment values, and the other saved scope remain unchanged. Non-interactive deletion requires --yes. reset clears only session overrides, revealing environment and saved layers.

Both files use schema version 1:

{
  "schemaVersion": 1,
  "apiUrl": "http://localhost:3002/v2",
  "timeoutMs": 120000
}

URL and timeout are optional; timeout must be an integer from 1 to 86,400,000 ms. Unsupported fields or schema versions are rejected. Missing files are normal and are not created on read. Project/global load failures are reported independently. Settings read/save/clear failures name the selected scope and destination without exposing file contents or sensitive input values. Failed saves or clears leave the loaded saved layer unchanged. Saving never silently replaces invalid or newer-schema files; explicit confirmed clear can remove them. Run /reload after manual edits.

Writes use a private temporary file and atomic replacement, with mode 0600 where supported. Mutations are serialized per path within a process; separate Pi processes use last-writer-wins behavior. Existing agent-directory permissions are not changed. Request artifacts remain temporary; this feature does not add response caching or history.

Search and scrape are available immediately; firecrawl_load is only needed before using deferred map or crawl tools. Request and argument failures remain sanitized failed tool executions.

Tools

Always active:

  • firecrawl_search — web search with a query, optional result limit, advanced JSON options, and concise/raw output.
  • firecrawl_scrape — scrape a URL in markdown or other Firecrawl formats.
  • firecrawl_load — load deferred tools without removing Pi or other extension tools.

Deferred tools:

  • firecrawl_map — discover URLs on a site.
  • firecrawl_crawl — start a crawl job.
  • firecrawl_crawl_status — inspect progress and returned documents.

Every model-facing schema is flat. options_json accepts an object for advanced Firecrawl features such as browser actions, extraction, headers, location, tags, search filters, webhooks, and nested scrape options. It rejects malformed JSON, arrays, primitive values, and prototype-pollution keys. Explicit top-level arguments always override advanced options. For scrape formats, options_json.formats is preserved when the flat formats argument is omitted; markdown is the default only when neither source specifies formats. Crawl max_depth maps to Firecrawl v2's maxDiscoveryDepth.

Responses are concise by default and are rendered with readable labels and width-aware Markdown in Pi. Set response_mode to raw when debugging or passing the full Firecrawl response to another step. Output is bounded at 50 KiB and 2,000 lines. A truncated response is saved in a private, session-specific temporary directory with restrictive permissions, and the tool reports its exact path and limits. Artifacts are removed when the session shuts down or the extension reloads.

Development

This repository uses Bun:

bun install
bun run validate

The validation command runs unit/integration-style mocked-fetch tests, strict typechecking, and the exact npm tarball allowlist check. The package intentionally publishes only package.json, README.md, LICENSE, the root index.ts entry point, and the four runtime files under src/; tests, workflows, lockfiles, PLAN.md, and local artifacts are excluded.

Releases

Publishing is deliberately release-driven. Maintainers should create and publish a Gitea release whose tag exactly matches the package version, for example v0.1.0. The repository workflow then verifies tests, types, package contents, npm availability, and the tag/version match before publishing:

  • stable releases use npm dist-tag latest;
  • SemVer prereleases use npm dist-tag next.

Do not run npm publish manually. A tag push by itself does not publish anything.

License

MIT © James Lindfors