pi-webveil
Pi extension: web_search and web_fetch tools backed by webveil. A drop-in, anonymity-capable replacement for Ollama's web_search/web_fetch.
Package details
Install pi-webveil from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-webveil- Package
pi-webveil- Version
0.7.1- Published
- Sep 30, 2026
- Downloads
- 249/mo · 40/wk
- Author
- wighawag
- License
- AGPL-3.0-or-later
- Types
- extension
- Size
- 139.6 KB
- Dependencies
- 1 dependency · 0 peers
Pi manifest JSON
{
"extensions": [
"./src/index.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
Anonymous-capable, self-hosted, account-free web search + fetch for AI agents.
webveil replaces account-bound tools (notably Ollama's web_search / web_fetch, which
proxy a hosted service and sign every request with your account identity) with a
self-hosted path that has no account, no API key, and an egress you control
(direct, HTTP proxy, or SOCKS5/Tor) so searches and fetches can be anonymous. It also
works perfectly well non-anonymously (direct egress).
Packages
webveil is a pnpm workspace monorepo. The core (search() / fetch()) is plain,
framework-agnostic. Two thin frontends wrap that same core:
webveil, an incur-based CLI + MCP server (--mcp, skills,--llms, TOON output). Pi-agnostic; usable by any agent (pi via pi-mcp-adapter, Claude Code, Cursor, Codex, bash). Has awebveilbin.pi-webveil, a pi extension registeringweb_searchandweb_fetchtools that call the core in-process. A drop-in replacement for Ollama's tools (same names), which is the original motivation. Depends onwebveilviaworkspace:*.
Quick start
npm install -g webveil # or, in pi: pi install npm:pi-webveil
webveil search "hello world"
That is the whole setup (Node 22 or later, on Linux x64 and arm64 with glibc, macOS x64 and arm64, and Windows x64): no service to run, no account, nothing else to install. The same holds for the pi extension: pi install npm:pi-webveil and its web_search and web_fetch tools work at once, with the same default and the same config. The default backend is searchcast: webveil sends the search itself, over HTTP with a real browser's TLS fingerprint (the native library comes with webveil on those platforms), from your own IP (direct egress) unless you configure an egress.
The default is a floor, not the product. It is there so that search works after one install. What you add with recipes is where webveil earns its keep: see what recipes add, just below.
The default engines, and what recipes add
With no engines configured, webveil tries two engines in order, ["mwmbl", "marginalia"]: Mwmbl first, and Marginalia Search when Mwmbl fails. Both are small, independent, non-commercial indexes, queried through their public keyless APIs (Mwmbl's, Marginalia's) by two code recipes that ship in the webveil package (recipes/mwmbl.mjs and recipes/marginalia.mjs, copies of searchcast's examples/recipes/, each naming the commit it was copied from). Know their limits:
- Per-IP quotas. Without a key, Mwmbl allows 1,000 requests a month at 1 request per second, counted per IP address; Marginalia's shared
publickey has one rate limit for everyone who uses it (on the evening of 2026-09-30 it answered HTTP 504 after 60 seconds to every request, which is why Mwmbl comes first). A Tor exit or VPN shares a quota with everyone else on it. Over quota an engine answers HTTP 429 or 503, which webveil reports as the engine being blocked, and the chain moves on. - More quota with a key. Ask Mwmbl for a personal key and set
MWMBL_API_KEY, and ask Marginalia for a free personal key (non-commercial use) and setMARGINALIA_API_KEY, in webveil's environment. Mwmbl's key goes in its request URL (theapi_keyquery parameter), so it can appear in error messages. Marginalia's key selects its current API (api2.marginalia-search.com), which receives it in anAPI-Keyheader, never in the URL; without it, Marginalia is asked through its older URL-keyed API with the sharedpublickey. - The results licence. Both provide their results under CC BY-NC-SA 4.0.
- Small indexes. Neither is a general-purpose engine: expect fewer results, and different ones, than a big engine returns. Each sees your IP (or your egress's) and your queries.
webveil is much more capable with recipes. A recipe is a small file that turns a site's search into an engine, and the chain of engines is yours to choose:
- Any site's search as an engine. A declarative recipe (JSON) gives the search URL and CSS selectors for the results, the "no results" page and a challenge page.
- Code recipes (JavaScript modules) handle what a template cannot: a site's JavaScript result flow, a JSON API, or a challenge to answer before the results.
- The engine chain falls through. An engine that is blocked, times out or no longer fits its recipe hands over to the next one, and the answer names the engines that failed first.
- A real browser as the last fallback. A browser engine runs a recipe in
@searchcast/browser, with a real browser's fingerprint and the page's JavaScript, for when the HTTP engines are blocked. - One pinned command installs a recipe set:
webveil install-recipes <url> --sha256 <hex>(the checksum is your trust decision), then name it withset:<name>.
The shape of a declarative recipe (placeholders; the format is documented in @searchcast/recipe):
{
"name": "example",
"navigate": {"url": "https://example.test/search?q={query}"},
"ready": ".result a",
"empty": ".no-results",
"blocked": ["#challenge"],
"results": {
"item": ".result",
"fields": {
"title": {"selector": "a"},
"url": {"selector": "a", "attr": "href"},
"content": {"selector": ".snippet"}
}
}
}
And of a code recipe (placeholders; see searchcast's code recipes and the two bundled ones for real examples):
export default {
name: 'example-api',
async search(query, ctx) {
const data = await ctx.http.json(
`https://api.example.test/search?q=${encodeURIComponent(query)}`,
{kind: 'document'},
);
if (!Array.isArray(data?.results)) ctx.recipeError('no results array');
return data.results.map((r) => ({title: r.title, url: r.url}));
},
};
Then list them in the global config, ~/.config/webveil/config.json, most wanted first, with the browser engine last:
{
"searchcast": {
"recipes": ["recipes"],
"codeRecipes": ["recipes/example-api.mjs"],
"engines": ["example-api", "example", "searchcast:example"]
}
}
Any engines, recipes or code recipes you configure replace the default chain entirely (nothing is merged with it; to keep Mwmbl or Marginalia, copy its file next to your recipes and list it). Code recipes run with full Node access, so they are accepted only from the global config or env, never from a project webveil.json. The full walkthrough is The searchcast backend; webveil doctor shows the chain in use (its defaultChain notice says when it is the default one).
Other backends, and upgrading from the SearXNG default
web_fetch (webveil fetch <url>) uses the browser fingerprint too with the searchcast backend. Where the native library is missing, search and fetch fail with the fix (webveil install-libcurl, or "fetchTransport": "plain" for fetch), never a silent fallback to Node's fingerprint. webveil doctor checks the whole setup and makes no network request.
Besides searchcast recipes, the options are:
searxng: a SearXNG metasearch service you run yourself ("backend": "searxng"). See With a local SearXNG.tavily-compat(an account or key) andcustom(your own command): see How it works.
No option is zero-setup, account-free and full web results at once (see work/notes/ideas/default-backend-policy-account-vs-origin.md). Zero setup gives you two small independent indexes through keyless APIs with per-IP quotas, with no account (each engine sees your IP, or your egress's, and your queries); recipe sets need recipes you trust, SearXNG needs a service you run, and tavily-compat needs an account or a key.
Upgrading from the SearXNG default. Up to webveil 0.10 the default backend was a local SearXNG at http://127.0.0.1:8080. webveil does not look for one (that would be a guess, and a request you did not ask for): set "backend": "searxng" in the global config (or WEBVEIL_BACKEND=searxng) to keep using it; baseUrl still defaults to http://127.0.0.1:8080. While the backend is the built-in default, a failed search and webveil doctor (its defaultBackend notice) remind you of this.
The searchcast backend: recipes and engine chains
Requires Node 22 or later. The native library comes with webveil on Linux x64 and arm64 (glibc), macOS x64 and arm64, and Windows x64; elsewhere, use your own build (step 2).
webveil is the only thing you install. To run engines of your own choosing instead of the default chain, in short:
npm i -g webveil # brings the native library on supported platforms
webveil install-recipes <url|file> --sha256 <hex> # a recipe set you trust (or write your own, step 3)
webveil doctor # library, config, engines and recipe files, in one report
Install webveil. It brings the
searchcastlibrary (formerlyserpcast) with it and, on the platforms above, the pinned libcurl-impersonate as searchcast's optional dependency@searchcast/libcurl-<platform>, with no download at run time.npm install -g webveilCheck libcurl-impersonate, the native library every engine request goes through, so that the TLS and HTTP/2 fingerprint are Chrome's. On the platforms above it came with step 1:
webveil doctorshows it withsource: platform package.webveil install-libcurlis only needed where that optional dependency was skipped (optional dependencies turned off, for examplenpm install --omit=optional) or the platform has no package: it downloads the pinned release, verifies its sha256 and puts it in~/.local/share/searchcast/($XDG_DATA_HOME/searchcast/), where it is found with no configuration. It is a download of code webveil will load, so it happens only when you run it.webveil doctor # libcurl.impersonating: true, and no problems webveil install-libcurl # only when doctor finds no library; with a proxy egress configured: add --egress to download through itA library
serpcast install-libcurlput in~/.local/share/serpcast/is still found for one release when the new directory has none;webveil doctorthen says so and prints themvcommand that moves it.The download never silently uses your configured
egress, because where a download goes is your decision when you type the command (and a per-folder egress picked up from the cwd could be the wrong one). With adirectegress it is direct unless you pass--proxy(http://,socks5://orsocks5h://;socks5h://resolves host names at the proxy). With a proxy egress (egressorfetchEgressset tohttporsocks5, from any config layer or env) it is not guessed either way: a direct download would tell GitHub that your own IP installed it, so the command refuses (exit 2, before any request) until you choose the route. Usually that is--egress: the download goes through the egress webveil already uses, resolved assearchresolves it (cwd, project, global config, env) and mapped as for searchcast (SOCKS assocks5h://, so host names resolve at the proxy), with its credentials used but never printed. It takesegress, orfetchEgresswhen only that hop is a proxy; when both are proxies and differ it takesegressand says so. With every hopdirect,--egressis simply a direct download. An egress the downloader cannot use (anhttps://proxy, a malformed URL) fails before any request. The other choices are--proxy <url>(the error shows the value matching your egress, credentials shown as***for you to put back) and--direct. Only one of the three may be given. An identical library already installed is left alone; a different one is replaced only with--force.Already have libcurl-impersonate 2.1.1 or later (Nix, a distro package, the one your SearXNG uses)? Skip the install and point webveil at it:
"searchcast": {"libcurlPath": "/path/to/libcurl-impersonate.so"}in the global config, orWEBVEIL_SEARCHCAST_LIBCURL_PATH. webveil loads that library, so a projectwebveil.jsoncannot set it (see What a project config can and cannot do).Drop a recipe in a recipes directory. A recipe is a JSON file that says how to query one engine and read its results page: the URL, and CSS selectors for the results, for "no results" and for a challenge page. webveil and searchcast ship no recipe for a real site: write one for an engine whose terms allow automated access (the format is documented in
@searchcast/recipe, formerlyserpcast-recipe). For example~/.config/webveil/recipes/web.json:{ "name": "web", "navigate": {"url": "https://search.example/?q={query}"}, "ready": "article.result a.title", "empty": ".no-results", "blocked": ["#captcha"], "results": { "item": "article.result", "fields": { "title": {"selector": "a.title"}, "url": {"selector": "a.title", "attr": "href"}, "content": {"selector": ".snippet"} } } }Keep recipes in a directory of their own. Every
*.jsonfile in a recipes directory is loaded as a recipe, so any other JSON file there (for example awebveil.json, via"recipes": ["."]) fails every search with a recipe error naming that file. Try a recipe on its own withnpx searchcast query --recipe ~/.config/webveil/recipes/web.json "hello world".Select the backend in the global config,
~/.config/webveil/config.json:{ "backend": "searchcast", "searchcast": {"recipes": ["recipes"], "engines": ["web"]} }recipeslists recipe files or directories; a relative path is relative to the config file that sets it, never to the cwd.enginesis the engine chain: recipe names, tried in order, and the first engine that answers wins. Setting any ofengines,recipesorcodeRecipesreplaces the default chain entirely (nothing is merged with it): a missing or emptyenginesis then an error, and so is a name that no loaded recipe has. (backendis alreadysearchcastby default; stating it keeps the config working if the default ever changes.WEBVEIL_BACKEND=searchcastworks too.)Search.
webveil search "hello world"
What to expect:
- A failing engine hands over to the next one. An engine that is blocked, times out, or returns a page that no longer fits its recipe is skipped for the next engine in the chain, and the results then carry
unresponsiveEngines, naming the engines that failed first, so a fallback answer never passes for the first choice. An engine that answeredblockedcools down: it is skipped forsearchcast.cooldownMs(default 5 minutes). - Decoy guard (for Bing and the like). Some engines, Bing above all, answer many queries (about half of realistic ones, measured in September 2026) with a well-formed page of results unrelated to the query: a decoy. searchcast checks the answers of a guarded engine: a page where at most one of the top 5 results mentions two of the query's content words counts as a decoy, which fails that engine (named in
unresponsiveEngines, and listed asdecoyif every engine fails) and hands over to the next one, with no cooldown. An engine is guarded when its recipe declares"decoyProne": true(the recipe's author knows the site serves decoys, so nothing is needed in your config), or when you list it insearchcast.decoyGuard(for example"decoyGuard": ["bing"], any config layer, orWEBVEIL_SEARCHCAST_DECOY_GUARD=bing, comma-separated).decoyGuardis therefore only needed for a recipe that does not declare it; it adds to the recipes' own declarations and cannot switch a declared one off. An engine that is neither is never judged. The rule is searchcast'sisDecoy: see searchcast's decoy guard. - Every engine failed: an error, never an empty list. The error lists each engine's failure. An empty result list means an engine's
emptyselector matched: the engine really found nothing. - No fingerprint, no search. Impersonation is always strict: if libcurl-impersonate is missing or not the right library, webveil refuses to search and says how to fix it. It never sends a request with a non-browser fingerprint.
- Sessions persist. Engine cookies (such as a challenge clearance) and cooldowns are kept between calls in
~/.local/state/webveil/, one directory per identity (egress plus searchcast settings), and expire aftersearchcast.sessionIdleMs(default 10 minutes) of disuse."searchcast": {"state": {"persist": false}}keeps them in memory instead, for the life of the process, and writes nothing there.webveil state cleardrops the current identity's state,--allevery identity's. See searchcast state on disk under How it works. - Anonymity. With this backend
egressgoverns the requests that reach the search engines: see Where does anonymity live?. web_fetchgets the browser fingerprint too. With this backendfetchTransportdefaults tosearchcast, soweb_fetchalso goes through libcurl-impersonate (and so also needs it). Set"fetchTransport": "plain"to keep the plain Node transport: see Theweb_fetchtransport.
webveil doctor reports, for the current folder's config: the libcurl-impersonate library webveil would load (where it was found, usually platform package with the package @searchcast/libcurl-<platform>, else the setting or directory that named it; its version; whether impersonation is active), the backend, both egress hops (proxy credentials shown as ***), the web_fetch transport, each engine of the chain with its kind and recipe file, and whether each configured set: is installed. When something is still read from serpcast's old data directory (~/.local/share/serpcast: the library or a recipe set), it adds oldDataDir, listing those items with the command that moves them to searchcast's; that is a notice, not a problem. While the default chain is in use (the searchcast backend with no engines, recipes or code recipes configured), it adds defaultChain, one line saying so and pointing to recipes; a notice too. It exits 1 with the list of problems when something is wrong (the library only counts when the searchcast backend or the searchcast fetch transport uses it). It makes no network request unless you add --remote, which asks the fingerprint echo service https://tls.browserleaks.com/json once, through the searchcast backend's egress (else the fetch hop's), and reports the JA3, JA4 and HTTP/2 values it saw.
The four setup commands (install-libcurl, install-recipes, recipes, doctor) are CLI only: they are not MCP tools (webveil --mcp serves search, fetch and state_clear) and not pi tools. The installers download code webveil runs, and the checksum pin is your trust decision, not an agent's. They load searchcast's install code only when they run: search and fetch never load it.
Private recipes (code recipes, or recipes you would rather not publish) go outside any repository, for example in ~/.config/webveil/recipes/, named from the global config: see Private searchcast recipes under How it works. The default chain runs searchcast's examples/recipes/mwmbl.mjs and marginalia.mjs, bundled in webveil, but only while you configure no chain of your own. To keep either in your own chain, copy its file next to your recipes and list it under searchcast.codeRecipes in the global config (a code recipe, so never from a project webveil.json), with mwmbl or marginalia in engines. For more quota, set MWMBL_API_KEY (sent in Mwmbl's query) or MARGINALIA_API_KEY (selects Marginalia's current API, key in a header) (see The default engines).
Installing recipes
A recipe repository can publish its recipes as one release archive (a .tar.gz of *.json and *.mjs files, optionally with a manifest.json naming the set). webveil installs it (through searchcast's installer), only when you type the command:
webveil install-recipes https://example.com/my-set-1.2.0.tar.gz --sha256 <hex> # with a proxy egress configured: add --egress to download through it
webveil recipes # the installed sets, their source, sha256 and files
- Why the pin is required. A set may hold code recipes, which are code with full Node access, so
--sha256(the archive's checksum, required for a URL and a local file alike) is your trust decision: exactly these bytes, which you reviewed or whose publisher you trust. Get it from the publisher through a channel you trust, or compute it yourself (sha256sum my-set-1.2.0.tar.gz) after reviewing the archive. It is checked before anything is unpacked. - Where it goes. The set is written to
~/.local/share/searchcast/recipes/<set>/($XDG_DATA_HOME/searchcast/recipes/<set>/), named after--nameor the archive'smanifest.json, with a.source.jsonrecording where it came from. An identical set already there is left alone; a different one is replaced only with--force. There is no--dir: webveil looks sets up there only, and, for one release, in serpcast's old~/.local/share/serpcast/recipes/when the new directory has no set of that name (nothing is ever written there;webveil recipeslists those sets asoldSets, andwebveil doctorprints the command that moves them). - How the download leaves, as for
install-libcurl: direct with adirectegress unless you pass--egressor--proxy; with a proxy egress (egressorfetchEgress), refused until you pass--egress(through the configured egress, credentials included),--proxy <url>or--direct. A local file makes no request, so it needs none of them (--egressand--directare accepted and ignored). - A private GitHub release asset needs authentication, which
install-recipesdoes not do. Download it with the GitHub CLI first, then install the file:gh release download v1.2.0 --repo owner/recipes --pattern 'my-set-1.2.0.tar.gz', thenwebveil install-recipes ./my-set-1.2.0.tar.gz --sha256 <hex>. webveil does not routegh: if the download must not come from your own IP, run it through the proxy yourself (torsocks gh release download ..., orALL_PROXY/HTTPS_PROXYforgh). Under anonctl the kernel forces it through the tunnel anyway.
Point webveil at an installed set with set:<name> in searchcast.recipes (its declarative recipes) and searchcast.codeRecipes (its code recipes), in the global config:
{
"backend": "searchcast",
"searchcast": {
"recipes": ["set:my-set"],
"codeRecipes": ["set:my-set"],
"engines": ["api", "web"]
}
}
set:my-set is the whole set: every *.json file except its manifest.json for recipes, every *.js and *.mjs file for codeRecipes. set:my-set/api.mjs names one file, for a set that also holds helper modules that are not engines. The set is looked up in searchcast's recipes directory, then serpcast's old one (following XDG_DATA_HOME in webveil's environment), never in the cwd; a set that is not installed is an error saying how to install it. codeRecipes stays a global-config-or-env setting whatever its form (WEBVEIL_SEARCHCAST_CODE_RECIPES=set:my-set works too): a project webveil.json naming set:... there is refused like any other code recipe path. The plain path (~/.local/share/searchcast/recipes/my-set) works as well, but ignores XDG_DATA_HOME.
A browser engine as the fallback (searchcast)
A browser engine runs a declarative recipe in searchcast's browser runner @searchcast/browser (formerly the searchcast 0.1 package), a real browser, so it has a real browser's fingerprint and runs the page's JavaScript. It is the heaviest engine, so it usually goes last in the chain, to answer when the HTTP engines are blocked. Name it searchcast:<recipe> in searchcast.engines, so one chain mixes HTTP and browser engines and one recipe file can run both ways:
{
"backend": "searchcast",
"searchcast": {
"recipes": ["recipes"],
"engines": ["web", "searchcast:web"],
"browser": {"mode": "library", "xvfb": "/usr/bin/Xvfb"}
}
}
searchcast.browser.mode says how webveil reaches the browser:
library(the default): webveil starts the browser runner in-process, withegressas the browser's proxy. The recipe must be a loaded declarative recipe (a code recipe cannot run in a browser).@searchcast/browseris not installed with webveil: install it next to webveil (npm install -g @searchcast/browserfor a global webveil), with a Chromium or Chrome executable, found by the browser runner ($SEARCHCAST_CHROME, thenPATH) or set assearchcast.browser.chrome. Optional keys:xvfb(an Xvfb executable: the browser then runs headful on a private virtual display),chromeArgs(extra Chromium arguments) andpersistProfile(truekeeps the browser profile, its cookies included, in the identity's state directory, deleted once idle pastsessionIdleMs; the default is a temporary profile per process).chrome,xvfbandchromeArgsmake webveil run a program, so they are accepted only from the global config or env (WEBVEIL_SEARCHCAST_BROWSER_CHROME,WEBVEIL_SEARCHCAST_BROWSER_XVFB, andWEBVEIL_SEARCHCAST_BROWSER_CHROME_ARGS, split on whitespace), and a Chromium argument that sets the browser's proxy or host resolution is refused (see What a project config can and cannot do).endpoint: asearchcast serveyou run yourself,"browser": {"mode": "endpoint", "endpoint": "/run/searchcast/searchcast.sock"}(a Unix socket path or anhttp://URL);<recipe>is then the recipe's name on that server and is not loaded locally. webveil does not run that browser, so it cannot put it on its egress: endpoint mode requiresegress: direct(see Where does anonymity live?).
Two limits of library mode. A persistent profile (persistProfile) must not be used by two webveil processes of the same identity at once (two parallel CLI calls, or the CLI next to a running MCP server or pi extension): Chromium's profile lock refuses the second browser. And @searchcast/browser is resolved from webveil's own install location, so it must be installed alongside webveil: webveil declares it as an optional peer dependency (@searchcast/browser >=0.1.0 <0.2.0), so a package manager does not install it for you but, once you install it next to webveil (in the same project, or globally next to a global webveil), exposes it to webveil even under a strict pnpm layout. Without it a library-mode engine fails before any search with an error saying to install it. If you installed the old searchcast 0.1.x package next to webveil for this, install @searchcast/browser instead (for a global webveil: npm uninstall -g searchcast && npm install -g @searchcast/browser; webveil brings its own searchcast library).
On NixOS
Everything below runs webveil's searchcast backend (then named serpcast) on NixOS without anything else. Verified on 2026-09-29 (NixOS, Node 24.19.0, webveil 0.5.0, serpcast 0.1.1), except where a point says otherwise. The setup commands became webveil's own on 2026-09-30 (webveil install-libcurl, webveil install-recipes, webveil recipes, webveil doctor); they run searchcast's installers (serpcast's, renamed), the ones verified here as npx serpcast ..., which now write to ~/.local/share/searchcast.
libcurl-impersonate. Installing webveil brings searchcast's linux platform package, which holds the same pinned prebuilt library as
webveil install-libcurl(not yet checked on NixOS itself;webveil doctortells).webveil install-libcurlworks (with the Tor egress below,webveil install-libcurl --egress, so the download goes through Tor too): the prebuilt linux-x64 library needs only libc and uses NixOS's CA bundle at/etc/ssl/certs/ca-certificates.crt. nixpkgs'curl-impersonateworks too (the same pinned 2.1.1 release, identical JA4 and HTTP/2 fingerprint). Do not put a/nix/store/...path in a config file (it breaks after an upgrade or garbage collection): link the library where searchcast looks by default instead, for example with home-managerxdg.dataFile."searchcast/libcurl-impersonate.so".source = "${pkgs.curl-impersonate}/lib/libcurl-impersonate.so";, or setWEBVEIL_SEARCHCAST_LIBCURL_PATHin a dev shell. Check withwebveil doctor.Recipes.
webveil install-recipes <archive> --sha256 <hex>works as is on NixOS, into~/.local/share/searchcast/recipes/<set>/(see Installing recipes). With the Tor egress below, both install commands require a route for a URL:--egressdownloads through that same Tor egress (assocks5h://127.0.0.1:9050), or give--proxy <url>or--direct, and a private asset fetched withgh release downloadgoes through Tor only if you routeghyourself (torsocks gh ..., orALL_PROXY/HTTPS_PROXY); under anonctl the kernel forces both anyway. For a declarative setup, home-manager can place a pinned, unpacked archive there instead, which webveil reads the same way (set:<name>):# a public archive: fetchzip pins the UNPACKED content (a Nix hash, not the archive's sha256) xdg.dataFile."searchcast/recipes/my-set".source = pkgs.fetchzip { url = "https://example.com/my-set-1.2.0.tar.gz"; hash = "sha256-..."; # leave empty once; Nix prints the right one # stripRoot = false; # when the files sit at the archive's root, not under one directory }; # a private archive (downloaded with gh, see above): a local, already unpacked directory # xdg.dataFile."searchcast/recipes/my-set".source = ./recipes/my-set;The
fetchzipform was built with Nix 2.34.8 and nixpkgs 26.11 on 2026-09-29 (a set under one top directory with the defaultstripRoot, and a root-level set withstripRoot = false; a root-level set with the default fails with "zip file must contain a single file or directory"). Thexdg.dataFileline is home-manager's documented option shape, not run here. A set placed this way has no.source.json, sowebveil recipesshows no source for it; webveil loads it anyway, andwebveil doctorshows it installed and lists its engines' files.Where the global config lives.
~/.config/webveil/config.json($XDG_CONFIG_HOME/webveil/config.json), as on any Linux. With home-manager, write it from Nix:xdg.configFile."webveil/config.json".text = builtins.toJSON { ... };(the file is then a read-only link into the store; edit the Nix expression instead).A complete minimal config: the searchcast backend over Tor, recipes from an installed set, and
web_fetchover searchcast too:services.tor.client.enable = true; # NixOS: SOCKS on 127.0.0.1:9050 # home-manager: xdg.configFile."webveil/config.json".text = builtins.toJSON { backend = "searchcast"; egress = { mode = "socks5"; url = "socks5://127.0.0.1:9050"; }; fetchTransport = "searchcast"; # the default with this backend; stated for clarity searchcast = { recipes = [ "set:my-set" ]; codeRecipes = [ "set:my-set" ]; engines = [ "api" "web" ]; # recipe names from the set, in order }; }; xdg.dataFile."searchcast/libcurl-impersonate.so".source = "${pkgs.curl-impersonate}/lib/libcurl-impersonate.so";The
builtins.toJSONexpression was evaluated withnix eval(it gives{"backend":"searchcast","egress":{"mode":"socks5","url":"socks5://127.0.0.1:9050"},"fetchTransport":"searchcast","searchcast":{"codeRecipes":["set:my-set"],"engines":["api","web"],"recipes":["set:my-set"]}}), and webveil's tests load an installed set throughset:<name>; the whole block was not applied with home-manager here. Replacemy-set,apiandwebwith your set and its recipe names, then check withwebveil doctorandwebveil search "hello world".npm's allow-scripts warning for
koffi(npm 11 blocks its install script) is harmless: koffi's prebuilt binary loads without it.Browser engines (library mode). Set
searchcast.browser.chrometo nixpkgs Chromium (${pkgs.chromium}/bin/chromium, or/run/current-system/sw/bin/chromiumwhen installed system-wide); browsers downloaded by other tools (such as Playwright's) run only withprograms.nix-ldenabled. Without a display (a server, an SSH session) setsearchcast.browser.xvfbto nixpkgs' Xvfb (${pkgs.xvfb}/bin/Xvfb,pkgs.xorg.xvfbon older nixpkgs), or pass"chromeArgs": ["--headless=new"](easier for sites to detect). Both verified.Tor.
services.tor.client.enable = true;gives a SOCKS proxy on127.0.0.1:9050: use"egress": {"mode": "socks5", "url": "socks5://127.0.0.1:9050"}(webveil hands searchcastsocks5h, so DNS stays at Tor). Search andweb_fetch(both transports) exit through Tor (check.torproject.org/api/ipreportsIsTor: true) with the same fingerprint as direct.
Upgrading from serpcast (webveil 0.10)
Up to webveil 0.10 this backend was called serpcast, after the library it ran (now searchcast). Every spelling below is renamed; the old one still works for one release, exactly like the new one, and prints one warning per spelling naming the new one (on stderr for the CLI and the MCP server, in the tool result and pi's notifications for pi-webveil, and under deprecations in webveil doctor). The next minor release removes the old spellings.
| old | new |
|---|---|
backend: "serpcast" |
backend: "searchcast" |
config section serpcast.* |
searchcast.* |
serpcast.searchcast.* (the browser runner's options) |
searchcast.browser.* |
fetchTransport: "serpcast" |
fetchTransport: "searchcast" |
fetchSerpcast.* |
fetchSearchcast.* |
WEBVEIL_SERPCAST_* |
WEBVEIL_SEARCHCAST_* |
WEBVEIL_SERPCAST_SEARCHCAST_* |
WEBVEIL_SEARCHCAST_BROWSER_* |
WEBVEIL_FETCH_SERPCAST_* |
WEBVEIL_FETCH_SEARCHCAST_* |
WEBVEIL_BACKEND=serpcast |
WEBVEIL_BACKEND=searchcast |
- Both spellings of one setting in one place is an error naming both (for example
serpcast.enginesandsearchcast.enginesin the same file, orWEBVEIL_SERPCAST_LIBCURL_PATHandWEBVEIL_SEARCHCAST_LIBCURL_PATH), never a silent pick. Different settings may mix spellings, and across files and env the usual precedence (env > project > global) applies to the renamed setting. - The trust rule does not depend on the spelling: a project
webveil.jsoncannot setserpcast.codeRecipes,serpcast.libcurlPathorserpcast.searchcast.chromeany more than their new names. - The rename resets nothing: the identity key of a config (its state directory and browser profile) is the same in either spelling, and the rename does not change it.
- The browser engine names
searchcast:<recipe>inenginesare unchanged.
With a local SearXNG
A local SearXNG at http://127.0.0.1:8080 on direct egress (non-anonymous) was the default backend up to webveil 0.10. Select it in the global config, ~/.config/webveil/config.json:
{"backend": "searxng"}
(or WEBVEIL_BACKEND=searxng). Run one with Docker:
# The container binds 8080 internally; map host 8080 -> 8080 to match the default.
docker run -d --name searxng -p 8080:8080 searxng/searxng
Then searches and fetches work with no other setting:
webveil search "hello world"
That's the whole happy path. Two things to know:
- Enable SearXNG's JSON API. A fresh install may serve only HTML and ship with the rate
limiter on, giving webveil
429or an HTML page. The stock Docker image above works out of the box; other installs needjsoninsearch.formatsandserver.limiter: falsefor local use. See SearXNG setup. - Different port? Point webveil at it (a non-Docker install often defaults to 8888):or set
export WEBVEIL_BASE_URL=http://127.0.0.1:8888 # or wherever your instance listensbaseUrlinwebveil.json.
For any non-Docker topology (install script, Unix sockets, reverse proxy, the
uwsgi-vs-http-socket catch), see SearXNG setup (detailed).
Where does anonymity live? (read before turning on egress)
webveil's egress only anonymizes webveil's OWN outbound hop (webveil → backend, and
web_fetch → the target URL). It does NOT anonymize what a backend does next. This has a
load-bearing consequence for SearXNG:
- A local SearXNG makes its actual search-engine requests (→ Google/Bing/…) from
its own process, on your machine, with your real IP. That hop is OUTSIDE webveil's
egress. So setting
WEBVEIL_EGRESS=socks5whilebaseUrlis127.0.0.1does NOT make your searches anonymous, webveil would just be proxying a pointless localhost call, while SearXNG crawls the web from your real IP. That is false confidence, the worst outcome. - webveil refuses this combo (fail-loud): a non-
directegress (http/socks5) with a loopbackbaseUrlis rejected with an error, rather than silently giving you fake anonymity. (A remote SearXNG over SOCKS is legitimate and allowed, the guard keys on loopback specifically.)
The searchcast backend has no such split: the search hop is webveil's own. webveil itself sends every engine request (declarative recipes, code recipes, and the library-mode searchcast browser, which gets egress as its proxy), so egress governs exactly the traffic that reaches the search engines, and egress: socks5 really anonymizes search. A SOCKS proxy is always handed over as socks5h, so host names are resolved at the proxy. The loopback guard above does not apply (searchcast has no baseUrl hop); it concerns SearXNG. searchcast has one case of the same trap, and webveil refuses it the same way:
- An external searchcast endpoint (
searchcast.browser.mode: "endpoint") is a browser webveil does not run, so it reaches the engines from wherever it runs, outside webveil's egress. With a non-directegress and asearchcast:engine in the chain, webveil refuses to search. Use library mode, oregress: directwith the searchcast server proxied by other means (anonctl, below). - Also refused, with a library-mode browser engine: a SOCKS
egressURL with credentials. Chromium has no SOCKS authentication, so the credentials (and any Tor circuit isolation they select) would be silently dropped.
Under anonctl (every connection of one Unix account forced through an anonymizer such as Tor by the kernel, fail-closed), run webveil as that account with egress: direct, the default. searchcast's engine requests, the library-mode browser and web_fetch all leave through the forced tunnel with no further configuration, and endpoint mode is allowed too (direct), provided the searchcast server runs under the forced account as well.
webveil splits egress into two independent hops so you can set them differently (see Per-hop egress below):
egressgoverns the BACKEND hop (webveil -> backendbaseUrl).fetchEgressgoverns theweb_fetchhop (webveil -> arbitrary URL). It DEFAULTS to inheritingegresswhen unset, so a singleegressstill drives both hops exactly as before.
So the correct setups:
| Goal | backend hop (egress) |
web_fetch hop (fetchEgress) |
backend | Who anonymizes the web hop |
|---|---|---|---|---|
| Local SearXNG, anonymous searches | direct |
direct |
local SearXNG | SearXNG itself, set its outgoing.proxies (Tor/SOCKS) in settings.yml |
Local SearXNG + anonymous web_fetch (the common one) |
direct |
socks5 |
local SearXNG | SearXNG's outgoing.proxies for SEARCH; webveil's hop for FETCH |
| Remote SearXNG, hide your IP from it | socks5 |
(inherits) | the remote SearXNG url | webveil's hop (Mullvad/Tor) |
Anonymous web_fetch of arbitrary URLs |
direct/(any) |
socks5 |
(any) | webveil's hop |
| Non-anonymous everyday use | direct |
direct |
local SearXNG | nobody (honest) |
| searchcast, anonymous searches (no SearXNG) | socks5 |
(inherits) | searchcast |
webveil's hop, for SEARCH and FETCH |
| searchcast under anonctl | direct |
direct |
searchcast |
anonctl (the whole account is forced) |
| searchcast, non-anonymous | direct |
direct |
searchcast |
nobody (honest) |
Rule of thumb: proxy the hop that actually reaches the public internet. For SEARCH
via a self-hosted SearXNG that hop is SearXNG's, so the search proxy goes on SearXNG
(outgoing.proxies) and webveil's backend hop stays direct. For web_fetch the public
hop is webveil's own, so fetchEgress: socks5 is correct even while the backend stays
local. webveil's backend socks5 mode is for remote backends. See
work/notes/findings/webveil-anonymity-boundary.md
and docs/adr/0003.
With the searchcast backend the hop that reaches the public internet for SEARCH is webveil's own, so the search proxy goes on egress, and fetchEgress inherits it unless you set it. See docs/adr/0004.
The web_fetch transport (fetchTransport)
web_fetch downloads the page itself (distilly then turns it into markdown), and fetchTransport says how:
| key | values | default | env |
|---|---|---|---|
fetchTransport |
plain: undici over the fetch-hop egress, a Node TLS fingerprint. searchcast: searchcast's libcurl-impersonate transport, Chrome's TLS and HTTP/2 fingerprint and the header set of a Chrome page load |
searchcast when backend is searchcast, plain otherwise |
WEBVEIL_FETCH_TRANSPORT |
An explicit value (from any config file or env) always wins over the default, so "fetchTransport": "searchcast" works with SearXNG too, and "plain" keeps the Node transport with the searchcast backend. Any other value is an error. Sites that gate on the TLS fingerprint answer a Node fetch with a block or challenge page; searchcast gets the page a browser would, as long as the page does not need JavaScript (no script runs).
With searchcast:
- Same egress. Requests leave through the fetch-hop egress (
fetchEgress, elseegress), mapped as for the searchcast backend:directis a direct connection, anhttpproxy is used as is, and a SOCKS proxy is always used assocks5h(host names resolved at the proxy). - Same SSRF guard, on every hop. The transport follows no redirects, so webveil follows them (at most
fetchMaxRedirects, default 20, http and https only) and checks each target before sending it: ondirectegress a redirect to a private or loopback address is refused. - No fingerprint, no fetch. Impersonation is strict, as for search: without libcurl-impersonate
web_fetchfails with the fix in the error (webveil install-libcurl, orsearchcast.libcurlPath). It never falls back toplain. The library path is the samesearchcast.libcurlPath, with the same trust rule (never from a projectwebveil.json);searchcast.enginesis not needed. - No cookies between fetches. Each fetch starts with an empty cookie jar: cookies set during one fetch's redirects are sent on its later hops, then dropped. Nothing is written to the searchcast state directory: a fetch is not a search identity, and cookies carried from page to page would link the fetches.
- Connections are reused between fetches (of the same egress). A new TCP + TLS setup per fetch costs about 0.5 s over Tor, a reused connection about 0.09 s, so webveil keeps a small pool of searchcast sessions per fetch-hop egress: a fetch takes an idle one (its cookies cleared), uses it alone, and returns it. Concurrent fetches get different sessions, a different egress never shares one, at most 4 idle sessions are kept (
fetchSearchcast.maxIdleSessions), and an idle session is closed after 10 minutes (fetchSearchcast.sessionIdleMs) or when webveil exits.fetchSearchcast.reuseConnections: falseopens one connection per request instead. The trade-off: a site sees repeated fetches to it arrive on one TLS connection, which links them at the connection level. Through one proxy or Tor circuit they already share an exit IP, and no cookie or other state is carried from one fetch to the next. - GET only, and only the Chrome page-load headers are sent. The time limit (15 s,
fetchSearchcast.timeoutMs) and the largest body (16 MiB,fetchSearchcast.maxBodyBytes) default to searchcast's; exceeding either is an error. The redirect limit (20) isfetchMaxRedirects, shared with theplaintransport.
A backend with its own /extract (tavily-compat) still fetches through that endpoint whatever fetchTransport says.
Chained egress: Mullvad base → rotating-residential exit → site
For the strongest anonymity stack, a single commercial proxy is a single point of legal exposure: a rotating-residential provider is an adversary that knows who paid (and sees the traffic), while Mullvad sees the traffic but is deliberately decoupled from your identity. Chain them so neither adversary sees both ends:
app → Mullvad (WireGuard base) → rotating-residential exit → site
The whole chain can live inside ONE local SOCKS5 listener (e.g. a sing-box unit whose SOCKS inbound detours through its WireGuard outbound: SOCKS → Mullvad → residential exit), so the chain looks like any other SOCKS endpoint. webveil's own hops are unchanged by this — a chain is just another SOCKS5 listener to webveil, it needs no special mode:
fetchEgresspoints at the chain's port like anysocks5h://URL (socks5h://127.0.0.1:1080), exactly like the single-proxy examples above.- A local SearXNG backend hop stays
directexactly as in the table above (the backend-hop guard still rejects proxying it); SearXNG carries its own engine crawl through the same chain via itsoutgoing.proxies.
Tor cannot be a chain base. Tor's exit path is fixed to Tor relays — there is no "Tor exits into a paid SOCKS" direction — so Tor can carry the whole path or nothing. A chain that needs a controlled exit identity starts from a WireGuard base (Mullvad), not Tor.
A chain (or any egress rotation) is also the repair when engines flag your direct IP and
serve irrelevant-but-confident results — that fix lives in SearXNG's outgoing.proxies,
never in webveil: see
Troubleshooting: irrelevant-but-confident results.
How it works (seams)
- core, the framework-agnostic
search(query, opts)andfetch(url, opts)functions. Both frontends call the same core. - backend seam, where results/content come from:
searxng(keyless self-hosted metasearch),tavily-compat(a generic Tavily-shaped/search+/extract),custom(a local command via a JSON stdin/stdout contract), andsearchcast(keyless search engines from recipes over libcurl-impersonate, no SearXNG; webveil'segressis its search egress; the default backend, with the bundled Mwmbl and Marginalia chain). The backend is handed a proxiedhttphelper so it cannot bypass egress. The searxng backend also surfaces engine degradation from the response'sunresponsive_engines— partial failures annotate the results (unresponsiveEngines, so a degraded answer never masquerades as a clean one), a probable full outage fails loud. Honest limit: it cannot detect junk results — a 200 "decoy SERP" parses like a real one at this layer — see SearXNG troubleshooting. - egress seam, how outbound HTTP leaves the machine:
direct,http(undiciProxyAgent), orsocks5(Tor127.0.0.1:9050, Mullvad10.64.0.1:1080). SOCKS5 is the mode that matters for anonymity. Fail-loud if a configured proxy cannot be built. Egress is per-request and scoped to webveil ONLY, it is NOT a system-wide proxy. It governs webveil's own search/fetch traffic (and thefetchit injects into distilly), and nothing else: your shell,git push, the browser, and the OS are untouched. So webveil onsocks5does NOT route yourgit pushthrough the proxy. See Anonymous egress andwork/notes/findings/mullvad-socks5-egress-mechanics.md. - config seam, per-folder resolution: env > nearest
webveil.jsonwalking up from cwd > global$XDG_CONFIG_HOME/webveil/config.json(default~/.config/webveil/config.json) > defaults. Per folder = per account/egress. The project file is a frontend-neutralwebveil.jsonread identically by the CLI and the pi extension. Seedocs/adr/0002.- Merge rule. Layers merge key by key, and so do config sections (plain objects): a project
webveil.jsonthat sets one key of a section keeps the global config's other keys in that section. Scalars and arrays are replaced whole by the highest layer that sets them.egressandfetchEgressare also replaced whole (a project{"mode": "direct"}over a global SOCKS5 egress is exactlydirect, neverdirectplus a stray url). - What a project config can and cannot do. A
webveil.jsonis read automatically from any checkout you run webveil in, so a cloned repository must never be able to make webveil run code. A project config may choose the backend, its URL, egress, fetch size and other plain settings, but a setting that makes webveil run code (today: thecustombackend's command, i.e. itsbaseUrlwhenbackendiscustom, the searchcast backend'ssearchcast.libcurlPath, a native library webveil loads, itssearchcast.codeRecipes, JS modules webveil imports and so runs, and the searchcast browser'ssearchcast.browser.chrome,searchcast.browser.xvfbandsearchcast.browser.chromeArgs, programs and arguments webveil launches) is refused when it comes from a projectwebveil.json, with an error naming the file and the key: put it in the global config or env instead. Even from a trusted layer, a searchcast chrome argument that controls the browser's proxy or host resolution (--proxy-server,--no-proxy-server,--proxy-bypass-list,--proxy-pac-url,--proxy-auto-detect,--winhttp-proxy-resolver,--host-resolver-rules,--host-rules, in any case, with one or two dashes) is refused with an error naming it: the browser's proxy comes fromegress, and such an argument would silently take the browser off it. A project may still say"backend": "custom"when the command itself comes from the global config or env. The check runs only where the setting is used, soweb_fetchkeeps working in such a folder. Paths in executable settings never resolve against the cwd:~/is your home directory, a relative path is relative to the config file that set it, a value from env must be absolute or~/-prefixed, and a bare command name (no slash; on Windows, no/, no\and no drive prefix likeC:) is looked up onPATHby webveil itself, searching only absolutePATHentries: empty,.and other relative entries are ignored, so a file in the cwd is never picked up, and a name found in no absolute entry is an error before anything runs. On Windows, only a drive path (C:\...) or a UNC path (\\server\share\...) counts as absolute;\tools\x.exeis relative to the config file (so it takes that file's drive), and a drive-relativeC:x.exeis refused. Seedocs/adr/0004. - Private searchcast recipes. A code recipe (a JS module for an engine that needs challenge handling, or a site whose terms you would rather not publish in a repo) is code with full Node access, and loading it runs it. Keep such recipes outside any repository, for example in
~/.config/webveil/recipes/, and list them in the global config as"searchcast": {"codeRecipes": ["recipes"]}(a module file, or a directory whose*.jsand*.mjsfiles are all loaded; a relative path is relative to the global config file), or in env asWEBVEIL_SEARCHCAST_CODE_RECIPES(absolute or~/paths, separated by:, or;on Windows). A set installed withwebveil install-recipesis namedset:<name>(orset:<name>/<file>) in either form: see Installing recipes. A projectwebveil.jsoncannot name them: it is read automatically from any checkout, so a cloned repository could otherwise make webveil import its own module. The project config may still list those recipes in itssearchcast.enginesorder, becausesearchcastmerges key by key and keeps the globalcodeRecipes. Code and declarative recipes share one name space; a name used twice is an error. - searchcast state on disk, per identity. The searchcast backend keeps its sessions (an engine's cookies, such as a challenge clearance, and a code recipe's own JSON state) and its engine cooldowns in
$XDG_STATE_HOME/webveil/(default~/.local/state/webveil/), so the one-shot CLI behaves like the long-lived MCP server and pi extension. Nothing else is stored (no queries, no results), except the searchcast browser profile when you opt intosearchcast.browser.persistProfile: it lives in the identity's directory as<hash>/browser-profile, holds that browser's own cookies and history, and is deleted before the next use once the identity has been idle pastsearchcast.sessionIdleMs. State is partitioned per identity, the backend-hopegressplus the resolvedsearchcastsection, hashed: each identity gets its own directory<hash>/state.json, so no proxy URL, credential or path appears in a file name, and a session obtained on one egress is never replayed on another (which would link the two). A session keeps its idle expiry (searchcast.sessionIdleMs, default 10 minutes) on disk: it is checked on every read and an expired entry is dropped, so saving state never lets cookies link searches across days. Concurrentwebveil searchcalls share the file safely (a lock file around each update, writes replaced atomically). Files are0600, directories0700. Clear it withwebveil state clear(the identity of the current folder's config) orwebveil state clear --all(every identity); deleting the directory by hand is also fine.
- Merge rule. Layers merge key by key, and so do config sections (plain objects): a project
- extractor seam,
urlToMarkdownviadistilly/fetchby default, injected with webveil's egress-boundfetch; a backend's own/extract(Tavily-compat) may override it. Owns the context-friendly markdown + size presets (s/m/l/f). Seedocs/adr/0001. The injectedfetchis the plain egress fetch, or withfetchTransport: searchcastsearchcast's impersonated transport (see Theweb_fetchtransport). - security, an SSRF guard lives in the egress fetch, so it covers distilly's rule-rewritten requests too. The searchcast fetch transport runs it on every redirect hop.
Configuration reference
Every key, in the global config (~/.config/webveil/config.json), a project webveil.json or env, resolved env > project > global > default (see the config seam above). A section (searchcast, searchcast.browser, searchcast.state, searchcast.decoyRule, fetchSearchcast) merges key by key across layers; a dotted key below is a key inside its section ("searchcast": {"state": {"persist": false}}). Every value is checked where it is used and a bad one fails loud naming the key (an env switch is true or false, anything else is an error). Unset, each key keeps the behaviour webveil had before it was a setting. Keys marked executable make webveil run code and are refused from a project webveil.json.
| key | default | env | meaning |
|---|---|---|---|
backend |
searchcast (searxng up to 0.10) |
WEBVEIL_BACKEND |
searxng, tavily-compat, custom or searchcast. |
baseUrl |
http://127.0.0.1:8080 |
WEBVEIL_BASE_URL |
The backend's URL (unix: socket paths too); for custom, its command (executable). |
apiKey |
none | WEBVEIL_API_KEY |
The tavily-compat key. |
egress |
{"mode": "direct"} |
WEBVEIL_EGRESS, WEBVEIL_EGRESS_URL |
The backend hop's egress: direct, http or socks5 with a url. With searchcast, the search egress. |
fetchEgress |
egress |
WEBVEIL_FETCH_EGRESS, WEBVEIL_FETCH_EGRESS_URL |
The web_fetch hop's egress. |
fetchSize |
m |
WEBVEIL_FETCH_SIZE |
The web_fetch page budget: s, m, l, f. |
fetchTransport |
searchcast with that backend, else plain |
WEBVEIL_FETCH_TRANSPORT |
How web_fetch sends requests. |
fetchMaxRedirects |
20 | WEBVEIL_FETCH_MAX_REDIRECTS |
Redirects web_fetch follows (both transports; 0 follows none). |
fetchSearchcast.maxIdleSessions |
4 | WEBVEIL_FETCH_SEARCHCAST_MAX_IDLE_SESSIONS |
Idle searchcast fetch sessions kept per fetch egress (0 keeps none). |
fetchSearchcast.sessionIdleMs |
600000 (10 min) | WEBVEIL_FETCH_SEARCHCAST_SESSION_IDLE_MS |
An idle searchcast fetch session is closed after this long. |
fetchSearchcast.reuseConnections |
true |
WEBVEIL_FETCH_SEARCHCAST_REUSE_CONNECTIONS |
Reuse connections across fetches; false: one connection per request. |
fetchSearchcast.timeoutMs |
15000 | WEBVEIL_FETCH_SEARCHCAST_TIMEOUT_MS |
The searchcast fetch transport's per-request time limit. |
fetchSearchcast.maxBodyBytes |
16777216 (16 MiB) | WEBVEIL_FETCH_SEARCHCAST_MAX_BODY_BYTES |
The largest page the searchcast fetch transport accepts. |
maxResults |
10 | WEBVEIL_MAX_RESULTS |
Results a search returns when the caller passes none (a caller's maxResults wins). |
httpTimeoutMs |
30000 | WEBVEIL_HTTP_TIMEOUT_MS |
The per-request timeout of webveil's requests to a backend (searxng, tavily-compat). |
searchcast.engines |
the default chain ["mwmbl", "marginalia"] (bundled) while engines, recipes and codeRecipes are all unset; else required |
none | The engine chain, in order. |
searchcast.recipes |
none | none | Declarative recipe files or directories, or set:<name>[/<file>]. |
searchcast.codeRecipes |
none | WEBVEIL_SEARCHCAST_CODE_RECIPES |
Code recipe modules or directories, or set:... (executable). |
searchcast.libcurlPath |
searchcast's lookup (SEARCHCAST_LIBCURL_PATH, the data directory, then the platform package) |
WEBVEIL_SEARCHCAST_LIBCURL_PATH |
The libcurl-impersonate library (executable). |
searchcast.sessionIdleMs |
600000 (10 min) | WEBVEIL_SEARCHCAST_SESSION_IDLE_MS |
An engine's session is dropped after this long unused. |
searchcast.cooldownMs |
300000 (5 min) | WEBVEIL_SEARCHCAST_COOLDOWN_MS |
How long a blocked engine is skipped (0: no cooldown). |
searchcast.decoyGuard |
none | WEBVEIL_SEARCHCAST_DECOY_GUARD, WEBVEIL_SEARCHCAST_DECOY_GUARD_EXCLUDE |
Engines checked for decoys, or {"include": [...], "exclude": [...]} (exclude switches the guard off even for a decoyProne recipe). Replaced whole across layers. Env: comma-separated names. |
searchcast.decoyRule.top, .maxRelevant, .prefix |
5, 1, 5 | WEBVEIL_SEARCHCAST_DECOY_RULE_TOP, _MAX_RELEVANT, _PREFIX |
The decoy rule's thresholds. The defaults are the measured values: other values are at your own risk. |
searchcast.timeoutMs |
15000 | WEBVEIL_SEARCHCAST_TIMEOUT_MS |
Per-request time limit of engine requests. |
searchcast.maxBodyBytes |
16777216 (16 MiB) | WEBVEIL_SEARCHCAST_MAX_BODY_BYTES |
The largest engine response. |
searchcast.reuseConnections |
true |
WEBVEIL_SEARCHCAST_REUSE_CONNECTIONS |
Keep an engine session's connections open between its requests. |
searchcast.keepSessions |
true |
WEBVEIL_SEARCHCAST_KEEP_SESSIONS |
Keep each engine's connections between searches; false: new connections per search. |
searchcast.idlePollMs |
5 | WEBVEIL_SEARCHCAST_IDLE_POLL_MS |
How long an in-flight request waits between looks at idle sockets. |
searchcast.maxRequestBodyBytes |
1048576 (1 MiB) | WEBVEIL_SEARCHCAST_MAX_REQUEST_BODY_BYTES |
The largest POST body a code recipe may send. |
searchcast.preflightCache |
true |
WEBVEIL_SEARCHCAST_PREFLIGHT_CACHE |
Remember an allowed CORS preflight, as Chrome does. |
searchcast.maxPreflightAgeS |
7200 | WEBVEIL_SEARCHCAST_MAX_PREFLIGHT_AGE_S |
The longest a preflight is remembered, in seconds. |
searchcast.maxRedirects |
20 | WEBVEIL_SEARCHCAST_MAX_REDIRECTS |
Redirects a declarative recipe follows (0 follows none). |
searchcast.state.persist |
true |
WEBVEIL_SEARCHCAST_STATE_PERSIST |
Keep sessions and cooldowns on disk; false: in memory only, nothing written to $XDG_STATE_HOME. |
searchcast.state.lockStaleMs |
10000 | WEBVEIL_SEARCHCAST_STATE_LOCK_STALE_MS |
A state lock older than this is taken over (a crashed writer). |
searchcast.state.lockWaitMs |
30000 | WEBVEIL_SEARCHCAST_STATE_LOCK_WAIT_MS |
Waiting longer than this for the state lock is an error. |
searchcast.browser.mode |
library |
none | library or endpoint. |
searchcast.browser.endpoint |
none | none | Endpoint mode: the searchcast serve URL or socket path. |
searchcast.browser.timeoutMs |
30000 | WEBVEIL_SEARCHCAST_BROWSER_TIMEOUT_MS |
Endpoint mode: the whole request's time limit. |
searchcast.browser.maxBodyBytes |
16777216 (16 MiB) | WEBVEIL_SEARCHCAST_BROWSER_MAX_BODY_BYTES |
Endpoint mode: the largest answer. |
searchcast.browser.chrome |
searchcast's lookup | WEBVEIL_SEARCHCAST_BROWSER_CHROME |
Library mode: the browser (executable). |
searchcast.browser.xvfb |
none | WEBVEIL_SEARCHCAST_BROWSER_XVFB |
Library mode: an Xvfb for a private display (executable). |
searchcast.browser.chromeArgs |
none | WEBVEIL_SEARCHCAST_BROWSER_CHROME_ARGS |
Library mode: extra Chromium arguments, whitespace-separated in env (executable). |
searchcast.browser.persistProfile |
false |
none | Library mode: keep the browser profile in the identity's state directory (not with state.persist: false). |
Every key of the searchcast section is part of the search identity (the state partition and the cached instance), so changing any of them, a timeout included, starts a fresh identity with new engine sessions; an unset key does not count, so existing state keeps its place.
Not configurable, on purpose. Strict impersonation (no fingerprint, no request: there is no strict key), the SSRF guard, the trust rule for executable settings (never from a project webveil.json), and the checksum pins of install-libcurl (pinned in searchcast's source) and install-recipes (--sha256 is required). These are the guarantees webveil exists to give, not tuning.
Anonymous egress (Mullvad / Tor)
By default webveil uses direct egress (your real IP, non-anonymous). Anonymity is
opt-in: it is enabled ONLY when you set it in config/env. webveil never auto-enables a
proxy (silent anonymity would be a footgun in the other direction).
Enable SOCKS5 egress for webveil:
export WEBVEIL_EGRESS=socks5
export WEBVEIL_EGRESS_URL=socks5://10.64.0.1:1080 # Mullvad
# or socks5://127.0.0.1:9050 # Tor
or per folder in webveil.json:
{ "egress": { "mode": "socks5", "url": "socks5://10.64.0.1:1080" } }
egress: socks5is for a REMOTE backend, NOT a local SearXNG. If yourbaseUrlis a local SearXNG (unix:or loopback like127.0.0.1/localhost),WEBVEIL_EGRESS=socks5is rejected (fail-loud), because webveil → local-SearXNG is a local hop; proxying it would give fake anonymity while SearXNG still crawls the web from your real IP. The hop that needs proxying is SearXNG's own, so you put the search proxy on SearXNG (outgoing.proxies) and keep the backend hopdirect. To proxyweb_fetchwhile keeping the local backend, setfetchEgress(the FETCH hop), NOTegress(see Per-hop egress). See Where does anonymity live? for the full table; the same SOCKS5 listener (Mullvad/Tor/wireproxy below) plugs into either side.
With the searchcast backend there is no local backend hop to protect: egress: socks5 is exactly the setting that anonymizes your searches, since webveil itself sends the engine requests.
Per-hop egress: local SearXNG + proxied web_fetch
The most common self-hosted topology is a local SearXNG (loopback TCP or a unix:
socket) for search, with web_fetch of arbitrary URLs going out through a SOCKS5
proxy (e.g. wireproxy → ProtonVPN at socks5h://127.0.0.1:1080). The two hops are
independent, so set them independently:
- backend hop (
egress):direct. The local SearXNG anonymizes its OWN engine crawl viaoutgoing.proxiesinsettings.yml. web_fetchhop (fetchEgress):socks5. The fetch target is a real public URL, so proxying it genuinely anonymizes that hop.
Concrete ~/.config/webveil/config.json:
{
"baseUrl": "http://127.0.0.1:8080",
"egress": { "mode": "direct" },
"fetchEgress": { "mode": "socks5", "url": "socks5h://127.0.0.1:1080" }
}
Or a unix:-socket SearXNG with the same proxied fetch:
{
"baseUrl": "unix:/usr/local/searxng/run/socket",
"fetchEgress": { "mode": "socks5", "url": "socks5h://127.0.0.1:1080" }
}
(egress omitted defaults to direct; fetchEgress does NOT inherit it here, it is set
explicitly to socks5.) Or via env:
export WEBVEIL_BASE_URL=http://127.0.0.1:8080 # local SearXNG, backend hop stays direct
export WEBVEIL_FETCH_EGRESS=socks5
export WEBVEIL_FETCH_EGRESS_URL=socks5h://127.0.0.1:1080
Prefer socks5h:// for the fetch hop (the h means the proxy does DNS resolution):
webveil does not resolve hostnames locally under proxy egress, and socks5h makes the
remote-DNS intent explicit so a target hostname is never leaked to your local resolver.
(fetchEgress defaults to inheriting egress when unset, so existing single-egress
configs are completely unchanged. The backend-hop guard still rejects egress: socks5
with a local baseUrl; it does NOT block this fetchEgress setup, because the fetch hop
is a different, genuinely-public hop. See
docs/adr/0003.)
Two layers keep your git push (and everything else) off the proxy
A common worry: "if I route through Mullvad, will my git push to GitHub leak under the
VPN exit IP?" With webveil, no, for two independent reasons:
- webveil's egress is per-request and webveil-only. It applies the SOCKS5 dispatcher
inside its own search/fetch code; it does not install a system proxy.
git, your shell, and the OS are never touched. webveil onsocks5proxies webveil's traffic and nothing else. - You configure split routing (below) so that even at the OS level, only the proxy IP goes through the tunnel.
Mullvad: use the SOCKS5 proxy WITHOUT tunnelling all your traffic
Mullvad's SOCKS5 proxy at 10.64.0.1:1080 only exists while a Mullvad WireGuard tunnel
is up (it is reachable only through the tunnel). The trick is to keep the tunnel up but
tell WireGuard NOT to route your normal traffic through it, only the proxy IP. Add this to
your Mullvad WireGuard .conf ([Interface] section):
Table = off
PostUp = ip -4 route add 10.64.0.1/32 dev %i; ip -4 route add 10.124.0.0/22 dev %i
PreDown = ip -4 route delete 10.64.0.1/32 dev %i; ip -4 route delete 10.124.0.0/22 dev %i
Table = off stops WireGuard from grabbing the default route; the manual routes send ONLY
Mullvad's SOCKS5 proxy IPs through the tunnel (10.124.0.0/22 is the multihop range).
Result: webveil's SOCKS5 requests exit via Mullvad; all other traffic (git, browser, OS)
uses your normal ISP connection. (Simpler alternative: leave WireGuard's routing alone and
rely on layer 1, but split routing is the belt-and-braces version.)
Verify the proxy works: curl https://ipv4.am.i.mullvad.net --socks5-hostname 10.64.0.1
should return a Mullvad exit IP; a plain curl https://am.i.mullvad.net should return your
real IP (proving only the proxy is tunnelled).
"Different exit identity for webveil than for the rest of the machine"
If you want webveil to exit somewhere different from your system, you have options, but be
clear on what is and isn't possible (see
work/notes/findings/mullvad-socks5-egress-mechanics.md):
- Different exit LOCATION, same account (easy). Point webveil at a specific multihop
SOCKS5 host so it exits elsewhere than your tunnel's entry:
WEBVEIL_EGRESS_URL=socks5://us-nyc-wg-socks5-001.relays.mullvad.net:1080. Your tunnel enters where your Mullvad app is connected; webveil's traffic exits in NYC. Same Mullvad account, unlinkable-by-location. - Two DIFFERENT Mullvad ACCOUNTS at once (hard, not a webveil feature). Mullvad's SOCKS5 proxy is a property of the ONE active WireGuard tunnel, which is tied to ONE account's key. SOCKS5 multihop changes exit location, NOT account. To run account A system-wide AND account B for webveil simultaneously, you must isolate them at the OS level: run webveil inside its own network namespace / VM / container that has its own WireGuard tunnel on account B, while the host runs account A. That is infrastructure work outside webveil. For most people, "don't link my searches to my git" is already solved by split routing above (searches exit via Mullvad, git stays on your real IP, not correlated by exit IP), without needing a second account.
Tor
WEBVEIL_EGRESS_URL=socks5://127.0.0.1:9050 with the Tor daemon running. Same per-request,
webveil-only scoping applies.
Other SOCKS5 providers
webveil's socks5 egress is generic: it builds a SOCKS5 dispatcher from any
socks5://host:port URL. Mullvad and Tor are just the documented examples. Anything that
exposes a SOCKS5 endpoint works, e.g. an SSH dynamic forward (ssh -D 1080 user@host, then
socks5://127.0.0.1:1080), or a local shadowsocks/sing-box listener. Use the socks5://
scheme (remote DNS, no leak; webveil does not resolve hostnames locally under proxy
egress). Verify any provider with
curl https://ipv4.am.i.mullvad.net --socks5-hostname <host>:<port>.
ProtonVPN (via wireproxy)
ProtonVPN does not offer a native SOCKS5 proxy (and says it never
will), unlike Mullvad's built-in 10.64.0.1:1080.
There is no socks5:// endpoint Proton hands you. But you can wrap a Proton WireGuard
tunnel in a local SOCKS5 listener and point webveil at that, exactly like the Tor case.
wireproxy is a userspace WireGuard client that
exposes a SOCKS5 port, which suits webveil's design (a local 127.0.0.1 listener,
webveil-only scope, no system-wide tunnel). You do not need the Proton app or CLI
running: wireproxy speaks the WireGuard protocol itself in userspace, so the .conf's
keys + endpoint are all it needs to establish the tunnel (no wg/wg-quick, no network
interface, no root). Proton's dashboard is just where you generate the config once. (One
limit: wireproxy proxies TCP via SOCKS5 CONNECT, which is all webveil needs; it is not a
UDP path.)
- Download a WireGuard config from Proton's account dashboard (Downloads → WireGuard configuration).
- Add a
[Socks5]block and run wireproxy:# proton.conf (from Proton's WireGuard download, plus the [Socks5] block) [Interface] PrivateKey = <from Proton> Address = 10.2.0.2/32 DNS = 10.2.0.1 [Peer] PublicKey = <from Proton> Endpoint = <proton-server>:51820 AllowedIPs = 0.0.0.0/0 [Socks5] BindAddress = 127.0.0.1:1080wireproxy -c proton.conf - Point the SOCKS5 endpoint (
socks5://127.0.0.1:1080) at the hop that reaches the public web (see the warning above):- Remote backend (search via a remote SearXNG) → webveil's BACKEND egress:
export WEBVEIL_EGRESS=socks5 export WEBVEIL_EGRESS_URL=socks5h://127.0.0.1:1080 web_fetch(arbitrary URLs, with ANY backend incl. a LOCAL SearXNG) → webveil's FETCH egress:This proxiesexport WEBVEIL_FETCH_EGRESS=socks5 export WEBVEIL_FETCH_EGRESS_URL=socks5h://127.0.0.1:1080web_fetchwhile the local-SearXNG backend hop staysdirect(see Per-hop egress).- Local SearXNG SEARCH → SearXNG's own outbound, in its
settings.yml(keep the backend hopdirect):This routes SearXNG's engine requests (→ Google/Bing/…) through Proton; webveil's local hop to SearXNG stays direct. (outgoing: proxies: all://: - socks5://127.0.0.1:1080WEBVEIL_EGRESS=socks5with a localbaseUrlis rejected; proxy SEARCH on SearXNG, and proxyweb_fetchviaWEBVEIL_FETCH_EGRESS.)
- Remote backend (search via a remote SearXNG) → webveil's BACKEND egress:
Because wireproxy is userspace, only the traffic you point at it exits via Proton; your
git push, shell, and OS stay on your real IP, with no system tunnel. A ready-made Docker
wrapper that does the same (Proton WireGuard creds in, SOCKS5 on 1080 out) is
SamuelMoraesF/protonvpn-proxy. Verify
the proxy with curl https://ipv4.am.i.mullvad.net --socks5-hostname 127.0.0.1:1080 (a
Proton exit IP means traffic through it exits via Proton); webveil fails loud if a
configured proxy is unbuildable, so it never silently falls back to your real IP.
Caveat: webveil's
socks5mode is NOT a whole-machine VPN. Do not assume enabling it anonymizes anything other than webveil. Conversely, a system-wide full-tunnel VPN under your logged-in identity is the thing that CAN deanonymize agit push; webveil's scoped egress deliberately avoids that.
License
AGPL-3.0-or-later. webveil depends on distilly (MIT, the local HTML-to-markdown
extractor; webveil uses its networked distilly/fetch entrypoint with an injected egress
fetch) and incur (MIT). MIT code may be used by AGPL software; distilly stays
GPL/AGPL-free so it remains cleanly reusable under MIT. See LICENSE and
COPYRIGHT.
Size discipline (per-module LOC)
Every module stays small with one responsibility. Per-module LOC is tracked here as a
first-class quality signal. target is the rough ceiling from CONTEXT.md (a ceiling, not
a promise); LOC is the actual line count of the built file.
packages/webveil (core + CLI/MCP frontend)
| module | LOC | target |
|---|---|---|
| src/index.ts (barrel) | 141 | - |
| src/cli.ts (incur frontend) | 326 | ~80 |
| src/setup.ts (setup commands) | 549 | - |
| src/core/search.ts | 150 | ~90 |
| src/core/fetch.ts | 176 | ~90 |
| src/core/fetch-transport.ts | 295 | - |
| src/core/config.ts | 505 | ~80 |
| src/core/tunables.ts (tuning rules) | 204 | - |
| src/core/layers.ts (merge + prov.) | 129 | - |
| src/core/spellings.ts (old spellings) | 329 | - |
| src/core/trust.ts (exec. settings) | 230 | - |
| src/core/identity.ts | 29 | ~30 |
| src/core/state.ts (per-id. store) | 397 | ~120 |
| src/core/egress.ts | 175 | ~70 |
| src/core/http.ts | 72 | ~60 |
| src/core/extract.ts | 82 | ~60 |
| src/core/security.ts (SSRF guard) | 314 | - |
| src/core/baseurl.ts (transport) | 104 | - |
| src/core/backends/types.ts | 70 | ~40 |
| src/core/backends/registry.ts | 57 | ~60 |
| src/core/backends/searxng.ts | 116 | ~90 |
| src/core/backends/tavily-compat.ts | 156 | ~90 |
| src/core/backends/custom.ts | 201 | ~70 |
| src/core/backends/searchcast.ts | 818 | ~90 |
| subtotal | 5625 |
packages/pi-webveil (pi extension frontend)
| module | LOC | target |
|---|---|---|
| src/index.ts | 246 | ~90 |
Total own source: 5871 LOC (excluding deps).
Reality vs. target: several modules currently exceed their
CONTEXT.mdceilings (notablybackends/searchcast.ts, which also carries the trust, browser and egress-guard policy of that backend,config.ts,tavily-compat.ts,custom.tsandpi-webveil/src/index.ts), and nine built modules (theindex.tsbarrel,setup.ts(the CLI-only setup commands),tunables.ts(every tuning value's rule and default), thesecurity.tsSSRF guard, thebaseurl.tsbackend transport,layers.tsandtrust.tsbehind the trust rule,spellings.ts(the oldserpcastspellings, for one release), andfetch-transport.ts, theweb_fetchtransport choice and searchcast adapter) have no ceiling of their own. The table above reflects the modules as actually built. For calibration, comparable pi web-search extensions:pi-searxng-search350 LOC (1 backend, no egress, no fetch),leing2021/pi-search1714,pi-search-hub9047,pi-web-providers18961. webveil delivers a 4-backend + egress + fetch + per-folder-config tool by leaning onincur(CLI/MCP/skills),distilly(extraction) andsearchcast(the search engines of thesearchcastbackend).
Develop
pnpm install
pnpm build
pnpm test
pnpm format:check
The test workflow (.github/workflows/test.yml) runs the same gate on every push to main and every pull request.
Release
Both packages are released with changesets, each with its own version. A PR that should ship adds a changeset (pnpm changeset, pick the packages and the bump). On main, the release workflow (.github/workflows/release.yml) opens or updates a "Version Packages" PR from the pending changesets; merging it publishes the bumped packages to npm through npm Trusted Publishing (OIDC, no token), with provenance. There is no local publish script: every release goes through that workflow. Both published packages carry this README and the AGPL LICENSE, copied in at pack time by scripts/copy-publish-assets.mjs.