@yceachan/pi-vision-helper
pi extension + skill: visual understanding via a configurable vision model (pi-registry reuse or custom responses API; default gpt-5.6-luna) when the main model cannot see images — pure TypeScript, single runtime
Package details
Install @yceachan/pi-vision-helper from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:@yceachan/pi-vision-helper- Package
@yceachan/pi-vision-helper- Version
0.1.1- Published
- Aug 16, 2026
- Downloads
- 480/mo · 19/wk
- Author
- yceachan
- License
- MIT
- Types
- extension, skill
- Size
- 84.2 KB
- Dependencies
- 0 dependencies · 2 peers
Pi manifest JSON
{
"extensions": [
"./extensions"
],
"skills": [
"./skills"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
pi-vision-helper
Visual understanding for pi: when the main model has no vision capability, this package routes image recognition / description / analysis to a configurable vision model via the OpenAI Responses API. Pure TypeScript — single runtime, no python.
- Default (no config): searches the pi-registry — the first vision-capable
lunamodel in~/.pi/agent/models-store.json(currentlygpt-5.6-lunaviaopencode-go), API key from~/.pi/agent/auth.json. - Config-driven:
vision.models[]mixespi-registryentries (fuzzy provider/model match, pi auth) andresponsesentries (custom baseUrl +$ENV_VAR/literal apiKey);vision.activepicks the model,enabled/forceVisionBridge/maxTokens/timeoutMs/systemPrompttune behavior.
中文文档:README.zh.md
Ships as a full pi package bundling:
pi-vision-helper/
├── package.json # pi manifest: extensions + skills
├── lib/ # core (shared by tool + CLI): config, registry, vision
│ ├── config.ts # config resolution + schema normalization
│ ├── registry.ts # models-store / auth.json access + fuzzy matching
│ └── vision.ts # target resolution + Responses API round trip
├── extensions/
│ └── pi-vision-helper.ts # registers the `pi-vision-helper` custom tool (in-process)
├── skills/
│ └── pi-vision-helper/
│ ├── SKILL.md # trigger conditions, config schema, troubleshooting
│ └── scripts/
│ └── vision.ts # CLI for manual debugging (bun), same lib/
├── README.md # this file (en)
└── README.zh.md # 中文文档
Install
pi install npm:@yceachan/pi-vision-helper
Or as a local path in ~/.pi/agent/settings.json:
"~/work/pi-agent-harness/extensions/mono/packages/pi-vision-helper"
After changes, run /reload in pi (or restart pi) to hot-reload.
Usage
Prefer the pi-vision-helper tool (schema requires images and prompt):
pi-vision-helper images=[path1, path2] prompt="Describe this image in detail" effort=high max_tokens=4096
images: image paths; Windows (C:\...) and WSL (/mnt/...) paths supported; at least one image (the tool schema rejects an empty array — a pure-text request would still be billed)prompt: required; must be explicitly constructed from the user's intent (describe / transcribe / compare, etc.), never omittedeffort: thinking depth (default from configdefaultEffort, else legacydefaults.effort, elsehigh;mediumis prone to speculative hallucinations;xhigh/maxneedmax_tokensraised to 8000+;off/minimalomit the reasoning field — the zen gateway rejects them verbatim with HTTP 400)max_tokens: output budget including reasoning tokens (default from configmaxTokens, else 4096; enforced range 256–32768 — config and CLI values are validated)
Manual CLI (same core, for debugging):
bun packages/pi-vision-helper/skills/pi-vision-helper/scripts/vision.ts img.png \
--prompt "describe" --model luna-customer --effort high --max-tokens 4096
Configuration
Config files, first found wins:
| Priority | Path | Scope |
|---|---|---|
| 1 | --config <path> (CLI flag) |
explicit, missing file = hard error |
| 2 | $PI_VISION_HELPER_CONFIG |
explicit env override, missing file = hard error |
| 3 | $CWD/.pi/vision-helper.json |
project-level (committable, shared per repo) |
| 4 | ~/.pi/agent/pi-vision-helper.json |
user-level (global default) |
No config file at all = default behavior: search the pi-registry and use the first matching
luna model (with image in its input list) found in the first provider that has one.
Full config example
{
// ── Global switches ────────────────────────────────────────────────────
"enabled": true, // master switch; false = the tool/CLI refuses to run
"forceVisionBridge": false, // true = delegation allowed even when the main model
// is itself a VLM (default: only for blind models)
"defaultEffort": "high", // default reasoning effort; the tool/CLI effort
// parameter overrides it (see the effort field table)
"maxTokens": 4096, // default max output tokens incl. reasoning;
// the tool/CLI parameter overrides it
"timeoutMs": 60000, // default timeout per vision call (ms);
// the pi tool falls back to 300000 when unset
"systemPrompt": "", // custom system prompt for the vision model;
// empty = built-in default (verbatim transcription etc.)
// ── Models & APIs ──────────────────────────────────────────────────────
"vision": {
"active": "luna", // active model name; absent = first entry in models[];
// models[] empty = legacy flat fields / default luna search
"models": [
{
// ① pi-registry entry — reuse pi's model catalog & pi auth
"name": "luna", // unique name (used by vision.active / CLI --model)
"type": "pi-registry", // "pi-registry" | "responses"
"Provider": "opencode-go", // fuzzy-match a provider in models-store.json
// (case-insensitive; no match = error listing providers)
"Model": "gpt-5.6-luna", // fuzzy-match a model under that provider; prefers
// entries with "image" in input; absent = first
// vision-capable (luna-first) model
// optional overrides (pi-registry):
// "cost": { "input": 0.1, "output": 0.6,
// "cacheRead": 0.01, "cacheWrite": 0.125 }, // USD per M tokens;
// // default = store entry cost
// "headers": { "X-Foo": "bar" } // extra request headers
},
{
// ② responses entry — any OpenAI-Responses-compatible endpoint
"name": "luna-customer",
"type": "responses",
"baseUrl": "https://opencode.ai/zen/go/v1",
// REQUIRED: endpoint root ("/responses" is appended)
"apiKey": "$VISION_API_KEY",// REQUIRED: "$ENV_VAR" reference or a literal key
"model": "gpt-5.6-luna", // REQUIRED: literal model id (no registry lookup)
// optional (responses):
// "cost": { "input": 0.1, "output": 0.6 }, // default = 0 (uncounted)
// "headers": { "X-Foo": "bar" } // extra request headers
}
]
}
}
Field reference
Top level
| Field | Type | Default | Meaning |
|---|---|---|---|
enabled |
bool | true |
Master switch. false = the tool returns a "disabled" text and the CLI exits with an error. |
forceVisionBridge |
bool | false |
Tool-only gate: allow delegation even when the main model is itself a VLM. Without it the tool refuses to run when the active main model's input includes "image" (the CLI always delegates — it has no main model). |
defaultEffort |
string | high |
Default reasoning effort for the tool/CLI when none is passed. Must be one of off/low/medium/high/xhigh/max (invalid values = config error). |
maxTokens |
number | 4096 | Default max_output_tokens (incl. reasoning) when the tool/CLI passes none. Enforced range 256–32768 (out-of-range = config error). |
timeoutMs |
number | 60000 | Default per-call HTTP timeout in milliseconds (the pi tool uses 300000 when unset). |
systemPrompt |
string | "" |
Replaces the built-in vision-assistant instructions; "" keeps the default. |
vision |
object | — | Model selection block (below). Absent = legacy flat fields / default luna search. |
vision
| Field | Type | Default | Meaning |
|---|---|---|---|
active |
string | first entry | Name of the active model entry. Unknown name = error listing the available names (no silent fallback). |
models |
array | [] |
Model entries in priority order. Empty = legacy flat fields / default luna search. |
Model entry
Fields depend on type:
| Field | pi-registry |
responses |
Meaning |
|---|---|---|---|
name |
✓ | ✓ | Unique entry name (active / CLI --model). Defaults to <Provider>/<Model> or <model>. |
type |
✓ | ✓ | "pi-registry" = reuse pi's catalog + auth; "responses" = custom endpoint. |
Provider |
✓ | — | Fuzzy provider match in models-store.json (case-insensitive; provider also accepted). No match = error listing available providers. |
Model |
✓ | — | Fuzzy model match under that provider, preferring vision-capable entries (model also accepted). Absent = first vision-capable (luna-first) model. |
baseUrl |
— | ✓ | Endpoint root; /responses is appended. |
apiKey |
— | ✓ | $ENV_VAR reference (expanded at call time; unset env = error) or a literal key. |
model |
— | ✓ | Literal model id — no registry lookup. |
cost |
✓ | ✓ | USD per M tokens {input, output, cacheRead, cacheWrite}. pi-registry: overrides the store entry cost (default = store cost). responses: default = 0 (uncounted). |
headers |
✓ | ✓ | Extra HTTP headers merged into the request (default {}). |
Fuzzy matching
Deterministic order: exact (case-insensitive) > candidate contains needle > needle contains candidate > no match = error (never silently picks the first candidate — that would route unknown providers to the wrong backend and bill it).
Legacy flat fields (still supported)
Without a vision block, the old single-model fields keep working:
{
"provider": "opencode-go", // registry provider (default: first provider)
"model": "gpt-5.6-luna", // registry model id (fuzzy; default: luna search)
"modelMatch": "exact", // "exact" | "substring"
"baseUrl": "https://...", // override / custom endpoint
"apiKey": "$VISION_API_KEY", // key override ($ENV_VAR or literal)
"apiKeyEnv": "VISION_API_KEY", // alternative: env var name holding the key
"cost": { "input": 0.1, "output": 0.6 },
"headers": { "X-Foo": "bar" },
"defaults": { "effort": "high", "maxTokens": 4096 }
}
Resolution precedence
Tool/CLI parameters > config file > pi-registry (models-store.json + auth.json) > built-in
defaults (luna / high / 4096). Effort specifically: tool/CLI effort > defaultEffort >
legacy defaults.effort > high. For keys: entry apiKey > entry env var >
auth.json[provider].key.
How it works
- Request: OpenAI Responses API —
POST {baseUrl}/responseswithinput_text(your prompt) and oneinput_imageper image (data:<mime>;base64,...). The MIME type is derived from the file extension (png/jpg/jpeg/jfif/gif/webp/bmp/heic/avif); unknown extensions fail loudly instead of being mislabeled. The base64 bytes exist only in memory and in that HTTPS request — they never enter the main model's context or the session history. A model entry'sapifield in models-store.json is advisory only: the zen gateway serves/responsesfor models declared as openai-completions too (verified with kimi-k2.7-code). - Effort:
off/minimalomit thereasoningfield entirely (the zen gateway maps them to null; sending them verbatim produces HTTP 400invalid_prompt). - Usage & cost: the tool reports tokens and USD cost through pi's Usage accounting —
registry entry cost by default, per-entry
costoverride, 0 when unknown. - Errors: config/registry problems fail loudly with actionable messages (available providers/models listed, env-var names, which config file was read) — no silent fallbacks.