pi-multivision
Give text-only pi models vision — a native vision tool with a bundled vision script and a multi-backend model chain (auto-fallback, timeout, retry). Configure models via .env, JSON, or a custom script.
Package details
Install pi-multivision from npm and Pi will load the resources declared by the package manifest.
$ pi install npm:pi-multivision- Package
pi-multivision- Version
0.3.0- Published
- Aug 19, 2026
- Downloads
- 1,140/mo · 292/wk
- Author
- likeattract
- License
- MIT
- Types
- extension
- Size
- 27 KB
- Dependencies
- 0 dependencies · 1 peer
Pi manifest JSON
{
"extensions": [
"./multivision.ts"
]
}Security note
Pi packages can execute code and influence agent behavior. Review the source before installing third-party packages.
README
👁️ pi-multivision
Give text-only pi models vision
One native tool call — a multi-backend vision model chain with automatic fallback.
The Problem
Some of the best coding models are blind. You paste a screenshot, a UI mock, or a diagram into pi — and a text-only model (DeepSeek, etc.) simply cannot see it. The read tool omits the image, and describing it yourself is a chore.
The Solution
pi-multivision registers a native tool (multivision) that any text-only model can call directly — no remembering to use a skill, no shelling out manually. The extension hands the image to a vision-capable model and returns the text description as the tool result.
- 🔌 Multi-backend with auto-fallback — providers are tried in config order (e.g. Step-3.7-Flash → GLM-4.6V-Flash → Qwen-Chat). If a backend is rate-limited, times out, or returns an empty response, the next one is tried automatically.
- ⏱️ Timeout & retry — per-request timeout (240s default, configurable), rate-limit backoff, and empty-response fallback. Slow models fail fast with a clear message instead of hanging.
- 🖼️ Single or multiple images —
imagePathfor one,imagePathsfor comparison across several. - 🧠 Model-driven — the tool is described in the model's tool list, so the agent picks it automatically whenever it needs to "see" an image.
- 📦 Bundled vision script — the package ships a ready-to-use
vision.js(OpenAI-compatible, multi-provider fallback). No custom script needed.
Installation
From npm (recommended):
pi install npm:pi-multivision
From GitHub:
pi install git:github.com/like-attract/pi-multivision
Then /reload (or restart pi).
Configuration(仅 .env)
First use requires at least one vision model. The extension ships with a bundled script but no default API keys.
1. Create ~/.pi/agent/pi-multivision.env(或设置 VISION_ENV 指向其他路径)
VISION_MODEL_1_NAME=glm
VISION_MODEL_1_URL=https://open.bigmodel.cn/api/paas/v4
VISION_MODEL_1_MODEL=glm-4.6v-flash
VISION_MODEL_1_KEY=your_api_key_here
- Up to 10 models:
VISION_MODEL_1_*,VISION_MODEL_2_*, … — the number is the fallback order (failed providers are skipped automatically). - Optional
VISION_TIMEOUT=<seconds>(default 240) — per-request timeout. - A template with common providers is shipped as
pi-multivision.env.exampleinside the package.
2. Call it
That's it — the extension auto-uses the bundled vision.js, so no JSON config, no custom script, no /reload after setup.
If no .env is found, the tool returns a step-by-step setup guide instead of guessing.
Free/low-cost vision models that work out of the box:
- GLM-4.6V-Flash —
https://open.bigmodel.cn/api/paas/v4 - Step-3.7-Flash —
https://api-inference.modelscope.cn/v1(model idstepfun-ai/Step-3.7-Flash)
Usage
Just ask. The model calls the tool itself:
> 看一下 login-page.png 里有什么文字
> 对比这两张截图有什么区别 (pass imagePaths)
Or force it in a session with any image path:
multivision(imagePath="screenshot.png", prompt="识别图中所有文字")
Progress display
While the analysis is running, the extension shows live progress via the extension UI protocol (ctx.ui.setWidget / ctx.ui.setStatus): a widget panel and a status-bar entry in both pi TUI and pi-web:
◆ 视觉分析中
图片: screenshot.png
模型链请求中(超时 240s,失败自动切换)
The widget/status is cleared automatically when the tool finishes.
Agent Usage Guidelines(模型行为准则)
The tool registers a promptSnippet, so it appears in the system prompt's Available tools list automatically — models see it without the user naming it explicitly.
Text-only models must call multivision immediately, without waiting for the user to ask, when any of these triggers occur:
- Image omitted from context — the message shows
image omitted: model does not support images,(image omitted),[图片已省略], or similar markers. This means the image was stripped because the model cannot see it natively. - User supplied an image — an image path, an image URL, or phrasing like "look at this screenshot / 看图 / 识图 / OCR / 看这个界面".
- A "seeing" task is involved — UI screenshots, error dialogs, flowcharts, diagrams, scanned documents, multi-image comparison.
- User asks "what is this?" while an image attachment is present in the conversation.
If the image path is unknown: search common locations first (%TEMP%, session dir, cwd) for png/jpg/webp files, then call the tool. If still not found, ask the user for the path rather than skipping the analysis.
This guideline is also mirrored in the
visionskill (SKILL.md) which is the canonical reference.
Why not pi-vision-handoff?
pi-vision-handoff intercepts image blocks at the context event and swaps them for descriptions automatically — great if you want zero-visible-tool behavior. pi-multivision takes the opposite approach: an explicit native tool with a hardened multi-backend chain (timeout, retry, empty-response fallback) that the model calls on demand. Trade-off: you see the tool call in the transcript (auditable), and you keep full control over which backend is used.
License
MIT