Instructions to use divinetribe/Vision-Narrator-0.8B-4bit-mlx with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use divinetribe/Vision-Narrator-0.8B-4bit-mlx with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("divinetribe/Vision-Narrator-0.8B-4bit-mlx") config = load_config("divinetribe/Vision-Narrator-0.8B-4bit-mlx") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use divinetribe/Vision-Narrator-0.8B-4bit-mlx with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Vision-Narrator-0.8B-4bit-mlx"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "divinetribe/Vision-Narrator-0.8B-4bit-mlx" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent
How to use divinetribe/Vision-Narrator-0.8B-4bit-mlx with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Vision-Narrator-0.8B-4bit-mlx"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default divinetribe/Vision-Narrator-0.8B-4bit-mlx
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use divinetribe/Vision-Narrator-0.8B-4bit-mlx with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "divinetribe/Vision-Narrator-0.8B-4bit-mlx"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "divinetribe/Vision-Narrator-0.8B-4bit-mlx" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Vision Narrator 0.8B (4-bit MLX)
A vision-language model small enough to live on a phone, trained to answer the questions blind people actually ask about what's in front of them.
This is the model inside RealTime AI Cam, a free camera app for blind and low-vision people from Nice Dreamz. You point the camera at a room, a letter, or the coffee maker, ask it a question, and it answers out loud. Everything runs offline, on the phone.
Get the app: App Store (iPhone) · Google Play (Android) · Source on GitHub
New in v09: it answers the question first
Earlier versions described the whole photo no matter what you asked. Ask "is this light on?" and you'd get a paragraph about the lamp, the table and the window blinds, and never hear the answer.
v09 was retrained on questions real blind people asked about photos they took themselves, from the VizWiz dataset. Now the answer comes first:
| Question | v08 (before) | v09 (now) |
|---|---|---|
| Is this light on? | "A lamp is on a small wooden table in front of the window, which has closed blinds above it..." | "Yes, the lamp is on, casting a warm glow across the room." |
| Is this cup full? | "This is a clear Starbucks Coffee tumbler with a green logo on it, and it is sitting on a white surface..." | "No, it's empty, with just a green straw sticking straight down into the clear plastic." |
| Does the screen say anything now? | "The photo is too blurry and dark to read the screen..." | "No, the screen is dark and blank, with just a faint reflection of light on the glass." |
| Is there a little light on under the Add Water, or is that off? | "The photo shows a machine's control panel with three round buttons..." | "No, there is no light on under the add water; only a green power light and blue icons are visible." |
Those are unedited outputs. The first three are VizWiz photos the model never saw in training. The last one is a photo of our own Keurig.
Scores
Tested on 295 VizWiz questions whose photos were kept out of training, scored with the standard VizWiz accuracy (agreement with 10 human answers):
| v08 | v09 | |
|---|---|---|
| Overall | 46.1% | 66.4% |
| Yes/no questions | 10.9% | 62.5% |
| Number questions | 47.3% | 61.3% |
| Other questions | 50.2% | 60.7% |
| Correctly says "I can't tell" on unanswerable photos | 78.0% | 92.0% |
| Puts the answer first | 34.6% | 65.8% |
The size didn't change (about 620 MB). Reading mail and describing rooms held steady: 192/200 facts right on the describe test, vs 194/200 before.
Still weak: elevator direction. Asked "is the elevator going up or down?", it usually says it can't tell. Ask "which arrow is lit?" instead and it gets most of them right. Direction indicators are what the next round adds.
What makes it different
It's trained to be useful, not just descriptive. A generic captioner tells a blind user "a piece of paper with writing on it." This model answers the four things that matter on a letter: what it is, who sent it, how much, and what happens if you ignore it. On a scene, it says what's in front of you and where.
The part that sees isn't compressed. The language half runs at 4-bit (affine, group size 64). The vision tower stays at full precision, because compressing it made the model misread dollar amounts. On this model, 8-bit vision dropped accuracy to 96.2% and 4-bit vision dropped it to 91.5%.
It says when it can't tell. It's trained to say something like "I can't read this page, the photo is too dark. Add some light or turn on the flash and try again" instead of inventing a plausible bill. For this audience that's a feature.
It runs on a 2022 phone. Tested on an iPhone 14 Pro Max, fully offline.
Real outputs (describe mode)
Scene: "A white mattress with a star pattern is right in front of you, so close it blocks most of the room. A person in purple pants is lying on a patterned rug to your left, with a bright orange heater sitting on the floor just ahead of them."
Bill: "This is a utility bill for $81.13, due October 8th, 2026. If it goes unpaid they can shut off service 15 days after that."
Bad photo: "The top of the page is cut off, so I can't tell who sent this. Move the phone back a little so the whole page is in view and try again."
Use it
Two system prompts, one per mode. For a question about a photo (what v09 was trained for):
from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template
ASK = ("You are the eyes of a blind person. They are asking a question about this photo. "
"Answer it in one short spoken sentence, and start with the answer itself: yes or no, "
"a number, a colour, a name, or the words that are printed. Then, only if it helps, a few "
"words about where or why. If the photo does not show the answer, say you can't tell for "
"sure and say what you can see. Plain words, no lists, no markdown.")
model, processor = load("divinetribe/Vision-Narrator-0.8B-4bit-mlx")
messages = [{"role": "system", "content": ASK},
{"role": "user", "content": "Is this light on?"}]
prompt = apply_chat_template(processor, model.config, messages, num_images=1)
print(generate(model, processor, prompt, image="photo.jpg", max_tokens=96).text)
For "what's in front of me" or reading mail, use the describe prompt from the app source. On iPhone it runs through MLX Swift. The Android app runs a GGUF conversion of these weights through llama.cpp.
Architecture
| Base | Qwen3.5 0.8B (Qwen3_5ForConditionalGeneration) |
| Language model | 1024 hidden, 24 layers, 8 heads, 248,320 vocab |
| Vision tower | 768 hidden, depth 12, projects to 1024 |
| Quantization | 4-bit affine, group size 64, language model only; vision tower full precision |
| On disk | about 620 MB |
| Runtime | MLX / MLX Swift |
Training
LoRA fine-tune of Qwen3.5 0.8B, rank 16, on all 24 language-model layers, with the vision tower frozen, then fused and quantized. The model learned what to say about what it already sees.
- Describe mode (v08 and earlier): rendered and photographed household documents (utility bills, past-due notices, statements), plus scene descriptions of VizWiz-Captions photos written by a larger teacher model.
- Ask mode (v09): 3,552 question and answer pairs built from VizWiz-VQA questions, answers written in answer-first style by a Qwen 27B teacher, plus 2,500 describe-mode examples replayed so the old skills weren't lost. Continued from v08, 2,300 steps.
The previous version is still available under the v08 tag of this repo.
Limits
English only. Tuned for US household documents, indoor scenes and everyday objects. It will sometimes misread a brand name or a sender on a badly lit page. It's a 0.8B model, not a human reader. Don't use it as the only source of truth for anything financial or medical. It's a fast first look for someone who would otherwise get nothing.
Credits
VizWiz. The v09 training questions and the scene photos come from the VizWiz datasets: photos taken and questions asked by blind people, collected by Danna Gurari and colleagues at the University of Colorado Boulder, licensed CC BY 4.0. Thank you for making this possible.
- Gurari, Li, Stangl, Guo, Lin, Grauman, Luo, Bigham. VizWiz Grand Challenge: Answering Visual Questions from Blind People. CVPR 2018. arXiv:1802.08218
- Gurari, Zhao, Zhang, Bhattacharya. Captioning Images Taken by People Who Are Blind. ECCV 2020. arXiv:2002.08565
Qwen. Built on Qwen3.5 0.8B by the Qwen team at Alibaba (Apache 2.0).
MLX. Quantized and served with MLX and mlx-vlm.
Trained by Matt Macosko and released through Nice Dreamz as divinetribe.
- Downloads last month
- 96
4-bit