Vision Narrator 0.8B (4-bit MLX)

A vision-language model small enough to live on a phone, trained to answer the questions blind people actually ask about what's in front of them.

This is the model inside RealTime AI Cam, a free camera app for blind and low-vision people from Nice Dreamz. You point the camera at a room, a letter, or the coffee maker, ask it a question, and it answers out loud. Everything runs offline, on the phone.

Get the app: App Store (iPhone) · Google Play (Android) · Source on GitHub


New in v09: it answers the question first

Earlier versions described the whole photo no matter what you asked. Ask "is this light on?" and you'd get a paragraph about the lamp, the table and the window blinds, and never hear the answer.

v09 was retrained on questions real blind people asked about photos they took themselves, from the VizWiz dataset. Now the answer comes first:

Question v08 (before) v09 (now)
Is this light on? "A lamp is on a small wooden table in front of the window, which has closed blinds above it..." "Yes, the lamp is on, casting a warm glow across the room."
Is this cup full? "This is a clear Starbucks Coffee tumbler with a green logo on it, and it is sitting on a white surface..." "No, it's empty, with just a green straw sticking straight down into the clear plastic."
Does the screen say anything now? "The photo is too blurry and dark to read the screen..." "No, the screen is dark and blank, with just a faint reflection of light on the glass."
Is there a little light on under the Add Water, or is that off? "The photo shows a machine's control panel with three round buttons..." "No, there is no light on under the add water; only a green power light and blue icons are visible."

Those are unedited outputs. The first three are VizWiz photos the model never saw in training. The last one is a photo of our own Keurig.

Scores

Tested on 295 VizWiz questions whose photos were kept out of training, scored with the standard VizWiz accuracy (agreement with 10 human answers):

v08 v09
Overall 46.1% 66.4%
Yes/no questions 10.9% 62.5%
Number questions 47.3% 61.3%
Other questions 50.2% 60.7%
Correctly says "I can't tell" on unanswerable photos 78.0% 92.0%
Puts the answer first 34.6% 65.8%

The size didn't change (about 620 MB). Reading mail and describing rooms held steady: 192/200 facts right on the describe test, vs 194/200 before.

Still weak: elevator direction. Asked "is the elevator going up or down?", it usually says it can't tell. Ask "which arrow is lit?" instead and it gets most of them right. Direction indicators are what the next round adds.


What makes it different

It's trained to be useful, not just descriptive. A generic captioner tells a blind user "a piece of paper with writing on it." This model answers the four things that matter on a letter: what it is, who sent it, how much, and what happens if you ignore it. On a scene, it says what's in front of you and where.

The part that sees isn't compressed. The language half runs at 4-bit (affine, group size 64). The vision tower stays at full precision, because compressing it made the model misread dollar amounts. On this model, 8-bit vision dropped accuracy to 96.2% and 4-bit vision dropped it to 91.5%.

It says when it can't tell. It's trained to say something like "I can't read this page, the photo is too dark. Add some light or turn on the flash and try again" instead of inventing a plausible bill. For this audience that's a feature.

It runs on a 2022 phone. Tested on an iPhone 14 Pro Max, fully offline.

Real outputs (describe mode)

Scene: "A white mattress with a star pattern is right in front of you, so close it blocks most of the room. A person in purple pants is lying on a patterned rug to your left, with a bright orange heater sitting on the floor just ahead of them."

Bill: "This is a utility bill for $81.13, due October 8th, 2026. If it goes unpaid they can shut off service 15 days after that."

Bad photo: "The top of the page is cut off, so I can't tell who sent this. Move the phone back a little so the whole page is in view and try again."


Use it

Two system prompts, one per mode. For a question about a photo (what v09 was trained for):

from mlx_vlm import load, generate
from mlx_vlm.prompt_utils import apply_chat_template

ASK = ("You are the eyes of a blind person. They are asking a question about this photo. "
       "Answer it in one short spoken sentence, and start with the answer itself: yes or no, "
       "a number, a colour, a name, or the words that are printed. Then, only if it helps, a few "
       "words about where or why. If the photo does not show the answer, say you can't tell for "
       "sure and say what you can see. Plain words, no lists, no markdown.")

model, processor = load("divinetribe/Vision-Narrator-0.8B-4bit-mlx")
messages = [{"role": "system", "content": ASK},
            {"role": "user", "content": "Is this light on?"}]
prompt = apply_chat_template(processor, model.config, messages, num_images=1)
print(generate(model, processor, prompt, image="photo.jpg", max_tokens=96).text)

For "what's in front of me" or reading mail, use the describe prompt from the app source. On iPhone it runs through MLX Swift. The Android app runs a GGUF conversion of these weights through llama.cpp.

Architecture

Base Qwen3.5 0.8B (Qwen3_5ForConditionalGeneration)
Language model 1024 hidden, 24 layers, 8 heads, 248,320 vocab
Vision tower 768 hidden, depth 12, projects to 1024
Quantization 4-bit affine, group size 64, language model only; vision tower full precision
On disk about 620 MB
Runtime MLX / MLX Swift

Training

LoRA fine-tune of Qwen3.5 0.8B, rank 16, on all 24 language-model layers, with the vision tower frozen, then fused and quantized. The model learned what to say about what it already sees.

  • Describe mode (v08 and earlier): rendered and photographed household documents (utility bills, past-due notices, statements), plus scene descriptions of VizWiz-Captions photos written by a larger teacher model.
  • Ask mode (v09): 3,552 question and answer pairs built from VizWiz-VQA questions, answers written in answer-first style by a Qwen 27B teacher, plus 2,500 describe-mode examples replayed so the old skills weren't lost. Continued from v08, 2,300 steps.

The previous version is still available under the v08 tag of this repo.

Limits

English only. Tuned for US household documents, indoor scenes and everyday objects. It will sometimes misread a brand name or a sender on a badly lit page. It's a 0.8B model, not a human reader. Don't use it as the only source of truth for anything financial or medical. It's a fast first look for someone who would otherwise get nothing.

Credits

VizWiz. The v09 training questions and the scene photos come from the VizWiz datasets: photos taken and questions asked by blind people, collected by Danna Gurari and colleagues at the University of Colorado Boulder, licensed CC BY 4.0. Thank you for making this possible.

  • Gurari, Li, Stangl, Guo, Lin, Grauman, Luo, Bigham. VizWiz Grand Challenge: Answering Visual Questions from Blind People. CVPR 2018. arXiv:1802.08218
  • Gurari, Zhao, Zhang, Bhattacharya. Captioning Images Taken by People Who Are Blind. ECCV 2020. arXiv:2002.08565

Qwen. Built on Qwen3.5 0.8B by the Qwen team at Alibaba (Apache 2.0).

MLX. Quantized and served with MLX and mlx-vlm.

Trained by Matt Macosko and released through Nice Dreamz as divinetribe.

Downloads last month
96
Safetensors
Model size
0.9B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for divinetribe/Vision-Narrator-0.8B-4bit-mlx

Finetuned
(423)
this model

Papers for divinetribe/Vision-Narrator-0.8B-4bit-mlx