3-bit VLM base + 4-bit MTP drafter. Pair both for accelerated instruct decode. reasoning_effort baked to low.
Jorge Leon
leonsarmiento
AI & ML interests
AI for Environment
Recent Activity
updated a model about 9 hours ago
leonsarmiento/Orion-26B-A4B-v1-6bit-XL-mlx published a model about 9 hours ago
leonsarmiento/Orion-26B-A4B-v1-6bit-XL-mlx updated a model about 9 hours ago
leonsarmiento/Tiger-Gemma-12B-v3-4bit-mlx