North Micro Vision Instruct
AvailableCohere's compact document-focused vision-language model, published Aug 12 2026 under Apache 2.0. A 2.4B-parameter VLM combining a custom 400M native-resolution vision encoder, a 2B language model on the Command A+ architecture, and a projector; it preserves aspect ratio for images up to 1654x2339px (an A4 page at 200 dpi). Multilingual visual understanding across documents, charts, and natural images, text output. Vendor-reported: 0.921 DocVQA and 0.808 ChartQA on document tasks, 0.732 RefCOCO on visual grounding, and 0.687 MMBench on general VQA; text-only capability lags larger models. Part of Cohere's North product family alongside North Mini Code. Open weights on Hugging Face.
Specifications
- License
- Open source ยท Apache-2.0
- Weights
- Downloadable
- Architecture
- unknown
- Parameters
- 2.4B
- Context window
- โ tokens
- Max output
- โ
- Knowledge cutoff
- โ
- Price (in / out, $/M)
- โ
- Modalities
- TextVision
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.