LLM Releases
โ† Catalog

North Micro Vision Instruct

Available
CohereOpen source

Cohere's compact document-focused vision-language model, published Aug 12 2026 under Apache 2.0. A 2.4B-parameter VLM combining a custom 400M native-resolution vision encoder, a 2B language model on the Command A+ architecture, and a projector; it preserves aspect ratio for images up to 1654x2339px (an A4 page at 200 dpi). Multilingual visual understanding across documents, charts, and natural images, text output. Vendor-reported: 0.921 DocVQA and 0.808 ChartQA on document tasks, 0.732 RefCOCO on visual grounding, and 0.687 MMBench on general VQA; text-only capability lags larger models. Part of Cohere's North product family alongside North Mini Code. Open weights on Hugging Face.

Specifications

License
Open source ยท Apache-2.0
Weights
Downloadable
Architecture
unknown
Parameters
2.4B
Context window
โ€” tokens
Max output
โ€”
Knowledge cutoff
โ€”
Price (in / out, $/M)
โ€”
Modalities
TextVision

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.