Parse 5
AvailableCohere's enterprise document-intelligence vision-language model, released Aug 27 2026 (parse-v5.0). A compact 2.3B-parameter VLM that converts PDFs, slides, and images into structured Markdown at scale โ detecting and interpreting tables, forms, diagrams, and embedded images, extracting semantic context rather than a flat character stream, and returning bounding boxes for visual elements to support retrieval, grounding, and citation. Trained on business documents across finance, insurance, and scientific research, with support for nine major world languages, an 8,192-token context window, and a ~4.6GB footprint. Priced at $1.50 per 1,000 pages via the Cohere API, with Model Vault for higher-volume/managed deployment and availability on Amazon SageMaker and Microsoft Azure; it can also be deployed privately. Cohere's own comparison places Parse 5 behind larger general-purpose frontier models (GPT-5.5, Opus 4.8, Gemini 3.5 Flash) on raw accuracy, positioning it on price-to-performance and cost per page. Part of Cohere's document-AI line alongside North Micro Vision. Vendor figures, unverified independently at launch.
Specifications
- License
- Proprietary
- Weights
- Not released
- Architecture
- unknown
- Parameters
- 2.3B
- Context window
- 8K tokens
- Max output
- โ
- Knowledge cutoff
- โ
- Price (in / out, $/M)
- โ
- Modalities
- TextVision
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.