Skip to content

Latest commit

 

History

History
71 lines (58 loc) · 3.59 KB

File metadata and controls

71 lines (58 loc) · 3.59 KB
title Standard Vault Pricing
slug docs/model-vault/standard/pricing
hidden false
description Standard Vault pricing models (Fixed and Flex) and per-model performance tiers and rates.
keywords standard vault, model vault, pricing, performance tiers, fixed, flex, cohere

Standard Vault is billed as a Cohere-managed service. Pricing depends on the models you select and each model's performance tier. Cohere manages the underlying infrastructure and scaling, and customers can choose between two pricing models:

Feature Fixed Flex
Commitment Monthly or annual Monthly or annual
Capacity Fixed number of instances (no autoscaling) Minimum baseline instances, plus autoscaling
Sizing Determined through a sizing exercise or a production trial (for example, based on expected load) --
Autoscaling -- Scales up/down based on request rate and agreed latency SLOs
Pause/resume You can pause a model or restart a paused model to save on costs. You can pause a model or restart a paused model to save on costs.
Overages -- Additional capacity billed per instance-hour
Max capacity -- Maximum instance cap per model

The following table summarizes the available models and their rates. All rates are per instance.

Model Performance Tier Hourly rate Monthly rate Annual rate
Embed 4 Small $4.00 $2,500 $25,000
Embed 4 Medium $5.00 $3,250 $32,500
Rerank 3.5 Medium $5.00 $3,250 $32,500
Rerank 4 Fast Medium $5.00 $3,250 $32,500
Rerank 4 Pro Medium $5.00 $3,250 $32,500
Rerank 4 Pro Large $10.00 $6,500 $65,000
Parse 5 Medium $4.00 $2,500 $25,000
Parse 5 XL $7.00 $4,300 $43,000

You may also want to compare Standard Vault pricing against the operational and capacity costs of running inference directly in your cloud provider account (for example, AWS), where cloud-provider credits may apply.

Generative models

The rates above cover Embed and Rerank, which are available self-serve. Generative models (the Command A family and North) are also available on Standard Vault, typically through a waitlist. All rates are per instance, per hour.

Model Hourly rate XL hourly rate
Command A $40.00 $48.00
Command A Vision $40.00 $48.00
Command A Translate $40.00 $48.00
Command A Reasoning $48.00 $57.50
Command A+ -- $57.50
North Mini Code -- $57.50

Monthly and annual commitment pricing follows the same Fixed and Flex plans described above; contact Cohere for those rates and for access. Generative access typically requires a waitlist, so check the model dropdown when creating a vault, or contact Cohere. Bundles and customized models are also available; see Supported Models for the full list.

For the pricing of the encrypted, confidential-computing product, see Model Vault Encrypted Pricing.

Performance Tiers

Each model has a performance tier based on latency requirements and throughput service level objectives (SLOs) per instance. You can see the tiers listed in the model dropdown selection as a size letter (e.g., S, M, L). The tiers follow an instance/hour pricing, which is then incorporated into your payment plan. We recommend selecting the model-performance tier combination that matches the nature of your workflow and the required measures of performance.