Skip to content
DMarketer Tayeeb – Digital Marketing Expert in Bangalore | SEO, SEM & SMM Expert
Contact

Cohere Embed 5 Pro vs Fast: Code, Costs and Limits

Quick choice: Use Embed 5 Pro when retrieval quality matters most for complex documents. If you index a corpus once and query it often, test Pro for documents and Fast for live queries. Cohere reports a normalized mean nDCG@10 score of 98.4 for this pairing, compared with 100 for Pro on both sides across 40 development datasets. Cohere says the two Embed 5 models share an embedding space, so Pro vectors can be queried with Fast when both sides use the same output dimension. Cohere’s launch results and compatibility note.

Which model belongs on each side of retrieval?

Cohere positions Pro for quality-first indexing and retrieval over complex, financial, multilingual, or visually rich material. Fast targets latency-sensitive and cost-sensitive workloads such as interactive search and high-volume query traffic. Those are product positions, not a guarantee that one wins on your corpus. See Cohere’s workload guidance.

Two model IDs in the Embed 5 shared space. Cohere launch details.
ChoiceAPI model IDStarting pointCohere API text rate
Embed 5 Proembed-v5.0-proOffline indexing; quality-critical or complex-document retrieval$0.12 per 1M input tokens
Embed 5 Fastembed-v5.0-fastLive queries; latency- or volume-sensitive retrieval$0.08 per 1M input tokens

The text-rate difference is $4 for every 100 million tokens embedded: $12 with Pro or $8 with Fast. This is a token calculation, not a per-request charge. It excludes image input, vector storage, your database, reranking, and generation. Cohere lists image input at $0.40 per million image tokens for both models, so the lower Fast text rate does not change the published image-token rate. See Cohere’s price table.

The shared space separates indexing from query choice

Cohere says Pro and Fast vectors can be compared directly, so you can choose the corpus encoder and query encoder independently. Its evaluation covers all four combinations below. The score is a mean nDCG@10 normalized to Pro corpus plus Pro query = 100. It is not an absolute score or the share of searches your system will get right.

Cohere-reported mean nDCG at 10 across 40 development datasets, normalized so Pro corpus plus Pro query equals 100. Scores: Pro/Pro 100.0; Pro/Fast 98.4; Fast/Pro 97.3; Fast/Fast 96.6.
Cohere-reported mean nDCG@10 across 40 development datasets spanning text, image, fused, and parsed-document retrieval. Scores are normalized to Pro corpus + Pro query = 100. These are vendor results, not an independent test. See Cohere’s pairing results.
Exact values from Cohere’s cross-model test; deltas are against Pro/Pro. Source: Cohere launch results.
Corpus encoderQuery encoderNormalized mean nDCG@10Delta vs. Pro/Pro
ProPro100.0Reference
ProFast98.4−1.6
FastPro97.3−2.7
FastFast96.6−3.4

For a quality-first index with frequent live queries, Pro corpus plus Fast query is Cohere’s recommended mixed pairing. In Cohere’s aggregate, Fast corpus plus Pro query scores 97.3, which is 0.7 normalized points higher than Fast/Fast at 96.6. Pro corpus plus Fast query scores 98.4, 1.1 points higher than the reverse mix. These small aggregate differences do not establish a universal deployment winner. Cohere’s footnote requires the same output dimension on both sides and reports that the pattern persists with Matryoshka truncation and int8 quantization. See the benchmark setup and conditions.

The 40-dataset matrix is a vendor-reported mean. Cohere does not publish per-dataset scores or confidence intervals for this matrix, or the test hardware and latency settings. Treat 98.4 as a trial hypothesis, not a forecast for a different vector store, corpus, or query mix.

Hypothetical Cohere API text-token costs only. Rates from Cohere’s launch price table; no measurement period is assumed.
Corpus modelQuery modelCorpus stage: 100M tokensQuery stage: 1B tokensCombined text cost
ProPro$12$120$132
ProFast$12$80$92
FastPro$8$120$128
FastFast$8$80$88

Formula: (corpus tokens ÷ 1,000,000 × corpus rate) + (query tokens ÷ 1,000,000 × query rate). For example, Pro/Fast is (100 × $0.12) + (1,000 × $0.08) = $92. This scenario is hypothetical and shows text charges only. It excludes image tokens, storage, database costs, reranking, and generation, and it is not a forecast of real workload spend.

Check the API contract before changing your index

Cohere’s current model reference lists both models with a 128K-token context window, 100+ languages, and text, image, and fused text-plus-image inputs. Both list output dimensions of 256, 512, 768, 1,024, 1,536, and 2,048; the default is 2,048. See Cohere’s model table.

For retrieval, set input_type="search_document" for corpus content and input_type="search_query" for queries. Cohere’s v2 endpoint accepts up to 96 texts or mixed input objects per call. Image inputs are data URIs for JPEG, PNG, WebP, or GIF files; the reference caps combined image payload at 20 MB. Cohere’s mixed-content example sends text and image components together in one input object. The default truncation mode is END, which drops excess tokens from the end of an overlong input. Set truncate="NONE" when oversized inputs should fail rather than lose their tail. Check the v2 Embed API reference.

The Embed 5 launch and model table highlight float, int8, and binary outputs. The current API reference also lists uint8, ubinary, and base64 representations for supported Embed models. The default response is float. Choose an output type and dimension that your vector database can store, and keep corpus and query vectors compatible. If your index schema is fixed at 1,536 dimensions, set that dimension explicitly instead of relying on Embed 5’s 2,048 default. API reference and model table list the current options.

Cohere’s direct shared-space statement applies to Embed 5 Pro and Fast. It does not establish compatibility with Embed 4 vectors or embeddings from another model. Matching dimensions alone does not prove semantic compatibility. For an older-model migration, use a separate index or namespace, re-embed documents with the selected Embed 5 corpus encoder and embed new queries with a compatible Embed 5 query encoder, then evaluate before routing production traffic to the new vectors. Cohere’s compatibility statement covers the two Embed 5 variants.

Cohere model-table defaults and parameters compared with the published AWS SageMaker parameter tables. Cohere model table; AWS Pro listing; AWS Fast listing.
Published interfaceDefault output dimensionListed output dimensionsWhat to verify
Cohere Embed 5 model table2,048 for Pro and Fast256, 512, 768, 1,024, 1,536, 2,048Use explicit model, dimension, and output type.
AWS SageMaker Pro parameters1,536256, 512, 1,024, 1,536Pro listing appears out of sync with Cohere’s Embed 5 table.
AWS SageMaker Fast parameters1,024256, 512, 768, 1,024, 1,536, 2,048Fast listing has a different stated default.

The AWS Pro parameter table appears out of sync with Cohere’s Embed 5 model table, and AWS lists a different default for Fast. That is an inference from the published-source conflict, not a confirmed endpoint defect. Before moving an index to SageMaker, confirm the selected offer and endpoint’s current model version, allowed parameters, and defaults. Do not assume Cohere API SDK calls or defaults work unchanged on SageMaker or Microsoft Foundry. AWS Pro listing, AWS Fast listing, and Cohere’s current model table.

The Embed endpoint’s documented inputs are text strings, image data URIs, and mixed content; the current reference does not list a PDF-file field. If you need to turn PDFs into structured text or page content first, see our Cohere Parse guide to PDF ingestion and structured outputs. Keep parsing and embedding as separate steps in your estimate.

Compare the billing unit for your deployment

Cohere API, Model Vault, Microsoft Foundry, and Amazon SageMaker expose different cost units. Do not carry the Cohere API token rate over to an instance-based deployment.

Published billing examples checked October 3–4, 2026. Provider units differ. See Cohere pricing, the AWS Pro listing, and the AWS Fast listing.
Deployment routeEmbed 5 ProEmbed 5 FastUnit and caveat
Cohere API$0.12 per 1M text tokens; $0.40 per 1M image tokens$0.08 per 1M text tokens; $0.40 per 1M image tokensToken-based. The accessed Cohere API pricing and API references do not provide a conversion from image files or page count to billed image tokens.
Cohere Model Vault, Small$3/hour or $2,000/month$3/hour or $2,000/monthPer instance.
Cohere Model Vault, Medium$5/hour or $3,250/month$5/hour or $3,250/monthPer instance.
Amazon SageMaker, ml.g5.xlarge real-time example$2.39 per host-hour$2.39 per host-hourMarketplace software charge; AWS infrastructure costs may apply. Other instance types and modes have different rates.
Microsoft FoundryPay-as-you-go token billingPay-as-you-go token billingCohere Azure guidance confirms the billing basis but does not state a current rate; check the Foundry offer before budgeting.

Cohere lists Embed 5 on its API, Model Vault, Microsoft Foundry, and Amazon SageMaker. It also documents vLLM for private deployments in a customer’s own VPC or on premises; private-deployment pricing is custom. Cohere’s launch announcement lists these availability and private-deployment routes. Cohere’s pricing page distinguishes per-instance Model Vault rates from token billing. The SageMaker listing price is a software usage charge, with AWS infrastructure billed separately. Azure guidance confirms token billing but does not provide the Foundry rate in the accessed documentation. See Cohere’s Azure guidance.

For image-heavy workloads, do not multiply the number of page images by $0.40. Cohere quotes that rate per million image tokens, but the sources checked here do not publish a token conversion for a page or file. Check usage reporting or price a small approved sample before estimating a large image corpus.

Run a small evaluation on your corpus

  1. Choose representative documents and real query types, including cases that matter to your team: tables, scanned pages, multilingual terms, and short interactive searches.
  2. Hold parsing, chunking, filters, vector-store settings, output dimension, and precision constant. Compare Pro/Pro, Pro/Fast, Fast/Pro, and Fast/Fast so the encoder pairing is the change.
  3. Measure relevance with labels from your own search tasks. Track query latency separately from indexing throughput, token charges, image-token usage, and vector storage. Decide in advance what quality or latency change you will accept.

Cohere’s launch also reports ViDoRe V3 results with RCP-nDCG@10, a separate measure from the 40-dataset nDCG@10 matrix. Cohere describes RCP-nDCG@10 as evaluating the ordering of a fixed candidate set, so those results do not by themselves show first-stage retrieval from a full corpus. Read Cohere’s explanation of RCP-nDCG@10 and keep candidate generation in mind when comparing its scores.

Cohere API example: Pro documents and Fast queries

Prerequisites: install the official SDK and NumPy with pip install cohere numpy, then set CO_API_KEY in the process environment. This Cohere v2 example uses Pro for search_document, Fast for search_query, and the same 1,024-dimensional float output on both sides. It ranks an illustrative three-document corpus by cosine similarity. The method follows Cohere’s launch example and v2 API reference. Cohere launch sample; v2 Embed API.

import os
import cohere
import numpy as np

client = cohere.ClientV2(api_key=os.environ["CO_API_KEY"])

documents = [
    "Annual report: operating expenses rose 8 percent.",
    "Maintenance guide: inspect the compressor every 90 days.",
    "Leave policy: employees accrue 1.5 paid days per month.",
]

doc_embeddings = client.embed(
    model="embed-v5.0-pro",
    input_type="search_document",
    texts=documents,
    output_dimension=1024,
    embedding_types=["float"],
    truncate="NONE",
).embeddings.float_

query = "How often should the compressor be inspected?"
query_embedding = client.embed(
    model="embed-v5.0-fast",
    input_type="search_query",
    texts=[query],
    output_dimension=1024,
    embedding_types=["float"],
    truncate="NONE",
).embeddings.float_[0]

doc_vectors = np.asarray(doc_embeddings)
query_vector = np.asarray(query_embedding)
cosine_scores = doc_vectors @ query_vector / (
    np.linalg.norm(doc_vectors, axis=1) * np.linalg.norm(query_vector)
)
for index in np.argsort(cosine_scores)[::-1][:2]:
    print(f"{cosine_scores[index]:.3f} {documents[index]}")

Execution status: This API example was not executed by DMT. The document set is illustrative; no observed match, returned score, throughput, or latency is claimed.

AI assistance disclosure: AI tools assisted with source review and drafting. DMT did not run the Cohere API example or independently reproduce Cohere’s benchmark. All benchmark scores above are labeled as Cohere-reported.

Share this article

Published by

Tayeeb Khan

Tayeeb Khan is the founder of DMarketer Tayeeb, covering digital marketing, SEO and AI. Articles may draw on professional experience, source-based research and AI-assisted or automated production. Firsthand tests are identified in the relevant article; a byline does not imply personal testing or human review of every claim.

Leave a Comment

Your email address will not be published. Required fields are marked *

Stay ahead of the curve

Get actionable digital marketing, SEO, and AI insights delivered to your inbox. No fluff, just value.

No spam. Unsubscribe anytime.