Compass 即将登陆云端 | Cohere

Cohere Blog ·

一个介绍 Compass(Cohere 的检索系统)的来源,涵盖其云端部署的理由、精准检索带来的成本优势、面向 agentic 工作流的设计要求,以及在金融 RAG 任务上基准测试所体现的准确率提升。 阅读 4 条观点,查看支持证据与原始来源。

Cohere Team

理解这篇

4 个要点

综合解读

  1. 客户需求推动了托管产品的推出

    团队持续要求提供托管版 Compass——希望获得其一流的检索能力,同时避免自托管带来的运维开销。

    支持这项说法 1

    客户对托管方案的需求一直明确且持续:团队希望获得 Compass 一流的检索能力,但许多人不愿承担自托管带来的运维开销。

    Cohere Team · 段落 5

    原始摘录
    Customer demand for a managed option has been clear and consistent: teams want Compass' best-in-class retrieval capabilities, but many do not want the operational overhead that comes with self-hosting.
    回到原文语境 →
  2. 精准检索可降低 token 成本

    更精准的检索可为生成式模型提供更小、更高质量的输入,从而降低推理成本并缩短任务完成时间。检索是当今企业可用的最有效成本杠杆之一。

    支持这项说法 1

    代币经济学:每个传给模型的无关结果都会消耗代币,并占用有限的上下文空间。更精准的检索能生成更小、质量更高的输入,从而降低推理成本,缩短任务完成时间。检索是当前企业可用的最有效的成本调控手段之一。

    Cohere Team · 段落 10

    原始摘录
    Token economics: Every irrelevant result passed to a model consumes tokens and occupies limited context space. More precise retrieval creates smaller, higher-quality inputs, reducing inference costs and cutting the time needed to complete a task. Retrieval is one of the most effective cost levers available to businesses today.
    回到原文语境 →

    继续探索

    成本优化 →
  3. Agentic 工作流需要可靠的多查询检索

    智能体在每项任务中可能会发出数十次查询。因此,检索系统必须在一系列机器生成的查询中可靠运行,而不仅限于单次搜索,同时要保持相关性、在每一步执行权限控制,并避免延迟累积。

    支持这项说法 1

    智能体访问模式:智能体在完成单个任务过程中可能发出数十次查询,不断重构请求并遍历多个数据源。在此类多跳循环中,延迟持续累积,相关性可能逐渐偏移,且权限控制必须在每一步都严格执行。因此,检索系统必须能在一系列由机器生成的查询序列中稳定可靠地运行,而不仅限于单次查询场景。

    Cohere Team · 段落 11

    原始摘录
    Agentic access patterns: Agents may issue dozens of queries while completing one task, reformulating requests and traversing multiple sources. In these multi-hop loops, latency accumulates, relevance can drift, and permissions must be enforced at every step. Retrieval must therefore perform reliably across sequences of machine-generated queries, not only single-shot searches.
    回到原文语境 →
  4. Compass 在金融 RAG 工作负载上的准确率提高了 14–16 个百分点

    在 High Finance 基准测试(一项由 Cohere 构建的投资银行 RAG 工作负载)中,Compass 的准确率比 Azure Search 提高了 14–16 个百分点(从 64.8 提升至 81.1)。这一差距可能决定了最终用户得到的是不令人满意的答案还是出色的答案。

    支持这项说法 1

    下图展示了一个典型金融行业RAG工作负载中,相较于传统或独立搜索基础设施所实现的准确率提升:即对查询进行嵌入(embed),从索引中检索演示材料,并对顶部结果进行评分。在“高阶金融”(High Finance)——一个由Cohere构建的投资银行业务基准测试中,Compass在Azure Search上的表现提升了14至16个百分点(从64.8提升至81.1)。如此幅度的差距,可能直接决定终端用户获得的是令人不满的答案,还是出色的答案。

    Cohere Team · 段落 25

    原始摘录
    The figure below is one instance of the accuracy gain over traditional or standalone search infrastructure on a representative financial-industry RAG workload: embed a query, retrieve presentation materials from an index, and score the top results. On High Finance, a Cohere-built investment-banking benchmark, Compass achieved a 14-16 point improvement on Azure Search (from 64.8 to 81.1). A gap of this size can be the difference between an unsatisfactory answer and a great one for the end user.
    回到原文语境 →

关键段落4

带明确归属与语境的原文片段。打开原始文本核查出处。

产品部署策略

客户需求推动了托管产品的推出

客户对托管方案的需求一直明确且持续:团队希望获得 Compass 一流的检索能力,但许多人不愿承担自托管带来的运维开销。

原始摘录
Customer demand for a managed option has been clear and consistent: teams want Compass' best-in-class retrieval capabilities, but many do not want the operational overhead that comes with self-hosting.
成本优化

精准检索可降低 token 成本

代币经济学:每个传给模型的无关结果都会消耗代币,并占用有限的上下文空间。更精准的检索能生成更小、质量更高的输入,从而降低推理成本,缩短任务完成时间。检索是当前企业可用的最有效的成本调控手段之一。

原始摘录
Token economics: Every irrelevant result passed to a model consumes tokens and occupies limited context space. More precise retrieval creates smaller, higher-quality inputs, reducing inference costs and cutting the time needed to complete a task. Retrieval is one of the most effective cost levers available to businesses today.
性能基准测试

Compass 在金融 RAG 工作负载上的准确率提高了 14–16 个百分点

下图展示了一个典型金融行业RAG工作负载中,相较于传统或独立搜索基础设施所实现的准确率提升:即对查询进行嵌入(embed),从索引中检索演示材料,并对顶部结果进行评分。在“高阶金融”(High Finance)——一个由Cohere构建的投资银行业务基准测试中,Compass在Azure Search上的表现提升了14至16个百分点(从64.8提升至81.1)。如此幅度的差距,可能直接决定终端用户获得的是令人不满的答案,还是出色的答案。

原始摘录
The figure below is one instance of the accuracy gain over traditional or standalone search infrastructure on a representative financial-industry RAG workload: embed a query, retrieve presentation materials from an index, and score the top results. On High Finance, a Cohere-built investment-banking benchmark, Compass achieved a 14-16 point improvement on Azure Search (from 64.8 to 81.1). A gap of this size can be the difference between an unsatisfactory answer and a great one for the end user.
检索系统设计

Agentic 工作流需要可靠的多查询检索

智能体访问模式:智能体在完成单个任务过程中可能发出数十次查询,不断重构请求并遍历多个数据源。在此类多跳循环中,延迟持续累积,相关性可能逐渐偏移,且权限控制必须在每一步都严格执行。因此,检索系统必须能在一系列由机器生成的查询序列中稳定可靠地运行,而不仅限于单次查询场景。

原始摘录
Agentic access patterns: Agents may issue dozens of queries while completing one task, reformulating requests and traversing multiple sources. In these multi-hop loops, latency accumulates, relevance can drift, and permissions must be enforced at every step. Retrieval must therefore perform reliably across sequences of machine-generated queries, not only single-shot searches.

这里提到的

全部提及对象

Azure Search

仅提及

Cohere 团队将 Azure Search 用作 High Finance 基准测试中 Compass 的基准对照,仅报告其得分(64.8),未对其本身作出超出相对性能的评价。

查看支持证据 · Cohere Team

Compass

支持

Cohere 团队报告称,Compass 在 High Finance 基准测试中相较 Azure Search 准确率提升了 14–16 分,并指出这一差距对终端用户答案质量具有决定性影响。

查看支持证据 · Cohere Team

来源与研究方法

这些观点均关联原始来源。转述已明确标注,不作为逐字原话展示。

打开转录或来源材料 (在新标签页中打开)报告问题

继续研究相关话题

以下资料涉及本文的话题;讨论相同话题不代表观点一致。