> ## Documentation Index
> Fetch the complete documentation index at: https://docs.prysm1.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Orchestrate

> Run a prompt across several models — cascade, ensemble, decompose, or debate — and get one synthesized answer with a PrysmProof v2.

<Note>
  **POST** `https://api.prysm1.com/v2/orchestrate` · Requires authentication
</Note>

Where [`/v1/chat/completions`](/api-reference/chat-completions) routes a prompt to the
single best model, **orchestrate** plans and executes it across *several* models, then
returns one synthesized answer plus a [PrysmProof v2](/concepts/prysmproof) attesting to
how robustly it was produced — which models ran and how strongly they agreed.

You pick the objective with a **policy**; PRYSM picks the **strategy** (or you force one).
See [How orchestration works](/concepts/orchestration) for the full model.

<Warning>
  This endpoint lives under **`/v2`**, not `/v1`. The SDKs target it automatically with
  `client.orchestrate(...)`.
</Warning>

## Authorization

<ParamField header="Authorization" type="string" required>
  Your secret key as a bearer token: `Bearer prysm_sk_...`
</ParamField>

## Body

<ParamField body="messages" type="array" required>
  The conversation, OpenAI-style: a list of `{ "role": "user" | "assistant" | "system", "content": "..." }`.
</ParamField>

<ParamField body="policy" type="string" default="balanced">
  The objective dial: `efficiency` (cheapest path that clears a confidence bar),
  `depth` (cross several models in parallel for robustness), or `balanced`.
</ParamField>

<ParamField body="strategy" type="string">
  Force an execution shape instead of auto-planning: `single`, `cascade`, `ensemble_moa`,
  `rank_fuse`, `decompose_and_route`, `self_consistency`, or `debate`. Omit to let PRYSM
  choose from the policy and prompt.
</ParamField>

<ParamField body="k" type="integer">
  Ensemble / sample width — how many models or samples to cross for `ensemble_moa`,
  `rank_fuse`, and `self_consistency`. Defaults to a policy-appropriate value.
</ParamField>

<ParamField body="max_tokens" type="integer" default="1024">
  Maximum tokens per underlying model call.
</ParamField>

<ParamField body="temperature" type="number" default="0.7">
  Sampling temperature passed to the underlying models.
</ParamField>

<ParamField body="max_cost_usd" type="number">
  A soft budget hint, in USD. Cascades stop escalating to pricier models once the
  estimated spend approaches this cap.
</ParamField>

<ParamField body="judge_model" type="string">
  Preferred aggregator/fuser model for strategies that synthesize a final answer
  (`ensemble_moa`, `rank_fuse`, `debate`). Ignored if it isn't a known catalog model.
</ParamField>

<ParamField body="compliance" type="object">
  A [Policy-as-Code](/concepts/compliance) spec that confines the run to approved
  providers/models. Non-compliant models are filtered out **before** scoring, so the
  engine cannot select one. Same fields as
  [`/v2/compliance/preview`](/api-reference/compliance-preview) (`provider_allowlist`,
  `jurisdiction`, `frameworks`, `certifications`, `data_residency`, `block_data_classes`,
  `require_zero_retention`). When set, the response's `prysm.compliance` carries the
  decision and `prysm.proof.compliance` carries the attestation.
</ParamField>

<ParamField body="brain_config" type="object">
  A [BRAIN.md](/concepts/brain-md) config whose `compliance:` block applies if `compliance`
  is omitted.
</ParamField>

<ParamField body="include_trace" type="boolean" default="true">
  Include the per-stage execution trace in `prysm.stages`. Set `false` for a leaner
  response.
</ParamField>

## Response

<ResponseField name="id" type="string">Unique orchestration id, e.g. `prysm-a1b2c3d4`.</ResponseField>
<ResponseField name="object" type="string">Always `orchestration`.</ResponseField>
<ResponseField name="created" type="integer">Unix timestamp (seconds).</ResponseField>
<ResponseField name="policy" type="string">The policy that ran: `efficiency`, `balanced`, or `depth`.</ResponseField>
<ResponseField name="strategy" type="string">The strategy that ran (auto-planned or forced).</ResponseField>
<ResponseField name="reason" type="string">Plain-English explanation of why this policy/strategy was chosen.</ResponseField>

<ResponseField name="choices" type="array">
  OpenAI-compatible choices. The synthesized answer is `choices[0].message.content`.

  <Expandable title="choices[]">
    <ResponseField name="index" type="integer">Choice index (always `0`).</ResponseField>
    <ResponseField name="message" type="object">`{ "role": "assistant", "content": "..." }`.</ResponseField>
    <ResponseField name="finish_reason" type="string">Always `stop`.</ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="usage" type="object">
  Aggregate token usage across every model call.

  <Expandable title="usage">
    <ResponseField name="prompt_tokens" type="integer">Total input tokens.</ResponseField>
    <ResponseField name="completion_tokens" type="integer">Total output tokens.</ResponseField>
    <ResponseField name="total_tokens" type="integer">Sum of the two.</ResponseField>
  </Expandable>
</ResponseField>

<ResponseField name="prysm" type="object">
  The orchestration extension block.

  <Expandable title="prysm">
    <ResponseField name="orchestration" type="object">
      <Expandable title="orchestration">
        <ResponseField name="policy" type="string">The policy that ran.</ResponseField>
        <ResponseField name="strategy" type="string">The strategy that ran.</ResponseField>
        <ResponseField name="reason" type="string">Why this plan was chosen.</ResponseField>
        <ResponseField name="models_used" type="string[]">Every model that contributed.</ResponseField>
        <ResponseField name="confidence" type="number">Confidence in the final answer, `0`–`1`.</ResponseField>
        <ResponseField name="agreement" type="number">How strongly the models agreed, `0`–`1`.</ResponseField>
        <ResponseField name="escalated" type="boolean">Whether a cascade escalated to a stronger model.</ResponseField>
        <ResponseField name="latency_ms" type="integer">Wall-clock latency in milliseconds.</ResponseField>
      </Expandable>
    </ResponseField>

    <ResponseField name="cost" type="object">
      <Expandable title="cost">
        <ResponseField name="total_usd" type="number">Total cost across all model calls, USD.</ResponseField>
        <ResponseField name="estimated" type="boolean">Always `true` — v2 cost is estimated from text length (\~4 chars/token).</ResponseField>
      </Expandable>
    </ResponseField>

    <ResponseField name="proof" type="object">
      A verifiable [PrysmProof v2](/concepts/prysmproof). Verify it later via
      [`GET /v1/proof/{request_id}`](/api-reference/proof).

      <Expandable title="proof">
        <ResponseField name="request_id" type="string">The id to verify against.</ResponseField>
        <ResponseField name="timestamp" type="string">ISO-8601 UTC timestamp.</ResponseField>
        <ResponseField name="proof_hash" type="string">SHA-256 over the execution stages, e.g. `sha256:a1b2c3d4e5f6...`.</ResponseField>
        <ResponseField name="policy" type="string">The policy that ran.</ResponseField>
        <ResponseField name="strategy" type="string">The strategy that ran.</ResponseField>
        <ResponseField name="models_used" type="string[]">Every model that contributed.</ResponseField>
        <ResponseField name="confidence" type="number">Confidence in the final answer.</ResponseField>
        <ResponseField name="agreement" type="number">How strongly the models agreed.</ResponseField>
        <ResponseField name="verifiable" type="boolean">Always `true`.</ResponseField>
      </Expandable>
    </ResponseField>

    <ResponseField name="stages" type="array">
      Per-stage execution trace (present when `include_trace` is `true`). Each stage has a
      `name` (e.g. `propose` → `aggregate`, or `subtasks` → `synthesize`), a `results`
      array with one entry per model call (model id, ok, token counts, latency, cost,
      confidence, role — the raw text is omitted), and a strategy-specific `detail` object.
    </ResponseField>
  </Expandable>
</ResponseField>

## Errors

| Status | `error`                 | Meaning                                              |
| ------ | ----------------------- | ---------------------------------------------------- |
| `400`  | `no_messages`           | `messages[]` was empty.                              |
| `401`  | —                       | Missing or invalid API key.                          |
| `502`  | `all_models_failed`     | Keys are configured but every model call failed.     |
| `502`  | `orchestration_error`   | The orchestrator raised while planning or executing. |
| `503`  | `no_provider_available` | No provider API keys are configured on the server.   |

<RequestExample>
  ```python Python theme={null}
  from prysm import Prysm

  client = Prysm()
  r = client.orchestrate(
      "Compare three database designs for a 40-person team",
      policy="depth",
  )
  print(r["choices"][0]["message"]["content"])
  print(r["prysm"]["orchestration"]["models_used"])
  ```

  ```typescript Node theme={null}
  import { Prysm } from "@prysmai/sdk";

  const client = new Prysm();
  const r = await client.orchestrate(
    "Compare three database designs for a 40-person team",
    { policy: "depth" },
  );
  console.log(r.choices[0].message.content);
  console.log(r.prysm.orchestration.models_used);
  ```

  ```bash cURL theme={null}
  curl https://api.prysm1.com/v2/orchestrate \
    -H "Authorization: Bearer $PRYSM_API_KEY" \
    -H "Content-Type: application/json" \
    -d '{
      "messages": [
        { "role": "user", "content": "Compare three database designs for a 40-person team" }
      ],
      "policy": "depth"
    }'
  ```
</RequestExample>

<ResponseExample>
  ```json 200 theme={null}
  {
    "id": "prysm-a1b2c3d4",
    "object": "orchestration",
    "created": 1767312000,
    "policy": "depth",
    "strategy": "ensemble_moa",
    "reason": "depth policy on an analysis-heavy prompt: proposed across diverse models, then aggregated.",
    "choices": [
      {
        "index": 0,
        "message": {
          "role": "assistant",
          "content": "Across the three designs, the trade-offs cluster around consistency vs. operational cost..."
        },
        "finish_reason": "stop"
      }
    ],
    "usage": {
      "prompt_tokens": 412,
      "completion_tokens": 933,
      "total_tokens": 1345
    },
    "prysm": {
      "orchestration": {
        "policy": "depth",
        "strategy": "ensemble_moa",
        "reason": "depth policy on an analysis-heavy prompt: proposed across diverse models, then aggregated.",
        "models_used": ["claude-sonnet-4.5", "gpt-5.2", "gemini-3.1-pro"],
        "confidence": 0.88,
        "agreement": 0.81,
        "escalated": false,
        "latency_ms": 4120
      },
      "cost": { "total_usd": 0.004812, "estimated": true },
      "proof": {
        "request_id": "a1b2c3d4-...-...",
        "timestamp": "2026-06-03T00:00:00+00:00",
        "proof_hash": "sha256:a1b2c3d4e5f60718",
        "policy": "depth",
        "strategy": "ensemble_moa",
        "models_used": ["claude-sonnet-4.5", "gpt-5.2", "gemini-3.1-pro"],
        "confidence": 0.88,
        "agreement": 0.81,
        "verifiable": true
      },
      "stages": [
        {
          "name": "propose",
          "results": [
            { "model": "claude-sonnet-4.5", "ok": true, "in_tokens": 96, "out_tokens": 280, "latency_ms": 2010, "cost_usd": 0.00128, "confidence": 0.86, "role": "proposer" },
            { "model": "gpt-5.2", "ok": true, "in_tokens": 96, "out_tokens": 305, "latency_ms": 2240, "cost_usd": 0.00161, "confidence": 0.84, "role": "proposer" },
            { "model": "gemini-3.1-pro", "ok": true, "in_tokens": 96, "out_tokens": 268, "latency_ms": 1980, "cost_usd": 0.00098, "confidence": 0.82, "role": "proposer" }
          ],
          "detail": { "k": 3 }
        },
        {
          "name": "aggregate",
          "results": [
            { "model": "claude-sonnet-4.5", "ok": true, "in_tokens": 124, "out_tokens": 80, "latency_ms": 1890, "cost_usd": 0.00094, "confidence": 0.88, "role": "aggregator" }
          ],
          "detail": { "aggregator": "claude-sonnet-4.5" }
        }
      ]
    }
  }
  ```
</ResponseExample>
