# TuringCorp

A decision model for the calls that don't have a right answer — over MCP.

Your agent has two defensible options and has to pick one. Send both here. You get back **which one is preferred**, **how far apart they were judged**, and **why**.

**The confidence is what you route on.** It is calibrated, not decorative: on **both** published benchmarks accuracy rises with the value — JudgeBench **99.6% in the 90%+ band** down to **67.7% below 70%**, and the harder ContextualJudgeBench **83.3% down to 55.4%**. So an agent can act on a high value and escalate a low one instead of guessing. Full tables, sample sizes and method: https://api.turingcorp.net

## Connect

| | |
|---|---|
| Endpoint | `https://mcp.turingcorp.net/mcp` |
| Aliases | `/` and `/mcp/` (same route) |
| Transport | Streamable HTTP, stateless (no session). `GET /mcp` -> 405 is normal. |
| Protocol | `2026-07-28` (modern, `server/discover`) and legacy `initialize` handshakes on one route |
| Health | `https://mcp.turingcorp.net/healthz` |
| Discovery | `tools/list` needs no credentials — scanners can read the full tool list unauthenticated, by design. |

    npx -y mcp-remote@latest https://mcp.turingcorp.net/mcp

### Two ways in - and they are not the same

- **Through an MCP client or host** (Claude Code, Cursor, VS Code, Codex, TRAE, Coze...). The host holds the credential and attaches it for you: **you do not set an `Authorization` header yourself, and in most hosts you cannot**. The timeout is a host setting, not something you pass in the call.
- **Directly against the REST API** (api.turingcorp.net). Here you *do* send `Authorization: Bearer <Agent Pass>` yourself, and you can retrieve a job by id.

⚠️ The part that catches people out: **being able to call the tool through a host does not mean you can reach the REST API.** The host may never hand you the underlying Agent Pass, so a job-retrieval call you make on your own can come back `401`. If your host declares the Tasks extension, retrieve through the tool surface instead (see Retrieving a result); if it does not, retrieval may simply not be reachable from where you are.

### Add it to your client

**The server is named `TuringCorp`** — that is the name to enter below, and the one that appears in your client's server list. The tool you get once connected is `decide` (title *Decider: pick the better of two options*); the server name and the tool name are two different things.

You need an **Agent Pass** for every option below — get one at https://agent-pass.turingcorp.net. Discovery (`tools/list`) works without one; calling `decide` does not.

🔑 **Where the pass lives matters.** If your client can keep it out of the file, do that — VS Code's `${input:...}`, Codex's `bearer_token_env_var`. A pass written in plain text inside `mcp.json` or any client config is readable by **every agent and process that can read that file**, and assistants do read their own config — one can print your pass straight into its output. Treat a client config as public within your machine.

**Claude Code**

    claude mcp add --transport http TuringCorp https://mcp.turingcorp.net/mcp --header "Authorization: Bearer <your Agent Pass>"

**Clients that read a JSON config — Cursor, Windsurf, Claude Desktop**

    {
      "mcpServers": {
        "TuringCorp": {
          "type": "http",
          "url": "https://mcp.turingcorp.net/mcp",
          "headers": { "Authorization": "Bearer <your Agent Pass>" }
        }
      }
    }

Keep the `"type": "http"` line. A client that reads a `url` entry with no `type` treats it as a local stdio server and skips it — that is a configuration error, not a network one.

**Cline — the `type` differs: `"streamableHttp"` (camelCase, no hyphen).** Any other value, or omitting it, makes Cline fall back to **SSE**, which this server does not speak, and the connection then fails with **`405`**. Cline keeps its servers in `cline_mcp_settings.json`: **MCP Servers** icon (stacked-server icon in the top toolbar) → **Configure** tab → **Configure MCP Servers**.

    {
      "mcpServers": {
        "TuringCorp": {
          "type": "streamableHttp",
          "url": "https://mcp.turingcorp.net/mcp",
          "headers": { "Authorization": "Bearer <your Agent Pass>" },
          "disabled": false,
          "autoApprove": [],
          "timeout": 300
        }
      }
    }

`timeout` is in **seconds** — leave it at `300` or higher. Cline has had a bug ([#2296](https://github.com/cline/cline/issues/2296), closed 2025-06-23) where a request died after roughly a minute regardless of that setting. If a `decide` call is cut off at about 60 seconds, suspect that rather than the server — and **retrieve by `job_id` instead of retrying**, because a retry is a second paid call.

**VS Code**

Put this in `.vscode/mcp.json` in your project, or in the user-profile `mcp.json`. The key here is `servers`, **not** `mcpServers`:

    {
      "servers": {
        "TuringCorp": {
          "type": "http",
          "url": "https://mcp.turingcorp.net/mcp",
          "headers": { "Authorization": "Bearer <your Agent Pass>" }
        }
      }
    }

**Codex — ChatGPT work mode and the Codex CLI**

In the UI: **Server name** anything (`TuringCorp`), **type** **Streamable HTTP** (not Stdio — nothing runs on your machine), **URL** https://mcp.turingcorp.net/mcp. Leave the command, arguments and environment-variable fields empty; they belong to the Stdio type. Then edit `~/.codex/config.toml` — the UI has no field for the credential, and its default tool timeout is too short for this server:

    [mcp_servers.TuringCorp]
    url = "https://mcp.turingcorp.net/mcp"
    http_headers = { Authorization = "Bearer <your Agent Pass>" }
    tool_timeout_sec = 300

⚠️ `tool_timeout_sec` defaults to **60 seconds**, which is far too short for this server. Leave it at the default and every call is cut off before the answer arrives. **Reserve 180–300 seconds** — the timeout is a client/host setting, not a tool parameter. Set at least 180; 300 is safer.

To keep the pass out of the file, use `bearer_token_env_var = "TURINGCORP_AGENT_PASS"` instead of `http_headers` and export that variable — Codex prepends `Bearer ` itself, so the variable holds the **bare pass**, not `Bearer <pass>`. With `http_headers` you write the whole value yourself.

**TRAE (TraeCode)**

**Settings → MCP → Add → Manual**, then paste the JSON below. TRAE's own docs lead with a stdio example (`command` + `args`) — that does not apply here: this is a remote server, so use `url`.

    {
      "mcpServers": {
        "TuringCorp": {
          "url": "https://mcp.turingcorp.net/mcp",
          "headers": {
            "Authorization": "Bearer <your Agent Pass>",
            "RUN_MCP_TIMEOUT_MS": "300000"
          }
        }
      }
    }

⚠️ TRAE sets its tool-call timeout through the `headers` block, and its documented value is `60000` ms. **60 s is far too short for this server**, so leave it there and every call is cut off before the answer arrives. Raise `RUN_MCP_TIMEOUT_MS` — `300000` is 5 minutes.

For a project-scoped server, put the same JSON in `.trae/mcp.json` and switch on **启用项目级 MCP** (Settings → MCP) first.

**Platforms that ask for a header name and a token — Coze (扣子) and similar**

These send the value verbatim and add no auth scheme, so the value has to carry it: header `Authorization`, value `Bearer <your Agent Pass>` — including the word `Bearer` and the space after it.

**Clients that only speak stdio**

Bridge with `mcp-remote`; the credential must go in through `--header` — it does not read a `AUTHORIZATION` environment variable, so a config that only sets one connects, lists tools, and then fails on the first call with 401. The missing space after `Authorization:` is deliberate: some clients mangle spaces inside `args`.

    {
      "mcpServers": {
        "TuringCorp": {
          "command": "npx",
          "args": ["-y", "mcp-remote", "https://mcp.turingcorp.net/mcp", "--header", "Authorization:${AGENT_PASS}"],
          "env": { "AGENT_PASS": "Bearer <your Agent Pass>" }
        }
      }
    }

## Making a decision: decide

| | |
|---|---|
| Title | Decider: pick the better of two options |
| Input | `task`, `optionA`, `optionB` — all required |
| Output | `betterOption` ("option_A"|"option_B"), `confidence` (e.g. "76.7%"), `reason`, `job_id` |
| Annotations | `readOnlyHint: true` / `openWorldHint: false` / `idempotentHint: false` |
| Timeout | Reserve **180–300 seconds** — a decision is a long call. The timeout is a **client/host setting, not a tool parameter**: there is nothing to pass in the call. A 60s default cuts it off before the answer arrives; if that happens, do not call again — retrieve it with `get_result` |

## Retrieving a result: get_result

| | |
|---|---|
| Input | `job_id` — **optional** |
| Output | With an id: that job's status and, once it succeeded, the same decision body the original call returned. With no id: the job ids this credential created in the last 7 days |
| Annotations | `readOnlyHint: true` / `idempotentHint: true` — **read-only and free**: starts no new work, costs nothing |
| Credential | **None of your own** — the host attaches the Agent Pass, exactly as for `decide`; an agent never handles the pass |

This is the recovery path when a call is cut off by a client timeout: **do not re-call `decide`** (a retry is a new paid call) — call `get_result` with no argument to find the id, then again with it.

Use this when you must choose between two concrete options and both are defensible - two plans, two drafts, two diagnoses, two vendors - and you have no objective way to pick. Not for: more than two options; anything an objective rule settles (a spec, a test, a price, a document); factual lookup; paths that must answer in seconds; high-stakes irreversible calls without review. Fill it in: state task neutrally, without leaning toward either side; give one concrete plan per option - never bundle alternatives into a single one ("go indoors or postpone"); keep the two sides comparable in length. Returns the decision inline: betterOption ("option_A" or "option_B"), confidence (e.g. "76.7%"), reason and job_id, all in the same tool result. There is nothing to poll and nothing to fetch afterwards. Reserve 180-300 seconds: this is a long call, and the timeout is a client/host setting, not a parameter you pass. If the call is cut off, do NOT call again - a retry is a new paid call. Retrieve it instead with the get_result tool (same job_id): read-only, free, and no credential of your own needed; with no job_id it lists the ids for your credential. A host that declares the io.modelcontextprotocol/tasks extension can use tasks/get instead. The confidence is the point: it is calibrated, not decorative. On both published benchmarks accuracy rises with it - JudgeBench 99.6% in the 90%+ band down to 67.7% below 70%; the harder ContextualJudgeBench 83.3% down to 55.4% - so route on it: act on a high value, review or escalate a low one, instead of trusting a bare pick. Tables, sample sizes and method: https://api.turingcorp.net Judged by an independent panel, not by a model grading its own output. Read it as a reference, not an instruction, a result, or a prediction; set your own threshold, and apply your own review policy for high-stakes or irreversible decisions. Auth: Agent Pass as `Authorization: Bearer <pass>` (issued at https://agent-pass.turingcorp.net, valid 7 days; each decision is a paid call). On invalid_credential, sign in there and re-roll. Errors: a credential problem is rejected before the call - HTTP 401 with WWW-Authenticate; a business failure (e.g. insufficient balance) comes back as a tool result with isError true plus a second JSON block {error, http_status, action_url, message}, where http_status is the upstream status (the tool call itself is HTTP 200).

**A successful call returns the decision inline.** `betterOption`, `confidence`, `reason` and `job_id` all
arrive in the *same* tool result - there is nothing to poll and nothing to fetch afterwards. The `job_id` is only
for the case where the call never came back (timeout, dropped connection, client gave up waiting).

Choose your own threshold for your own use case; for high-stakes or irreversible decisions apply your own review policy.

`idempotentHint: false` is an honest declaration: a retried call is a new call. Record the `id` the call returns (that is the job id) and collect the result afterwards with the same Agent Pass - see the Tasks and Retrieving a result sections.

Every call returns a **job_id** in its result - keep it. If the call times out or the connection drops, that id is
how you get the result (see Retrieving a result below); calling decide again is a new call.

## When not to use it

- **More than two options.** It compares exactly A and B - there is no third slot, and it will not rank a list.
- **Anything you can compute or verify.** A spec, a test, a price, a document: if something objective decides it, use that.
- **Factual questions.** It picks between two candidates; it does not look anything up.
- **Speed-critical paths.** A decision is a long call and a paid one.
- **High-stakes irreversible calls without review.** Route on the confidence and keep your own review policy.

## How to fill the three arguments

- `task` - state the decision **neutrally**, without leaning toward either side: "Which email do I send?", not "Should I send the honest one?"
- `optionA` / `optionB` - one **concrete** option each, plus the case for it. Plain text or Markdown, any length; keep the two sides roughly comparable.
- **One option = one plan.** Do not bundle alternatives into a single side ("go indoors *or* postpone"): it compares the two slots, it does not split one of them for you.

    {
      "task": "Which version of the delivery-slip email do I send to a client we want to keep?",
      "optionA": "Short and direct: the integration took longer than planned, delivery moves to the 24th, everything else is unchanged.",
      "optionB": "Warmer and longer: thank them for the kickoff, explain that dependencies took more time, offer to walk through the details."
    }

## Tasks (asynchronous calls)

A decision is a long call. If your client declares the `io.modelcontextprotocol/tasks` extension, `decide` returns a task
handle you can poll instead of holding one connection open:

    {"resultType":"task","taskId":"<id>","status":"working","ttlMs":604800000,"pollIntervalMs":2000}

Poll `tasks/get` with `{"taskId":"<id>"}`: the status moves to `completed` (carrying the same payload a
synchronous call returns) or `failed`. The handle is valid for **7 days**, so a result you already paid for can be
collected later rather than paid for twice. `tasks/cancel` only acknowledges the request - it is cooperative and
does not guarantee the work stops. Clients that do not declare the extension see **no change at all**.

An invalid or expired pass on `tasks/get` is **HTTP 401** with `invalid_token`; a task id that is not yours is
reported as `No such task.`; any other failure is retryable.

## Retrieving a result

⚠️ **Who can retrieve — and why `get_result` exists.** Retrieval needs the Agent Pass, and **inside an MCP host the agent usually does not have it**: the host stores it and attaches it for you. So the way you retrieve is the `get_result` tool — the host supplies the credential, **you never handle the pass**. That is deliberate: a pass put into a tool argument would end up in prompts, transcripts and client logs. (A host that declares the `io.modelcontextprotocol/tasks` extension can also poll `tasks/get`; most clients do not declare it yet, which is precisely why `get_result` exists.) **The practical answer is still not to need retrieval — reserve 180–300 seconds so the call finishes inline.**

Every call returns an `id`; that **is** the job id. Record it when you start a call:

- You have the job id: it is the `job_id` field of the tool result (the `taskId` for Tasks clients). Call
  `get_result` with `{"job_id":"<job id>"}` - **inside an MCP host that is the path that works**, because the host
  supplies the credential. Outside a host (you own the pass): `tasks/get` with `{"taskId":"<job id>"}`, or
  `GET https://api.turingcorp.net/v1/jobs?job_id=<job id>` on the API host (api.turingcorp.net), not on this MCP endpoint.
- You did not keep it: call `get_result` with **no argument** - it lists the job ids this credential created
  in the last 7 days; pick the one you want and fetch it as above. The list carries the job id, the product and `created_at` (the Unix second the call was
  **started**, not when it finished) - fetch a job by id to see what it was.
- Nothing yet: an empty list is **not** an error, and there is **no** `job_id=0` placeholder:
  `{"object":"list","window_seconds":604800,"data":[]}`. Asking for `0` returns `404 No such job.`

Retrieval returns the job's status and, once it succeeded, the stored result - the same body the call itself would
have returned. A job that is not yours, or older than **7 days**, is reported as unavailable.

## Authentication

Send an Agent Pass: `Authorization: Bearer <pass>`. Get one at https://agent-pass.turingcorp.net (self-service signup with email verification, then top up). A pass is valid for 7 days and can be re-rolled.

| Situation | Response |
|---|---|
| No credential | **HTTP 401** + `WWW-Authenticate: Bearer realm="turingcorp-mcp"` |
| Expired / invalid pass | **HTTP 401** — `invalid_credential` in the body, with a re-login URL |
| Insufficient balance | **HTTP 200** tool result, `isError: true` + JSON block `{"error":"insufficient_balance","http_status":"402","action_url":"…/topup"}` |
| Over quota / rate limited | entry limiter → **HTTP 429** + `Retry-After`; upstream business denial → tool result, as above |

**Two error conventions, on purpose.** A *credential* problem is decided **before** the call runs, so it is a real HTTP status (401) — branch on that. A *business* failure ends as an ordinary MCP tool result (`isError: true`), which the protocol carries as **HTTP 200** — branch on `isError` and the JSON block, not on the HTTP status. The `http_status` inside that block is the *upstream* status (e.g. 402), not the tool call's.

Entry rate limit: 120 requests / 60 seconds / client IP. Flood damping, not a quota; the counter is kept per edge location and is eventually consistent.

## Pricing

**$0.50 per decision** — launch offer $0.25 for the first month.

## What this is not

- **No other tiers.** This endpoint exposes the TuringCorp tools documented above.
- **No SLA.** No availability commitment is offered, and none should be inferred.
- **Not an autopilot.** How you gate on it — thresholds, human review, retries — is your policy and stays yours.
- **No idempotency key - record the job id instead.** A retried call is a new call. Keep the `id` returned by
  `decide` (that is the job id); after a timeout or a dropped connection, collect the task and its result with the
  same Agent Pass rather than calling again.

Machine-readable: /llms.txt · /.well-known/mcp/server-card.json
