POST /v1/responses is the core call. You make it against your instance, not against api.agent37.com: every instance serves its own chat API at https://{instanceId}.agent37.app, the url of the default port in the create response. This page uses https://ab12cd34ef.agent37.app. Authenticate with the same sk_live_ key you use on the hosting API, sent as the X-Agent37-Key header.
The call is agentic by default: the agent can browse, run code, use a terminal, read and write files, call connected tools, and reason across many steps before answering.
Request body
Request bodies are capped at 2 MB; anything larger returns413 payload_too_large.
string
required
The message or task, a plain string. There is no image field; to attach files, upload them first and list their paths in
files. See Sessions for how history carries across turns.string[]
Paths of files on the instance to attach to this turn, typically the
path returned by PUT /v1/files/content. Each must name an existing file on the instance, or the call returns 400 validation_error. The paths are appended to the input, and the agent reads them from disk.string
Continue an existing conversation. Omit it to start a new one; the response returns the new session’s id. The harness owns sessions and creates one on first use, so an id it has not seen simply starts a fresh thread under that id rather than erroring. The exception is a harness that mints its own session ids:
codex and opencode reject a first turn on an id they did not issue with 400 validation_error (param: session_id), so on those two always start a thread by omitting session_id and reuse the id the response returns.boolean
default:"false"
true returns a Server-Sent Events stream; false returns the finished response as one JSON body. See Streaming.string
The LLM to run this turn on, an id from
GET /v1/models (on OpenClaw the ids are provider/model). On Hermes it applies to this turn only: omit it and the turn runs on the instance’s default model, whatever earlier turns chose. On OpenClaw a named model is written to the session and sticks for the turns that follow; a turn that omits it leaves the session’s model as is. A model OpenClaw does not allow (an id outside GET /v1/models) never starts the turn: the response settles with status: "failed" and error.code agent_error, with OpenClaw’s reason in message, and the session keeps its previous model. See Models.string
The model’s provider, for example
anthropic. Hermes only, and per turn like model: it tells Hermes which configured provider to run the named model on. OpenClaw ignores it (the response still echoes it); there the provider is the prefix of the provider/model id you pass as model.string
How hard the model thinks:
none, minimal, low, medium, high, xhigh, max, or ultra. max is the strongest plain thinking level; ultra is the harness’s own Ultra mode, top thinking plus its agentic orchestration (on OpenClaw it also turns on proactive subagent orchestration). It applies to this turn only on both harnesses. Omit it and Hermes uses the instance’s configured effort, while OpenClaw uses the session’s own thinking level. On OpenClaw the values map one to one onto its thinking levels, with none sent as off. max and ultra need gateway 0.10.0: older instances reject them with validation_error until updated.object
Up to 16 key/value pairs, at most 64 KB serialized. Echoed back on the response object, never interpreted.
string
Which agent harness runs the turn,
hermes, openclaw, claude-code, codex, grok, or opencode. Omit it to use the instance’s configured default (hermes on agent37-hermes, openclaw on agent37-openclaw, claude-code on agent37-claude-code, codex on agent37-codex, grok on agent37-grok, opencode on agent37-opencode). Routing is per request: the gateway keeps no session-to-agent binding, so if a session runs on a non-default harness, send agent on every turn of it. Targeting a harness the instance was not provisioned with does not reject the call: the turn settles with status: "failed" and error.code agent_unavailable.string
default:"chat"
chat runs one turn and replies. goal is reserved: sending it returns 400 validation_error today.instance_id in the body is accepted and ignored. The URL names the instance: one gateway per instance, so there is nothing to route.Response
The response object. Ids are 32-character hex strings; timestamps from the gateway are epoch milliseconds.string
The response id. Use it to reconnect or cancel the turn.
string
The conversation this turn belongs to. Reuse it on the next call to continue the thread.
string
in_progress, then a terminal completed, failed, or cancelled.string
The agent that ran the turn,
hermes, openclaw, claude-code, codex, or opencode.string | null
The
model you sent on this request, echoed back. null when you omitted it: the turn still ran on the harness’s default or session model, but the gateway does not report which model the harness resolved.string | null
The
provider you sent on this request, echoed back; null when you omitted it.string
The agent’s final answer. Always a string, empty if the turn produced none.
object | null
Token counts and cost for the turn:
{ input_tokens, output_tokens, cost_usd }. cost_usd is absent or null when the provider did not report a cost.object | null
The session’s context window as of this turn:
{ used_tokens, window_tokens }, the tokens occupying the model’s window against the window’s size. null when the harness didn’t report a measurement; a cancelled turn reports none on either harness. Hermes reports it on every completed turn. OpenClaw reports none, so the gateway derives it from the turn: used_tokens is the turn’s token total (input, output, and cache tokens) and window_tokens is the model’s context window from OpenClaw’s model catalog, or 128000 (OpenClaw’s own default budget) when the catalog lists no window for the model. On OpenClaw it is null when the turn reported no usage, no model, a zero token total, or a model that is not in the catalog.object | null
Your request metadata, echoed back verbatim.
number
When the turn started, epoch milliseconds.
A failed turn does not reject the HTTP call. The POST still returns 200 with
status: "failed" and error set. Branch on status, not on the HTTP code.On a non-streaming call the gateway sends the
200 and headers as soon as the turn starts, then keeps the connection alive with a whitespace tick every 25 seconds while the agent works. The JSON body arrives when the turn finishes, prefixed by that whitespace. Leading whitespace is valid JSON, so standard parsers handle it unchanged. Don’t treat the early headers as the response being ready.Example
Continue a conversation
The first message omitssession_id and starts a session. The reply returns a session_id; pass it on the next message to continue the same thread. The session holds the full history, so you never resend a transcript: you send only the new input.
One active turn per session. A session runs one response at a time. Sending new input while one is in flight returns
409 session_busy, normally with the running response’s id in error.response_id. Use another session, reattach to the running turn, or cancel it first.When the agent needs a decision from you
The agent asks by ending its turn: the response comes backcompleted with the question (and the options, if any) as output_text, and the turn is over. Answer by sending the reply as the next input on the same session_id. There is no in-turn prompt to answer while a response is in_progress, so a chat client needs nothing special: render the question like any other reply, and send the user’s answer like any other message.
Follow up on a response
Every response has an id you can use after the call returns.GET /v1/responses/{id}/stream replays every event so far in order, then stays attached live, so a dropped connection never loses the answer, including after the turn has finished, while the response is still retained. See Streaming for the replay window. Lost the id (page reload, new device)? GET /v1/sessions/{id} returns the running response as active_response_id.
POST /v1/responses/{id}/cancel takes no body and stops a running turn (on OpenClaw it aborts the run inside OpenClaw). It returns 200 with the current response object as soon as the stop is requested, not once the agent has actually stopped, so the body normally still reads status: "in_progress"; the response settles to cancelled when the turn unwinds. To see it end, read GET /v1/responses/{id}/stream: a cancelled turn still closes with response.completed, carrying the output_text accumulated so far. Cancelling a finished response is a no-op that returns its terminal state, still 200.
Status values
A response moves fromin_progress to exactly one terminal status.