API reference
Endpoints, model ids, the memtro extension, streaming and tools.
Memtro exposes the OpenAI chat-completions shape at https://platform.memtro.com/v1, so any OpenAI-compatible client works unchanged. Authenticate with Authorization: Bearer sk-memtro-….
Endpoints
| Method and path | Purpose |
|---|---|
POST /v1/chat/completions | Chat, streaming or not, with tools. The main endpoint. |
GET /v1/models | The model ids available to you: memtro, auto, auto/<provider>, and provider/model for every provider you have a key for. |
POST /v1/embeddings | Embeddings through your OpenAI or Google key. |
POST /api/mcp | Memtro as an MCP server. |
Model ids
| Id | Meaning |
|---|---|
auto | Complexity routing across the configured ladder. See Model routing. |
auto/anthropic, auto/openai, auto/google, auto/claude-code | Routing within one provider. |
memtro | The organisation's default model (set under Providers). brain is accepted as an older alias. |
anthropic/claude-opus-5, openai/gpt-5.4, google/gemini-3.8-flash | A specific model through your key for that provider. Bare ids such as claude-sonnet-5 are inferred. |
claude-code/opus, claude-code/sonnet, claude-code/haiku, claude-code/fable | Run through the Claude Code CLI with your Claude Code token or Anthropic key. |
compat:<label>/<model> | An OpenAI-compatible endpoint you configured under that label. |
Any id with :agent or :chat | Forces agent mode or single-reply mode for this request, e.g. auto:agent. |
The memtro request field
Add a memtro object to the request body to control how Memtro behaves. Every field is optional. (brain is still accepted as the field name for older clients.)
{
"model": "auto",
"messages": [{"role": "user", "content": "Summarise open VIP tickets"}],
"memtro": {
"mode": "agent",
"memory": "org",
"retrieve": true,
"tools": true,
"show_progress": true,
"engine": "native",
"cache": true
}
}
| Field | Values | Default | What it does |
|---|---|---|---|
mode | chat, agent | chat, or the organisation's default | agent runs the multi-step harness: plan, query every relevant system, cross-reference, answer once. See Agent mode. |
memory | org, personal, off | org when you are in an organisation | Where facts learned from this exchange are stored. Acts as a ceiling: within an org request, facts about you personally are still classified as personal. |
retrieve | boolean | true | Inject relevant memories and graph context into the request. |
tools | boolean | true | Expose Memtro's memory tools and your connector tools to the model. |
show_progress | boolean | true | In agent mode, stream tool activity as reasoning_content deltas. |
engine | native, opencode, claude-code | the organisation's setting | Which harness runs agent mode. |
cache | boolean | true | Serve identical connector calls from the short-lived result cache. |
What comes back
Responses are standard chat completions plus a memtro field:
{
"memtro": {
"mode": "agent",
"engine": "native",
"routing": {"requested": "auto", "model": "anthropic/claude-opus-5", "complexity": 3, "method": "heuristic", "reason": "…"},
"tools": [{"tool": "helpscout_get_conversation", "input": {"id": 2450}, "ok": true, "summary": "…"}]
}
}
When streaming, the first chunk carries memtro with the mode, engine and routing decision; tool activity arrives as reasoning_content deltas while the model works, and a final chunk lists the tools used. Open WebUI, Cherry Studio and most chat UIs render the reasoning as a collapsible "thinking" block.
Tools
- Server tools run inside Memtro:
memtro_search,memtro_remember,memtro_explore, the job tools,memtro_delegatein agent mode, and one set per connector you have access to (hubspot_search,helpscout_get_conversation,postgres_query, and so on). You do not declare them; they are always there unlesstools: false. - Client tools declared in the request's
toolsarray work the OpenAI way: the response returnstool_callsfor you to execute. Client tools are not available withclaude-code/*models, and requests that declare them fall back to the native harness.
Streaming
Set stream: true for server-sent events in the OpenAI chunk format. Add stream_options: {"include_usage": true} for a final usage chunk. The API also accepts temperature, top_p, max_tokens or max_completion_tokens, stop and image content parts.
Errors
Errors use the OpenAI error envelope ({"error": {"message", "type", "code"}}). Common codes: invalid_api_key, provider_not_configured (add a key under Providers), model_not_found, routing_unavailable, unsupported_tools (client tools with a claude-code/* model), engine_unavailable.
Limits
The API is rate limited per client address and long requests (agent runs) may take minutes; keep client timeouts generous. Each agent request has a step budget (default 30) set under Dashboard → Routing.