> ## Documentation Index
> Fetch the complete documentation index at: https://docs.gmicloud.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# GMI Router Overview

The GMI Router turns a plain-language request into the right model for the job. Send a chat, and the router picks the model and returns a completion, automatically.

Access the playground here: [https://console.gmicloud.ai/user-console/ie/gmi-router/test-routing](https://console.gmicloud.ai/user-console/ie/gmi-router/test-routing)

<img src="https://mintcdn.com/gmicloud/sjCgJNK4c7I1nIHE/images/GMI-Router-ui.gif?s=37e4eebb535e5e56e7c1f841ada6c33a" alt="GMI Router Ui" width="1280" height="720" data-path="images/GMI-Router-ui.gif" />

## **How routing works**

Autoroute treats model selection as a process, not a single guess. Every Autoroute request moves through the same sequence:

1. **Understand.** The router reads your latest message and classifies the task and its requirements.
2. **Compare.** It scores the allowed model pool across quality, cost, and latency.
3. **Validate.** The top candidate is checked before it is committed. A high score alone does not win the request.
4. **Select.** The router locks in the winning model.
5. **Respond.** The selected model generates the completion, returned inline or streamed token by token.

The flow, at a glance: Prompt -> Understand -> Compare -> Validate -> Select -> Respond.

**/autoroute** picks the model and returns the completion in one call. Streaming is the default.

## **Authentication**

Base URL: `https://console.gmicloud.ai/api/v1/ie/recommendation`

Send a bearer token in the Authorization header:

`Authorization: Bearer YOUR_TOKEN`

Most routed responses include an X-Recommendation-ID header. It is absent on auth and early validation errors.

## **Quick start: Autoroute**

Send an OpenAI-style messages\[] array. The router reads the latest user message, ranks the model pool, and generates with the top candidate.

```text theme={null}
curl -X POST https://console.gmicloud.ai/api/v1/ie/recommendation/autoroute \
  -H "Authorization: Bearer YOUR_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "messages": [
      { "role": "user", "content": "Summarize this contract in three bullet points." }
    ],
    "mode": "balanced",
    "stream": false
  }'
```

Non-streaming response:

```text theme={null}
{
  "model": "provider/model-name",
  "message": { "role": "assistant", "content": "..." },
  "routing_metadata": {
    "selected_model": "provider/model-name",
    "task_type": "summarization",
    "selected_mode": "balanced",
    "fallback_models": ["provider/backup-a"],
    "recommendation_id": "bab0a8bb-195f-49c5-8f9c-016da0d89cc8"
  }
}
```

### **Request fields**

| **Field**  | **Type** | **Required** | **Description**                                                                                  |
| :--------- | :------- | :----------- | :----------------------------------------------------------------------------------------------- |
| `messages` | array    | Yes          | OpenAI-style conversation. The full conversation influences routing, not only the latest message |
| `mode`     | string   | No           | Accepted values are cost, balanced, or quality. Unknown request fields are ignored.              |
| `stream`   | boolean  | No           | Streaming is the default. Set false for one-shot JSON.                                           |

### **Streaming vs non-streaming**

* **Streaming (default).** Returns text/event-stream. Chunks are proxied as they arrive, followed by a routing\_metadata event, then \[DONE]. One backup model may be tried before the first token on a 5xx, provider error, 429, or a 10-second first-token timeout. Once tokens start flowing, no further fallback happens; a mid-stream failure keeps the partial output, sends an error event, and closes the stream.
* **Non-streaming (stream: false).** Returns one JSON object, trying the primary model plus one backup on a transient failure (5xx, provider error, 429), with a combined timeout of 10 minutes for primary and backup.

The router is stateless. Resend the full conversation on every request.

## **Status codes**

| **Code** | **Meaning**                                      |
| :------- | :----------------------------------------------- |
| `200`    | Success.                                         |
| `401`    | Missing or invalid credentials.                  |
| `4xx`    | Invalid request, such as a missing user message. |
| `5xx`    | Routing or model generation failed.              |
