> ## Documentation Index
> Fetch the complete documentation index at: https://docs.shareofmodel.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Rate limits

> Request quotas for the Core API, the Search API and the MCP server, and how to stay within them

Share Of Model limits how many requests a client can send within a sliding time window. These rate limits apply to every public surface: the **Core API**, the **Search API** and the **MCP server**.

<Info>
  **At a glance.** As a guideline, unauthenticated requests are currently limited to around **30 requests per minute**, and authenticated requests to around **120 requests per minute**. Limits are counted per **IP address and user combined**, over a **sliding window**. These values are indicative and may change. See [Adaptive and burst protection](#adaptive-and-burst-protection).
</Info>

## Why rate limits exist

Share Of Model takes the security of its services and of your data very seriously. Rate limits are one of several layers of protection. They are not there to slow down legitimate integrations. They keep the platform reliable, fair and secure for every customer.

<CardGroup cols={2}>
  <Card title="Service stability" icon="server">
    Capping request bursts protects the infrastructure behind the APIs from sudden spikes, so response times stay predictable for everyone, including during peak hours.
  </Card>

  <Card title="Data protection" icon="shield-halved">
    Limits make large-scale scraping, credential stuffing and brute-force attempts impractical, which helps protect your brand data, your analyses and your account.
  </Card>

  <Card title="Fair usage" icon="scale-balanced">
    Shared capacity is spread evenly, so a single runaway script or misconfigured agent cannot degrade the experience of other users and organisations.
  </Card>

  <Card title="Abuse prevention" icon="user-shield">
    Unauthenticated traffic gets a lower limit, which reduces the attack surface of public endpoints such as token exchange.
  </Card>
</CardGroup>

## Limits

| Request type | Indicative limit | Applies to |
| - | - | - |
| Unauthenticated | **\~30 requests / minute** | Requests sent without a valid Bearer token, for example calls to `/v1/auth/token` |
| Authenticated | **\~120 requests / minute** | Requests sent with a valid Bearer JWT, across the Core API, the Search API and the MCP server |

<Tabs>
  <Tab title="Unauthenticated" icon="lock-open">
    A request is unauthenticated when it carries no valid token, for example when you exchange an API key for a JWT. These requests are subject to the lower limit of around **30 requests per minute**.

    <Tip>
      Tokens can be reused until they expire. Exchange your API key once, cache the JWT, and refresh it only when you receive a `401`. Requesting a new token before every call wastes your unauthenticated quota.
    </Tip>
  </Tab>

  <Tab title="Authenticated" icon="lock">
    A request is authenticated when it carries a valid Bearer JWT. These requests get the higher limit of around **120 requests per minute**. This covers every endpoint of the [Core API and the Search API](/api-reference/introduction). See [Authentication](/api-reference/authentication) for how to get a token.
  </Tab>

  <Tab title="MCP server" icon="plug">
    The [MCP server](/api-reference/mcp-integration) at `https://mcp.shareofmodel.ai/mcp` follows the same rules. Your LLM client, whether Claude, Microsoft Copilot Studio or another MCP-compatible tool, is subject to the same quotas as a direct API integration.

    <Note>
      Agents can chain many tool calls to answer a single question. If a conversation stops mid-answer because of rate limiting, wait a moment and ask the assistant to continue, or ask a narrower question so it needs fewer calls.
    </Note>
  </Tab>
</Tabs>

## Adaptive and burst protection

The per-minute values above are not the only safeguard. The platform also applies **burst protection** that limits how many requests can arrive within a very short period, even when the per-minute limit has not been reached. The thresholds for burst protection are not published.

<Warning>
  All the limits described on this page are **indicative and subject to change at any time, without notice**. Share Of Model continuously adjusts its protections based on security signals, traffic patterns and suspected abuse. A client may therefore be limited sooner than the published values suggest, temporarily or permanently.
</Warning>

<Tip>
  Do not design your integration to use the full published quota. Keep a comfortable margin below it, and always handle rejected requests gracefully, as described in [Handle rate limiting in your code](#handle-rate-limiting-in-your-code).
</Tip>

## How limits are counted

Limits are tracked per **combination of IP address and user**. Each pair of client IP address and authenticated identity has its own counter.

The counter uses a **sliding window**: at any moment, the platform looks at the requests you sent during the preceding minute. There is no fixed reset time. Capacity frees up gradually as your older requests move out of the window.

```mermaid theme={null}
flowchart LR
    A[Incoming request] --> B{Valid Bearer token?}
    B -- No --> C[Unauthenticated counter<br/>~30 requests / sliding minute]
    B -- Yes --> D[Authenticated counter<br/>~120 requests / sliding minute]
    C --> E{Under the limit?}
    D --> E
    E -- Yes --> F[Request processed]
    E -- No --> G[Request rejected]
```

What this means in practice:

<AccordionGroup>
  <Accordion title="Two users behind the same IP address" icon="users">
    Each user has their own quota. Colleagues who share an office network or a VPN exit node do not use up each other's requests.
  </Accordion>

  <Accordion title="One user on several IP addresses" icon="network-wired">
    Each IP address counts separately for the same user. For example, a laptop and a server both using the same account are tracked separately.
  </Accordion>

  <Accordion title="API calls and MCP calls from the same user" icon="plug">
    The same limits apply to the APIs and the MCP server. Plan for both when you run scripts and agents on the same account.
  </Accordion>

  <Accordion title="Core API and Search API" icon="layer-group">
    Both surfaces are subject to the limits described on this page. Treat them as one budget when you design high-volume jobs.
  </Accordion>
</AccordionGroup>

## When you exceed the limit

Requests sent after a limit is reached are rejected with HTTP `429 Too Many Requests`. Because the window slides, requests are accepted again as your earlier requests move out of the last minute. Your data, analyses and tokens are not affected. The rejected request is simply not processed, and you can send it again after waiting.

<Warning>
  Retrying immediately in a tight loop does not help. Every retry counts against your quota and keeps you over the limit for longer. Always wait before retrying.
</Warning>

## Handle rate limiting in your code

Follow these steps to build an integration that stays within the limits and recovers cleanly when it does not.

<Steps>
  <Step title="Reuse your access token">
    Exchange your API key once and keep the JWT in memory until it expires. This keeps you well below the unauthenticated limit.
  </Step>

  <Step title="Pace your requests">
    Spread calls evenly over time rather than sending them in bursts. Aim to stay well below the published per-minute values, since burst protection can reject a quick series of requests on its own.
  </Step>

  <Step title="Retry with exponential backoff">
    When a request is rejected for rate limiting, wait before retrying, and double the wait on each consecutive failure. If the response includes a `Retry-After` header, use its value instead.
  </Step>

  <Step title="Cap the number of retries">
    Give up after a few attempts and log the failure, so a persistent problem surfaces instead of looping forever.
  </Step>
</Steps>

The examples below apply these steps. They retry on HTTP `429 Too Many Requests` and honour the `Retry-After` header when it is present.

<CodeGroup>
  ```python Python theme={null}
  import time

  import requests

  API_URL = "https://api.shareofmodel.ai/v1"
  MAX_RETRIES = 5


  def get_with_backoff(path: str, token: str) -> requests.Response:
      delay = 1.0
      for _ in range(MAX_RETRIES):
          response = requests.get(
              f"{API_URL}{path}",
              headers={"Authorization": f"Bearer {token}"},
              timeout=30,
          )
          if response.status_code != 429:
              return response
          retry_after = response.headers.get("Retry-After")
          time.sleep(float(retry_after) if retry_after else delay)
          delay *= 2
      response.raise_for_status()
      return response
  ```

  ```javascript JavaScript theme={null}
  const API_URL = "https://api.shareofmodel.ai/v1";
  const MAX_RETRIES = 5;

  const sleep = (ms) => new Promise((resolve) => setTimeout(resolve, ms));

  async function getWithBackoff(path, token) {
    let delay = 1000;
    let response;
    for (let attempt = 0; attempt < MAX_RETRIES; attempt += 1) {
      response = await fetch(`${API_URL}${path}`, {
        headers: { Authorization: `Bearer ${token}` },
      });
      if (response.status !== 429) {
        return response;
      }
      const retryAfter = response.headers.get("Retry-After");
      await sleep(retryAfter ? Number(retryAfter) * 1000 : delay);
      delay *= 2;
    }
    throw new Error(`Rate limited after ${MAX_RETRIES} attempts: ${path}`);
  }
  ```
</CodeGroup>

## Best practices

<Columns cols={2}>
  <Card title="Cache responses" icon="database">
    Metrics are produced per [collect](/concepts/analyses-and-collects#collects), so the data behind a finished collect does not change between calls. Store results locally instead of calling the same endpoint repeatedly.
  </Card>

  <Card title="Request only what you need" icon="filter">
    Pass several IDs in `collect_ids` to get aggregated metrics in one call, rather than requesting each collect separately and combining the results yourself.
  </Card>

  <Card title="Run jobs sequentially" icon="arrow-right-arrow-left">
    Avoid firing dozens of parallel requests from the same account and IP address. A small worker pool with pacing is faster overall than a burst that gets rejected.
  </Card>

  <Card title="Keep agents focused" icon="robot">
    With the MCP server, specific questions scoped to one brand, analysis or period need fewer tool calls than broad, open-ended ones.
  </Card>
</Columns>

<Check>
  An integration that reuses its token, paces its calls, keeps a margin below the published values and backs off on rejection is much less likely to be limited.
</Check>

## Related pages

<CardGroup cols={3}>
  <Card title="Authentication" icon="key" href="/api-reference/authentication">
    Get and refresh a Bearer JWT.
  </Card>

  <Card title="MCP Integration" icon="plug" href="/api-reference/mcp-integration">
    Connect an LLM agent to your data.
  </Card>

  <Card title="API keys" icon="lock" href="/platform/getting-started/creating-and-managing-api-keys">
    Create, scope and rotate API keys.
  </Card>
</CardGroup>
