↩ BLOG
/
Tutorial

Understanding RPM, ITPM, OTPM: API Rate Limits Explained

Learn how Neosantara's 3D rate limiting works β€” RPM, ITPM, and OTPM per tier. Optimize your AI app to stay within limits and scale efficiently.
Er Rickow
Er Rickow
Tutorial
July 5, 2026Β·6 min read
Understanding RPM, ITPM, OTPM: API Rate Limits Explained
Share
Tutorial

Most AI APIs enforce a single rate limit: requests per minute. Neosantara uses a 3-dimensional rate limiting system (RPM, ITPM, OTPM) that gives you finer control over throughput, prevents accidental cost spikes, and ensures fair resource allocation across all users.

If you've ever hit a 429 error and wondered why, this guide explains exactly what each limit means, how they interact, and how to maximize your throughput within them.

Key Takeaways

  • RPM (Requests Per Minute) limits how many API calls you can make per minute, regardless of size
  • ITPM (Input Tokens Per Minute) limits how much context you can send to models each minute
  • OTPM (Output Tokens Per Minute) limits how much models can generate for you each minute
  • All three are enforced simultaneously. Hitting any one triggers a 429 response
  • Neosantara refunds your balance automatically when a request is rate-limited

What Are RPM, ITPM, and OTPM?

Traditional API providers give you one number: "500 requests per minute." That's simple but crude. A one-sentence classification request and a 100K-token document analysis both count as one request, even though they consume vastly different resources.

Neosantara splits rate limiting into three independent dimensions:

LimitFull NameWhat It ControlsWhy It Exists
RPMRequests Per MinuteNumber of API callsPrevents connection flooding
ITPMInput Tokens Per MinuteTotal input tokens sentPrevents context-window abuse
OTPMOutput Tokens Per MinuteTotal output tokens generatedPrevents compute-heavy generation spikes

Each limit uses a sliding window (RPM) or token bucket (ITPM/OTPM) algorithm, enforced per-user via Redis. When any single dimension hits its limit, the request returns HTTP 429 with headers telling you exactly which limit was hit and when it resets.

How the Three Limits Interact

Think of it like a highway with three toll gates. Your request must pass all three to proceed:

Your Request
    β”‚
    β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”     β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  RPM    β”‚ ──▢ β”‚  ITPM   β”‚ ──▢ β”‚  OTPM*  β”‚ ──▢ Model
β”‚ Check   β”‚     β”‚ Check   β”‚     β”‚ Settle  β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜     β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
                                 * settled post-response

RPM is checked first: have you exceeded your request count this minute?

ITPM is checked second: does your input (system prompt + messages + tools) exceed your token budget for this minute?

OTPM is special: it's settled after the response completes, because the output length isn't known upfront. If your cumulative output this minute exceeds the OTPM limit, your next request will be blocked until the window resets.

Rate Limits by Tier

Neosantara offers multiple tiers, each with progressively higher limits. All PAYG (Pay-As-You-Go) tiers use balance-based billing in Rupiah.

TierRPMITPMOTPMBest For
Free1530K8KTesting and prototyping
Basic50500K80KSide projects and MVPs
Standard1,0002M320KProduction applications
Pro2,0005M800KHigh-traffic SaaS products
Enterprise4,00010M1.6MMission-critical workloads

Coding Plan Tiers Work Differently

The flat-fee Coding Plan subscription (for tools like Claude Code, Cline, Continue) does not use RPM/ITPM/OTPM at all. Instead, it enforces two separate mechanisms:

  1. Concurrency limit β€” the maximum number of simultaneous in-flight requests
  2. Prompt quota β€” a rolling 5-hour and weekly budget of prompts, where each prompt can consume 1-3x quota depending on model load
TierMax Concurrent Requests5-Hour QuotaWeekly QuotaMonthly Price
CodingHemat4~80 prompts~400 promptsRp 149.000
CodingReguler10~400 prompts~2.000 promptsRp 499.000
CodingEksklusif20~1.600 prompts~8.000 promptsRp 1.290.000

This mirrors GLM's approach: instead of a hard per-minute request cap, the plan throttles how many requests can run at the same time, then hard-stops once your rolling prompt budget is exhausted. No overage, no balance deduction β€” the Coding Plan never draws on your PAYG balance.

What Happens When You Hit a Limit?

When any limit is exceeded, Neosantara returns a 429 Too Many Requests response with detailed headers:

HTTP/1.1 429 Too Many Requests
x-neosantara-ratelimit-requests-limit: 50
x-neosantara-ratelimit-requests-remaining: 0
x-neosantara-ratelimit-requests-reset: 2026-07-06T02:16:00.000Z
x-neosantara-ratelimit-input-tokens-limit: 500000
x-neosantara-ratelimit-input-tokens-remaining: 487231
x-neosantara-ratelimit-output-tokens-limit: 80000
x-neosantara-ratelimit-output-tokens-remaining: 72104

In this example, RPM hit zero while ITPM and OTPM still have budget. You know exactly which limit was the bottleneck.

Automatic Balance Refund

If you're on PAYG billing and a request is rate-limited, Neosantara automatically refunds any reserved balance. You never pay for requests that didn't complete. The refund happens instantly with a description "Refund: Rate Limit Blocked" in your transaction history.

How to Optimize Your Throughput

1. Batch smaller requests together

If you're making many small requests (classifications, extractions), you'll hit RPM before ITPM. Consider using the Batch API to submit hundreds of requests in one call.

2. Reduce input token waste

Large system prompts eat ITPM budget on every request. Strategies:

  • Use prompt caching to avoid re-sending repeated context
  • Trim conversation history to only relevant messages
  • Move static instructions to a shorter system prompt

3. Control output length

Set max_tokens appropriately. If you only need a yes/no classification, set max_tokens: 10 instead of leaving it at the model default (4,096+). This conserves your OTPM budget.

4. Use model routing strategically

Simple tasks routed to DeepSeek V4 Flash are typically faster, meaning you get responses back quicker and your per-minute utilization stays lower.

5. Implement exponential backoff

When you receive a 429, read the reset header and wait until that timestamp. Don't retry immediately.

import time
from openai import OpenAI, RateLimitError

client = OpenAI(
    base_url="https://api.neosantara.xyz/v1",
    api_key="nst-your-key"
)

def call_with_backoff(messages, max_retries=3):
    for attempt in range(max_retries):
        try:
            return client.chat.completions.create(
                model="deepseek-v4-flash",
                messages=messages
            )
        except RateLimitError as e:
            wait = 2 ** attempt  # 1s, 2s, 4s
            print(f"Rate limited. Waiting {wait}s...")
            time.sleep(wait)
    raise Exception("Max retries exceeded")

Monitoring Your Usage

Neosantara's dashboard shows real-time rate limit consumption. Every response also includes the remaining budget in headers, so you can build client-side awareness:

response = client.chat.completions.create(
    model="claude-sonnet-4-6",
    messages=[{"role": "user", "content": "Hello"}]
)

# Check remaining budget from response headers
# Available in the raw HTTP response
print(f"RPM remaining: {response.headers.get('x-neosantara-ratelimit-requests-remaining')}")
print(f"ITPM remaining: {response.headers.get('x-neosantara-ratelimit-input-tokens-remaining')}")

Frequently Asked Questions

Do rate limits reset at the start of each minute?

No. Neosantara uses a sliding window for RPM and token bucket for ITPM/OTPM. This means limits refill continuously rather than resetting at clock boundaries. A burst at 12:00:30 won't suddenly unlock at 12:01:00.

Are limits per API key or per user?

For the standard PAYG API, limits are enforced per API key β€” each key gets its own independent RPM/ITPM/OTPM budget at your account's tier. If you have 3 keys and your tier allows 50 RPM, each key gets its own 50 RPM, not a shared pool. The exception is MCP endpoints (/v1/mcp), where a user-level aggregate limit applies on top of the per-key limit.

Can I get higher limits?

Yes. Upgrade your tier through the Neosantara dashboard. Enterprise limits (4,000 RPM, 10M ITPM) are available for high-volume applications.

Why is OTPM settled after the response?

Because the model hasn't generated its output yet when the request arrives. Neosantara can check RPM and estimate ITPM upfront, but OTPM is only known after generation completes. If your OTPM is exhausted, the current response still completes (you're never cut off mid-stream), but subsequent requests will be blocked until the window refills.

Do Coding Plan limits stack with PAYG?

No. Coding Plan users operate under their subscription's concurrency limit and rolling prompt quota (not RPM/ITPM/OTPM). If you also have PAYG balance, it's used for models outside your plan's allowed list, under your base tier's standard rate limits.

Conclusion

Neosantara's 3-dimensional rate limiting gives you predictable, fair access to AI models:

  • RPM prevents connection storms
  • ITPM prevents context-window abuse
  • OTPM prevents compute-heavy generation spikes

Together, they ensure your application scales smoothly from prototype to production. Monitor your usage via response headers and the dashboard, implement backoff on 429s, and use the Batch API for workloads that don't need real-time responses.

Ready to build? Get your API key β†’


Sources: