Understanding RPM, ITPM, OTPM: API Rate Limits Explained

Most AI APIs enforce a single rate limit: requests per minute. Neosantara uses a 3-dimensional rate limiting system (RPM, ITPM, OTPM) that gives you finer control over throughput, prevents accidental cost spikes, and ensures fair resource allocation across all users.
If you've ever hit a 429 error and wondered why, this guide explains exactly what each limit means, how they interact, and how to maximize your throughput within them.
Key Takeaways
- RPM (Requests Per Minute) limits how many API calls you can make per minute, regardless of size
- ITPM (Input Tokens Per Minute) limits how much context you can send to models each minute
- OTPM (Output Tokens Per Minute) limits how much models can generate for you each minute
- All three are enforced simultaneously. Hitting any one triggers a 429 response
- Neosantara refunds your balance automatically when a request is rate-limited
What Are RPM, ITPM, and OTPM?
Traditional API providers give you one number: "500 requests per minute." That's simple but crude. A one-sentence classification request and a 100K-token document analysis both count as one request, even though they consume vastly different resources.
Neosantara splits rate limiting into three independent dimensions:
| Limit | Full Name | What It Controls | Why It Exists |
|---|---|---|---|
| RPM | Requests Per Minute | Number of API calls | Prevents connection flooding |
| ITPM | Input Tokens Per Minute | Total input tokens sent | Prevents context-window abuse |
| OTPM | Output Tokens Per Minute | Total output tokens generated | Prevents compute-heavy generation spikes |
Each limit uses a sliding window (RPM) or token bucket (ITPM/OTPM) algorithm, enforced per-user via Redis. When any single dimension hits its limit, the request returns HTTP 429 with headers telling you exactly which limit was hit and when it resets.
How the Three Limits Interact
Think of it like a highway with three toll gates. Your request must pass all three to proceed:
Your Request
β
βΌ
βββββββββββ βββββββββββ βββββββββββ
β RPM β βββΆ β ITPM β βββΆ β OTPM* β βββΆ Model
β Check β β Check β β Settle β
βββββββββββ βββββββββββ βββββββββββ
* settled post-response
RPM is checked first: have you exceeded your request count this minute?
ITPM is checked second: does your input (system prompt + messages + tools) exceed your token budget for this minute?
OTPM is special: it's settled after the response completes, because the output length isn't known upfront. If your cumulative output this minute exceeds the OTPM limit, your next request will be blocked until the window resets.
Rate Limits by Tier
Neosantara offers multiple tiers, each with progressively higher limits. All PAYG (Pay-As-You-Go) tiers use balance-based billing in Rupiah.
| Tier | RPM | ITPM | OTPM | Best For |
|---|---|---|---|---|
| Free | 15 | 30K | 8K | Testing and prototyping |
| Basic | 50 | 500K | 80K | Side projects and MVPs |
| Standard | 1,000 | 2M | 320K | Production applications |
| Pro | 2,000 | 5M | 800K | High-traffic SaaS products |
| Enterprise | 4,000 | 10M | 1.6M | Mission-critical workloads |
Coding Plan Tiers Work Differently
The flat-fee Coding Plan subscription (for tools like Claude Code, Cline, Continue) does not use RPM/ITPM/OTPM at all. Instead, it enforces two separate mechanisms:
- Concurrency limit β the maximum number of simultaneous in-flight requests
- Prompt quota β a rolling 5-hour and weekly budget of prompts, where each prompt can consume 1-3x quota depending on model load
| Tier | Max Concurrent Requests | 5-Hour Quota | Weekly Quota | Monthly Price |
|---|---|---|---|---|
| CodingHemat | 4 | ~80 prompts | ~400 prompts | Rp 149.000 |
| CodingReguler | 10 | ~400 prompts | ~2.000 prompts | Rp 499.000 |
| CodingEksklusif | 20 | ~1.600 prompts | ~8.000 prompts | Rp 1.290.000 |
This mirrors GLM's approach: instead of a hard per-minute request cap, the plan throttles how many requests can run at the same time, then hard-stops once your rolling prompt budget is exhausted. No overage, no balance deduction β the Coding Plan never draws on your PAYG balance.
What Happens When You Hit a Limit?
When any limit is exceeded, Neosantara returns a 429 Too Many Requests response with detailed headers:
HTTP/1.1 429 Too Many Requests
x-neosantara-ratelimit-requests-limit: 50
x-neosantara-ratelimit-requests-remaining: 0
x-neosantara-ratelimit-requests-reset: 2026-07-06T02:16:00.000Z
x-neosantara-ratelimit-input-tokens-limit: 500000
x-neosantara-ratelimit-input-tokens-remaining: 487231
x-neosantara-ratelimit-output-tokens-limit: 80000
x-neosantara-ratelimit-output-tokens-remaining: 72104In this example, RPM hit zero while ITPM and OTPM still have budget. You know exactly which limit was the bottleneck.
Automatic Balance Refund
If you're on PAYG billing and a request is rate-limited, Neosantara automatically refunds any reserved balance. You never pay for requests that didn't complete. The refund happens instantly with a description "Refund: Rate Limit Blocked" in your transaction history.
How to Optimize Your Throughput
1. Batch smaller requests together
If you're making many small requests (classifications, extractions), you'll hit RPM before ITPM. Consider using the Batch API to submit hundreds of requests in one call.
2. Reduce input token waste
Large system prompts eat ITPM budget on every request. Strategies:
- Use prompt caching to avoid re-sending repeated context
- Trim conversation history to only relevant messages
- Move static instructions to a shorter system prompt
3. Control output length
Set max_tokens appropriately. If you only need a yes/no classification, set max_tokens: 10 instead of leaving it at the model default (4,096+). This conserves your OTPM budget.
4. Use model routing strategically
Simple tasks routed to DeepSeek V4 Flash are typically faster, meaning you get responses back quicker and your per-minute utilization stays lower.
5. Implement exponential backoff
When you receive a 429, read the reset header and wait until that timestamp. Don't retry immediately.
import time
from openai import OpenAI, RateLimitError
client = OpenAI(
base_url="https://api.neosantara.xyz/v1",
api_key="nst-your-key"
)
def call_with_backoff(messages, max_retries=3):
for attempt in range(max_retries):
try:
return client.chat.completions.create(
model="deepseek-v4-flash",
messages=messages
)
except RateLimitError as e:
wait = 2 ** attempt # 1s, 2s, 4s
print(f"Rate limited. Waiting {wait}s...")
time.sleep(wait)
raise Exception("Max retries exceeded")Monitoring Your Usage
Neosantara's dashboard shows real-time rate limit consumption. Every response also includes the remaining budget in headers, so you can build client-side awareness:
response = client.chat.completions.create(
model="claude-sonnet-4-6",
messages=[{"role": "user", "content": "Hello"}]
)
# Check remaining budget from response headers
# Available in the raw HTTP response
print(f"RPM remaining: {response.headers.get('x-neosantara-ratelimit-requests-remaining')}")
print(f"ITPM remaining: {response.headers.get('x-neosantara-ratelimit-input-tokens-remaining')}")Frequently Asked Questions
Do rate limits reset at the start of each minute?
No. Neosantara uses a sliding window for RPM and token bucket for ITPM/OTPM. This means limits refill continuously rather than resetting at clock boundaries. A burst at 12:00:30 won't suddenly unlock at 12:01:00.
Are limits per API key or per user?
For the standard PAYG API, limits are enforced per API key β each key gets its own independent RPM/ITPM/OTPM budget at your account's tier. If you have 3 keys and your tier allows 50 RPM, each key gets its own 50 RPM, not a shared pool. The exception is MCP endpoints (/v1/mcp), where a user-level aggregate limit applies on top of the per-key limit.
Can I get higher limits?
Yes. Upgrade your tier through the Neosantara dashboard. Enterprise limits (4,000 RPM, 10M ITPM) are available for high-volume applications.
Why is OTPM settled after the response?
Because the model hasn't generated its output yet when the request arrives. Neosantara can check RPM and estimate ITPM upfront, but OTPM is only known after generation completes. If your OTPM is exhausted, the current response still completes (you're never cut off mid-stream), but subsequent requests will be blocked until the window refills.
Do Coding Plan limits stack with PAYG?
No. Coding Plan users operate under their subscription's concurrency limit and rolling prompt quota (not RPM/ITPM/OTPM). If you also have PAYG balance, it's used for models outside your plan's allowed list, under your base tier's standard rate limits.
Conclusion
Neosantara's 3-dimensional rate limiting gives you predictable, fair access to AI models:
- RPM prevents connection storms
- ITPM prevents context-window abuse
- OTPM prevents compute-heavy generation spikes
Together, they ensure your application scales smoothly from prototype to production. Monitor your usage via response headers and the dashboard, implement backoff on 429s, and use the Batch API for workloads that don't need real-time responses.
Ready to build? Get your API key β
Sources:
- Neosantara API Documentation, "Rate Limits", https://docs.neosantara.xyz/en/rate-limits
- Anthropic, "Rate Limits (May 2026 update)", https://docs.anthropic.com/en/api/rate-limits
- Upstash, "Rate Limiting with Redis", https://upstash.com/docs/oss/sdks/ts/ratelimit



