Introducing Bonsai-27B on Neosantara: 262k Context

Neosantara now offers direct API access to PrismML Bonsai-27B, the multimodal flagship of the Bonsai family. One API call gets you 27-billion-parameter reasoning, image understanding, and function calling behind a 262,144-token context window, all through our unified OpenAI-compatible endpoint with Rupiah billing.
Key Takeaways
- PrismML compresses Bonsai-27B with end-to-end 1.58-bit ternary quantization, making it one of the most efficient models in its class.
- Published benchmarks retain 94.6% of FP16 baseline quality (Hugging Face, July 2026) [2].
- 262K-token context window, native vision, and structured tool calling out of the box.
- Available today, free for a limited time, for every Neosantara account.
Why Bonsai-27B?
Most capable models demand serious hardware. Bonsai-27B takes a different path: PrismML applies its end-to-end ternary {-1, 0, +1} quantization to the Qwen3.6-27B backbone, covering embeddings, attention, MLPs, and the LM head with no higher-precision escape hatches, so a 27-billion-parameter model fits where larger setups cannot go. It is the first model in its capability class small enough to run on a phone (PrismML, July 2026) [1], and it ships under the Apache 2.0 license.
What you get on Neosantara:
- 262K-token context: entire codebases, contract stacks, or thousands of log lines in one prompt.
- Native vision: diagram OCR and screenshot understanding through a built-in vision tower.
- Structured tool calling: multi-step agent loops without extra scaffolding.
- OpenAI compatibility: drop-in for LangChain, LlamaIndex, Cursor, or raw SDK workflows.
During the launch window, Bonsai-27B is free for everyone on Neosantara for a limited time (3 RPM on pure Free tier, unlocked to full 10 RPM with an initial deposit of Rp 15.000 or higher), so you can test reasoning, vision, and agent loops before committing a budget. Current rates after the preview are on the pricing page.
The numbers back it up. Across PrismML's 15-benchmark suite, Ternary Bonsai 27B scores 80.5 overall against 85.0 for the full-precision Qwen3.6 baseline, a 95% retention rate with math and coding nearly untouched (PrismML, July 2026) [1]:

Fig: Intelligence density (per GB) of Bonsai 27B versus other models in its parameter class. (Source: PrismML, 2026)
Getting Started
If you have used the OpenAI SDK, integration takes minutes:
from openai import OpenAI
client = OpenAI(
base_url="https://api.neosantara.xyz/v1",
api_key="YOUR_NEOSANTARA_API_KEY"
)
# Text & Deep Reasoning
response = client.chat.completions.create(
model="bonsai-27b",
messages=[
{"role": "system", "content": "You are a helpful and rigorous analytical AI assistant."},
{"role": "user", "content": "Analyze the time complexity trade-offs of B-Tree vs LSM-Tree for write-heavy storage engines."}
],
temperature=0.6,
max_tokens=1000
)
print(response.choices[0].message.content)Browse the model gallery for the full lineup, or see how Kimi K2 handles trillion-parameter agentic workloads.
How Do You Plug Bonsai-27B into Coding Agents?
Because the endpoint speaks both OpenAI and Anthropic dialects, terminal coding agents work without wrappers.
OpenCode picks up Neosantara as a custom provider in opencode.json (custom provider docs):
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"neosantara": {
"npm": "@ai-sdk/openai-compatible",
"name": "Neosantara",
"options": {
"baseURL": "https://api.neosantara.xyz/v1"
},
"models": {
"bonsai-27b": { "name": "Bonsai 27B" }
}
}
}
}Add your key with the /connect command or export OPENAI_API_KEY=..., then pick the model with /models.
Claude Code routes through the Anthropic-compatible base URL:
export ANTHROPIC_BASE_URL="https://api.neosantara.xyz/anthropic"
export ANTHROPIC_AUTH_TOKEN="YOUR_NEOSANTARA_API_KEY"
export ANTHROPIC_API_KEY=""
export ANTHROPIC_MODEL="bonsai-27b"The empty ANTHROPIC_API_KEY matters: Claude Code checks it first and would otherwise ignore your auth token (Anthropic-compatible API, Neosantara docs). Per-tool setup for Cursor, Cline, and friends lives in the AI Coding Tools guide.
Agent frameworks with native Neosantara support need only a few lines each:
# LiteLLM (native provider, reads NEOSANTARA_API_KEY)
import litellm
resp = litellm.completion(
model="neosantara/bonsai-27b",
messages=[{"role": "user", "content": "Design a rate limiter for a multi-tenant API."}],
)
# Agno
from agno.agent import Agent
from agno.models.neosantara import Neosantara
agent = Agent(model=Neosantara(id="bonsai-27b"))
agent.print_response("Summarize this stack trace and propose a fix.", stream=True)
# any-llm
from any_llm import completion
resp = completion(
"neosantara/bonsai-27b",
messages=[{"role": "user", "content": "Review this diff for race conditions."}],
)How Fast Is Bonsai-27B on Neosantara?
We ran our own benchmark through the production gateway on August 21, 2026, five runs per scenario. Warm requests produce first tokens in about 6 seconds for short prompts, rising to about 9 seconds with a 9,700-token input, while sustaining roughly 60 to 80 output tokens per second end to end:
| Scenario | Median first token | Effective throughput |
|---|---|---|
| Short prompt (~50 input tokens) | ~6.1 s | ~64 tok/s |
| Medium prompt (~500 input tokens) | ~7.9 s | ~76 tok/s |
| Long prompt (~9,700 input tokens) | ~9.3 s | ~64 tok/s |
Most of that waiting is not idle time. Bonsai-27B is a thinking model: it spends 400 to 800 hidden reasoning tokens working through the problem before the visible answer streams, which is exactly where its benchmark strength comes from. You can also tune or constrain reasoning intensity using reasoning_effort ("low", "medium", or "high") or thinking_budget_tokens in your payload:
# Fine-tune thinking depth & budget (Reasoning Control)
response = client.chat.completions.create(
model="bonsai-27b",
messages=[{"role": "user", "content": "Explain event-driven microservices architecture."}],
extra_body={
"reasoning_effort": "low", # Options: 'low', 'medium', 'high'
# OR specify explicit token budget:
# "thinking_budget_tokens": 512 # 512 (Low), 2048 (Medium), -1 (Unconstrained)
},
max_tokens=1500 # Recommend >= 1000 so text output is not truncated during reasoning
)💡 max_tokens Tip: Because Bonsai-27B requires token budget for its internal reasoning trace (Chain-of-Thought) before writing the answer, always allocate
max_tokens >= 1000. Setting it too low (e.g.max_tokens: 100) may consume all tokens during thinking, yieldingcontent: null.
Under the hood, Neosantara's live Bonsai-27B inference and production benchmarks are powered by Daytona Spot GPU sandboxes on NVIDIA RTX PRO 6000 (96GB VRAM) hardware with automated scale-to-zero orchestration, achieving enterprise-grade throughput with zero persistent idle compute costs. Plan for the reasoning budget when setting max_tokens, and see our guide to understanding rate limits if you are sizing throughput for production. Expect one cold-start request of around 74 seconds after a fresh deploy.
Vision and Tool Calling
Image input works through the same chat interface, supporting both public URLs and Base64 encoded strings (data:image/png;base64,...). Pass an image alongside your prompt and Bonsai-27B reads diagrams, screenshots, and documents [3]:
# Multimodal Image Analysis
response = client.chat.completions.create(
model="bonsai-27b",
messages=[{
"role": "user",
"content": [
{"type": "text", "text": "Extract and explain the architecture diagram in this image:"},
{"type": "image_url", "image_url": {"url": "https://example.com/system-architecture.png"}}
]
}]
)Tool calling uses the same OpenAI function-calling schema you already know:
# Structured Tool Calling
response = client.chat.completions.create(
model="bonsai-27b",
messages=[{"role": "user", "content": "What is the weather in Jakarta right now?"}],
tools=[{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get current weather for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"]
}
}
}]
)
call = response.choices[0].message.tool_calls[0]
print(call.function.name, call.function.arguments)Frequently Asked Questions
A standard 27B model in FP16 format requires more than 54 GB of memory. Ternary Bonsai-27B stores weights at an effective 1.71 bits per weight, coming in at roughly 7.2 GB when deployed, roughly a 9x reduction.
PrismML's evaluation scores Ternary Bonsai 27B at 80.49 against 85.07 for the FP16 baseline, preserving 94.6% of baseline quality. Standard 2-bit quantization on the same backbone retains only 85.5%.
Yes. Weights are published under Apache 2.0 on Hugging Face in GGUF and MLX formats. Note that the ternary kernels require PrismML's llama.cpp fork rather than a stock build, alongside MLX and CUDA options, while the Neosantara API needs no local hardware.
Source References
- [1] PrismML, "Announcing Bonsai 27B: The First 27B-Class Model to Run on a Phone," July 2026. URL: https://prismml.com/news/bonsai-27b (retrieved August 21, 2026)
- [2] PrismML, "prism-ml/Ternary-Bonsai-27B-gguf Model Card," Hugging Face, July 2026. URL: https://huggingface.co/prism-ml/Ternary-Bonsai-27B-gguf (retrieved August 21, 2026)
- [3] PrismML, "Bonsai 27B: Multimodal & Tool Use Docs," PrismML Docs, 2026. URL: https://docs.prismml.com/models/bonsai-27b (retrieved August 21, 2026)
Try Bonsai-27B Free for a Limited Time
Sign up for a free Neosantara AI account and put Bonsai-27B's 262k context, vision, and tool calling behind one OpenAI-compatible endpoint at no cost during the launch window.



