Product · Tinfield 1 + inference gateway
Tinfield 1 on Runtime
Our open-weight flagship for terminal work and long-horizon software engineering.
The BF16 reference scored 33.0 on Terminal-Bench 4.0 and 62.0 on DeepSWE v1.1 across the full task sets at k=5. Runtime serves Tinfield 1 Compact at $0.15/M input tokens and $0.60/M output tokens. The served Compact build has not yet been rerun on those full benchmark suites.
from openai import OpenAI
client = OpenAI(
base_url="https://api.rntm.sh/v1",
api_key=BTL_KEY,
)
client.chat.completions.create(
model="tinfield-1",
messages=[{"role": "user",
"content": "ship it"}],
)How it cuts spend
A token-efficiency layer,
not just a router.
Routing to a cheaper equivalent upstream is only half of it. The runtime also sends fewer billable tokens, reuses the ones it must send, and avoids doing the same work twice. You keep the model boundary you chose; we cut the waste before the request reaches it.
What teams get
Switch the base URL,
keep the product.
Best fit for teams already shipping AI products and feeling real spend or latency pressure. No exact-vendor lock-in — ask for a specific provider when you need it, let the gateway choose when you don't.
Launch Runtime →Customer API surface
The routes that
actually matter.
Most traffic only ever touches two of these. The rest are for keys, usage, and the catalog. No /v1/admin/* or ops-only auth in the customer path.
Stop paying for token waste.
Tell us your stack, providers, traffic, and constraints. We'll get you a key.