developer preview · api.earthruntime.com

Qwen and GPT-OSS inference without marketplace overhead.

We operate the hardware ourselves. You get direct pricing, predictable performance, and full visibility into the model you run.

100,000 free tokens. No card required. A 6-digit code is emailed; your key is shown here once.

Code sent to

Didn't get it? Send a new code — it expires in 15 minutes.

Generating your key…
This usually takes about 30 seconds. Keep this page open.
Check your inbox
This address already joined the beta, so your key is on its way to by email.
Your key, shown once

It's already pasted into the request on the right. Run it.

first request · bash
curl https://api.earthruntime.com/v1/chat/completions \
  -H "Authorization: Bearer $EARTHRUNTIME_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "qwen3.6-35b",
    "messages": [{"role": "user", "content": "Explain why tail latency matters for production inference in three sentences."}]
  }'
Trouble pasting? Use the single-line version.
curl https://api.earthruntime.com/v1/chat/completions -H "Authorization: Bearer $EARTHRUNTIME_KEY" -H "Content-Type: application/json" -d '{"model":"qwen3.6-35b","messages":[{"role":"user","content":"Explain why tail latency matters for production inference in three sentences."}]}'

Works with your existing OpenAI client — change the base URL and the key, keep everything else. Read the docs →

01 models

One key, one endpoint. Change the model id to switch.

Every model is served from hardware we own and operate. What you see on the card is what runs — precision, context window, and whether reasoning is on.

qwen3.6-35b

Qwen 3.6 35B

Parameters
35B
Context
262K
Precision
Qwen official FP8
Reasoning
off by default, toggle per request
Price / 1M
$0.10 in · $0.90 out

Served from Qwen's own official FP8 checkpoint. We do not requantize it ourselves. Reasoning is disabled by default for lower latency; enable it when the workload benefits.

qwen3.8-27b

Qwen 3.8 27B

Parameters
27B
Context
262K
Precision
Qwen official FP8, FP8 KV
Reasoning
off by default, toggle per request
Price / 1M
$0.24 in · $2.20 out

The smaller Qwen for cost-sensitive, high-volume work.

gpt-oss-120b

GPT-OSS 120B

Parameters
120B (MoE)
Context
131K
Precision
checkpoint-native MXFP4, FP8 KV
Reasoning
on by default, reasoning_effort
Price / 1M
$0.03 in · $0.17 out

OpenAI's open-weight model, served without a middleman.

coming soon deepseek-v4-flash

DeepSeek V4 Flash

Context
262K
Precision
FP8 dense, MXFP4 experts
Price / 1M
$0.30 in · $0.90 out

Not yet available; the rate above applies at launch. Email contact@earthruntime.com to be told when it opens.

02 benchmarks

Every request served. No 429s.

Measured from a neutral vantage outside our network, at 64 concurrent requests, on 2026-09-24. The exact script, workload and error definition are on the methodology page. Run it against us and against anyone else.

Our fleet at 64 concurrent requests. 256 requests per run, median of 3 runs, max_tokens 128, from an AWS EC2 t3.xlarge in us-east-2c on 2026-09-24.
ModelThroughputTTFT p50 at c=64Latency p50Latency p95Errors
qwen3.6-35b2,464 tok/s765 ms2.47 s3.56 s0%
gpt-oss-120b2,276 tok/s537 ms3.11 s4.57 s0%
qwen3.8-27b1,608 tok/s508 ms4.70 s8.59 s0%
Zero errors across all 3,104 requests. TTFT is quoted at 64 concurrent streams, which is where it was measured; at a single stream it is 73 ms. Why the gap.
Head to head on qwen3.6-35b, same run, same vantage, same workload. Higher throughput and lower error rate are better.
EndpointThroughputLatency p50Latency p95Errors
earthruntime2,464 tok/s2.47 s3.56 s0%
OpenRouter, default routing1,534 tok/s1.65 s4.68 s0%
OpenRouter to AtlasCloud1,327 tok/s1.87 s3.18 s0%
OpenRouter to Parasail3,042 tok/s2.18 s2.74 s17.6%
Read those two right-hand columns together. Parasail posts the highest raw throughput because it returns 429 to 17.6% of the load and finishes a smaller job sooner. OpenRouter's default routing is genuinely faster than us on the median; we are ahead on throughput, on the p95 tail, and on serving every request we were sent.

Read the methodology and reproduce it

03 how it works

Three steps. No SDK to learn.

If your code already talks to OpenAI, it already talks to us.

1

Get a key

Enter your email at the top of this page. A 6-digit code arrives, you paste it, and the key is shown once — with 100,000 free tokens already on it.

2

Point your OpenAI client at the base URL

Set base_url to https://api.earthruntime.com/v1 and pass the key. Request and response shapes are unchanged, streaming included.

3

Swap the model id

Use qwen3.6-35b, qwen3.8-27b or gpt-oss-120b. Nothing else in the call changes.

python · openai ≥ 1.0
from openai import OpenAI

client = OpenAI(
    base_url="https://api.earthruntime.com/v1",   # step 2
    api_key=os.environ["EARTHRUNTIME_KEY"],
)

stream = client.chat.completions.create(
    model="qwen3.6-35b",                          # step 3
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
node · openai ≥ 4
import OpenAI from "openai";

const client = new OpenAI({
  baseURL: "https://api.earthruntime.com/v1",
  apiKey: process.env.EARTHRUNTIME_KEY,
});

const r = await client.chat.completions.create({
  model: "gpt-oss-120b",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(r.choices[0].message.content);
04 pricing

Pay per token. No subscription, no minimum.

Every rate on this page is plain HTML — nothing is loaded by script. The same numbers are published as /pricing.md for machines.

Per-model rates, USD per 1M tokens. Exclusive of applicable taxes.
ModelInput / 1MOutput / 1MContextNotes
qwen3.6-35b$0.10$0.90262KQwen official FP8; reasoning off by default
qwen3.8-27b$0.24$2.20262KQwen official FP8; reasoning off by default
gpt-oss-120b$0.03$0.17131KMXFP4; reasoning on by default
deepseek-v4-flash$0.30$0.90262KNot yet available; rate applies at launch
Rates are per million tokens and are billed on the usage block of each response. Notice periods for rate changes are in the terms.

Free tier

100,000 free tokens for new users. No card required. Get a key at the top of the page — it arrives with the tokens already on it.

Custom volume

Need more than the packs cover, or a committed rate? contact@earthruntime.com.

Credit packs

Checkout cancelled — nothing was charged. Pick a pack whenever you're ready.

Credits are added to the key tied to the email you enter. No subscription, no account setup.

Checkout is handled by Stripe. Token estimates assume a typical input/output mix.

05 infrastructure

Direct infrastructure changes the economics.

A marketplace resells someone else's GPUs and takes a margin on every token. We operate inference hardware in cities and look for local uses for its physical outputs — so the price you pay is ours, and the performance you see is ours to keep.

01 · compute

Inference hardware sited in cities, run by us, no intermediary.

02 · capture

The first system pairs that compute with direct-air carbon capture.

03 · CO2

Recovered CO2 is bottled and supplied to nearby hospitality businesses.

04 · the city

Faster inference is the product today. A more useful relationship between cities and compute is what we're building toward.

Yes: the carbonation in your soda at a local bar can come from the machine that answered your API call.

06 tested on real workloads

What developers found

“There were no latency issues using Qwen.”
Agentic coding

A developer tested Qwen on a Rust DSP mastering tool containing a deliberately misleading hardcoded ID. The model attempted the obvious fix, saw its tests fail, backtracked, identified the underlying cause, and produced a patch that two other models independently reviewed.

From a private evaluation. We have not published the name.
“Much better than the OpenRouter options we tested.”
Production spam-call classification

After benchmarking earthruntime against OpenRouter-hosted alternatives in production, the team made earthruntime its primary provider and retained the marketplace as an outage fallback.

From a private evaluation. We have not published the name.
07 faq

Questions developers actually ask.

Which models does earthruntime serve?

Qwen 3.6 35B (qwen3.6-35b), Qwen 3.8 27B (qwen3.8-27b), and GPT-OSS 120B (gpt-oss-120b). DeepSeek V4 Flash is coming soon.

Is the API OpenAI-compatible?

Yes. Point your existing OpenAI client at https://api.earthruntime.com/v1 with your earthruntime API key; request and response formats stay the same.

Is there a free tier?

Yes — 100,000 free tokens to start, no card required.

How is earthruntime different from a model marketplace?

We operate the inference hardware ourselves, so you get direct pricing, predictable performance, and full visibility into the model you are running — no marketplace overhead.

Are the models quantized?

Qwen 3.6 35B is served from Qwen's own official FP8 checkpoint, at its full 262K context. Nothing is requantized by us, and the precision each model is served at is listed on its model card.

How do I add credit, and is there a subscription?

No subscription. Buy a credit pack ($1, $5 or $20) with the email address your key is tied to; credits land within a minute. For higher volumes, email contact@earthruntime.com.

Where does the hardware run?

In cities, on hardware we operate ourselves. Our first system pairs the compute with direct-air carbon capture and supplies the recovered CO2 to nearby hospitality businesses.