This usually takes about 30 seconds. Keep this page open.
Check your inbox
This address already joined the beta, so your key is on its way to by email.
Your key, shown once
It's already pasted into the request on the right. Run it.
first request · bash
curl https://api.earthruntime.com/v1/chat/completions \
-H "Authorization: Bearer $EARTHRUNTIME_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.6-35b",
"messages": [{"role": "user", "content": "Explain why tail latency matters for production inference in three sentences."}]
}'
Trouble pasting? Use the single-line version.
curl https://api.earthruntime.com/v1/chat/completions -H "Authorization: Bearer $EARTHRUNTIME_KEY" -H "Content-Type: application/json" -d '{"model":"qwen3.6-35b","messages":[{"role":"user","content":"Explain why tail latency matters for production inference in three sentences."}]}'
Works with your existing OpenAI client — change the base URL and the key, keep everything else. Read the docs →
01 models
One key, one endpoint. Change the model id to switch.
Every model is served from hardware we own and operate. What you see on the card is what runs — precision, context window, and whether reasoning is on.
qwen3.6-35b
Qwen 3.6 35B
Parameters
35B
Context
262K
Precision
Qwen official FP8
Reasoning
off by default, toggle per request
Price / 1M
$0.10 in · $0.90 out
Served from Qwen's own official FP8 checkpoint. We do not requantize it ourselves. Reasoning is disabled by default for lower latency; enable it when the workload benefits.
qwen3.8-27b
Qwen 3.8 27B
Parameters
27B
Context
262K
Precision
Qwen official FP8, FP8 KV
Reasoning
off by default, toggle per request
Price / 1M
$0.24 in · $2.20 out
The smaller Qwen for cost-sensitive, high-volume work.
gpt-oss-120b
GPT-OSS 120B
Parameters
120B (MoE)
Context
131K
Precision
checkpoint-native MXFP4, FP8 KV
Reasoning
on by default, reasoning_effort
Price / 1M
$0.03 in · $0.17 out
OpenAI's open-weight model, served without a middleman.
coming soondeepseek-v4-flash
DeepSeek V4 Flash
Context
262K
Precision
FP8 dense, MXFP4 experts
Price / 1M
$0.30 in · $0.90 out
Not yet available; the rate above applies at launch. Email contact@earthruntime.com to be told when it opens.
02 benchmarks
Every request served. No 429s.
Measured from a neutral vantage outside our network, at 64 concurrent requests, on 2026-09-24. The exact script, workload and error definition are on the methodology page. Run it against us and against anyone else.
Our fleet at 64 concurrent requests. 256 requests per run, median of 3 runs, max_tokens 128, from an AWS EC2 t3.xlarge in us-east-2c on 2026-09-24.
Model
Throughput
TTFT p50 at c=64
Latency p50
Latency p95
Errors
qwen3.6-35b
2,464 tok/s
765 ms
2.47 s
3.56 s
0%
gpt-oss-120b
2,276 tok/s
537 ms
3.11 s
4.57 s
0%
qwen3.8-27b
1,608 tok/s
508 ms
4.70 s
8.59 s
0%
Zero errors across all 3,104 requests. TTFT is quoted at 64 concurrent streams, which is where it was measured; at a single stream it is 73 ms. Why the gap.
Head to head on qwen3.6-35b, same run, same vantage, same workload. Higher throughput and lower error rate are better.
Endpoint
Throughput
Latency p50
Latency p95
Errors
earthruntime
2,464 tok/s
2.47 s
3.56 s
0%
OpenRouter, default routing
1,534 tok/s
1.65 s
4.68 s
0%
OpenRouter to AtlasCloud
1,327 tok/s
1.87 s
3.18 s
0%
OpenRouter to Parasail
3,042 tok/s
2.18 s
2.74 s
17.6%
Read those two right-hand columns together. Parasail posts the highest raw throughput because it returns 429 to 17.6% of the load and finishes a smaller job sooner. OpenRouter's default routing is genuinely faster than us on the median; we are ahead on throughput, on the p95 tail, and on serving every request we were sent.
Your credits land within a minute, on the email address you paid with. Charged by Provocative Science Holdings, Inc.
Checkout cancelled — nothing was charged. Pick a pack whenever you're ready.
Credits are added to the key tied to the email you enter. No subscription, no account setup.
Checkout is handled by Stripe. Token estimates assume a typical input/output mix.
05 infrastructure
Direct infrastructure changes the economics.
A marketplace resells someone else's GPUs and takes a margin on every token. We operate inference hardware in cities and look for local uses for its physical outputs — so the price you pay is ours, and the performance you see is ours to keep.
01 · compute
Inference hardware sited in cities, run by us, no intermediary.
02 · capture
The first system pairs that compute with direct-air carbon capture.
03 · CO2
Recovered CO2 is bottled and supplied to nearby hospitality businesses.
04 · the city
Faster inference is the product today. A more useful relationship between cities and compute is what we're building toward.
Yes: the carbonation in your soda at a local bar can come from the machine that answered your API call.
06 tested on real workloads
What developers found
“There were no latency issues using Qwen.”
Agentic coding
A developer tested Qwen on a Rust DSP mastering tool containing a deliberately misleading hardcoded ID. The model attempted the obvious fix, saw its tests fail, backtracked, identified the underlying cause, and produced a patch that two other models independently reviewed.
“Much better than the OpenRouter options we tested.”
Production spam-call classification
After benchmarking earthruntime against OpenRouter-hosted alternatives in production, the team made earthruntime its primary provider and retained the marketplace as an outage fallback.
07 faq
Questions developers actually ask.
Which models does earthruntime serve?
Qwen 3.6 35B (qwen3.6-35b), Qwen 3.8 27B (qwen3.8-27b), and GPT-OSS 120B (gpt-oss-120b). DeepSeek V4 Flash is coming soon.
Is the API OpenAI-compatible?
Yes. Point your existing OpenAI client at https://api.earthruntime.com/v1 with your earthruntime API key; request and response formats stay the same.
Is there a free tier?
Yes — 100,000 free tokens to start, no card required.
How is earthruntime different from a model marketplace?
We operate the inference hardware ourselves, so you get direct pricing, predictable performance, and full visibility into the model you are running — no marketplace overhead.
Are the models quantized?
Qwen 3.6 35B is served from Qwen's own official FP8 checkpoint, at its full 262K context. Nothing is requantized by us, and the precision each model is served at is listed on its model card.
How do I add credit, and is there a subscription?
No subscription. Buy a credit pack ($1, $5 or $20) with the email address your key is tied to; credits land within a minute. For higher volumes, email contact@earthruntime.com.
Where does the hardware run?
In cities, on hardware we operate ourselves. Our first system pairs the compute with direct-air carbon capture and supplies the recovered CO2 to nearby hospitality businesses.