Local Deployment
Running the local binary has no per-token cost. Once you download the binary, you own it — no phone-home, no metering, no internet connection required.- Zero per-token cost. Every token you generate is free.
- No usage-based billing. There is no meter running in the background.
- Pay for your GPU once; run unlimited tokens. A single A10 delivers 312 tokens/sec on Llama 3.3 70B.
- No internet required after download. The binary has weights embedded and operates fully offline.
Cloud API Pricing
The table below shows current cloud API rates alongside the lowest price available for the same model elsewhere.Prices decrease as the compiler finds new efficiencies and those savings are passed on to you directly. Prices never increase without at least 30 days of advance notice.
Free Credits
Every new account receives $20 in credits — approximately 1 million tokens on Llama 3.3 70B Instruct. No credit card is required to sign up and start using the API. Credits are applied automatically to your first requests.Billing
Cloud API usage is billed per token for text models and per minute of audio for Whisper. You can monitor your usage and remaining balance in the billing dashboard at any time. A few things to keep in mind:- Usage is tallied in real time and visible in your dashboard immediately.
- Accounts with an outstanding balance may be suspended until payment is made.
- You can set a monthly spend cap in the dashboard to avoid unexpected charges.