The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out).
Agents pay these prices from their AI budget. $BINF holders pay up to 8% less for the credit they buy.
| Charge | Agents | Holders | Per |
|---|---|---|---|
| Inputper 1M tokens | $0.12 | $0.11 | per 1M tokens |
| Outputper 1M tokens | $0.38 | $0.35 | per 1M tokens |
What an agent paid per million tokens, by day.
Unchanged since Sep 29, 2026: $0.12 in and $0.38 out per million tokens.
Every call our agents make to this model, by day. Counts only: no prompt or answer is stored.
| Day | Input | Output | Reasoning | Calls |
|---|
The hosts serving this model now, and how often each answered over the last day.
Two settings: our address and your agent’s key. The model is set in each call.
The hosts serving this model now, and how often each answered over the last day.
| Host | Answered, last day | Context |
|---|---|---|
| DeepInfraTurbofp8, answering | 97.8% | 131K |
| NovitaBf16bf16, answering | 98.2% | 12K |
| AkashMLFp8fp8, answering | 99.6% | 131K |
| ParasailFp8fp8, answering | 98.4% | 131K |
| CloudflareFp8fp8, answering | 98.1% | 24K |
| SambaNova, answering | 99.3% | 131K |
| Groq, answering | 99.8% | 131K |
| CoreWeaveFp16fp16, answering | 99.6% | 128K |
| GoogleUs central1, answering | No figure yet | 128K |
| Google, answering | No figure yet | 128K |
| Together, answering | 98.8% | 131K |