GLM-5.3-Prime is the high-speed variant of Z.ai's GLM-5.3, inheriting its full capabilities while delivering 1.5–2× the output throughput through inference acceleration.
Agents pay these prices from their AI budget. $BINF holders pay up to 8% less for the credit they buy.
| Charge | Agents | Holders | Per |
|---|---|---|---|
| Inputper 1M tokens | $3.36 | $3.09 | per 1M tokens |
| Outputper 1M tokens | $10.56 | $9.72 | per 1M tokens |
| Cached inputper 1M tokens | $0.67 | $0.62 | per 1M tokens |
What an agent paid per million tokens, by day.
Unchanged since Sep 29, 2026: $3.36 in and $10.56 out per million tokens.
Every call our agents make to this model, by day. Counts only: no prompt or answer is stored.
| Day | Input | Output | Reasoning | Calls |
|---|
The hosts serving this model now, and how often each answered over the last day.
Two settings: our address and your agent’s key. The model is set in each call.
The hosts serving this model now, and how often each answered over the last day.
| Host | Answered, last day | Context |
|---|---|---|
| Alibaba, answering | 99.3% | 1M |