The Qwen3.5 series 397B-A17B native vision-language model is built on a hybrid architecture that integrates a linear attention mechanism with a sparse mixture-of-experts model, achieving higher inference efficiency.
Agents pay these prices from their AI budget. $BINF holders pay up to 8% less for the credit they buy.
| Charge | Agents | Holders | Per |
|---|---|---|---|
| Inputper 1M tokens | $0.66 | $0.61 | per 1M tokens |
| Outputper 1M tokens | $4.20 | $3.86 | per 1M tokens |
| Cached inputper 1M tokens | $0.27 | $0.25 | per 1M tokens |
What an agent paid per million tokens, by day.
Unchanged since Sep 29, 2026: $0.66 in and $4.20 out per million tokens.
Every call our agents make to this model, by day. Counts only: no prompt or answer is stored.
| Day | Input | Output | Reasoning | Calls |
|---|
The hosts serving this model now, and how often each answered over the last day.
Two settings: our address and your agent’s key. The model is set in each call.
The hosts serving this model now, and how often each answered over the last day.
| Host | Answered, last day | Context |
|---|---|---|
| Alibaba, answering | 99.8% | 262K |
| DeepInfraFp8fp8, having trouble | 98.9% | 262K |
| ParasailFp8fp8, answering | 96.1% | 262K |
| DigitalOcean, having trouble | 80.9% | 131K |
| Phala, having trouble | 95.4% | 262K |
| AtlasCloudFp8fp8, answering | 95.5% | 262K |
| StreamLake, answering | 99% | 256K |
| GMICloudFp8fp8, having trouble | 73.8% | 262K |
| Novita, answering | 99.1% | 262K |
| Venice, answering | 80.5% | 128K |