API Model List and Pricing
| Model Type | Model Name (API Call Parameter Value) | Backend Model Version | Model Source | Prompt (1M token) | Completion (1M token) | Context | Brief Model Capability Description |
|---|---|---|---|---|---|---|---|
| Chat | gpt-5.6-luna | gpt-5.6-luna | External | 1.4662 | 8.7996 | 1M | High cost-effective chat model, suitable for general-purpose scenarios |
| Chat | gpt-5.6-terra | gpt-5.6-terra | External | 14.662 | 87.996 | 1M | High-performance chat model, suitable for complex reasoning tasks |
| Chat | gpt-4o | gpt-4o-2024-08-06 | External | 18.328 | 73.310 | 128K | Multimodal flagship model, supports image-text understanding |
| Chat | gpt-4o-mini | gpt-4o-mini | External | 1.110 | 4.399 | 128K | Lightweight and fast chat model, suitable for high-frequency calls |
| Embedding | text-embedding-3-large | text-embedding-3-large | External | 0.95310 | / | 8K | High-precision text vectorization, suitable for semantic retrieval |
| Embedding | text-embedding-3-small | text-embedding-3-small | External | 0.14663 | / | 8K | Lightweight text vectorization, suitable for large-scale embedding |
| Chat | GLM-5.2 | GLM-5.2 | On-premises | 4.00 | 14.00 | 256K | Bilingual Chinese-English chat model, suitable for Chinese-language scenarios |
| Chat | Kimi-K3 | Kimi-K3 | On-premises | 10.00 | 50.00 | 128K | Strong long-text processing capability, suitable for document analysis |
| Chat | Qwen3.6-35B-A3B | Qwen3.6-35B-A3B | On-premises | 0.90 | 5.40 | 256K | Lightweight and efficient chat model, suitable for resource-constrained scenarios |
| Chat | DeepSeek-V4-Pro | DeepSeek-V4-Pro | On-premises | 2.25 | 6.75 | 64K | Professional reasoning model, suitable for coding and math tasks |
| Chat | DeepSeek-V4-Flash | DeepSeek-V4-Flash | On-premises | 0.75 | 2.25 | 1M | Ultra-fast lightweight chat model, suitable for real-time interaction |