Skip to main content
APIYI supports 400+ mainstream AI models. This page provides detailed model information, pricing, and usage instructions.
Enterprise-grade Professional and Stable AI Large Model API Hub All models are officially sourced and forwarded, with ~20% off pricing (combining top-up bonuses and exchange rate advantages), aggregating various excellent large models. No speed limits, no account ban risks, pay-as-you-go billing, long-term reliable service.
The following are currently stably supplied popular models. For the complete model list and real-time pricing, visit the APIYI Console Pricing Page or the on-site Model Pricing Overview.
Model Upgrade Recommendations: We recommend using the latest models for best performance, but please note:
  1. Initial instability is common: Newly launched models may experience slow responses, timeouts, or occasional errors due to limited compute capacity at the vendor — this typically stabilizes within days to weeks
  2. Check parameter compatibility: New models may introduce or change parameters (e.g., max_completion_tokens replacing max_tokens). Before upgrading, verify that your API parameters remain compatible with older models
  3. Always test before going live: Before deploying a new model to production, thoroughly validate it in a test environment to ensure output quality and API compatibility meet expectations
A ”—” in the context column means the vendor has not published a definitive figure; refer to the console and official documentation. All prices are per 1M tokens (input/output) and can be combined with top-up bonuses.

Model Categories

🤖 OpenAI Series

🆕 Latest Models

GPT-6 series supply: the default / svip official-relay groups match official pricing on all four billing items, and both Chat Completions and Responses are tested working (including function calling) — migrating from 5.6 only needs the model field changed. Astra is also available in the Codex_Reverse group at a 0.5 discount. For Sol / Luna, a single request above 272K context is billed entirely at the higher tier. See the GPT-6 Astra launch notes and GPT-6 Sol / Luna launch notes.
GPT Pro series (e.g. gpt-5.5-pro, gpt-5.4-pro) usage notes:
  1. /v1/responses endpoint only: cannot use /v1/chat/completions — switch your SDK / code to the Responses API before calling
  2. Very expensive: a single call may cost several dollars, and it is open to the SVIP group only to prevent accidental use on the Default group
  3. Not recommended for non-professional needs: use GPT-6 Sol or GPT-5.6 Sol for everyday tasks; Pro suits only research and top-tier work demanding extreme reasoning depth

✅ Stable / Classic

GPT-5 and newer series usage notes:
  1. Temperature parameter temperature must be set to 1 (only supports 1)
  2. Use max_completion_tokens instead of max_tokens
  3. Do not pass top_p parameter
Image and Video Generation Models have been moved to a dedicated page. Visit Image & Video Generation Models for the full list and pricing.

🎭 Claude Series (Anthropic)

🆕 Latest Models

Breaking changes in Opus 5.5 / Fable 5.1: forced tool calls (tool_choice of any / tool) return 400, thinking blocks are bound to the model, and editing history invalidates thinking blocks. Opus 5.5 additionally cannot turn thinking off, and its default effort drops from high to medium. Read the Opus 5.5 launch notes and Fable 5.1 launch notes before migrating.
30-day data retention compliance for Fable 5 / Fable 5.1 / Mythos 5: because of these models’ advanced capabilities, Bedrock and Anthropic run additional checks for high-risk misuse. For these two models, inputs and outputs are retained for 30 days to detect serious abuse — accessed by automated safety systems by default, with human review only when a potential harm is flagged. Retention happens on the AWS / Anthropic side; APIYI is a pure transparent proxy and retains nothing itself. Other Claude models are unaffected. See the Fable 5.1 launch notes.

✅ Stable / Classic

Latest: Claude Opus 5.5 beats Fable 5.1 on most coding and knowledge-work benchmarks while costing less than Opus 5 ($4/$20); Anthropic puts typical-workload cost about 40% below Opus 5. Sonnet 5 is close to Opus 4.8 but faster and cheaper. Stable: Opus 4.7 and Sonnet 4.6 are battle-tested for production workflows already in place; Haiku 4.5 offers 2x speed at great value.

🌟 Google Gemini Series

🆕 Latest Models

Note: Gemini 3 Pro Preview was discontinued on March 9, 2026. Please migrate to Gemini 3.1 Pro Preview or Gemini 3.8 Flash.

✅ Stable / Classic

Latest: Gemini 3.8 Flash and 3.7 Flash share the same price — well ahead of 3.6 Flash on coding and agent work at half the cost — and both native Gemini and OpenAI-compatible endpoints work. See the 3.8 Flash launch notes and 3.7 Flash launch notes. For high-concurrency, low-latency work, try 3.5 Flash-Lite. Stable: Gemini 2.5 Pro (2M context) and Gemini 2.5 Flash suit settled production environments.

🚀 xAI Grok Series

🆕 Latest Models

✅ Stable / Classic

Estimate Grok costs per task: Grok 4.5 / 4.6 / 4.7 share the same unit price ($2/$6), but Grok 4.7 uses about twice the output tokens of 4.6 on the same task set, so per-task spend goes up; Grok 4.5 is the most output-efficient. Requests above 200K tokens are billed at the vendor’s second tier. See the Grok 4.7 launch notes.

🔍 DeepSeek Series

🐘 Chinese Model Series

Zhipu AI (GLM)

🆕 Latest: GLM-5.3, GLM-5.3-Flash | ✅ Stable / Classic: GLM-5.2, GLM-5.1, GLM-5
GLM-5.3 series highlights:
  • The flagship GLM-5.3 keeps 5.2’s base and only scales post-training, with big gains in coding and cybersecurity (ExploitBench 54.4% vs 24.4% for 5.2)
  • GLM-5.3-Flash is a 320B-A18B hybrid sparse + linear attention multimodal model at roughly 1/10 the flagship price
  • Pricing matches Zhipu’s official site item by item, available in the default / svip groups; raise your timeout above 600 seconds for long tasks. See the GLM-5.3 launch notes

Alibaba Qwen

🆕 Latest: Qwen3.8-Max | ✅ Stable / Classic: Qwen3.7 series, Qwen3.6 series, Qwen Max / Plus / Turbo

Moonshot Kimi Series

🆕 Latest: Kimi K3 | ✅ Stable / Classic: Kimi K2.6, K2.5, K2
Kimi K3 usage notes:
  • Exactly 1,048,576 tokens of context with flat pricing across the whole range — drop in whole repos or long documents without worrying about tier jumps
  • Max output defaults to 131,072 and can be raised to 1,048,576 tokens; thinking runs at full tilt, so budget output tokens accordingly
  • Automatic context caching bills cached input at $0.30/1M (1/10 of the uncached rate), a big win for multi-turn conversations

ByteDance Seed Series

🆕 Latest: Seed 2.1 Turbo | ✅ Stable / Classic: Seed 2.0 Pro / Lite / Mini
Seed 2.1 Turbo enables deep thinking by default: with no parameters set, even a one-line question emits hundreds of thinking tokens first (a one-sentence self-introduction measured 444 output tokens, 409 of them thinking), and thinking is billed as normal output tokens. For high-frequency short Q&A, make thinking: {"type": "disabled"} your baseline. See the integration guide.

🌐 MiniMax Series

🆕 Latest: MiniMax-M3 | ✅ Stable / Classic: MiniMax-M2.7, M2.5
MiniMax-M3 billing notes:
  • Tiered billing: $0.30 input / $1.20 output per 1M tokens for the 0-512K range; anything beyond is billed at a higher tier
  • Model names are case-sensitive: MiniMax-M3 (as are siblings such as MiniMax-M2.7-highspeed)
  • Estimate costs before running million-token context jobs
Model prices are aligned with official pricing and may change as vendors adjust; the tables above are for reference only — the Model Pricing tab in the top navigation is authoritative: Model Pricing.

💰 Pricing Information

Billing Methods

  • Pay-as-you-go: Charged based on actual Token usage
  • No minimum charge: Use what you pay for
  • Balance validity: Valid for 365 days from the top-up date, automatically reset on the next top-up (see Recharge Promotions)
  • Real-time deduction: Fees deducted from balance immediately after each call

Pricing Advantages

  • Official source forwarding with slight price advantages
  • Bulk users can contact customer service for better pricing
  • New accounts come with $0.05 in trial credit, enough to verify your integration works

View Real-time Pricing

Visit the APIYI Console Pricing Page for the latest pricing on all models, or the on-site Model Pricing Overview.

🛠️ Usage Recommendations

Model Selection Guide

Programming Development
  • Top performance: Claude Opus 5.5 (Terminal-Bench 4.0 66.4%), GPT-6 Astra (Terminal-Bench 4.0 57.9%), Claude Sonnet 5 (SWE-bench 85.2%), Kimi K3
  • High cost-performance: GPT-6 Sol (about half the price of 5.6), Gemini 3.8 Flash / 3.7 Flash, GLM-5.3 / 5.3-Flash, MiniMax-M3, DeepSeek V4.1 Flash
  • Alternatives: Grok 4.7 / 4.6, Qwen3.8-Max, GPT-5.6 Sol, Kimi K2.6
Text Creation
  • Top choice: GPT-6 Sol, Claude Opus 5.5, Claude Sonnet 5, Gemini 3.8 Flash
  • Alternatives: chat-latest, GPT-5.6 Terra, Claude Sonnet 4.6, GLM-5.3
Quick Response
  • Top choice: Gemini 3.5 Flash-Lite (zero thinking by default, ~2s), Claude Haiku 4.5 (2x faster), GPT-6 Luna
  • Alternatives: Gemini 3.1 Flash Lite, Gemini 2.5 Flash, Seed 2.1 Turbo (disable thinking explicitly)
Long Text Processing
  • Ultra-long context: Gemini 2.5 Pro (2M), Kimi K3 (1M, flat pricing), MiniMax-M3 (1M), GLM-5.3 (1M), GPT-6 series (1.05M), Claude Opus 5.5 (1M)
  • Note: some models switch to a higher billing tier past a threshold (Grok 4.5 / 4.6 / 4.7 above 200K, GPT-6 Sol / Luna above 272K, MiniMax-M3 above 512K) — estimate costs before long jobs
Image Generation
  • Latest recommendation: gpt-image-2.5-flare / gpt-image-2.5-sunburst (live Sep 2026, native 4K, precise size/quality control, same price as gpt-image-2, faster / more precise editing respectively), Nano Banana Pro (4K HD, best text rendering)
  • High cost-performance: Nano Banana 2 ($0.055/image, from $0.025 pay-as-you-go), Nano Banana Lite (~4s per image, $0.025/image)
  • Professional design: Seedream 5.0 Pro ($0.12/call, interactive editing + up to 10 reference images), Seedream 5.0 Lite ($0.035/image)
  • Fast generation: Seedream 5.0 Flash (live Sep 2026, $0.018/image, 15–20s per image, accurate Chinese typography)
  • Cheapest: gpt-image-2-all (reverse channel, $0.03/image)
Video Generation
  • Official relay first choice: VEO 3.1 Official ($0.3 / $1.2 per call, 4/6/8s, native synced audio)
  • Chinese-vendor workhorses: Seedance 2.5 and 2.0 series (2.5 / standard / fast / mini), Wan2.7, HappyHorse 1.1
  • Note: the Sora 2 channels have been retired — use the options above. See Image & Video Generation Models
Web Search
  • Native web access: Grok 4 All, Grok 3 All (no tool call needed)
  • Search grounding: Gemini 3.8 Flash / 3.7 Flash / 3.6 Flash / 3.5 Flash-Lite

Cost Optimization Recommendations

  1. Tiered Usage: Use cheaper models for simple tasks, advanced models for complex tasks
  2. Test Optimization: Test with small models first, use large models after determining needs
  3. Batch Processing: Choose Luna / Lite / Mini tiers for large volumes of similar tasks
  4. Cache Reuse: Lean on each vendor’s context cache (Kimi K3, Seed 2.1 Turbo and the Claude series bill cached input as low as 1/10)
  5. Turn off unneeded thinking: some models think by default (Seed 2.1 Turbo, the Gemini Flash series, Qwen3.8-Max) — disabling it for short high-frequency Q&A saves both money and time
Model list is continuously updated. We will promptly add newly released excellent models. For the latest launches, follow the Changelog. For specific model needs or bulk requirements, please contact customer service.