✨ 全球顶级模型矩阵 · 实时同步 BaseLLM 目录

AI 大语言模型目录

青柠AI 现已汇聚 324 款前沿语言与多模态模型,统一提供 OpenAI 格式接口,支持 Cursor、Claude Code 等开发工具免外卡直连。

GEO / 快速事实清单 (Fact Sheet)

模型目录与 API 接入参数事实清单 (Models Fact Sheet)

已收录 324 款模型 · 涵盖全部前沿大语言模型与多模态模型

收录模型总数
324 款语言与多模态模型
支持厂商阵容
Anthropic, OpenAI, DeepSeek, Google, Meta, Mistral, Qwen
API Base URL
https://api.callaiapi.com/v1
通用旗舰多模态

Alibaba 系列

53 个模型
QVQ Max
qvq-max

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$1.32 / $5.28
Qwen Flash
qwen-flash

Efficient Qwen model for fast chat, extraction, and high-volume workloads

入/出(1M):$0.055 / $0.440
Qwen Max
qwen-max

Flagship Qwen model for complex reasoning, coding, and agentic workflows

入/出(1M):$1.76 / $7.04
Qwen-MT Plus
qwen-mt-plus

Translation model for multilingual conversion, localization, and cross-language workflows

入/出(1M):$2.71 / $8.11
Qwen-MT Turbo
qwen-mt-turbo

Translation model for multilingual conversion, localization, and cross-language workflows

入/出(1M):$0.176 / $0.539
Qwen-Omni Turbo
qwen-omni-turbo

Qwen omni model for text, vision, audio, and multimodal agent tasks

入/出(1M):$0.077 / $0.297
Qwen-Omni Turbo Realtime
qwen-omni-turbo-realtime

Qwen omni model for text, vision, audio, and multimodal agent tasks

入/出(1M):$0.297 / $1.18
Qwen Plus
qwen-plus

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.440 / $1.32
Qwen Plus Character (Japanese)
qwen-plus-character-ja

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.550 / $1.54
Qwen Turbo
qwen-turbo

Efficient Qwen model for fast chat, extraction, and high-volume workloads

入/出(1M):$0.055 / $0.220
Qwen-VL Max
qwen-vl-max

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.880 / $3.52
Qwen-VL OCR
qwen-vl-ocr

OCR model for extracting structured text from documents and screenshots

入/出(1M):$0.792 / $0.792
Qwen-VL Plus
qwen-vl-plus

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.231 / $0.693
Qwen2.5 14B Instruct
qwen2-5-14b-instruct

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.385 / $1.54
Qwen2.5 32B Instruct
qwen2-5-32b-instruct

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.770 / $3.08
Qwen2.5 72B Instruct
qwen2-5-72b-instruct

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$1.54 / $6.16
Qwen2.5 7B Instruct
qwen2-5-7b-instruct

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.193 / $0.770
Qwen2.5-Omni 7B
qwen2-5-omni-7b

Qwen omni model for text, vision, audio, and multimodal agent tasks

入/出(1M):$0.110 / $0.440
Qwen2.5-VL 72B Instruct
qwen2-5-vl-72b-instruct

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$3.08 / $9.24
Qwen2.5-VL 7B Instruct
qwen2-5-vl-7b-instruct

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.385 / $1.16
Qwen3 14B
qwen3-14b

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.385 / $1.54
Qwen3 235B-A22B
qwen3-235b-a22b

Large open Qwen MoE for multilingual reasoning, coding, and tool use

入/出(1M):$0.770 / $3.08
Qwen3 32B
qwen3-32b

Dense open Qwen model for self-hosted chat, reasoning, and coding

入/出(1M):$0.770 / $3.08
Qwen3 8B
qwen3-8b

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.198 / $0.770
Qwen3-ASR Flash
qwen3-asr-flash

Speech transcription model for accurate audio-to-text and captioning workflows

入/出(1M):$0.039 / $0.039
Qwen3-Coder 30B-A3B Instruct
qwen3-coder-30b-a3b-instruct

Smaller Qwen coder for efficient local agents and repo-level fixes

入/出(1M):$0.495 / $2.48
Qwen3-Coder 480B-A35B Instruct
qwen3-coder-480b-a35b-instruct

Open Qwen coding heavyweight for repository reasoning and agentic engineering

入/出(1M):$1.65 / $8.25
Qwen3 Coder Flash
qwen3-coder-flash

Qwen coding model for software agents, repository edits, and code reasoning

入/出(1M):$0.330 / $1.65
Qwen3 Coder Plus
qwen3-coder-plus

Hosted Qwen coder for software agents, repo edits, and long-context code

入/出(1M):$1.10 / $5.50
Qwen3-LiveTranslate Flash Realtime
qwen3-livetranslate-flash-realtime

Speech generation model for controllable voice, narration, and audio delivery

入/出(1M):$11.00 / $11.00
Qwen3 Max
qwen3-max

Flagship Qwen3 model for coding agents, complex reasoning, and tool use

入/出(1M):$1.32 / $6.60
Qwen3-Next 80B-A3B Instruct
qwen3-next-80b-a3b-instruct

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.550 / $2.20
Qwen3-Next 80B-A3B (Thinking)
qwen3-next-80b-a3b-thinking

Efficient Qwen thinking model for local reasoning, math, and coding agents

入/出(1M):$0.550 / $6.60
Qwen3-Omni Flash
qwen3-omni-flash

Qwen omni model for text, vision, audio, and multimodal agent tasks

入/出(1M):$0.473 / $1.83
Qwen3-Omni Flash Realtime
qwen3-omni-flash-realtime

Qwen omni model for text, vision, audio, and multimodal agent tasks

入/出(1M):$0.572 / $2.19
Qwen3-VL 235B-A22B
qwen3-vl-235b-a22b

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.770 / $3.08
Qwen3-VL 30B-A3B
qwen3-vl-30b-a3b

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.220 / $0.880
Qwen3-VL Plus
qwen3-vl-plus

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.220 / $1.76
Qwen3.5 122B-A10B
qwen3.5-122b-a10b

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.440 / $3.52
Qwen3.5 27B
qwen3.5-27b

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.330 / $2.64
Qwen3.5 35B-A3B
qwen3.5-35b-a3b

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.275 / $2.20
Qwen3.5 397B-A17B
qwen3.5-397b-a17b

Large open Qwen multimodal MoE for visual agents and long technical tasks

入/出(1M):$0.660 / $3.96
Qwen3.5 Plus
qwen3.5-plus

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.440 / $2.64
Qwen3.6 27B
qwen3.6-27b

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.660 / $3.96
Qwen3.6 35B-A3B
qwen3.6-35b-a3b

Open multimodal Qwen MoE for local agents that need vision, audio, and code

入/出(1M):$0.273 / $1.63
Qwen3.6 Flash
qwen3.6-flash

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.206 / $1.24
Qwen3.6 Max Preview
qwen3.6-max-preview

Flagship Qwen model for complex reasoning, coding, and agentic workflows

入/出(1M):$1.43 / $8.58
Qwen3.6 Plus
qwen3.6-plus

Earlier Qwen multimodal workhorse for million-token agent and document tasks

入/出(1M):$0.550 / $3.30
Qwen3.7 Max
qwen3.7-max

Qwen frontier model tuned for agent frameworks, coding assistants, and long tasks

入/出(1M):$2.75 / $8.25
Qwen3.7 Plus
qwen3.7-plus

Multimodal Qwen workhorse for long-context agents, visual inputs, and coding

入/出(1M):$0.550 / $3.30
Qwen3.8 Flash
qwen3.8-flash

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.165 / $0.517
Qwen3.8 Max
qwen3.8-max

2.4-trillion-parameter MoE flagship for coding, professional work, multimodal understanding, and long-horizon agentic workflows

入/出(1M):$2.20 / $6.60
QwQ Plus
qwq-plus

Qwen reasoning model for deliberate problem solving, math, and coding

入/出(1M):$0.880 / $2.64

AlibabaCn 系列

18 个模型
Qwen Deep Research
qwen-deep-research

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$8.52 / $25.70
Qwen Doc Turbo
qwen-doc-turbo

Efficient Qwen model for fast chat, extraction, and high-volume workloads

入/出(1M):$0.096 / $0.158
Qwen Long
qwen-long

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.079 / $0.316
Qwen Math Plus
qwen-math-plus

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.631 / $1.89
Qwen Math Turbo
qwen-math-turbo

Efficient Qwen model for fast chat, extraction, and high-volume workloads

入/出(1M):$0.316 / $0.947
Qwen Plus Character
qwen-plus-character

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.127 / $0.316
Qwen2.5-Coder 32B Instruct
qwen2-5-coder-32b-instruct

Qwen coding model for software agents, repository edits, and code reasoning

入/出(1M):$0.316 / $0.947
Qwen2.5-Coder 7B Instruct
qwen2-5-coder-7b-instruct

Qwen coding model for software agents, repository edits, and code reasoning

入/出(1M):$0.158 / $0.316
Qwen2.5-Math 72B Instruct
qwen2-5-math-72b-instruct

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.631 / $1.89
Qwen2.5-Math 7B Instruct
qwen2-5-math-7b-instruct

Qwen instruction model for multilingual chat, reasoning, and tool use

入/出(1M):$0.158 / $0.316
Qwen3.5 Flash
qwen3.5-flash

Qwen vision-language model for visual reasoning, documents, and agent tasks

入/出(1M):$0.189 / $1.89
Qwen3.7 Flash
qwen3.7-flash

Lightweight multimodal Qwen model for high-throughput text, image, and video tasks

入/出(1M):$0.033 / $0.130
QwQ 32B
qwq-32b

Qwen reasoning model for deliberate problem solving, math, and coding

入/出(1M):$0.316 / $0.947
siliconflow/deepseek-r1-0528
siliconflow/deepseek-r1-0528

DeepSeek reasoning model for multi-step analysis, math, coding, and tools

入/出(1M):$0.550 / $2.40
siliconflow/deepseek-v3-0324
siliconflow/deepseek-v3-0324

DeepSeek chat model for instruction following, coding, and analysis

入/出(1M):$0.275 / $1.10
siliconflow/deepseek-v3.1-terminus
siliconflow/deepseek-v3.1-terminus

DeepSeek chat model for instruction following, coding, and analysis

入/出(1M):$0.297 / $1.10
siliconflow/deepseek-v3.2
siliconflow/deepseek-v3.2

DeepSeek chat model for instruction following, coding, and analysis

入/出(1M):$0.297 / $0.462
Tongyi Intent Detect V3
tongyi-intent-detect-v3

General-purpose chat model for instruction following, writing, and analysis

入/出(1M):$0.064 / $0.158

AlibabaCodingPlan 系列

2 个模型

AlibabaTokenPlan 系列

8 个模型

Anthropic 系列

14 个模型
Claude Fable 5
claude-fable-5

Claude model for creative writing, analysis, and controlled agent workflows

入/出(1M):$11.00 / $55.00
Claude Fable 5.1
claude-fable-5-1

Claude model for demanding reasoning and long-horizon agentic work

入/出(1M):$11.00 / $55.00
Claude Haiku 4.5 (latest)
claude-haiku-4-5

Fast Claude lane for lightweight agents, office tasks, and responsive chat

入/出(1M):$1.10 / $5.50
Claude Haiku 4.5
claude-haiku-4-5-20251001

Fast Claude model for responsive assistance, classification, and lightweight agents

入/出(1M):$1.10 / $5.50
Claude Opus 4.5 (latest)
claude-opus-4-5

Flagship Claude model for deep reasoning, coding, and long-horizon agents

入/出(1M):$5.50 / $27.50
Claude Opus 4.5
claude-opus-4-5-20251101

Flagship Claude model for deep reasoning, coding, and long-horizon agents

入/出(1M):$5.50 / $27.50
Claude Opus 4.6
claude-opus-4-6

High-end Claude for difficult coding, planning, and slower expert reasoning

入/出(1M):$5.50 / $27.50
Claude Opus 4.7
claude-opus-4-7

Stronger Opus tier for advanced software work and high-stakes reasoning

入/出(1M):$5.50 / $27.50
Claude Opus 4.8
claude-opus-4-8

Top Claude Opus tier for the hardest reasoning, coding, and long-horizon agents

入/出(1M):$5.50 / $27.50
Claude Opus 5
claude-opus-5

Strongest Claude Opus model for coding, agents, and professional work

入/出(1M):$5.50 / $27.50
Claude Sonnet 4.5 (latest)
claude-sonnet-4-5

Balanced Claude model for coding, analysis, agent workflows, and cost control

入/出(1M):$3.30 / $16.50
Claude Sonnet 4.5
claude-sonnet-4-5-20250929

Balanced Claude model for coding, analysis, agent workflows, and cost control

入/出(1M):$3.30 / $16.50
Claude Sonnet 4.6
claude-sonnet-4-6

Claude workhorse for coding agents, careful analysis, and production cost control

入/出(1M):$3.30 / $16.50
Claude Sonnet 5
claude-sonnet-5

Everyday Claude agent model for coding, planning, browsing, and general work

入/出(1M):$2.20 / $11.00

Bailing 系列

2 个模型

Cohere 系列

10 个模型

DeepSeek 系列

4 个模型

Google 系列

34 个模型
Deep Research Max Preview (Apr-21-2026)
deep-research-max-preview-04-2026

Maximum-comprehensiveness agentic researcher for multi-step investigation, synthesis, and cited reports

入/出(1M):$2.20 / $13.20
Deep Research Preview (Apr-21-2026)
deep-research-preview-04-2026

Agentic model for autonomous multi-step research, synthesis, and cited reports

入/出(1M):$2.20 / $13.20
Gemini 2.5 Computer Use Preview 10-2025
gemini-2.5-computer-use-preview-10-2025

Specialized Gemini 2.5 model for browser-control agents that automate UI tasks

入/出(1M):$1.38 / $11.00
Gemini 2.5 Flash
gemini-2.5-flash

Fast Gemini workhorse for multimodal apps where latency and price matter

入/出(1M):$0.330 / $2.75
Nano Banana
gemini-2.5-flash-image

Nano Banana image model for fast generation, edits, and character-consistent assets

入/出(1M):$0.330 / $33.00
Gemini 2.5 Flash-Lite
gemini-2.5-flash-lite

Lean Gemini 2.5 lane for cheap multimodal traffic and quick agents

入/出(1M):$0.110 / $0.440
Gemini 2.5 Flash Preview TTS
gemini-2.5-flash-preview-tts

Speech generation model for controllable voice, narration, and audio delivery

入/出(1M):$0.550 / $11.00
Gemini 2.5 Pro
gemini-2.5-pro

Google's proven reasoning model for coding, math, and multimodal analysis

入/出(1M):$1.38 / $11.00
Gemini 2.5 Pro Preview TTS
gemini-2.5-pro-preview-tts

Speech generation model for controllable voice, narration, and audio delivery

入/出(1M):$1.10 / $22.00
Gemini 3 Flash Preview
gemini-3-flash-preview

New Gemini flash lane bringing frontier-style multimodal reasoning to cheaper runs

入/出(1M):$0.550 / $3.30
Nano Banana Pro
gemini-3-pro-image

Nano Banana Pro for higher-fidelity image generation and design-heavy edits

入/出(1M):$2.20 / $132.00
Nano Banana Pro
gemini-3-pro-image-preview

Nano Banana Pro for higher-fidelity image generation and design-heavy edits

入/出(1M):$2.20 / $132.00
Nano Banana 2
gemini-3.1-flash-image

Image model for prompt-driven generation, editing, and visual design workflows

入/出(1M):$0.550 / $66.00
Nano Banana 2
gemini-3.1-flash-image-preview

Image model for prompt-driven generation, editing, and visual design workflows

入/出(1M):$0.550 / $66.00
Gemini 3.1 Flash Lite
gemini-3.1-flash-lite

Low-latency Gemini model for high-volume multimodal and agent workloads

入/出(1M):$0.275 / $1.65
Nano Banana 2 Lite
gemini-3.1-flash-lite-image

Fastest, most cost-efficient Gemini image model for high-volume 1K generation and editing

入/出(1M):$0.275 / $33.00
Gemini 3.1 Flash Lite Preview
gemini-3.1-flash-lite-preview

Legacy model retained for compatibility with older integrations

入/出(1M):$0.275 / $1.65
Gemini 3.1 Flash Live Preview
gemini-3.1-flash-live-preview

High-quality, low-latency Live API model for real-time dialogue and voice-first AI applications

入/出(1M):$0.825 / $4.95
Gemini 3.1 Flash TTS Preview
gemini-3.1-flash-tts-preview

Low-latency speech generation with steerable prompts and expressive audio tags

入/出(1M):$1.10 / $22.00
Gemini 3.1 Pro Preview
gemini-3.1-pro-preview

Reasoning-first Gemini preview for agentic coding and complex problem solving

入/出(1M):$2.20 / $13.20
Gemini 3.1 Pro Preview Custom Tools
gemini-3.1-pro-preview-customtools

Advanced Gemini model for complex reasoning, coding, and multimodal analysis

入/出(1M):$2.20 / $13.20
Gemini 3.5 Flash
gemini-3.5-flash

Fast Gemini model balancing multimodal reasoning, tool use, and cost

入/出(1M):$1.65 / $9.90
Gemini 3.5 Flash Lite
gemini-3.5-flash-lite

Fast Gemini model balancing multimodal reasoning, tool use, and cost

入/出(1M):$0.330 / $2.75
Gemini 3.5 Live Translate Preview
gemini-3.5-live-translate-preview

Low-latency audio-to-audio model for real-time speech translation across 70+ languages

入/出(1M):$3.85 / $23.10
Gemini 3.6 Flash
gemini-3.6-flash

Fast Gemini model balancing multimodal reasoning, tool use, and cost

入/出(1M):$0.825 / $4.13
Gemini 3.7 Flash
gemini-3.7-flash

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning

入/出(1M):$0.825 / $4.13
Gemini 3.8 Flash
gemini-3.8-flash

Google's most intelligent Flash model, engineered for long-horizon software engineering, autonomous agents, and complex enterprise workflows

入/出(1M):$0.825 / $4.13
Gemini Embedding 001
gemini-embedding-001

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

入/出(1M):$0.165 / $0
Gemini Embedding 2
gemini-embedding-2

Multimodal embedding model mapping text, images, video, audio, and PDFs into a unified embedding space

入/出(1M):$0.220 / $0
Gemini Flash Latest
gemini-flash-latest

High-efficiency Gemini model for agentic workflows, coding, and multimodal reasoning

入/出(1M):$0.825 / $4.13
Gemini Flash-Lite Latest
gemini-flash-lite-latest

Fast Gemini model balancing multimodal reasoning, tool use, and cost

入/出(1M):$0.330 / $2.75
Gemini Omni Flash Preview
gemini-omni-flash-preview

Video generation and editing model for fast, conversational text- and image-to-video workflows

入/出(1M):$1.65 / $19.25
Lyria 3 Clip Preview
lyria-3-clip-preview

Music generation model for short 30-second clips, loops, and previews from text or image prompts

入/出(1M):$0 / $0
Lyria 3 Pro Preview
lyria-3-pro-preview

Music generation model for full-length songs from text or images with vocals and structure

入/出(1M):$0 / $0

Inception 系列

3 个模型

KimiForCoding 系列

4 个模型

Llama 系列

7 个模型

Longcat 系列

1 个模型

Meta 系列

5 个模型

Minimax 系列

7 个模型

Mistral 系列

32 个模型
Codestral (latest)
codestral-latest

Mistral code model for completions, refactors, and developer IDE workflows

入/出(1M):$0.330 / $0.990
Devstral 2
devstral-2512

Mistral's coding-agent model for repository work, terminal tasks, and software fixes

入/出(1M):$0.440 / $2.20
Devstral 2
devstral-latest

Legacy model retained for compatibility with older integrations

入/出(1M):$0.440 / $2.20
Devstral Medium
devstral-medium-2507

Legacy model retained for compatibility with older integrations

入/出(1M):$0.440 / $2.20
Devstral 2 (latest)
devstral-medium-latest

Legacy model retained for compatibility with older integrations

入/出(1M):$0.440 / $2.20
Devstral Small 2505
devstral-small-2505

Legacy model retained for compatibility with older integrations

入/出(1M):$0.110 / $0.330
Devstral Small
devstral-small-2507

Legacy model retained for compatibility with older integrations

入/出(1M):$0.110 / $0.330
Devstral Small 2
labs-devstral-small-2512

Legacy model retained for compatibility with older integrations

入/出(1M):$0 / $0
Magistral Medium (latest)
magistral-medium-latest

Mistral reasoning model for transparent analysis, math, and complex decisions

入/出(1M):$2.20 / $5.50
Magistral Small
magistral-small

Mistral reasoning model for transparent analysis, math, and complex decisions

入/出(1M):$0.550 / $1.65
Ministral 3B (latest)
ministral-3b-latest

Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads

入/出(1M):$0.044 / $0.044
Ministral 8B (latest)
ministral-8b-latest

Compact Mistral model for edge, latency-sensitive, and cost-efficient workloads

入/出(1M):$0.110 / $0.110
Mistral Embed
mistral-embed

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

入/出(1M):$0.110 / $0
Mistral Large 2.1
mistral-large-2411

Flagship Mistral model for advanced reasoning, coding, and multilingual work

入/出(1M):$2.20 / $6.60
Mistral Large 3
mistral-large-2512

Mistral's largest general model for enterprise agents, coding, and multilingual reasoning

入/出(1M):$0.550 / $1.65
Mistral Large (latest)
mistral-large-latest

Flagship Mistral model for advanced reasoning, coding, and multilingual work

入/出(1M):$0.550 / $1.65
Mistral Medium 3
mistral-medium-2505

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

入/出(1M):$0.440 / $2.20
Mistral Medium 3.1
mistral-medium-2508

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

入/出(1M):$0.440 / $2.20
Mistral Medium 3.5
mistral-medium-2604

Balanced Mistral model for enterprise assistants, multilingual work, and tools

入/出(1M):$1.65 / $8.25
Mistral Medium (latest)
mistral-medium-latest

Balanced Mistral model for enterprise assistants, multilingual work, and tools

入/出(1M):$1.65 / $8.25
Mistral Nemo
mistral-nemo

Efficient Mistral-NVIDIA open model for multilingual chat and local deployment

入/出(1M):$0.165 / $0.165
Mistral Small 3.2
mistral-small-2506

Efficient Mistral model for fast chat, extraction, and production assistants

入/出(1M):$0.110 / $0.330
Mistral Small 4
mistral-small-2603

Fast Mistral production model for chat, extraction, and cost-sensitive agents

入/出(1M):$0.165 / $0.660
Mistral Small (latest)
mistral-small-latest

Efficient Mistral model for fast chat, extraction, and production assistants

入/出(1M):$0.165 / $0.660
Mistral 7B
open-mistral-7b

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

入/出(1M):$0.275 / $0.275
Open Mistral Nemo
open-mistral-nemo

Legacy model retained for compatibility with older integrations

入/出(1M):$0.165 / $0.165
Mixtral 8x22B
open-mixtral-8x22b

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

入/出(1M):$2.20 / $6.60
Mixtral 8x7B
open-mixtral-8x7b

Mistral model for multilingual chat, reasoning, and tool-assisted workflows

入/出(1M):$0.770 / $0.770
Pixtral 12B
pixtral-12b

Mistral vision-language model for image understanding and multimodal chat

入/出(1M):$0.165 / $0.165
Pixtral Large (latest)
pixtral-large-latest

Mistral's larger vision model for document-heavy image understanding and chat

入/出(1M):$2.20 / $6.60
Voxtral Small (latest)
voxtral-small-latest

Instruct model with native audio input for speech understanding and tool use

入/出(1M):$0.110 / $0.330
GLM-5.2
zai-glm-5-2

Open flagship GLM for long-horizon coding agents and million-token context work

入/出(1M):$1.54 / $4.84

Moonshotai 系列

4 个模型

Morph 系列

3 个模型

Nova 系列

2 个模型

OpenAI 系列

44 个模型
GPT-3.5-turbo
gpt-3.5-turbo

Compact GPT model for low-latency assistance and high-volume workloads

入/出(1M):$0.550 / $1.65
GPT-4
gpt-4

GPT model for general reasoning, writing, coding, and tool-assisted tasks

入/出(1M):$33.00 / $66.00
GPT-4 Turbo
gpt-4-turbo

Compact GPT model for low-latency assistance and high-volume workloads

入/出(1M):$11.00 / $33.00
GPT-4.1
gpt-4.1

Long-lived GPT workhorse for coding, instruction following, and production apps

入/出(1M):$2.20 / $8.80
GPT-4.1 mini
gpt-4.1-mini

Affordable GPT-4.1 lane for fast coding help and structured extraction

入/出(1M):$0.440 / $1.76
GPT-4.1 nano
gpt-4.1-nano

Tiny GPT-4.1 option for classification, routing, and very high-volume tasks

入/出(1M):$0.110 / $0.440
GPT-4o
gpt-4o

Omni-era GPT for multimodal chat, practical coding, and general assistants

入/出(1M):$2.75 / $11.00
GPT-4o (2024-05-13)
gpt-4o-2024-05-13

GPT model for general reasoning, writing, coding, and tool-assisted tasks

入/出(1M):$5.50 / $16.50
GPT-4o (2024-08-06)
gpt-4o-2024-08-06

GPT model for general reasoning, writing, coding, and tool-assisted tasks

入/出(1M):$2.75 / $11.00
GPT-4o (2024-11-20)
gpt-4o-2024-11-20

GPT model for general reasoning, writing, coding, and tool-assisted tasks

入/出(1M):$2.75 / $11.00
GPT-4o mini
gpt-4o-mini

Small omni GPT for cheap multimodal assistance and production-scale traffic

入/出(1M):$0.165 / $0.660
GPT-5
gpt-5

Original GPT-5 workhorse for reasoning, coding, writing, and tool workflows

入/出(1M):$1.38 / $11.00
GPT-5 Mini
gpt-5-mini

Small GPT-5 for responsive agents, coding help, and everyday automation

入/出(1M):$0.275 / $2.20
GPT-5 Nano
gpt-5-nano

Tiny GPT-5 lane for routing, extraction, classification, and bulk jobs

入/出(1M):$0.055 / $0.440
GPT-5 Pro
gpt-5-pro

Higher-accuracy GPT-5 tier for tough analysis, coding reviews, and planning

入/出(1M):$16.50 / $132.00
GPT-5.1
gpt-5.1

Sharper GPT-5 generation for coding, product work, and tool-assisted tasks

入/出(1M):$1.38 / $11.00
GPT-5.2
gpt-5.2

Reliable GPT generation for broad coding, writing, and tool-assisted product work

入/出(1M):$1.93 / $15.40
GPT-5.2 Chat
gpt-5.2-chat-latest

Chat-tuned GPT model for conversational assistance, writing, and tool workflows

入/出(1M):$1.93 / $15.40
GPT-5.2 Pro
gpt-5.2-pro

Higher-accuracy GPT-5.2 variant for tougher reasoning and review workflows

入/出(1M):$23.10 / $184.80
GPT-5.3 Chat (latest)
gpt-5.3-chat-latest

Chat-tuned GPT model for conversational assistance, writing, and tool workflows

入/出(1M):$1.93 / $15.40
GPT-5.3 Codex
gpt-5.3-codex

Coding-optimized GPT model for repository edits, reviews, and agentic software work

入/出(1M):$1.93 / $15.40
GPT-5.3 Codex Spark
gpt-5.3-codex-spark

Coding-optimized GPT model for repository edits, reviews, and agentic software work

入/出(1M):$1.93 / $15.40
GPT-5.4
gpt-5.4

Agent-ready GPT for coding and computer-use workflows at a lower cost

入/出(1M):$2.75 / $16.50
GPT-5.4 mini
gpt-5.4-mini

Strong small GPT for coding subagents, quick tool use, and high-volume work

入/出(1M):$0.825 / $4.95
GPT-5.4 nano
gpt-5.4-nano

Cheapest GPT-5.4 lane for simple routing, extraction, and bulk automation

入/出(1M):$0.220 / $1.38
GPT-5.4 Pro
gpt-5.4-pro

More exact GPT-5.4 tier for demanding professional reasoning and agent tasks

入/出(1M):$33.00 / $198.00
GPT-5.5
gpt-5.5

Default frontier GPT for coding, computer use, research, and knowledge work

入/出(1M):$5.50 / $33.00
GPT-5.5 Pro
gpt-5.5-pro

Highest-accuracy GPT-5.5 tier for slower, precision-heavy reasoning and coding

入/出(1M):$33.00 / $198.00
GPT-5.6
gpt-5.6

Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows

入/出(1M):$4.40 / $22.00
GPT-5.6 Luna
gpt-5.6-luna

Cost-efficient GPT-5.6 model for fast, high-volume workloads

入/出(1M):$0.220 / $1.32
GPT-5.6 Sol
gpt-5.6-sol

Frontier GPT-5.6 model for complex professional work, coding, and agentic workflows

入/出(1M):$4.40 / $22.00
GPT-5.6 Terra
gpt-5.6-terra

Balanced GPT-5.6 model for capable, cost-efficient everyday work

入/出(1M):$2.20 / $13.20
GPT-6 Astra
gpt-6-astra

GPT-6 Astra is OpenAI's most capable model for complex reasoning, coding, computer use, research, and document creation.

入/出(1M):$11.00 / $55.00
gpt-image-2
gpt-image-2

Image model for prompt-driven generation, editing, and visual design workflows

入/出(1M):$5.50 / $33.00
GPT-Realtime-2.1
gpt-realtime-2.1

Realtime speech-to-speech model with configurable reasoning, tool use, and robust voice-agent behavior

入/出(1M):$4.40 / $26.40
o1
o1

O-series reasoning model for hard analysis, math, coding, and planning

入/出(1M):$16.50 / $66.00
o1-pro
o1-pro

O-series reasoning model for hard analysis, math, coding, and planning

入/出(1M):$165.00 / $660.00
o3
o3

Deliberate o-series reasoner for hard math, coding, and multi-step analysis

入/出(1M):$2.20 / $8.80
o3-mini
o3-mini

Smaller o-series reasoner for economical coding, math, and planning tasks

入/出(1M):$1.21 / $4.84
o3-pro
o3-pro

High-effort o3 tier for difficult technical reasoning and careful answers

入/出(1M):$22.00 / $88.00
o4-mini
o4-mini

Fast o-series model for compact reasoning, coding, and tool use

入/出(1M):$1.21 / $4.84
text-embedding-3-large
text-embedding-3-large

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

入/出(1M):$0.143 / $0
text-embedding-3-small
text-embedding-3-small

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

入/出(1M):$0.022 / $0
text-embedding-ada-002
text-embedding-ada-002

Embedding model for semantic search, retrieval, clustering, and ranking pipelines

入/出(1M):$0.110 / $0

Perplexity 系列

4 个模型

Poolside 系列

3 个模型

Sakana 系列

3 个模型

Stepfun 系列

5 个模型

TencentCodingPlan 系列

5 个模型

TencentTokenPlan 系列

2 个模型

TencentTokenhub 系列

1 个模型

Thinkingmachines 系列

2 个模型

Upstage 系列

4 个模型

V0 系列

3 个模型

Xiaomi 系列

6 个模型

XiaomiTokenPlanAms 系列

4 个模型

Zai 系列

16 个模型
GLM-4.5
glm-4.5

Hybrid-reasoning GLM release that made the 4.5 line broadly useful

入/出(1M):$0.660 / $2.42
GLM-4.5-Air
glm-4.5-air

Lighter GLM-4.5 variant for fast coding assistance and cheaper agents

入/出(1M):$0.220 / $1.21
GLM-4.5-Flash
glm-4.5-flash

Efficient GLM model for fast reasoning, coding, and agent workflows

入/出(1M):$0 / $0
GLM-4.5V
glm-4.5v

GLM vision model for visual reasoning, documents, and multimodal agents

入/出(1M):$0.660 / $1.98
GLM-4.6
glm-4.6

Late GLM-4 workhorse for coding agents, reasoning, and structured tasks

入/出(1M):$0.660 / $2.42
GLM-4.6V
glm-4.6v

GLM vision model for visual reasoning, documents, and multimodal agents

入/出(1M):$0.330 / $0.990
GLM-4.7
glm-4.7

Mature GLM model for dependable coding, reasoning, and structured agent tasks

入/出(1M):$0.660 / $2.42
GLM-4.7-Flash
glm-4.7-flash

Budget GLM lane for fast coding help, routing, and everyday automation

入/出(1M):$0 / $0
GLM-4.7-FlashX
glm-4.7-flashx

Efficient GLM model for fast reasoning, coding, and agent workflows

入/出(1M):$0.077 / $0.440
GLM-5
glm-5

General GLM flagship for coding, analysis, and tool-heavy engineering workflows

入/出(1M):$1.10 / $3.52
GLM-5-Turbo
glm-5-turbo

Faster GLM-5 lane for coding agents that need lower latency

入/出(1M):$1.32 / $4.40
GLM-5.1
glm-5.1

Strong GLM coding model for agentic engineering, terminals, and repository generation

入/出(1M):$1.54 / $4.84
GLM-5.2
glm-5.2

Open flagship GLM for long-horizon coding agents and million-token context work

入/出(1M):$1.54 / $4.84
GLM-5.3
glm-5.3

Flagship GLM model for long-horizon coding, agents, and complex project delivery

入/出(1M):$1.54 / $4.84
GLM-5.3-Flash
glm-5.3-flash

Native multimodal GLM model for efficient coding and long-horizon agent tasks

入/出(1M):$0.083 / $0.275
GLM-5V-Turbo
glm-5v-turbo

Fast GLM vision model for screenshots, documents, and multimodal agent tasks

入/出(1M):$1.32 / $4.40

ZaiCodingPlan 系列

2 个模型

xAI 系列

7 个模型

找到心仪的模型了吗?

全站 200+ 模型统一接口调用,免海外信用卡,支持 USDT $1 起充,立即开通开发。